Benchmarks
80.9% of tasks come in at exactly the planned size.
Measured across 3,827 completed builds in 397 cycles on PAPI itself. When an estimate is off, it is usually too big: 15.2% of tasks came in smaller than planned and 3.9% came in bigger.
What the estimates are
- XS
- S
- M
- L
- XL
Effort sizes. Not story points or hours.
Plan sizes the task
When you plan a cycle, your AI tool gives each task a size from its scope and the files it will likely touch. It also sees how earlier estimates turned out, by module.
Build reports the actual size
When the build is done, the AI that built it reports the size the task turned out to be, on the same scale. The size is the AI’s own judgement. Nothing is timed.
This page compares the two
Sizes become points (XS=1, S=2, M=3, L=5, XL=8) only so they can be added up and compared.
How the actual size compared, across 3,827 tasks
- Same size as estimated· 3,097
- 80.9%
- Smaller than estimated· 580
- 15.2%
- Bigger than estimated· 150
- 3.9%
The headline counts only tasks that came in at exactly the planned size. Counting smaller tasks too gives 96.1% within estimate. When an estimate misses, it is usually too big: 580 tasks came in smaller and 150 tasks came in bigger.
Cycle by cycle
Show the numbers as a table
| Cycle | Tasks | Same size | Smaller | Bigger |
|---|---|---|---|---|
| 401 | 1 | 1 | 0 | 0 |
| 400 | 10 | 10 | 0 | 0 |
| 399 | 12 | 12 | 0 | 0 |
| 398 | 14 | 14 | 0 | 0 |
| 397 | 1 | 1 | 0 | 0 |
| 396 | 12 | 11 | 1 | 0 |
| 395 | 25 | 23 | 2 | 0 |
| 394 | 13 | 11 | 0 | 2 |
| 393 | 6 | 6 | 0 | 0 |
| 392 | 9 | 8 | 0 | 1 |
| 391 | 13 | 11 | 0 | 2 |
| 390 | 11 | 11 | 0 | 0 |
| 389 | 11 | 10 | 0 | 1 |
| 388 | 11 | 10 | 1 | 0 |
| 387 | 13 | 11 | 1 | 1 |
| 386 | 15 | 14 | 1 | 0 |
| 385 | 10 | 4 | 5 | 1 |
| 384 | 30 | 16 | 12 | 2 |
| 383 | 14 | 8 | 5 | 1 |
| 382 | 20 | 17 | 1 | 2 |
| 381 | 48 | 42 | 3 | 3 |
| 380 | 12 | 7 | 2 | 3 |
| 379 | 17 | 12 | 5 | 0 |
| 378 | 11 | 10 | 1 | 0 |
| 377 | 10 | 10 | 0 | 0 |
| 376 | 13 | 11 | 2 | 0 |
| 375 | 9 | 7 | 2 | 0 |
| 374 | 9 | 9 | 0 | 0 |
| 373 | 12 | 12 | 0 | 0 |
| 372 | 14 | 11 | 0 | 3 |
| 371 | 8 | 8 | 0 | 0 |
| 370 | 20 | 20 | 0 | 0 |
| 369 | 11 | 10 | 1 | 0 |
| 368 | 10 | 10 | 0 | 0 |
| 367 | 13 | 13 | 0 | 0 |
| 366 | 12 | 10 | 0 | 2 |
| 365 | 14 | 12 | 0 | 2 |
| 364 | 26 | 17 | 1 | 8 |
| 363 | 25 | 22 | 0 | 3 |
| 362 | 29 | 21 | 3 | 5 |
How it compares
McConnell's industry reference for software estimation puts unaided projects at roughly even odds of landing within estimate at all, with median schedule overruns of 30 to 50% (Software Estimation: Demystifying the Black Art, 2006). That baseline measures time against a schedule, while PAPI's figure compares a reported size with a planned size and counts smaller-than-planned tasks as hits, so treat it as a rough reference point. PAPI's figure covers 397 cycles of one project: PAPI itself.
Methodology
Source: PapiUI's own build_reports table in PAPI's own database. Each row is a completed build with the planner's estimated effort and the builder's reported actual effort.
- Cycle range: 0 to 401 inclusive (effort tracking began at cycle 17).
- Estimate: set during plan by the planner (the AI running in your AI tool), written into each task's build handoff. PAPI runs no AI of its own.
- Actual: reported by the builder's AI in the build report when the build completes. It is a size the AI judges, not a measured time.
- Effort sizes map to points: XS=1, S=2, M=3, L=5, XL=8. Not story points or hours; the points exist only to add sizes up and compare them.
- “Within estimate” means actual_pts is at or below estimated_pts, so a task that came in smaller than planned counts as within.
- Filter:
completed = 'Yes'. Cancelled, abandoned, or partially-completed builds are excluded. - Result: 3,677 of 3,827 = 96.1% within estimate. Same size: 3,097. Smaller: 580. Bigger: 150.
What this means in practice
These numbers show how often a task's planned size matched the size it turned out to be, and which way the misses went. They don't show whether the right thing got built; reviews check that. Most misses are tasks sized too big, so a plan tends to overstate the work rather than understate it. Both sizes are judgements made by AI, so read them as a consistent yardstick within one project rather than as timings.