Benchmarks

80.9% of tasks come in at exactly the planned size.

Measured across 3,827 completed builds in 397 cycles on PAPI itself. When an estimate is off, it is usually too big: 15.2% of tasks came in smaller than planned and 3.9% came in bigger.

See the live trend →

What the estimates are

  • XS
  • S
  • M
  • L
  • XL

Effort sizes. Not story points or hours.

  1. Plan sizes the task

    When you plan a cycle, your AI tool gives each task a size from its scope and the files it will likely touch. It also sees how earlier estimates turned out, by module.

  2. Build reports the actual size

    When the build is done, the AI that built it reports the size the task turned out to be, on the same scale. The size is the AI’s own judgement. Nothing is timed.

  3. This page compares the two

    Sizes become points (XS=1, S=2, M=3, L=5, XL=8) only so they can be added up and compared.

How the actual size compared, across 3,827 tasks

Same size as estimated· 3,097
80.9%
Smaller than estimated· 580
15.2%
Bigger than estimated· 150
3.9%

The headline counts only tasks that came in at exactly the planned size. Counting smaller tasks too gives 96.1% within estimate. When an estimate misses, it is usually too big: 580 tasks came in smaller and 150 tasks came in bigger.

Cycle by cycle

Each column is one cycle, split by how its tasks' actual sizes compared with the estimates. Across all 397 cycles: 80.9% same size, 15.2% smaller, 3.9% bigger.
Show the numbers as a table
CycleTasksSame sizeSmallerBigger
4011100
400101000
399121200
398141400
3971100
396121110
395252320
394131102
3936600
3929801
391131102
390111100
389111001
388111010
387131111
386151410
38510451
3843016122
38314851
382201712
381484233
38012723
379171250
378111010
377101000
376131120
3759720
3749900
373121200
372141103
3718800
370202000
369111010
368101000
367131300
366121002
365141202
364261718
363252203
362292135

How it compares

McConnell's industry reference for software estimation puts unaided projects at roughly even odds of landing within estimate at all, with median schedule overruns of 30 to 50% (Software Estimation: Demystifying the Black Art, 2006). That baseline measures time against a schedule, while PAPI's figure compares a reported size with a planned size and counts smaller-than-planned tasks as hits, so treat it as a rough reference point. PAPI's figure covers 397 cycles of one project: PAPI itself.

Methodology

Source: PapiUI's own build_reports table in PAPI's own database. Each row is a completed build with the planner's estimated effort and the builder's reported actual effort.

  • Cycle range: 0 to 401 inclusive (effort tracking began at cycle 17).
  • Estimate: set during plan by the planner (the AI running in your AI tool), written into each task's build handoff. PAPI runs no AI of its own.
  • Actual: reported by the builder's AI in the build report when the build completes. It is a size the AI judges, not a measured time.
  • Effort sizes map to points: XS=1, S=2, M=3, L=5, XL=8. Not story points or hours; the points exist only to add sizes up and compare them.
  • “Within estimate” means actual_pts is at or below estimated_pts, so a task that came in smaller than planned counts as within.
  • Filter: completed = 'Yes'. Cancelled, abandoned, or partially-completed builds are excluded.
  • Result: 3,677 of 3,827 = 96.1% within estimate. Same size: 3,097. Smaller: 580. Bigger: 150.

What this means in practice

These numbers show how often a task's planned size matched the size it turned out to be, and which way the misses went. They don't show whether the right thing got built; reviews check that. Most misses are tasks sized too big, so a plan tends to overstate the work rather than understate it. Both sizes are judgements made by AI, so read them as a consistent yardstick within one project rather than as timings.

Start your first cycleSee the live trendRead the manifesto