Do story points still work when agents do the building?

Stefan-Iulian Tesoi · · 6 min read

A two-pan balance on a weathered counter, one pan empty and the other holding a rusted nut standing in for a calibrated weight — sizing against whatever reference happened to be to hand

Story points still work, but not as a measure of effort. What they usefully track now is specification risk — the chance an item comes back needing a decision nobody made — and on that axis the expensive work is not the technically hard work. It is the vague work, which is almost always estimated low.

That inversion is the problem with carrying the old scale forward: it does not read high where the risk is.

What do story points measure now?

How likely an item is to be built twice. That is the quantity that still varies once a coding agent is doing the building, and it is not the quantity the scale was calibrated for.

A point was always relative rather than absolute. Martin Fowler's entry on story points records why they exist: teams estimating in staff-hours struggled to produce useful numbers, so the practice became sizing each story against stories already sized. The names people gave the unit were honest about its precision — Joseph Pelrine's gummi bears, Josh Kerievsky's Nebulous Units of Time.

What went into that comparison was never one thing. A five bundled how much code it was, how much of the system it touched, and how much nobody had decided yet. Teams read the total as effort because effort was the part they could feel. Take execution time down to minutes and the first term collapses, leaving the term that was quietly doing the forecasting.

Effort was the wrong axis all along

The scale's own top value admits it. Fowler's note on using the highest number in a series is blunt: an eight really means "eight or more", and forecasting with it is unwise because "eight can turn into all sorts of numbers when it finally gets broken down" — better to say a story is too big to estimate than to record a false marker number.

So the number tracks how much of the work is visible, and a big one means the edges cannot be seen — which is why splitting stories reliably makes the total go up rather than down.

When a person did the building, that invisibility was partly absorbed. A developer holding an ambiguous ticket asked someone, decided something reasonable, or noticed the contradiction at the keyboard, and the cost surfaced as an estimate wrong by half a day. Agile estimation worked tolerably because the estimate and its error shared an axis.

An agent does not absorb it. Handed the same item it makes a plausible choice, implements it completely, and returns something satisfying the sentence it was given rather than the intent behind it. That ambiguity is resolved at review, for the price of the whole item.

Why are bad estimates wrong by a factor?

Because effort error is additive and specification error is multiplicative. One overruns; the other repeats the entire item.

The arithmetic is unforgiving once building is cheap. An item specified properly costs about forty minutes to write executably, some agent time nobody has to watch, and ten to twenty minutes of review against criteria that name a command — an hour of human time.

The same item with one open decision comes back. Fifteen minutes of review establishes that it misses the criteria, twenty more settle whether the code or the criterion was wrong, and then it is re-specified, re-executed and re-reviewed: just under two hours. A second bounce, common when the first fix addressed the symptom, takes it past two and a half.

A one-point item with an unanswered question inside it costs more human time than a five-point item with none. On the old scale, the expensive one is labelled cheap and pulled into the sprint first.

That is the factor. Nothing in the effort estimate moved — the code really was small — and the item cost two and a half times its neighbour. Estimation with AI agents goes wrong not because the numbers are imprecise, but because they are precise about a term that no longer dominates the total.

What replaces the effort estimate?

Counting the decisions an item still contains. A smaller scale, faster to apply, predicting rework instead of duration.

The question is how many choices the item leaves unanswered — not technical choices an agent can reasonably make, but product choices where a wrong guess is wrong rather than different.

Decisions leftWhat it meansWhat to do with it
NoneEvery criterion names a command and an expected resultReady. Size it small and move on
One, answerable nowA person can settle it in the planning meetingSettle it, write the answer into the item, then Ready
One, needing someone absentWaiting on a person, not on capacityOff the sprint. It is blocked, not small
Two or more, or unknownNot an item yetSplit it, or run a spike to produce the decisions

Most story points alternatives keep a number and change its meaning, which leaves everyone quietly converting back to days. This scale is deliberately not a number: nothing to sum, no velocity to derive, nothing to negotiate down. Its only output is whether an item goes on the sprint, which is the decision sprint planning now exists to make.

Three cheap practices make it work:

Laimonade drafts items in this shape and checks returned work against the criteria the item was accepted on, with a person deciding what is done — the division described in what an AI product owner actually does, and the item format in a backlog an agent can read.

What to do with historical velocity

Keep it, stop forecasting with it, and mark the date the builder changed.

Deleting the history is the wrong instinct — those sprints did deliver those points, and the average of recent completed sprints stays a fair observation of what this team, with these agents, got through.

What it cannot do is bridge the discontinuity. A series running across the week agents arrived is two measurements sharing an axis, and an average across the join describes no period at all. The honest treatment is an annotation on that week and a fresh baseline after it, with the new series carrying no forecasting value for three sprints.

Frequently asked questions

Should you re-estimate the existing backlog?

No, and re-pointing forty items is a week nobody gets back. Re-estimate an item when it is about to go on a sprint, which is when the decision-count question is cheap and its answer is still true. A number attached six months ago describes an understanding of the work that has since changed.

Can an agent estimate its own work?

It can produce a number, and the number carries the item's defect. An agent estimating an ambiguous backlog item prices the interpretation it happened to choose, so the estimate is confident exactly where the risk is. Counting open decisions needs a person to judge whether an unmade choice matters, and that judgement is the thing being estimated.

Is velocity still worth tracking?

Yes, as an observation rather than a target. It answers what this team actually completed over recent sprints, which is useful for noticing a change and useless for committing to a date. The moment it becomes a target it stops measuring anything, because the cheapest way to raise it is to point work more generously.

Do points still help with splitting?

This is the one job they still do best. A large number reliably signals that an item's edges are not visible, so "this feels like an eight" remains a good reason to break it up. The difference is that the eight diagnoses vagueness rather than forecasting duration.