Do story points still work when agents do the building?
Stefan-Iulian Tesoi · · 6 min read

Story points still work, but not as a measure of effort. What they usefully track now is specification risk — the chance an item comes back needing a decision nobody made — and on that axis the expensive work is not the technically hard work. It is the vague work, which is almost always estimated low.
That inversion is the problem with carrying the old scale forward: it does not read high where the risk is.
What do story points measure now?
How likely an item is to be built twice. That is the quantity that still varies once a coding agent is doing the building, and it is not the quantity the scale was calibrated for.
A point was always relative rather than absolute. Martin Fowler's entry on story points records why they exist: teams estimating in staff-hours struggled to produce useful numbers, so the practice became sizing each story against stories already sized. The names people gave the unit were honest about its precision — Joseph Pelrine's gummi bears, Josh Kerievsky's Nebulous Units of Time.
What went into that comparison was never one thing. A five bundled how much code it was, how much of the system it touched, and how much nobody had decided yet. Teams read the total as effort because effort was the part they could feel. Take execution time down to minutes and the first term collapses, leaving the term that was quietly doing the forecasting.
Effort was the wrong axis all along
The scale's own top value admits it. Fowler's note on using the highest number in a series is blunt: an eight really means "eight or more", and forecasting with it is unwise because "eight can turn into all sorts of numbers when it finally gets broken down" — better to say a story is too big to estimate than to record a false marker number.
So the number tracks how much of the work is visible, and a big one means the edges cannot be seen — which is why splitting stories reliably makes the total go up rather than down.
When a person did the building, that invisibility was partly absorbed. A developer holding an ambiguous ticket asked someone, decided something reasonable, or noticed the contradiction at the keyboard, and the cost surfaced as an estimate wrong by half a day. Agile estimation worked tolerably because the estimate and its error shared an axis.
An agent does not absorb it. Handed the same item it makes a plausible choice, implements it completely, and returns something satisfying the sentence it was given rather than the intent behind it. That ambiguity is resolved at review, for the price of the whole item.
Why are bad estimates wrong by a factor?
Because effort error is additive and specification error is multiplicative. One overruns; the other repeats the entire item.
The arithmetic is unforgiving once building is cheap. An item specified properly costs about forty minutes to write executably, some agent time nobody has to watch, and ten to twenty minutes of review against criteria that name a command — an hour of human time.
The same item with one open decision comes back. Fifteen minutes of review establishes that it misses the criteria, twenty more settle whether the code or the criterion was wrong, and then it is re-specified, re-executed and re-reviewed: just under two hours. A second bounce, common when the first fix addressed the symptom, takes it past two and a half.
A one-point item with an unanswered question inside it costs more human time than a five-point item with none. On the old scale, the expensive one is labelled cheap and pulled into the sprint first.
That is the factor. Nothing in the effort estimate moved — the code really was small — and the item cost two and a half times its neighbour. Estimation with AI agents goes wrong not because the numbers are imprecise, but because they are precise about a term that no longer dominates the total.
What replaces the effort estimate?
Counting the decisions an item still contains. A smaller scale, faster to apply, predicting rework instead of duration.
The question is how many choices the item leaves unanswered — not technical choices an agent can reasonably make, but product choices where a wrong guess is wrong rather than different.
| Decisions left | What it means | What to do with it |
|---|---|---|
| None | Every criterion names a command and an expected result | Ready. Size it small and move on |
| One, answerable now | A person can settle it in the planning meeting | Settle it, write the answer into the item, then Ready |
| One, needing someone absent | Waiting on a person, not on capacity | Off the sprint. It is blocked, not small |
| Two or more, or unknown | Not an item yet | Split it, or run a spike to produce the decisions |
Most story points alternatives keep a number and change its meaning, which leaves everyone quietly converting back to days. This scale is deliberately not a number: nothing to sum, no velocity to derive, nothing to negotiate down. Its only output is whether an item goes on the sprint, which is the decision sprint planning now exists to make.
Three cheap practices make it work:
- Write the decision into the item, not a comment. An agent reads the item; what is in the thread may as well be in someone's head.
- Record why the rejected option was rejected, because the question returns in three weeks and gets answered differently.
- Treat "we'll figure it out during implementation" as one open decision. It always was, absorbed by whoever implemented it.
Laimonade drafts items in this shape and checks returned work against the criteria the item was accepted on, with a person deciding what is done — the division described in what an AI product owner actually does, and the item format in a backlog an agent can read.
What to do with historical velocity
Keep it, stop forecasting with it, and mark the date the builder changed.
Deleting the history is the wrong instinct — those sprints did deliver those points, and the average of recent completed sprints stays a fair observation of what this team, with these agents, got through.
What it cannot do is bridge the discontinuity. A series running across the week agents arrived is two measurements sharing an axis, and an average across the join describes no period at all. The honest treatment is an annotation on that week and a fresh baseline after it, with the new series carrying no forecasting value for three sprints.
Frequently asked questions
Should you re-estimate the existing backlog?
No, and re-pointing forty items is a week nobody gets back. Re-estimate an item when it is about to go on a sprint, which is when the decision-count question is cheap and its answer is still true. A number attached six months ago describes an understanding of the work that has since changed.
Can an agent estimate its own work?
It can produce a number, and the number carries the item's defect. An agent estimating an ambiguous backlog item prices the interpretation it happened to choose, so the estimate is confident exactly where the risk is. Counting open decisions needs a person to judge whether an unmade choice matters, and that judgement is the thing being estimated.
Is velocity still worth tracking?
Yes, as an observation rather than a target. It answers what this team actually completed over recent sprints, which is useful for noticing a change and useless for committing to a date. The moment it becomes a target it stops measuring anything, because the cheapest way to raise it is to point work more generously.
Do points still help with splitting?
This is the one job they still do best. A large number reliably signals that an item's edges are not visible, so "this feels like an eight" remains a good reason to break it up. The difference is that the eight diagnoses vagueness rather than forecasting duration.