What changes about sprint planning with AI agents

Stefan-Iulian Tesoi · · 9 min read

Two rusted gate-valve handwheels on old pipework, each with a rating plate at its hub — nothing moves down the line until somebody opens the gate, and the plate states what it is rated to carry

Sprint planning with AI agents stops being a capacity negotiation and becomes a readiness check. The question is no longer how many points the team can absorb, but how many items are specified well enough that something can execute them without asking a question — and that number is almost always smaller than the backlog suggests.

Everything else that changes follows from that one substitution. The meeting keeps its slot and loses its subject.

What does sprint planning with AI agents look like?

A review of which items are executable, in priority order, until the list runs out. That is the whole meeting, and it is shorter than the one it replaces.

The 2020 Scrum Guide gives sprint planning three topics: why this sprint is valuable, what can be done in it, and how the chosen work will get done. The third is where the time went. A team of five would spend an hour deciding whether a story was three points or five, because the answer determined what else fit.

With a coding agent doing the building, the third topic mostly answers itself and the second one changes meaning. "What can be done this sprint" no longer describes the team's appetite. It describes how much of the backlog has an outcome, acceptance criteria, paths and dependencies already written down. Work that does not is not a smaller commitment — it is not a commitment at all, because nothing can start it.

From capacity negotiation to readiness check

Capacity was the scarce thing, so planning rationed it. It is not the scarce thing any more, and planning has not caught up.

The old meeting had a real function. Engineering time was fixed, roughly known, and fully subscribed, so the negotiation was about which work got some of it. Estimates existed to make that trade visible. Everyone knew the numbers were soft and used them anyway, because a soft number still ranks two options against a fixed budget.

Take the budget away and the ritual keeps running with nothing underneath it. Teams adopting agile with coding agents report the same first sprint: the board is loaded to historical velocity, the agents clear it by Wednesday, and Thursday and Friday are spent either idle or building things nobody specified. The plan was not too ambitious. It was measuring the wrong resource.

Planning against velocity asks how much this team delivered last month. Planning against readiness asks how much work could start this morning. When a machine does the building, only the second question has an answer that holds.

The readiness check is unglamorous and takes about twenty minutes. Walk the top of the backlog, and for each item ask whether an agent could execute it with no further conversation — a stated outcome, criteria naming a command and an expected result, the repositories and paths, and dependencies as item ids rather than as prose. Items that pass go to Ready. Items that fail go back with the specific missing piece named, which is the actual output of the meeting: not a sprint, but a to-do list of specification work with the sprint as a by-product.

What happens to estimates and velocity?

Both survive as historical records and stop working as forecasts, because they measure a queue that is no longer the constraint.

Velocity is delivered points per completed sprint. It forecasts well when the same people who delivered last month's points are delivering this month's, at the same rate, against the same kind of work. Substitute the builder and the series breaks — not by drifting, but by measuring something else entirely. What a sprint now delivers is bounded by how many items were ready, so a velocity chart is a lagging record of specification throughput wearing the label of engineering throughput.

That does not make the number useless. In Laimonade, velocity is the delivered points of completed sprints with points carried in from the previous week subtracted, and the average of the last three completed sprints is the capacity the next sprint is planned against. It works as a smoothed observation of what this team, with these agents, actually got through. It stops working the moment it is read as a target, because the way to hit a points target is to point work optimistically.

Estimation is the weaker of the two. An agent's execution time has almost no relationship to a story point, and the variance that used to live in implementation has moved into whether the item was right. The useful estimate is no longer of effort. It is of specification risk: how likely this item is to come back needing a decision. That is a different scale with different anchors, and pretending the old one measures it is how a plan stays confidently wrong.

Which ceremonies survive, and which do not?

Planning and the retrospective survive with new subjects. The standup loses most of its content. The review is unchanged, and is the one that matters most.

CeremonyWhat it was forWhat it becomes
Sprint planningFitting work to capacityA readiness check, and a specification to-do list
StandupSurfacing blockers between peopleAgents report themselves; the time goes to review
Sprint reviewShowing what was builtUnchanged — a person still decides what is done
RetrospectiveWhy were we slower than plannedWhich specifications failed, and how
EstimationPredicting a delivery dateRanking specification risk, on a different scale

The retrospective is the one worth rewriting deliberately. Its old question assumed the constraint was execution. The productive version asks which items came back not meeting their criteria, and whether the code was wrong or the criterion was, because a bad criterion produces the same defect again next week and a bad diff does not. What that hour looks like is set out in a sprint retrospective with coding agents in the loop.

The sprint review gains importance structurally. When a person wrote the code, a second person reading it layered a quality measure on top of the author's own judgement. When an agent wrote it, that reading is the only judgement in the loop. This is why Laimonade has no tool that moves an item from In Review to Done: an agent can carry work to the edge of the board and no further, and the last step is a person accepting it against criteria written before the work started.

How long should a sprint be now?

Shorter, and for a different reason than before — a sprint is a batching interval for human review, and review is what the calendar now has to accommodate.

Two-week sprints were sized around the cost of planning and the time it took to finish something worth showing. Neither cost is what it was. When items are executable and the work comes back within hours, a fortnight's boundary means a two-week-old decision is still governing what the agents are building, and that stale plan is the expensive part.

One week is a defensible default: long enough that planning is not a daily tax, short enough that a wrong call costs days rather than a fortnight. The rollover matters more than the length — in Laimonade, sprints close on Monday in the project's timezone, unfinished items carry forward keeping their column, and completed ones are archived with their completion date, so the historical record survives a status that gets overwritten. The other honest answer is that a team with a steady specification supply and same-day review does not need the boundary at all, and is running continuous flow with a weekly ceremony bolted on.

A sprint of this shape, start to finish

What follows is arithmetic rather than a case study: the shape a one-week sprint takes for a team of four running two agents.

  1. Monday, twenty minutes. Walk the top forty backlog items. Eleven are executable as written; nine more are close but name no verifiable check; the rest are ideas. Ready gets eleven items, and the meeting's real output is the list of nine and what each is missing.
  2. Monday to Wednesday. Agents pull from Ready one item at a time, each in its own worktree, and hand work back as In Review. The eleven are gone by Wednesday afternoon, which is expected rather than a sign the sprint was too small.
  3. Throughout. A person reviews each returned item against the criteria it was accepted on, at ten to twenty minutes each when the criteria name a command — roughly three hours across the week, spread out.
  4. Thursday and Friday. The nine near-ready items get specified properly, which is next week's sprint. This is the work that sets the ceiling, and it is the work that gets skipped when the week is judged by how busy the agents looked.
  5. Friday, fifteen minutes. Retrospective on the two items that bounced. Both were rejected for the same missing criterion, so the fix is to the item template rather than to anyone's diligence.

The uncomfortable part is step four. It is the only step that raises next week's number, it produces nothing visible on the board, and every capacity-planning instinct says the time would be better spent shipping.

Laimonade exists to take the first and third steps off a person: drafting items with criteria, paths and dependencies already in them, then checking returned work against those criteria. Sprint planning automation is worth having only where it makes the readiness check cheap enough to do honestly — a person still decides what is worth building, and a coding agent cannot close its own item. The mechanics are in the sprint workflow, the wider shape in how Laimonade works, and the role this leaves for a person in what an AI product owner actually does.

An AI sprint workflow is therefore not the old ceremony set with faster building underneath it. It is the same calendar pointed at a different scarce resource, and teams that keep pointing it at capacity get a plan that is accurate about a constraint they no longer have. How many agents that calendar can feed is a separate question, answered in how many coding agents one team can run.

Frequently asked questions

Do story points still mean anything?

As a record, yes; as a forecast, much less. Points describe how much was delivered in sprints that have closed, which is a real fact about the past. They stop predicting because an agent's execution time barely tracks the effort a point was estimating, and because what actually varies now is whether an item was specified correctly — a risk the scale was never built to express.

Should agents attend standup?

No, because the meeting's content is already on the board. An agent reports what it did as it does it, so reading those reports aloud adds a delay rather than information. What justifies keeping the slot is the part a status update cannot carry: deciding what to do about work that came back wrong, and which of the near-ready items gets specified next.

What replaces the sprint commitment?

A readiness count, and it is honest in a way the commitment was not. Instead of promising eleven items, the team observes that eleven are executable and that the agents will finish roughly that many. The forward-looking promise moves to the other queue — how many items will be specified this week — because that is the number that decides what next week can contain.

Does this mean the backlog matters more than the sprint?

Yes, and that is the real change. The sprint is now a one-week window onto a backlog, and its contents are decided entirely by how much of that backlog was made executable beforehand. A team with a well-specified backlog can run almost any cadence; a team without one gets the same empty Thursday whatever length it picks.