The honest view from inside technical leadership
Analysis

Leadership

Field notes on how engineering leaders keep control of systems as AI changes how software gets built.

Leadership · Response

Error Budgets and AI: Why Healthy Dashboards Can Hide a Comprehension Gap

A service can meet its SLO while its team cannot explain a failed change. Here is how a pre-agreed comprehension hold extends error budgets for AI-era teams.

Line-art illustration of a dashboard gauge reading healthy beside a locked release gate, representing error budgets and AI comprehension holds

Error budgets and AI do not mix well, because an error budget measures failure and AI makes it cheaper to ship changes that nobody on the team fully understands. A service can stay inside its budget while the team loses the ability to explain what it runs. Paul LaPosta makes this case in a recent article, and this piece extends one claim from it.

Error budgets and AI: LaPosta says shipping now outruns understanding

LaPosta describes a service that meets its service level objective, handles a rollback smoothly, passes QA and approval, and still leaves its team unable to explain why a consequential change failed. His diagnosis is that "AI has split shipping a change from understanding it." Code, tests, documentation, and reviews have become cheaper to produce, and understanding has not become cheaper with them.

He points to early warning signs: merge requests that get reopened, rollbacks that turn into debates, and decisions that the team should settle being escalated to leadership. He advises against a comprehension score, which would suggest false precision. He suggests instead that leaders ask the owning team to explain recent material changes without the one engineer who knows the system and without AI-generated context, and that teams get the authority to delay releases while the service is still healthy.

An error budget can only trip after something has already failed

Google’s Site Reliability Engineering book defines the error budget as an objective measure of how unreliable a service is allowed to be within a quarter. When SLO violations use up the budget, releases are temporarily halted and the team invests in testing and development to make the system more resilient. The design has two parts. A meter counts failures, and a halt rule stops releases when the meter runs out.

The halt rule works because the organization agreed on it in advance. Nobody has to win an argument on the day the budget runs out. The meter, however, only moves when users feel an outage. LaPosta’s scenario is the case where the meter reads healthy at the moment the team is most fragile. The part of the error budget design worth keeping is the halt rule that was agreed before pressure arrived. The meter can stay as it is.

Naur’s 1985 essay explains why more AI output cannot supply the missing understanding

Peter Naur argued in his 1985 essay Programming as Theory Building that a program is a theory held in the minds of the people who build and maintain it. In his account, source code is a partial record of that theory. He wrote that a program can keep running after its team disperses and still be effectively dead, because nobody can answer demands for modification intelligently. He also held that rebuilding the theory from documentation alone is not possible.

Read next to LaPosta, Naur gives a reason why the gap he describes is structural. AI-generated code, tests, documentation, and review comments are all artifacts. If Naur is right, adding artifacts raises the volume of the record and adds nothing to the theory until a person absorbs it. There is a boundary here that I cannot resolve. Naur wrote about human teams and did not address AI tools. Whether an engineer who reads and interrogates an AI-generated explanation builds the same theory as one who wrote the code is an empirical question, and I have not seen a primary study that answers it.

A comprehension hold works if it copies the shape of a budget freeze

LaPosta asks leaders to give teams authority to delay healthy releases. Authority alone tends to go unused when delivery pressure is high, so a trigger needs to exist before the pressure does. A practical version borrows from the SRE design. First, the organization defines in advance which classes of change count as material. Second, before a material change ships, the owning team explains the change, the alternatives considered, and the risks, without the key engineer and without the prompt history. Third, a failed explanation triggers a hold that the team itself can invoke. Fourth, the outcome is recorded as pass or hold, with no score attached, which respects LaPosta’s warning about false precision. I call this a comprehension hold, and the name is mine.

The boundary conditions are real. Every hold costs delivery time. An explanation session can become a rehearsed ritual that anyone can pass. On a team of two, the key engineer is most of the team, so the test has little to measure. Low-risk changes would only collect cost. There is also a counter-case in favor of AI: a model that explains unfamiliar code well may help a team build understanding faster than before. The hold does not penalize that, because it tests the outcome, which is whether the humans can explain the change.

Why This Matters for Engineering Leadership

Leaders currently read reliability from dashboards, and LaPosta’s point is that the dashboards can be accurate and still silent about the thing that matters most before an incident. A second standing question helps: can the owning team explain the last material change without its most knowledgeable engineer? From an AI governance view, accountability requires a person who can explain a decision. A team that cannot explain a change it shipped has no accountable owner in any useful sense. Leaders who agree on the trigger for a hold in advance spare their teams the argument at the moment the pressure is highest.

FAQ

Why can’t error budgets detect an AI comprehension gap?

An error budget counts failures against a service level objective. A team can ship changes it poorly understands without consuming any budget, so the meter stays healthy until a failure happens.

What is a comprehension hold?

It is a pre-agreed pause on a material release when the owning team cannot explain the change, its alternatives, and its risks without the key engineer or AI-generated context. The term is this piece’s proposal, built on LaPosta’s suggestions.

Does Paul LaPosta recommend a comprehension score?

No. His article advises against a score because it would suggest false precision. He suggests leaders check whether the owning team can explain recent material changes.

What did Peter Naur argue about programs and teams?

In his 1985 essay Programming as Theory Building, Naur argued that a program is a theory held by the people who maintain it, and that a program is effectively dead once that team disperses.

What happens when an error budget is exhausted?

According to Google’s Site Reliability Engineering book, releases are temporarily halted and the team invests in testing and development to make the system more resilient.

Who should be allowed to delay a release on a healthy service?

LaPosta says teams should have the authority to delay releases while the service is healthy. This piece adds that the trigger for a delay should be agreed in advance.

Source: Your error budgets don’t know AI exists by Paul LaPosta, LeadDev, 30 September 2026.

Read more news & analysis