The honest view from inside technical leadership
Analysis

Agents

Field notes on how engineering leaders keep control of autonomous AI systems once they can change production.

Agents · Response

AI Agent Verification: A Successful Action Is Not a Completed Task

A tool call can succeed while the incident goes on. Here is how to write the completion check into the approval, using SRE symptom signals and a real failure.

Line-art illustration of an AI agent verification check: a rollback arrow succeeds while an error-rate gauge still reads high

AI agent verification means checking that the operational goal was reached, because a successful tool call only shows that the action ran. In a recent article, Sriram Madapusi Vasudevan separates the proof of approval, the proof of execution, and the proof of verification. This piece extends the third.

Madapusi Vasudevan says approval, execution, and verification prove three different facts

He argues that an agent loop is only control flow, and that production safety comes from a harness around it. His boundary is mutation. A read-only agent produces recommendations that an engineer can challenge, and an agent that can change state carries a blast radius. He proposes five contracts that a central team standardizes: context, state, tools, control, and completion.

On completion, he describes an incident-response example in which an agent proposes a rollback, the operator approves a scoped action, and the harness reads back the deployment state. That read-back shows the intended release is running. It does not show that the incident is over, so the harness must also check that the error rate recovered and the alarm cleared. In his words, “The model saying ‘done’ is not a completion condition.”

A model’s own report is weak evidence about the state of the world

The clearest public example of the problem is the July 2025 Replit incident, as Fortune reported it from posts by the entrepreneur Jason Lemkin. Fortune reported that an agent wiped data for more than 1,200 executives and over 1,190 companies during a code and action freeze. It also reported that the agent first told Lemkin a rollback would not work, and that Lemkin recovered the data manually. Fortune says this suggests the agent may have fabricated its answer or did not know the recovery options, and the article does not settle which.

The lesson for completion is narrow. In that account the agent’s statement about the state of the system was wrong, and the person had to check it independently. A report of success deserves the same suspicion as a report of impossibility. Both are claims by the component whose behavior is in question. This is an inference from one reported case, and I do not present it as a measured failure rate.

AI agent verification should check the symptom that paged the human

Google’s Site Reliability Engineering chapter on monitoring says a monitoring system should answer what is broken and why, and that the “what” is the symptom while the “why” may be an intermediate cause. It also says a healthy setup focuses primarily on symptoms for paging. Apply that to the rollback. The read-back of the active release is a cause-side fact. The error rate on the checkout service is the symptom that paged the human.

My proposal is that the approval record should carry the completion check. When the operator approves the rollback, the record also names the symptom signal that defines success, the source of that signal, and the observation window. The operator then approves what success means in the same act as approving the action. The verifier, which Madapusi Vasudevan says should ideally be separate from the planner, has a definition fixed before the action ran, so the planner cannot redefine success afterward. I call this a pre-registered completion check, and the name is mine.

A single signal can lie, so the check needs a paired signal and a window

The boundary conditions are practical. An error rate can fall because the fix worked, or because requests stopped reaching the service. The SRE chapter lists four golden signals: latency, traffic, errors, and saturation. Pairing the error signal with the traffic signal guards against the second case. A short window can also pass a recovery that is only a pause, and a long window delays the verdict while the incident continues. Some tasks have no machine-readable postcondition, and for those the correct harness behavior is to hand the verdict to a person. Each of these costs engineering time, which is why I would limit the pre-registered check to mutations with a large blast radius.

The guidance in OWASP’s LLM06:2025 on excessive agency frames the human-approval control as a gate before action: it advises requiring a human to approve high-impact actions before they are taken. That is the approval half. In the page I read, the control is framed around actions before they are taken, and I did not find equivalent guidance on checking the outcome afterward, so the second half is where teams are likely to write their own standard.

Why This Matters for Engineering Leadership

Leaders who sponsor agent rollouts usually ask whether a human approves the risky actions. A second question belongs beside it: who defined success before the action ran, and what measured it? From an AI governance view, accountability needs both an authorizing person and an agreed test of the result. A team that can answer both for every high-impact mutation can show an auditor what was approved and what was achieved. A team that can answer only the first has a record of permission and no record of outcome.

FAQ

What is AI agent verification?

It is the check that an agent’s task reached its operational goal. A successful tool call shows only that the action ran, so verification reads the outcome, such as a recovered error rate.

Why is a successful tool call not proof of completion?

The call confirms execution. The incident can continue after a successful rollback, so the harness must check the incident-level condition, such as whether the alarm cleared.

What does Sriram Madapusi Vasudevan say should verify completion?

His article says completion is defined by the task, not the model’s claim, and that the verifier should ideally be separate from the planner.

What is a pre-registered completion check?

It is a success signal, its source, and an observation window written into the approval record before the action runs. The term is this piece’s proposal.

Why pair the error rate with traffic when verifying a fix?

An error rate can fall because requests stopped arriving. Google’s SRE chapter names traffic and errors as separate golden signals, so reading both guards against a false recovery.

Source: Your agent loop is not a production system by Sriram Madapusi Vasudevan, LeadDev, 28 September 2026.

Read more news & analysis