In this fictional, unexecuted scenario, a coding agent works through an afternoon and leaves six commits. The tests appear green. Then a reviewer finds an unrelated edit, a weakened assertion and a bug in the output. Was that a successful afternoon?
The activity log cannot answer. Duration, tokens and commits describe work performed, not value accepted. For an engineering owner, the useful result is a change that meets its requirements—and an honest account of the human effort needed to get there.
This guide proposes a repeatable evaluation record for one bounded task. It separates the first result from the eventual result, tests continuation without confusing it with starting again, and keeps unrun observations explicitly unknown. No coding-agent test has been performed for this article.
Start here
Delegation decision: decide whether the assignment is worth handing over.
Claude Code product boundary: identify the product, model, tools and environment.
Repository preflight: check the repository before moving beyond synthetic work.
What the historical sources establish
What changed—and what the release does not prove distinguishes the 24 November 2025 release announcement from the separate 26 November 2025 engineering account. Neither supplies measured results for your team or the fictional tasks below.
This page turns that evidence limit into a local measurement method: define acceptance, preserve the conditions and count the effort. The method is an editorial recommendation, not an Anthropic benchmark.
Define acceptance before counting activity
An accepted change satisfies the specified user outcome, passes the agreed completion checks and receives independent human acceptance. State the scope of that acceptance: it does not promise the absence of every possible defect or authorise deployment.
Completion checks are the predefined tests and inspections used to judge the result. Include the requested behaviour and existing behaviour that must remain intact. A green test alone is insufficient if the diff weakens the test or misses the user’s actual problem.
Record two results:
- First-pass result: the artefact at the predefined first completion boundary, before corrective feedback or human code changes.
- Eventual result: the artefact accepted or rejected after any further attempts, corrections and review within the overall ceiling.
Define that first boundary before starting—for example, the first declaration of completion, terminal failure or ceiling reached. A planned interruption need not end the first attempt: if continuation is part of the protocol, say so.
Keep the original artefact and review findings. “Eventually accepted” must not erase the failed first pass. If the ceiling expires without acceptance, record not accepted within the ceiling, not a silently extended success.
Freeze the conditions that make the result interpretable
Before a future test, record:
- exact product, surface and installed version/build where observable;
- selected model/version, or the disclosed selection policy and any undisclosed details;
- repository starting snapshot and complete task text;
- tools, connections and effective permissions;
- starting context and harness—the surrounding execution and context-handling setup;
- test suite, commands and independent review criteria;
- elapsed-time, human-effort and cost ceilings, with a named reviewer.
The Claude Code product-boundary guide explains the identities; the repository preflight covers repository controls. Here, preserve their verified values as evaluation conditions rather than duplicating those procedures.
If the model version is undisclosed, write undisclosed. You may still learn something about that workflow, but cannot confidently attribute its result to a specific model version. Likewise, changed tests, added hints or extra tool access affect interpretation.
Record the baseline test result before the attempt. Without it, a failing check might be an existing defect rather than a regression. Preserve task wording and input artefacts so another engineer can recognise what was actually evaluated.
Two fictional tasks, with no invented results
The two task records below are fictional and unexecuted. Task A uses the parent guide’s sorting example; Task B adds a different requirement to illustrate a planned handoff. Both use invented appointment data only, with no client code, production credentials or deployment connection. Each would start from its own clean, recorded snapshot.
Task A: one contained sorting fix
Use the sorting requirement in A fictional decision: delegate the defect, not the overhaul, and copy its exact wording into the evaluation record before any authorised attempt.
For this record, completion checks would cover both ordering cases, the existing sorting suite and an independent inspection for hard-coded output or weakened assertions. The first-pass boundary would be the first completion declaration, terminal failure or ceiling reached.
Task A result cells
- First pass / eventual acceptance: Unknown — unrun
- Completion checks / independent review: Unknown — unrun
- Regressions / retries / human corrections: Unknown — unrun
- Human effort / runtime / attributable cost: Unknown — unrun
- Planned interruption: Not part of this task’s protocol
Task B: an export change across a handoff
The invented program exports a day’s appointments but gives no clear indication when the day is empty. The required change is an explicit empty-day result without changing the non-empty export.
The proposed protocol would interrupt work at a predefined checkpoint after initial inspection and before implementation, then continue in a fresh session with specified handoff material. Checks would cover empty and non-empty exports and the output a user receives.
Task B result cells
- First pass / eventual acceptance: Unknown — unrun
- Completion checks / independent review: Unknown — unrun
- Interruption / usable handoff / continuation outcome: Unknown — unrun
- Regressions / retries / human corrections: Unknown — unrun
- Human effort / runtime / attributable cost: Unknown — unrun
These tasks illustrate different evaluation questions. They are not a comparison of uninterrupted versus interrupted performance: their requirements differ. Unknown cells are not zeros, and two small tasks would not establish sustained-work reliability.
Test continuation separately from replay
A continuation resumes unfinished work from a recorded intermediate state. A replay starts the same task again from its original snapshot. They answer different questions.
For Task B, specify exactly what survives the interruption: repository state, original requirements, test output and a progress note, for example. Record whether conversation history is retained. Use only facilities the chosen setup actually supports; do not assume every harness offers identical session controls.
At the checkpoint, capture what is complete, what remains and which checks have run. On continuation, observe whether the next session identifies the unfinished requirement and proceeds without losing or repeating work. Count any human explanation needed to repair the handoff. If interruption cannot be performed as specified, record a protocol deviation, not a successful continuation test.
For a replay, retain the original result and start a separately identified attempt from the same initial snapshot. Hold the task, checks, context and limits constant where possible. A changed prompt is a new condition, not an exact replay.
A planned continuation is not automatically a retry. A retry is an additional corrective attempt after failure or rejection. Record unplanned restarts separately, including their cause, rather than hiding them inside an apparently continuous run.
Review the diff and the user outcome independently
Freeze the first-pass artefact before giving corrective feedback. A reviewer should compare the complete diff with the original requirements, independently run the agreed checks where possible, and inspect the resulting user behaviour.
Did the change solve the problem? Were unrelated files touched? Were tests weakened? For a visual task, inspect the relevant rendered states and interactions rather than relying only on automated assertions.
Prefer a reviewer who did not coach the attempt. If the operator must also review it, record that limitation; independently checking output is still useful, but it is not the same as a separate-person review.
A regression is newly broken existing behaviour. Use the baseline to distinguish it from a pre-existing failure; record uncertain causes as unknown. Keep regressions found before acceptance in the record even if later corrected.
Reviewer-requested fixes belong after the first-pass result. Record who made them, what changed and whether re-review was required.
Count the full effort without double-counting time
Review time includes inspection, verification, decision and re-review. Handoff effort includes preparing, reading and repairing progress records. Human setup, supervision and corrective coding also count.
Report three separate quantities:
- Human effort: person-minutes across those activities.
- Agent runtime: observed active runtime, with gaps or unavailable timing noted.
- Elapsed time: start to final decision, including waits and interruptions.
Do not add overlapping human time and runtime into one misleading “total hours” figure. Together, these measures describe total effort and turnaround.
Record attributable cost only when observed; otherwise mark it unknown. Tokens and commit counts may explain behaviour or resource use, but neither measures acceptance.
Compare setups only after actual controlled observations exist. Use the same task and snapshot, comparable review criteria and recorded limits; disclose remaining differences. Report first-pass and eventual results together, with human effort and regressions. One successful replay does not establish reliability, and an unknown cost cannot support a cost-saving claim.
Copyable scorecard: one task, one evaluation record
Record ID / date / operator / reviewer: ___
Product, surface, version/build: ___
Model/version or selection policy; undisclosed details: ___
Starting snapshot / task / expected user outcome: ___
Tools, connections, effective permissions: ___
Starting context / harness / permitted handoff material: ___
Baseline checks / completion checks / review criteria: ___
Elapsed-time, human-effort and cost ceilings: ___
First-pass boundary / frozen artefact / result: ___
Continuation, replay or retry IDs and conditions: ___
Independent diff and user-outcome findings: ___
Regressions / failures / protocol deviations: ___
Human corrections / eventual result: ___
Human effort by activity / runtime / elapsed time: ___
Observed cost or unknown / limitations: ___
Pilot or hold?
Prepare one synthetic task record and its acceptance checks first. Run it only with the required organisational authorisation, verified setup controls and a ready observation method; otherwise hold. Before any real-repository work, complete the repository preflight and obtain the authorised owner’s explicit, scoped access decision.
Afterwards, decide whether the evidence supports another controlled attempt—not whether the activity looked impressive. Until reviewed observations exist, hold performance comparisons and leave the result cells unknown.