Skip to main content
DigitalSanctum.
Master Guide AI Knowledge Hub

Longer-running coding agents: when to let them continue

A bounded delegation decision for coding work that continues across steps.

Longer-running coding agents

A coding agent has traced a bug, changed two files and started checking the result. It asks to continue. For an engineering lead, the important question is not whether it can keep working. It is whether the next stretch is still inside an agreed task—and whether someone can independently judge what it produces.

A longer run can be useful when several dependent steps belong together: reproduce a defect, trace its cause, make a bounded change and check the behaviour again. It becomes harder to justify when “continue” means discovering new requirements, widening access or accumulating changes nobody has time to review.

Delegate a bounded outcome, not an amount of uninterrupted activity. Before starting, establish the work contract, narrow authority, recoverable progress and independent acceptance. The framework below is an editorial recommendation for making that decision, not a report of a coding-agent trial.

Start here

What changed—and what the release does not prove

On 24 November 2025, Anthropic announced Claude Opus 4.5 alongside Claude Code updates and described new tools for longer-running agents. That supports a historical product change worth evaluating. It does not establish that Opus 4.5 could work autonomously for a particular number of hours, or outlast previous Opus models and competing flagships. The announcement’s quoted 30-minute coding-session observation is customer testimony, not Anthropic’s own general duration finding. Anthropic, Introducing Claude Opus 4.5

A separate engineering article, published 26 November 2025, discusses work spanning hours or days while explicitly describing consistent progress across many context windows as an open problem. It reports an experimental approach using initial setup, incremental coding and durable progress records. This later account explains why continuation needs careful design; it is not additional release-day evidence or a measured result for your team. Anthropic Engineering, Effective harnesses for long-running agents

The operational question therefore remains local: can this task produce a reviewable result within the authority, effort and uncertainty the team is prepared to accept?

Decide whether the task earns a longer run

Start with the outcome and work backwards. A suitable candidate has a behaviour someone can describe, a limited area of change, checks that expose likely mistakes and a reviewer who understands the result. Its steps benefit from continuity, but the task can still stop at a meaningful checkpoint.

A contained compatibility refactor might qualify. So might fixing a reproducible defect in a named component. Neither qualifies automatically: missing tests, sensitive inputs or difficult rollback can change the decision.

Use four questions:

Is the outcome clear? Proceed when: Observable behaviour and explicit exclusions. Narrow or hold when: “Improve the system” without a finish line.

Is the change bounded? Proceed when: Named components and understood dependencies. Narrow or hold when: Likely expansion across systems or owners.

Can authority stay narrow? Proceed when: Required actions can be limited and checked. Narrow or hold when: Broad credentials or external actions are necessary.

Can a person accept it? Proceed when: Reviewer, evidence and review time are available. Narrow or hold when: Nobody can independently check the result.

Do not turn this into an average score. A clear task does not compensate for uncontrolled access. Strong tests do not compensate for the absence of an accountable reviewer.

Also ask whether delegation is worthwhile at all. If preparing the task and checking its output would take more effort than a straightforward human fix, doing it directly may be the better choice. Conversely, an exploratory investigation may be useful if its deliverable is a bounded diagnosis—not an implied permission to implement whatever it discovers.

Write the work contract before “continue”

The work contract is a short agreement about the desired change and the limits around it. It should let a second engineer distinguish useful progress from drift without reconstructing the conversation.

State the user outcome, starting snapshot, permitted scope and behaviour that must remain unchanged. Name the evidence required for acceptance. Specify which decisions remain human: dependency changes, new destinations, broader repository access, merging or deployment.

Include a checkpoint and a ceiling on time or spend. A ceiling limits exposure; it does not define success. Reaching it with unfinished work means stopping with an honest handoff, not silently extending the run.

Copyable work-contract card

Owner / independent reviewer: ___
Desired user outcome: ___
Starting snapshot and permitted scope: ___
Must remain unchanged / explicitly out of scope: ___
Allowed actions; decisions reserved for a person: ___
Acceptance evidence: ___
Checkpoint / time and spend ceilings: ___
Stop triggers: ___
Progress record / recovery point / handoff destination: ___

If these entries cannot be completed meaningfully, narrow the assignment before starting. “Ask me if anything looks risky” is not a substitute for identifying foreseeable boundaries.

A fictional decision: delegate the defect, not the overhaul

Consider a fictional, unexecuted example. An engineering lead prepares a disposable appointment-list application containing only invented records. No client code, production credentials or real service connections are involved.

The good candidate is a sorting defect. Appointments with identical start times change order unexpectedly. The proposed contract asks for stable ordering in one named module, preserving the existing behaviour for different start times. It permits local inspection, relevant edits and specified local tests; no package installation, remote push or deployment. An engineer will independently inspect the diff and the displayed list.

This task has connected steps worth keeping together: reproduce the issue, understand the sorting rule, change it and verify both affected and unaffected cases. The proposed decision is a bounded synthetic pilot, subject to verifying the setup—not approval for real-repository access. Its result remains unknown because nothing has been run.

The poor candidate is the follow-on request: “While you are there, modernise scheduling, connect the client calendar and ship the improvement.” It combines an undefined outcome with new data access, integration choices and production consequences. More runtime cannot resolve who may authorise those decisions.

The lead should hold that request and separate the decisions. A smaller investigation could document requirements using synthetic inputs. A later implementation would need its own contract, access decision and repository preflight. Success on the sorting task would not authorise the overhaul.

Keep authority narrower than the task description

Reading code, editing files, running commands, installing packages and contacting external systems are different actions. An instruction to finish the task must not silently authorise all of them.

Claude Code provides one useful example of the distinction. Its documentation says permission rules are enforced by Claude Code, while prompts and CLAUDE.md instructions shape what the model tries to do. An instruction is not enforcement. An approval interface alone also does not establish the limits of the surrounding filesystem, credentials or network. Claude Code permissions documentation

Before a real run, require evidence that the chosen setup can enforce the intended limits. The detailed repository checks belong in Before staff use an agent on a real repository; this pillar’s decision is simpler: if the task needs authority the owner cannot confidently bound, do not extend the run.

External actions deserve an explicit decision. A local fix does not imply permission to open a pull request, send a message, change CI or deploy. Keep those decisions separate even when the next action seems convenient.

Preserve progress that another person can use

A longer task should leave understandable intermediate states, not just a growing transcript. At each agreed checkpoint, require a concise record of:

  • what changed and which requirement it addresses;
  • checks attempted, including failures;
  • unresolved questions and the next bounded step;
  • the code state from which work can resume or be recovered.

Anthropic’s later engineering account describes two relevant failure patterns: attempting too much and leaving undocumented partial work, and declaring completion too early. Its experimental approach uses incremental changes, Git history and progress notes to help subsequent sessions understand the state. Those findings motivate durable records; they do not prove that a particular continuation mechanism will work in every product or environment. Anthropic Engineering, 26 November 2025

Treat the progress note as a claim to check against the files and evidence. A commit preserves a state, not correctness. If a new session cannot reconcile the note with the repository, hand off the discrepancy rather than allowing it to guess.

Accept the outcome independently

The agent should not be the sole judge of its own completion. Reserve review time before starting, and have a human compare the entire diff with the original contract.

Review both implementation and user outcome. Tests provide evidence about the cases they cover; they do not establish that unrelated behaviour is intact or that a page is usable. For interface changes, include appropriate visual and interaction checks. Inspect changes to tests, scripts and dependencies rather than accepting a green result at face value.

Keep the first-pass result distinct from the eventual accepted result. Reviewer corrections and retries are part of the effort, not details to erase from a successful summary. What longer-running coding does and does not prove supplies the evaluation method; here, the decision is whether enough trustworthy evidence exists to accept, request a bounded correction or reject the change.

Acceptance of a diff is not permission to deploy it. Preserve the team’s separate release decision.

Go, stop or hand off

Go when the outcome is clear, the scope and authority are bounded, progress can be recovered, and a reviewer has both evidence requirements and time. Continue only while those conditions remain true.

Stop when the task seeks an unexpected credential, destination, dependency or licence change; expands beyond agreed systems; develops unexplained failures; loses a trustworthy recovery point; or reaches its ceiling. A stop is a control working as intended, not necessarily a failed experiment. Resolve the new decision before restarting.

Hand off at completion or when a person must decide. Supply the diff, checkpoint state, actions taken, check results, failures and explicit unknowns. Make clear what was not verified. A human should be able to resume or discard the work without relying on a polished completion message.

Longer-running capability is a reason to examine this workflow, not a reason to waive its boundaries. The practical decision is: can this task continue in reviewable increments, under narrow authority, with independent evidence of an acceptable result? If not, shorten it, change the deliverable or keep it with a person.

Timeline evidence · 24 Nov 2025 — Claude Code with Claude Opus 4.5

This milestone has its own dated source and review status; the timeline remains a selective timeline with incomplete coverage.

Milestone id
opus-45
Event date
2025-11-24
Issuer
Anthropic
Milestone title
Claude Code with Claude Opus 4.5
Evidence label
Announcement
Primary decision type
Capability
Secondary decision types
Place
Editorial decision implication
Decide when a longer coding-agent run is worth allowing, and which scope, stop points and checks must come first.
Historical source check
2026-09-26
Milestone review status
source-checked
Australian access
unresolved
Era
agents
Pillar
Decision pillar
Primary source
Anthropic — primary page

Capability — source support (close paraphrase): Effort control, context compaction and advanced tool use let Opus 4.5 run longer with less intervention. Location: New on the Claude Developer Platform and Product updates.

Place — source support (close paraphrase): Claude Code becomes available in the desktop app for local and remote sessions. Location: New on the Claude Developer Platform and Product updates.

Scope and limits: The announcement does not establish an hours-long run compared with earlier Opus models or every other flagship model. Its quoted 30-minute autonomous coding session is a customer statement, not Anthropic’s own duration finding.

View the timeline record

Digital Sanctum knowledge base

Search Digital Sanctum

Find services, processes, products, case studies, and strategic intelligence. Search stays in your browser.

Type at least two characters to search the knowledge base.