Skip to main content
DigitalSanctum.

AI Knowledge Hub · guide

What Claude Code is—and which decisions it owns

Identify the product, model, account, tools and environment before evaluating a coding agent.

Imagine a fictional demonstration: a team lead watches Claude Code open files, change a function and run tests, then says, “We tested Opus.” What would that demonstrate: a model’s reasoning, a coding product’s tools, or a particular account working inside a prepared environment?

For a small Australian engineering team, that distinction matters before anyone repeats the demonstration. Which model was selected? Where did the commands run? What else could the session reach? Watching a successful edit does not answer those questions.

Claude Code names the product, not the whole setup. To describe one useful pilot, name the product, model, surface, account, tools and repository environment separately. Then try one invented task in a disposable repository—not a convenient copy of client work.

Name the product and model separately

Anthropic describes Claude Code as a coding tool that can work with a codebase, edit files and run commands. Its documentation describes terminal, IDE, desktop and web surfaces, along with connections to external tools. These are documented capabilities, not evidence that every capability is available or enabled in your session. Claude Code overview

The model is a different part of the setup. On 24 November 2025, Anthropic announced Claude Opus 4.5 alongside Claude Code updates and new tools for longer-running agents. That historical announcement does not establish the model selected in a session today. Do not treat an old demonstration—or a historical model identifier—as current configuration guidance. Anthropic, Introducing Claude Opus 4.5

For the historical evidence and its duration limits, see What changed—and what the release does not prove. Here, verify the model or model-selection policy actually in use; do not infer it from a release-day example.

Start here

Delegation decision: when a longer coding assignment is worth handing over.
Evaluation: how to measure accepted work rather than activity.
Repository preflight: what must be checked before using real code.

Separate the surface from the account

The surface is where you interact with Claude Code: for example, a terminal or IDE extension. The account is the sign-in or provider arrangement through which you access it. Neither should be inferred from the other.

Anthropic’s overview describes several surfaces and different account or provider routes. That is a reason to verify your chosen route, not assume that a colleague’s desktop demonstration and your terminal session use identical access arrangements. Claude Code overview

Record where the session starts and where its actions execute. Then establish the exact account being used, applicable organisational controls and the model selection shown or documented for that session. Avoid recording credentials in the pilot note.

If the underlying model version is not disclosed, write that down rather than guessing. An undisclosed selection policy is a limitation on what you can later attribute to a model; it is not permission to ignore uncertainty about account ownership or effective access.

Multiple surfaces therefore do not imply one account setup, one permission configuration or a fixed model. A screenshot of a product name establishes none of those details.

Distinguish tools from their permissions

The model contributes proposed actions; Claude Code’s tools perform actions such as reading files, editing code and invoking commands. Connected tools can add other destinations. For example, Anthropic documents Model Context Protocol connections to external information and work systems. A capability listed in those docs does not prove that such a connection exists in your setup. Claude Code overview

For the pilot, ask two separate questions:

  • Which tools and connections are available?
  • Which actions do the effective controls permit?

Claude Code’s permission documentation explicitly distinguishes instructions from enforcement. Prompts and CLAUDE.md shape what the model tries to do; permission rules are enforced by Claude Code. Writing “only edit this function” does not itself change the tool’s access. Claude Code permissions

The surrounding environment also needs identifying. A test command may invoke a script; that script may use files, credentials or network access available to its process. Do not treat an approval prompt as evidence that those downstream resources are restricted.

The practical check is therefore the documented product behaviour, effective configuration and actual execution environment together. This is not a detailed security preflight—that belongs in Before staff use an agent on a real repository—but it is enough to reject “the prompt says no” as a description of an enforced boundary.

Copy this boundary card

Complete this before a proposed pilot. It is a planning record, not evidence of an approved account or a completed run. Keep unknowns visible.

Synthetic pilot boundary card

Product: Claude Code; version/build as observable, or explicitly unknown.
Model / selection policy: selected model as observable; any undisclosed version or automatic selection; who may change it.
Surface: chosen interface and where actions execute.
Account: verified sign-in/provider arrangement and applicable organisational controls; no credentials recorded.
Repository / environment: disposable checkout containing invented code, issue and tests; execution location; no client material or production credentials.
Tools / effective permissions: available tools and connections; permitted read, edit and local test actions; how those limits were checked.
External destinations: required product-service connections identified; no additional task destinations authorised.
Human owner: engineer who observes the pilot and independently reviews its output.
Unknowns / hold conditions: unresolved account or access questions, unchecked connections, or unavailable usage limits that prevent a bounded run.

Do not enter “no network” simply because the coding task is local. Distinguish connections needed to use the product from optional integrations or task-driven requests. If you cannot establish the relevant destinations and controls, hold the pilot rather than treating intention as observation.

Likewise, a new repository is only one part of the setup. It does not by itself establish what the account, shell or connected tools can access.

Apply the boundary card to the shared synthetic example

Use the fictional, unexecuted sorting task in A fictional decision: delegate the defect, not the overhaul. This is the same illustrative task, not a second pilot or a reported result.

Before any authorised attempt, complete the boundary card for the chosen setup. Identify the product and observable version, model or selection policy, surface, account, execution location, available tools and effective permissions. Distinguish required product-service connections from task destinations.

Verify that the setup can enforce the task’s recorded limits; the request alone cannot supply them. If account ownership, reachable resources, destinations or effective controls remain unresolved, hold the pilot.

During any future run, record requested and observed actions against that card. Afterwards, independently inspect the complete diff and test output. The question here is which setup performed the work and what it could reach; What longer-running coding does and does not prove covers how to evaluate the result.

Record what happened without enlarging the claim

Keep the completed boundary card with a short observation note: actions requested and observed, files changed, test output, human corrections and unexpected behaviour. Avoid secrets, personal information and unnecessary raw logs.

All results for this example remain unknown because it is unrun. A later passing test would show only that the observed change passed that check in the recorded setup. It would not establish broad model performance, sustained-work reliability or suitability for customer code.

For first-pass versus eventual results, continuation versus replay, and comparative evaluation, use What longer-running coding does and does not prove. Those are separate measurement questions, not conclusions this small identity check can supply.

Your next decision

Replace “we tested Opus” with a description another engineer can understand: this product and version, this model or selection policy, this surface and account, these tools and controls, this disposable environment.

First verify the exact account and effective settings. Run the synthetic pilot only with the required organisational authorisation and verified controls; otherwise hold it. Record only what is observed. Keep real-repository work on hold until the repository-specific preflight supplies the evidence for an authorised owner’s explicit access decision; a synthetic pilot does not grant that access. The useful outcome here is a named, bounded setup, not confidence borrowed from a demonstration.

Timeline evidence · 24 Nov 2025 — Claude Code with Claude Opus 4.5

This milestone has its own dated source and review status; the timeline remains a selective timeline with incomplete coverage.

Milestone id
opus-45
Event date
2025-11-24
Issuer
Anthropic
Milestone title
Claude Code with Claude Opus 4.5
Evidence label
Announcement
Primary decision type
Capability
Secondary decision types
Place
Editorial decision implication
Decide when a longer coding-agent run is worth allowing, and which scope, stop points and checks must come first.
Historical source check
2026-09-26
Milestone review status
source-checked
Australian access
unresolved
Era
agents
Pillar
Decision pillar
Primary source
Anthropic — primary page

Capability — source support (close paraphrase): Effort control, context compaction and advanced tool use let Opus 4.5 run longer with less intervention. Location: New on the Claude Developer Platform and Product updates.

Place — source support (close paraphrase): Claude Code becomes available in the desktop app for local and remote sessions. Location: New on the Claude Developer Platform and Product updates.

Scope and limits: The announcement does not establish an hours-long run compared with earlier Opus models or every other flagship model. Its quoted 30-minute autonomous coding session is a customer statement, not Anthropic’s own duration finding.

View the timeline record

Digital Sanctum knowledge base

Search Digital Sanctum

Find services, processes, products, case studies, and strategic intelligence. Search stays in your browser.

Type at least two characters to search the knowledge base.