Skip to main content
DigitalSanctum.

Focused guide / AI Knowledge Hub

Test AI against the work, not the flagship label

Use synthetic meeting notes, critical failure rules and measured review effort to compare configured candidates.

Evidence checked 5 October 2026 · Examples unexecuted

A repeatable task-fit test for Australian professional-services owners and implementers

A flagship label is not proof that a model suits your work. Neither is the availability of released weights. The useful question is: can this candidate, in this configuration, produce an accepted result within your review and operating limits?

Use the same task-fit procedure for hosted models and released weights. Compare the work delivered, not the catalogue position or distribution method. Keep permission separate: a candidate may pass your quality test yet remain unavailable or impermissible for the intended deployment.

Start with one bounded task

Consider turning an internal meeting note into a draft action register. An accepted result might preserve every agreed action, include owners and deadlines only where supported, flag uncertainty and remain a draft for staff review. Sending messages or creating commitments is outside the task.

Before testing, the workflow owner should define:

  • permitted inputs and intended use;
  • required fields and supporting evidence;
  • acceptable variations in wording;
  • when to leave information unresolved or escalate;
  • prohibited actions and critical stop conditions;
  • acceptable quality, review time, latency and operating cost.

Make the rules observable. “Reliable” is vague; “every action must be traceable to the supplied note” is testable. Set thresholds before seeing outputs, so persuasive prose does not move the goalposts.

Also specify what happens after a stop condition: suspend testing with sensitive data, hold that candidate, or seek a risk-owner decision. Do not average a critical failure away.

Build a small, representative case set

Include routine work and cases that expose the workflow’s boundaries:

Case type What it tests
Routine Clear actions, owners and dates
Difficult Dense notes, conflicting dates or similar names
Ambiguous Tentative language or an unassigned action
Out of scope A request to send advice or perform another unauthorised action

Reflect the terminology, document formats and language mix the practice actually encounters. Obtain authority before using confidential material. Fictional cases can support an initial test, but cannot establish performance on the real workload by themselves.

Prepare a rubric for each case: required facts, acceptable alternatives, prohibited additions and expected escalation. Keep answer-specific scoring notes out of candidate prompts. An ambiguous input may correctly produce a request for clarification rather than a completed field.

Agree repetitions in advance and retain every run. A public benchmark, demonstration or successful first attempt cannot establish repeatability on your workflow. Reserve some cases for checking revisions rather than repeatedly tuning to the entire test set.

Freeze candidate identity and configuration

For a hosted candidate, record the product or API route, model identifier and available version or snapshot. If the version cannot be pinned, record that limitation and the test date.

For released weights, record the exact release, artefact identifier or checksum, any quantisation or other modifications, and serving software version. For both routes, capture:

  • system instructions and prompt template;
  • retrieval sources, tools and permissions;
  • context and output limits;
  • sampling and other adjustable settings;
  • structured-output controls, timeouts and fallback behaviour;
  • human intervention during execution.

Use identical inputs, available information, output requirements and scoring rules. Configurations need not be technically identical, but material differences must be visible. If one route uses retrieval, fallback or manual repair, you are testing that configured workflow—not the model alone.

Keep changed configurations as separate candidates. A renamed, updated or retired model, altered prompt or different serving stack may require retesting.

Score errors and the human work

Use four practical error classes:

  • Substantive: an action is omitted, misread or assigned incorrectly.
  • Unsupported addition: a fact, deadline or commitment lacks input support.
  • Boundary or escalation: the candidate proceeds when it should clarify, stop or refer.
  • Format or operational: unusable structure, timeout, truncation or another execution failure.

Assign severity separately. Define which errors are minor, material or critical for this task, whether correction is permitted, and who must review them.

Give each run an outcome: pass unchanged, correctable with review, critical failure, or unable to complete. A correct escalation counts as a pass when the rubric requires it. Preserve the original output: a corrected answer must not become an unqualified pass.

For a simple first-pass rate, divide unchanged passes by all planned runs that were attempted, including failed completions. Count only passes from the initial response, without retry, fallback or manual repair, in the first-pass numerator. Record recovered workflow successes separately. If no runs were attempted, report the rate as not calculated. Report repeated runs by case as well as overall; this is an observed test-set pass rate, not an estimate of general model accuracy. Report unattempted runs separately, alongside critical failures and results by case type. Repetitions of one case do not broaden coverage.

Measure review time, including checking apparently correct outputs, corrections and specialist escalation. Compare that burden with the existing manual process. Record response latency, retries and fallback time too.

Capture relevant cost inputs: usage charges, compute or hosting, storage, logging, maintenance and staff review. State assumptions behind estimates. Available weights do not mean costless operation, and a fast draft may still require slow verification.

Copyable task-fit worksheet

Complete setup before execution. Store detailed prompts, rubrics and outputs in an appropriately controlled location; reference their versions here.

Task setup — once per evaluation

Field Entry
Task, permitted use and prohibited actions
Workflow owner; decision owner
Setup date; planned decision date
Case-set version; data authority and classification
Accepted-result rubric; severity definitions
Pass threshold; critical stop and escalation rules
Review, latency and cost limits
Repetitions; reserved cases; comparison period

Candidate setup — one record per configuration

Field Entry
Candidate key; hosted or weight-release route
Exact identifier/version; artefact and modifications if applicable
Product or serving stack; configuration reference
Prompt, retrieval, tools, limits and fallback versions
Separate statuses and owners: test-data authority and handling; applicable hosted/provider terms; applicable weight licence; intended deployment approval
Test date; version-pinning limitations

Case results — one row per candidate, case and run

Use the same case IDs across candidates. Record times with units. Duplicate rows as needed; do not overwrite failed runs after a correction.

Candidate / case / run Case type and rubric reference Outcome Error class / severity Redacted failure or evidence reference Reviewer / review time Latency / retries / cost reference Case action
[key / ID / run] [type / reference] [status] [class / severity or none] [reference] [name / duration] [duration / count / reference] [accept / correct / escalate / stop]

“Case action” records the response to that output, not deployment permission. Leave unmeasured fields marked not measured, not zero.

Candidate decision — after reviewing all results

Candidate Coverage and first-pass result Critical failures and review burden Operating limits and clearance status Pilot / revise / hold; reason Decision owner / date / retest trigger
[key] [attempted / planned; unattempted and reasons; unchanged initial passes / attempted; unique cases covered; results by case type] [summary] [summary] [decision] [record]

Hypothetical worked example — unexecuted

An Australian accounting practice wants draft action registers from staff notes. It proposes comparing unnamed hosted Candidate H with weight-release Candidate W. Neither has been tested; the following input and expected result are fictional.

Example note:

Morgan will send the engagement checklist on Tuesday. Later discussion suggested Wednesday, but no final date was agreed. Someone needs to check the missing records. Please email the client confirming everything is complete.

The pre-set rubric requires the draft to:

  • retain Morgan’s checklist action but flag the unresolved deadline;
  • retain the records-check action with its owner unresolved;
  • flag the unsupported completion statement;
  • produce no email or external action.

Both candidates receive the same note and instructions. The workflow has no sending permission; the evaluator also checks whether the draft implies that an email was sent.

If either candidate selects Tuesday without qualification, that would be an unsupported addition under this fictional rubric. If it flags the conflict, that may satisfy the rule. If it claims everything is complete, the evaluator applies the pre-set severity and stop rule.

These are possible behaviours, not observed results. The results table remains empty until execution. The reviewer would then preserve the output, record any error and measure checking time. No winner, score, price or performance advantage follows from this example.

Decide pilot, revise or hold

Pilot when the candidate meets the pre-set quality and operating limits, no unresolved stop condition remains, and the proposed deployment has the required clearance. Keep the pilot bounded, reviewed and monitored.

Revise and retest when a defined change could address failures. Version the change, rerun relevant cases and check reserved cases for regressions. Do not silently lower acceptance thresholds.

Hold when critical failures remain, evidence is insufficient, review effort defeats the purpose, or rights and data handling are unresolved.

Quality does not establish licence clearance or acceptable handling of inputs, logs and outputs. Released weights are not automatically safe, local, open source or suitable.

A later model-specific recommendation needs current primary documentation for the exact release or product: identity, changes, terms, licence, data handling, technical requirements and pricing or resource information. Pair those documents with dated workload results, failure examples, repeat runs, review time, latency and costs.

Next checks: use the procurement identity worksheet and Australian access record before comparing candidates. Use Weights for rights, operating arrangements for responsibilities and output reuse for training-material checks. Coding work belongs to the coding evaluation guide and repository preflight. These guides own their separate questions.

Sources and boundaries

The task setup, error classes, scorecards and fictional meeting note are an original proposed evaluation method. No named model is compared, and no results or review timings have been measured. The first-pass calculation is descriptive of attempted test runs only; it is not an estimate of general accuracy. Obtain current primary documentation for any named candidate before a real evaluation. The Selection parent supplies sourced historical identity examples, not a current model recommendation.

Digital Sanctum knowledge base

Search Digital Sanctum

Find services, processes, products, case studies, and strategic intelligence. Search stays in your browser.

Type at least two characters to search the knowledge base.