Claude Sonnet 5.5: calibrate effort without overreach

A guide to turning effort levels into an evaluable policy: segment tasks, measure complete outcomes, and escalate only when evidence supports it.

AI effortAgent evaluationScope controlClaude Sonnet 5.5
Wasyra AI Systems
Trust, copilots, and enterprise adoption
Published
September 29, 2026
min read
6 min read
Categoría
AI Systems
5gates for calibrating effort
Conceptual illustration of five effort levels: a balanced path passes an acceptance gate while other paths drift through excess or insufficiency.

Chapter 01

The maximum level can be the wrong setting

When an agent fails a complex task, raising effort looks like the obvious correction. The model reasons longer, explores more paths, and reviews more deeply. But a real operation does not reward activity; it rewards an acceptable outcome within the agreed scope, time, and budget. If the agent adds unrequested files, starts unnecessary review rounds, or exhausts the timeout, more work can reduce the chance of acceptance.

Anthropic released Claude Sonnet 5.5 on September 28, 2026 and documented five effort levels. The official release claims higher speed and lower cost per task than Sonnet 5, but these are vendor measurements we did not reproduce. The more useful engineering signal is different: on FrontierCode, max effort scored below xhigh because some runs started subagent reviews, timed out, or made out-of-scope changes. This article is published one Lima calendar day later.

Chapter 02

The mechanism: effort expands behavior, not just tokens

Effort controls how much reasoning and autonomy the model dedicates to a request. In Sonnet 5.5, low, medium, high, xhigh, and max are not fixed budgets and do not map directly to levels from an earlier version. With more effort, the agent can sustain a longer task, check its work, and resolve ambiguity; it can also interpret an open request as permission to add related documentation, tests, or refactors. The control changes the full trajectory, including tools, retries, and stopping decisions.

That is why the comparison unit should not be an isolated answer or token count. It should be the completed and accepted task: fixed input, allowed actions, expected artifact, required checks, and rejection criteria. Two levels may produce correct code, but only one may keep the diff scoped and finish before the limit. On another task, the lower level may stop too early. The right policy depends on the shape of the work, not on a universal intelligence hierarchy.

Chapter 03

1. Define the task envelope before sweeping levels

First separate workloads that genuinely require different behavior. Fixed-schema extraction, a sourced support reply, a one-file patch, and an ambiguous migration should not share the same default merely because they call the same model. For each class, record input size, available tools, reversibility, error cost, maximum duration, expected artifact, and whether a person can clarify uncertainty. That envelope turns the word 'complex' into observable constraints.

Then create representative and boundary cases inside each class. Keep the same inputs, tool versions, limits, and graders while comparing effort. Include incomplete requests, contradictory data, an unavailable dependency, and a case that must request approval. If you change the prompt, model, harness, and level at the same time, you will not know what caused an improvement. The sweep should isolate one decision and leave a reproducible baseline.

Chapter 04

2. Evaluate an acceptance vector, not an average score

Define independent gates: functional correctness, scope compliance, valid tool use, sufficient evidence, latency, total cost, and required human correction. An outcome fails if it breaks a critical contract even when its average is high. For code, for example, separate passing tests, allowed files, absence of side changes, review acceptance, and cycle time. For support, separate factual accuracy, citations, privacy, tone, and correct escalation.

Measure variability as well. Run each case multiple times when risk and budget allow, because a single trajectory can hide a problematic tail. Report complete acceptance, not merely the percentage of correct answers, and retain rejection reasons. The release question is 'Which level crosses every gate with enough stability?' rather than 'Which level achieved the highest aggregate score?' That distinction avoids buying reasoning that does not improve the operational outcome.

Chapter 05

3. Treat scope, tools, and time as budgets

Do not limit only tokens. Set maximum tool calls, duration, modifiable files or records, attempts per dependency, and self-review rounds. The agent should know the stopping condition and the next safe state: deliver the artifact, ask for clarification, escalate, or abort without effects. High effort without these edges can turn diligence into drift. Low effort without a completion criterion can return a plan when the contract required execution.

The Sonnet 5.5 guidance acknowledges that xhigh and max can initiate extra review and related actions. Anthropic suggests instructing the agent to stop when the task and its checks are complete and not to launch reviewers unless requested. That pattern generalizes: the prompt bounds initiative while the runtime enforces limits the prompt cannot guarantee. Keep both, because a sentence does not replace quotas, permissions, or effective cancellation.

Chapter 06

4. Escalate on observable signals and step back down

Start each class at the lowest level that clears its gates, not at the highest available level. Escalate when a defined signal appears: an incomplete plan, unresolved dependency, failed verification, material uncertainty, or a task crossing a length threshold. Do not escalate out of frustration after every error; some failures require better context, a repaired tool, or human approval, and more effort will only repeat the problem at higher cost.

Record the initial level, escalation reason, tools already used, reversible state, and final outcome. Escalation must not erase the first attempt or duplicate effects. Re-evaluate periodically and step effort down when the prompt, tools, or model improve. A policy that only permits increases turns every system improvement into permanent cost. The router should learn the complexity it needs rather than confusing age with criticality or volume with difficulty.

Chapter 07

5. Hypothetical example: resolving an invoice dispute

Suppose an agent receives a dispute, checks the contract and invoice, proposes a reply, and creates a task when it finds an inconsistency. Field extraction starts at low; analysis with two consistent documents starts at medium. If commercial terms conflict, the router moves to high with the same snapshot and requires a citation for every conclusion. No level may issue a credit note: that action sits outside the envelope and requires finance approval.

The team compares five gates: correct fields, traceable evidence, escalation decision, response within the SLA, and zero unauthorized effects. If high improves conflict handling but creates duplicate tasks, it is not approved until idempotency is fixed. If xhigh offers a more elegant explanation but takes too long without changing acceptance, high remains. If max searches unrequested attachments, the scope gate rejects it even when the answer is correct. This scenario is illustrative and does not represent a Wasyra customer or Wasyra results.

Chapter 08

Conclusion: buy acceptance, not activity

Anthropic's published results depend on its harnesses, graders, prompts, and pre-release deployments. We did not run Sonnet 5.5 or verify its speed, cost, or benchmark claims. Nor do we assume the anomaly between max and xhigh will appear in another workload. The transferable lesson is methodological: an effort control can change scope, tools, and stopping behavior, so it must be evaluated as a complete execution policy.

Approve a level when it clears every relevant gate with acceptable variability and sustainable total cost. Keep an explicit path to escalate and another to stop. Repeat the sweep when the model, prompt, tools, limits, or case distribution changes. Retain rejected runs as well; they often reveal whether the next step is more reasoning, better context, a runtime control, or a human decision.

A good router does not assume more is better. It selects the minimum effort that completes the task, preserves scope, and leaves reviewable evidence. If you need to turn that policy into a pilot with real cases, permissions, and release gates, Wasyra's custom agent service can design the path. It is an implementation service separate from the product available at agents.wasyra.com.

Written by

Wasyra AI Systems

Trust, copilots, and enterprise adoption

Wasyra AI Systems covers guardrails, suggestion-first modes, and review design so work assistants earn real adoption.

CopilotsTrustB2B AI
More from this author

Series

AI systems that actually reach production

A series on agents, copilots, and guardrails for bringing AI into real work without breaking trust or operations.

Posts in this series

Keep reading

Keep reading