HydraFusion: how to evaluate multi-model orchestration

A guide to turning multi-model routing into a verifiable policy: choose the smallest sufficient workflow, count every pass, and protect the final change.

Multi-model orchestrationAgent evaluationCoding agentsHydraFusion
Wasyra AI Systems
Trust, copilots, and enterprise adoption
Published
October 1, 2026
min read
7 min read
Categoría
AI Systems
5gates for evaluating orchestration
Conceptual illustration of a coding task branching into single, cascade, and critique paths before converging into one reviewable change.

Chapter 01

More models do not guarantee a better change

When an engineering task stalls, adding another model is tempting: one drafts, another critiques, and a third handles the hard parts. That composition can improve coverage, but it also duplicates context, adds latency, consumes more budget, and creates new intermediate states. If the critic reviews a different version from the one that ultimately changed the repository, or an escalation repeats effectful tools, apparent redundancy can reduce reliability rather than improve it.

On September 30, 2026, GitHub announced that HydraFusion, previously limited to Copilot CLI, is now available as a research preview in Visual Studio Code and the GitHub Copilot app. The system chooses among three patterns per request: Single, Cascade, and Critique. The central news is its expansion into everyday work surfaces, announced one Lima calendar day before this article. It is not a new model or a promise that every task will receive multiple passes.

Chapter 02

The mechanism: select a workflow, not only a model

Single gives the task to one model. Cascade starts with an efficient model and applies a quality gate that either accepts the result or escalates it to a stronger model. Critique separates roles: one model creates the draft, a tool-less critic from another family reviews it, and the first model performs one revision. HydraFusion uses signals for reasoning, code generation, debugging, and tool use to choose the pattern it expects will be sufficient to meet its quality bar.

The distinction from Auto matters. Auto selects one model per request; HydraFusion can also select the execution topology within the turn. That lightweight decision does not generate the response, but it determines what context is shared, how many passes are billed, and how long the outcome takes. The documentation says the model pool can change and users cannot select its members. An enterprise evaluation should therefore observe workflow behavior rather than depend on a named combination that may differ tomorrow.

Chapter 03

1. Freeze the task contract before evaluating the router

A router can only look intelligent when the task has a testable outcome. Define the starting commit, allowed files, tools, permissions, timeout, required checks, and rejection criteria. Separate simple tasks, cross-file bugs, migration changes, and ambiguous requests; their review needs are not equivalent. Include cases that should end without a patch, such as missing credentials, contradictory requirements, or an action requiring human approval. This lets you verify whether Single avoids unnecessary work and whether Cascade or Critique appear when they can add value.

Keep the prompt, repository, dependencies, and graders fixed when comparing policies. Record the chosen pattern and models as trajectory evidence, not as a permanent product definition. Run repetitions when risk warrants them, because one correct selection does not prove stability. The goal is not to force an ideal workflow distribution; it is to learn which task classes reach an accepted outcome, which escalate, and which should stop before touching the workspace.

Chapter 04

2. Evaluate the gate that decides whether to accept or escalate

In Cascade, the decisive component is not the stronger model but the gate that makes it avoidable. Measure false accepts: drafts that clear the gate and later fail tests, scope, or review. Also measure false escalations: correct solutions that consume another pass without changing the final decision. Segment these errors by task type and severity. A conservative gate can raise quality and cost together; a permissive gate can show good latency while allowing rare but critical failures through.

Do not reduce the gate to a textual self-assessment by the same system. Combine deterministic checks such as compilation, tests, and file lists with task-specific graders and sampled human review. For an API change, for example, require contract compatibility, negative tests, and no secrets; a persuasive explanation does not compensate for breaking them. Preserve the escalation reason and rejected artifact. Without that evidence, you only know that more compute was used, not whether the router detected a real deficiency.

Chapter 05

3. Count cost and latency across the complete workflow

GitHub bills every model HydraFusion uses at its standard rate, and the Auto discount does not apply. The correct cost includes routing, drafting, critique, revision, escalation, retries, and fallbacks. Latency should run from request to acceptable artifact, including tools and checks. Report medians and tails rather than averages alone: a Critique route may look reasonable in the center while still missing the SLA when the repository or conversation is large.

Context is part of the economics as well. HydraFusion keeps the primary model when possible to benefit from cached tokens and gives the critic only what it needs, but each pass follows its model's context limit. If the system requires conversation compaction, record what evidence was lost and whether the outcome changed. Compare against Auto and a fixed model using the same unit: accepted task. Token savings have little value if they increase human correction or repeat external effects.

Chapter 06

4. Verify critic isolation, independence, and diversity

Critique promises an independent perspective: the reviewer belongs to another model family and runs without tools. That isolation reduces the chance that it will modify the repository, but it does not guarantee a useful critique. Verify that it receives the contract, diff, test results, and only the necessary context; too little encourages invented problems, while the solver's full narrative can reproduce its assumptions. Classify valid, irrelevant, and missed findings on failures you seeded deliberately.

Also inspect how the solver uses feedback. Accepting every critique can degrade a correct solution; ignoring it makes the second pass decorative. The record should connect each finding to the decision to apply or reject it and the resulting change. Include permission, security, concurrency, and compatibility cases, not only style. Provider or family diversity is an architectural signal, but practical independence is demonstrated when the critic finds different failures and improves complete acceptance without introducing drift.

Chapter 07

5. Protect repository state and cancellation

The documentation warns about a key operational limit: when HydraFusion discards a draft, changes that draft already made in the workspace are not automatically undone. The final artifact therefore cannot be validated only against the visible response. Capture initial state, executed commands, touched files, and final diff. Run every case in an isolated checkout and define what happens on cancellation, timeout, or gate failure. A cancelled result must remain identifiable and never blend silently into the next attempt.

Extend the same principle to external tools. Use idempotency keys, test environments, and least privilege; a tool-less critic does not undo a ticket, deployment, or message created by the solver. Require orchestration to apply the patch only after the complete workflow validates, or to produce a reviewable change instead of mutating the target system. GitHub describes fail-safe application as an internal principle, but your process must still verify the material state left around the agent.

Chapter 08

Hypothetical example: a cross-file bug with a migration

Suppose a service loses events while retrying a schema migration. The contract fixes a commit, six allowed files, an ephemeral database, and three tests: duplicate, rollback, and concurrency. Single should solve straightforward cases. Cascade escalates when the first patch fails a deterministic test. Critique is useful when the patch passes tests but needs a separate reading of idempotency and compatibility. The critic receives the contract, diff, and results; it receives neither credentials nor workspace access.

The team approves the policy only when it clears five gates: complete correctness, respected scope, total cost within budget, latency within the SLA, and a clean workspace after failures or cancellations. It compares acceptance and cost against Auto and a fixed model. If Critique finds more problems but adds side changes when applying feedback, it fails scope. If Cascade escalates almost every case, the gate or task segmentation needs work. This example is illustrative; it does not describe a Wasyra implementation or result.

Chapter 09

Conclusion: adopt the router as a changing policy

GitHub published favorable offline results for selected HydraFusion configurations on three benchmarks, but the original research article is dated September 4 and is context here, not the news event. Results depend on benchmark revisions, model pool, tuned policies, and pricing assumptions. CheckpointBench is internal, and the vendor says the preview will test translation to real workloads. We did not reproduce those benchmarks or present their savings as a transferable expectation.

The useful decision is smaller: enable the preview for a bounded group, use substantial and well-defined tasks, capture pattern, passes, cost, latency, and state, and keep a control using Auto or a fixed model. Re-evaluate when HydraFusion's models or behavior change. The documentation states that there is no SLA and the research preview is not intended for production workloads. That makes it an evaluation target, not a default for critical workflows.

Multi-model orchestration is valuable when it buys verifiable acceptance, not when it merely accumulates opinions. If you need to design a pilot with task contracts, permissions, traces, and release gates, Wasyra's custom agent service can help build that layer. It is an implementation service separate from the product available at agents.wasyra.com.

Written by

Wasyra AI Systems

Trust, copilots, and enterprise adoption

Wasyra AI Systems covers guardrails, suggestion-first modes, and review design so work assistants earn real adoption.

CopilotsTrustB2B AI
More from this author

Series

AI systems that actually reach production

A series on agents, copilots, and guardrails for bringing AI into real work without breaking trust or operations.

Posts in this series

Keep reading

Keep reading