Local AI for coding agents: five boundaries you need to measure

Microsoft and GitHub announced local inference for coding agents and explicit local-model selection. The development opens a useful architecture for privacy, latency, and continuity, but it also exposes a critical distinction: where the model runs, where data travels, and what the agent can execute are separate decisions. This guide proposes five boundaries and an evaluation plan before adopting local AI.

Local AICoding agentsOn-device inferenceAI evaluation
Wasyra Engineering
Modernization, architecture, and reliable delivery
Published
min read
8 min read
Categoría
Engineering
5boundaries for a local agent
A local workstation, a cloud, and an isolated tool environment connect through a router with independent network and execution controls.

Chapter 01

Local is not a single property of the agent

On October 7, 2026, Microsoft and GitHub introduced an architecture in which GitHub Copilot will be able to decide whether a task uses on-device inference or cloud models. They also enabled Copilot CLI to discover models served by a local Ollama instance. The news matters because it moves the compute decision into the agent's real workflow: context, cache, tools, and a multi-turn session. But the term local can invite a false conclusion. Processing tokens on the machine does not prove that the client is offline, telemetry is disabled, no remote provider receives context, or commands execute with least privilege.

The decision for a CTO is not local versus cloud as two mutually exclusive products. It is which boundary each task class needs and what evidence will prove that the boundary holds. A small fix involving sensitive files may prioritize data residency. A complex migration may need a more capable remote model. An overnight automation may value continuity without internet while demanding strong isolation for shell and credentials. Before buying hardware or changing models, separate five contracts: inference, network, data, tools, and outcome. Otherwise, the team will optimize a label while leaving operational risk unchanged.

The central development is two calendar days old when this article is published. The memory, throughput, and benchmark figures disclosed by Microsoft come from its October 5 tests on specific hardware and configuration. They are provider results, not measurements reproduced by Wasyra or promises that transfer to other machines.

Chapter 02

The mechanism: inference, orchestration, and execution are separate layers

A coding agent is not just a model. The host gathers instructions and context; a router selects an endpoint; the model proposes text or tool calls; the runtime validates arguments and executes processes; finally, the host observes files, tests, and repository state. Local inference mainly changes the second step. It can reduce prompt transit and avoid network latency to the model, but it does not automatically redefine the other components. A remote MCP server, web search, telemetry event, or command that downloads dependencies can still leave the device.

Microsoft describes a quantized local model with 137 billion total parameters and 6.8 billion active parameters, plus speculative decoding to accelerate responses. Quantization reduces precision and memory; speculative decoding uses an auxiliary model or process to propose token blocks for the target model to verify. Both optimizations have system effects: the model becomes smaller, but context, KV cache, runtime, applications, and the operating system compete for the same memory. Keeping it loaded avoids part of the startup cost without guaranteeing constant latency as the session grows.

A hybrid router adds another variable. It can preserve cached work and move turns between edge and cloud, but every transition needs an explicit rule: which context is sent, which summary is generated, which tools remain available, and what happens if the selected model cannot support a call. Automatic selection is safe only when an external policy constrains destinations and data. If the conversation itself freely decides when to leave the device, the residency promise becomes subordinate to the very system it is meant to control.

Chapter 03

Five boundaries that must pass separately

Define each boundary as an observable claim, not a configuration preference. A selector showing a local model proves which endpoint was chosen for that turn; it does not prove the other limits. The following matrix prevents marketing, platform, and security from using the same word for different contracts.

Minimum contract for claiming a workflow is local
BoundaryQuestionPassing evidence
1. InferenceWhich endpoint processes every turn and fallback?Router traces identify model, location, and reason without exposing the prompt.
2. NetworkWhich destinations may receive traffic during the task?A test with egress blocked completes or fails explicitly, with no unexpected connections.
3. DataWhich files, excerpts, and metadata leave the boundary?The data inventory and redacted logs match the per-task policy.
4. ToolsWhich processes, paths, secrets, and services can it touch?Operating-system controls block ungranted access even when the model requests it.
5. OutcomeDoes it complete the task with operable quality and cost?Task tests measure acceptance, time, energy, memory, and failure recovery.
Local can accurately describe inference and still be false for the whole session. Always document the scope of the claim.

Chapter 04

Hypothetical example: fixing a financial SDK without exporting sensitive code

Imagine a team maintaining a payment SDK that needs to fix idempotency validation. The repository contains internal contracts, anonymized fixtures, and integration configuration. The task appears ideal for local inference: read three modules, propose a patch, and run tests. The policy allows read and write only in a disposable checkout, execution of the test chain without network access, and read access to local documentation. It denies the user's keychain, other folders, host sockets, and remote endpoints.

This example is hypothetical. The first trial uses an explicitly selected local model with egress blocked. If a dependency is missing, the agent cannot silently open the network: it returns a blocked state naming the required package. A second profile preinstalls verified dependencies. The team observes whether the patch compiles, whether new tests capture the defect, and whether the diff avoids out-of-scope files. The model's prose is secondary; repository state and tests are the evidence.

For complex cases, the team may allow remote fallback only after classifying and reducing context. The router sends a minimal reproduction without secrets or customer names, records the fallback reason, and requires confirmation before expanding the data. That workflow remains hybrid, not local. Naming it precisely lets product compare utility and security audit the exception without turning the whole session into a black box.

Chapter 05

How to implement without confusing privacy and performance

Start with a task matrix, not a model catalog. Classify short edits, repository search, migrations, test generation, and autonomous automation by sensitivity, context, tools, latency tolerance, and consequence of error. Then assign a profile: local required, hybrid with governed fallback, or remote allowed. Routing should be deterministic for sensitive tasks. A residency policy that depends on an opaque Auto heuristic is not a verifiable policy.

Also separate inference configuration from tool configuration. The local endpoint needs identity, version, streaming and tool-calling capabilities, context limits, and an update strategy. The runtime needs file-system, network, process, credential, and user-interface permissions. Microsoft Execution Containers formalizes this separation: policy remains outside the agent and maps to native controls. The documentation also clarifies that some built-in tools are checked inside the host and remote MCP servers sit outside the process sandbox; those exceptions belong in the threat model.

Prepare explicit degradation states. Insufficient memory, an unloaded model, an overlong context, a down endpoint, or an unsupported tool must not trigger the cloud by surprise. The system can summarize, split the task, ask permission for a fallback, or stop with an actionable reason. Record selected location, version, context size, peak memory, invoked tools, network destinations, and task outcome while redacting sensitive content. That telemetry improves routing without storing the code that was meant to remain local.

Chapter 06

The right evaluation covers a complete task, not tokens per second

Build a representative set with disposable repositories and verifiable outcomes. Include a trivial edit, a multi-file bug, a missing dependency, context that exceeds the expected memory budget, an unsupported tool, an attempt to read outside the checkout, and a network loss midway through the session. For each case define the allowed diff, required tests, forbidden accesses, time limit, and expected blocked-state output. Run enough repetitions to expose variation; one successful demonstration does not prove reliability.

Measure four layers. Quality: accepted tasks, tests that actually detect the defect, and regressions. System: time to first useful action, full duration, peak memory, energy, and stability as context grows. Boundaries: connections made, files read, denials, and fallbacks. Operations: time to install or update models, recover a session, and explain a failure. Tokens per second help diagnose performance, but they do not show whether the agent finished sooner or produced a correct change.

Compare local, hybrid, and cloud on the same task version and tools. Do not compare a quantized laptop model with a different remote model on different prompts. Report ranges and failures, not only averages. The promotion criterion may require zero egress in the local profile, zero out-of-scope access, a minimum accepted-task rate, and time and memory budgets. If local quality falls short, narrowing scope may be right; if cloud violates residency, that profile is disqualified even when it is faster.

Chapter 07

Limits and the adoption decision

Local AI shifts costs: less API consumption can mean more hardware, weight distribution, storage, energy, support, and security patching. Teams must also govern licenses, model provenance, and versions. An OpenAI-compatible endpoint simplifies integration, but protocol compatibility does not guarantee equivalent tool calling, context windows, or quality. A model available on the device can also become unusable when other applications compete for memory.

Adopt local when a stable task class benefits from residency, continuity, or latency and passes all five boundaries. Use hybrid when the value of a remote path justifies an explicit, minimized, auditable exception. Keep cloud when required capability dominates and data may leave under an accepted contract. A mature architecture does not promise that everything happens in one place: it shows, enforces, and tests where every part happens.

The practical next step is to choose a low-impact task, write its permission and data matrix, and run the same evaluation across all three profiles. Do not start by deploying the largest model that fits. Start by proving that the right workflow fits inside the right boundary. Wasyra can help turn that proof into an engineering slice with metrics, isolation, and reversible rollout, without confusing local inference with a fully isolated agent.

FAQ

Frequently asked questions

Does a local model make the agent work offline?

Not necessarily. The host, telemetry, tools, MCP servers, or configured provider may still use the network. Offline mode and egress must be configured and tested separately.

What should I measure before moving a coding agent on device?

Measure task outcome, end-to-end latency, memory and energy, network connections, file and tool access, fallback behavior, and recovery. Compare on the same tasks and versions.

Written by

Wasyra Engineering

Modernization, architecture, and reliable delivery

Wasyra Engineering documents patterns for moving legacy systems without freezing delivery or breaking ownership.

LegacyRefactorArchitecture
More from this author

Series

AI systems that actually reach production

A series on agents, copilots, and guardrails for bringing AI into real work without breaking trust or operations.

Posts in this series

Keep reading

Keep reading