Section 37 of 45 9 min read

Running Models Locally

The five reasons ranked by how often each is the real one, what is available, the hardware arithmetic, the serving stack, and the cost sheet.

Objective

Cover self-hosted open-weight models: when they are the right choice, the hardware arithmetic, the serving options, the licensing and provenance questions, and where the limits are.

This section closes Part V because tokenization, attention, sampling and grammar masking are all at their most controllable when you host the model yourself. Not exclusively: hosted APIs expose the sampling parameters, and caller-supplied grammars are available on some of them, including a GBNF grammar over final assistant text on an ordinary serverless endpoint (§32.5). What self-hosting adds is the masking implementation, the runtime, the absence of a provider-side complexity ceiling, and the pieces with no hosted analog at all—KV cache manipulation, custom sampling code. Local inference is where the mechanics stop being explanatory and become fully adjustable.

37.1 The Five Reasons, Ranked by How Often They Are the Real One #

Data residency and confidentiality

This is the strongest reason, and the one that survives scrutiny. Some code, some regulated data, and some air-gapped environments cannot leave the building. No contractual assurance changes that constraint. If this is your reason, the analysis is short: you self-host or you do not use a model.

Reproducibility

A model you host does not change under you. No silent serving-stack update, no model deprecation, no batch-dependent expert routing from other tenants’ traffic (§30.4). For anyone building evaluations or regulated systems, this is underrated.

Latency and offline operation

Local inference has no network round trip and works without connectivity. For completion-style workloads at high call frequency, this is a genuine experience difference.

Cost at sustained volume

This reason holds and is also the most commonly overstated. The break-even is discussed in §37.5 and it is further out than most estimates suggest.

Control over the mechanics

Arbitrary grammar constraints (§32.5), full logprob access, custom sampling, and KV cache manipulation are available locally and mostly not through hosted APIs. For a narrow set of applications this is decisive.

If your reason is only the fourth, work the arithmetic before committing.

37.2 What Is Actually Available #

The open-weight landscape changed materially with the shift to mixture-of-experts (§30). Coding-capable models now ship in sizes that fit a single workstation.

A representative current example: Qwen3.6-35B-A3B, released April 2026 under Apache 2.0, holding 35 billion parameters and activating 3 billion per token, reported by its authors at 73.4 percent on a verified software-engineering benchmark.66 That figure is vendor self-reported and should be treated as such—the general pattern of benchmark contamination and the specific practice of labs reporting their own numbers both apply, and §34 exists precisely because a public benchmark score is not a prediction about your repository.

The larger point is directional, and it survives the caveat: open-weight coding models moved from “small and weak” to “credible for a defined set of tasks” over roughly two years, and the mechanism was MoE rather than any breakthrough in dense scaling.

37.3 The Hardware Arithmetic #

The estimate that matters is simple enough to work out in your head:

memory for weights (GB) ≈ total parameters (B) × bits per weight ÷ 8

At 4-bit that is roughly half the parameter count in gigabytes. A 35B model needs about 17.5 GB for weights; the same model at 16-bit needs about 70 GB.69

Fig. 13—Weight MemoryWeight memory by parameter count and quantization

Weights, however, are not the whole requirement. The KV cache scales with context length and, as §29.1 showed, grows large quickly. Budget for both.

def vram_estimate(total_params_b: float, bits: int, ctx_tokens: int,
                  layers: int, kv_heads: int, head_dim: int,
                  kv_bits: int = 16) -> dict:
    """Rough planning estimate. Real usage varies by runtime and batch size."""
    weights = total_params_b * bits / 8                       # GB
    kv = (2 * layers * kv_heads * head_dim * ctx_tokens * kv_bits / 8) / 1e9
    overhead = 1.5                                            # activations, runtime
    return {"weights_gb": round(weights, 1),
            "kv_cache_gb": round(kv, 1),
            "total_gb": round(weights + kv + overhead, 1)}

# A hypothetical 48-layer full-attention/GQA model at 35B — NOT the
# Qwen3.6-35B-A3B card above, whose numbers differ sharply; see below.
print(vram_estimate(35, 4, 32_768, layers=48, kv_heads=8, head_dim=128))
# {'weights_gb': 17.5, 'kv_cache_gb': 6.4, 'total_gb': 25.4}

Those layer and head arguments are the load-bearing part, and they are a stand-in rather than a reading of any card. Substitute the actual Qwen3.6-35B-A3B configuration—40 layers of which only 10 are attention, 2 KV heads, head dimension 256—and the attention KV term at 32K falls to roughly 0.7 GB against the 6.4 GB above.66 That is not a deployed total either: the thirty gated-linear layers keep recurrent state, quantized weights carry block metadata the ideal 0.5 bytes per parameter omits, and allocator and runtime overhead are not free. The lesson is the shape of the calculation, not the 25.4. For a named model, read its config and measure.

Three levers exist when the number does not fit. Weight quantization, where 4-bit is the standard compromise and preserves most quality for coding work.69 KV cache quantization, where moving from 16-bit to 8-bit halves the quantized tensors—not necessarily the whole cache, since scales, unquantized layers and auxiliary state are not halved with them—and whose backend requirements are runtime-specific rather than a universal flash-attention dependency: vLLM, for instance, supports per-tensor and per-head FP8 schemes with different backend coverage.78 Name the runtime, the backend, the source and target dtypes, and whether you are quantizing K, V or both, then validate quality. Reducing context length, the lever people forget, though §12.1 argues you should be operating well below the maximum anyway.

Quantization is lossy, and the loss is not uniform across tasks. Treat a quantization change as a model change requiring your eval suite to run (§34.8).

37.4 The Serving Stack #

Three tools cover essentially all of it, and the choice follows from what you are doing rather than from preference.

RuntimeBest forTrade-off
OllamaGetting started, single-user workstationsLeast control; opinionated defaults
llama.cppHardware-level control, CPU and Apple SiliconMore configuration; you manage quantization
vLLMMulti-user serving, throughputGPU-oriented; server-shaped operational surface

All three expose an OpenAI-compatible endpoint, which is what makes them useful: most coding agents accept a base-URL override, so pointing an existing agent at a local model is a configuration change rather than an integration project. vLLM additionally runs in-process for offline batch work through LLM.generate() and LLM.chat(), which is the right shape for eval runs and dataset generation—useful to know before standing up a server to do something that does not need one.78

# llama.cpp, serving a quantized MoE model with an OpenAI-compatible API
./llama-server \
  --model Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf \
  --ctx-size 32768 \
  --flash-attn on \
  --cache-type-k q8_0 --cache-type-v q8_0 \
  --temp 0.6 --top-p 0.95 --top-k 20 \
  --port 8001
# endpoint: http://localhost:8001/v1
# Point an existing agent at it
export OPENAI_BASE_URL="http://localhost:8001/v1"
export OPENAI_API_KEY="local"

The sampling parameters in that command carry more weight than they appear to. Model authors publish recommended sampling settings, and unlike hosted APIs, where the provider has already tuned them, locally you get the defaults you type. Using a model at the wrong temperature is a common and invisible cause of “the local model is bad.”

Expect an integration gap rather than a clean substitution. Harness behavior is tuned per model (§6.2), and a harness whose system prompt, tool schemas, and conversation management were tuned against a frontier model will not automatically get the best out of a different one. This is the §6 argument applied in reverse, and it is the most common reason local model evaluations come out worse than the benchmark numbers suggest.

37.5 The Full Cost Comparison #

The comparison practitioners make sets hardware amortization against API spend. The comparison that matters, however, includes everything else.

CostHostedLocal
Per-tokenMeteredZero marginal
HardwareNoneCapital, or cloud GPU rental
Electricity and coolingNoneReal at sustained load
OperationsNoneSomeone owns uptime, upgrades, capacity
Model updatesAutomaticYou re-download, re-quantize, re-evaluate
Idle capacityNoneYou pay whether or not anyone is working
EvaluationOnce per vendor changeEvery model, quantization, and runtime change

The last two rows are what practitioners leave out. A workstation-class GPU is idle overnight and on weekends, and a self-hosted deployment inherits an evaluation obligation that a hosted API mostly absorbs on your behalf.

At low-to-moderate volume, hosted wins on total cost, and it is not close. At sustained high volume with an existing platform team, local can win. If the only argument is cost, run the numbers with all seven rows before committing, and note that the tier routing in §27 gets much of the same saving with none of the operational burden.

Where local wins outright is on the first two reasons in §37.1—confidentiality and reproducibility. Those are not cost arguments and they do not need one.

37.6 Weights, Licenses, and Training Data #

Open weights are not open source, and the distinction carries real obligations.

License classes

Apache 2.0 and MIT are permissive: commercial use, modification, and redistribution with essentially no field-of-use restriction. Custom “community” licenses frequently add restrictions—user thresholds, use-case carve-outs, naming requirements, and acceptable-use policies that bind downstream. Read the actual license, not the “open” label on the download page.

What “open weights” excludes

You get parameters. You almost never get the training data, the training code, or the full data-provenance record. This means you cannot independently audit what the model was trained on, cannot reproduce it, and cannot answer a data-provenance question with anything stronger than the publisher’s own statement.

Why this matters even though you are not training anything

Three questions land on any enterprise adopting an open-weight model, and none is answerable from the weights alone. Whether the training corpus included copyrighted or licensed code, and what that implies for output you ship. Whether it included personal data, and what that implies under privacy regimes. And whether the publisher’s provenance claims are auditable or merely asserted.

Litigation and regulation in this area are unsettled and moving, and this document does not offer a legal position. The practical point is procedural: the provenance question does not go away because you self-hosted; it moves from your vendor’s legal team to yours. A hosted provider typically offers contractual indemnification. A downloaded weights file offers a license file.

Fine-tuning is out of scope here

Adapting an open-weight model on your own data is a legitimate technique with its own literature, and it raises a further set of data-governance questions—what goes into the tuning set, whether it can be extracted from the resulting model, and who may use the artifact. Those questions belong to a governance framework rather than a prompting one.1

37.7 A Realistic Deployment Pattern #

Local and hosted are not mutually exclusive. The useful pattern splits work by sensitivity and difficulty rather than forcing an all-or-nothing choice.

def select_backend(task_class: str, data_class: str) -> str:
    # Confidentiality is a hard constraint, evaluated first and never traded
    # against cost or quality.
    if data_class in {"restricted", "regulated"}:
        return "local"
    if task_class in {"extract", "classify", "format", "summarize"}:
        return "local"          # mechanical work; a small local model suffices
    return "hosted"             # hard reasoning, long-horizon agentic work

That ordering is deliberate. Confidentiality is a constraint, not a preference, so it is evaluated first and never traded against cost. Everything after it is the ordinary routing decision from §27, with a local tier added below the small hosted tier.

37.8 Failure Modes #

  • Adopting local for cost without the full sheet. Idle capacity, operations, and evaluation obligation are the rows people omit.
  • Treating a vendor benchmark number as a prediction. Self-reported, and not measured on your repository.66
  • Using the runtime’s default sampling. Model authors publish recommended settings, and a hosted API applies them on your behalf. A local runtime gives you whatever you type.
  • Expecting a drop-in substitution. The harness was tuned for a different model (§6.2), and that alone can account for a large quality gap.
  • Changing quantization without re-evaluating. It is a model change.
  • Assuming open weights answer the provenance question. They relocate it.
  • Sizing memory on active parameters. Memory follows total, throughput follows active (§30.2).

References cited in this section

4 of 81 · numbering matches the PDF

  1. 66Qwen team model card for Qwen3.6-35B-A3B together with Unsloth's local-deployment documentation for Qwen3-Coder, https://unsloth.ai/docs/models/tutorials/qwen3-coder-how-to-run-locally, verified September 8, 2026. Vendor documentation and self-reported benchmarks. The parameter counts, Apache 2.0 licensing and reported benchmark figure come from the model card; the Unsloth page is deployment guidance and is not the source of any number here. Cited as product fact for the total-versus-active parameter naming convention, the Apache 2.0 licensing, and the availability of 35B-A3B and 480B-A35B configurations. Note that the card describes a mixed architecture, gated-linear layers interleaved with attention layers, which is why §29.1's KV formula is labeled as full-attention or grouped-query arithmetic rather than as a universal law. The 73.4 percent verified-software-engineering-benchmark figure is the model authors' own reported number, is not independently replicated, and should not be read as a prediction about any particular repository—see reference 55 on why public coding benchmarks travel badly.huggingface.co/Qwen/Qwen3.6-35B-A3B ↗
  2. 69Practitioner guidance on local inference hardware sizing, including the weights estimate of parameters × bits ÷ 8, 4-bit Q4_K_M as the standard quality-versus-memory compromise, and KV cache quantization requiring flash attention support. Community and blog-level sources verified September 8, 2026; the arithmetic is verifiable independently and the quantization-quality claim is a rule of thumb rather than a measurement. Two limits belong with it. The ideal 0.5 bytes per parameter at four bits is a floor: Q4_K-family files carry per-block scale and minimum metadata, so an actual GGUF is larger than the formula. And the flash-attention requirement for KV-cache quantization is a property of particular runtimes and build options rather than a universal one—name the runtime and check its flags. Reference 78 is the vLLM side of that question. Any quantization change should be treated as a model change and re-evaluated (§34.8).
  3. 78vLLM documentation: "Batch invariance," "Reproducibility," https://docs.vllm.ai/en/latest/usage/reproducibility/, and "Offline inference," https://docs.vllm.ai/en/latest/serving/offline_inference/, all verified September 25, 2026. Project documentation. Cited for two corrections. Batch-invariant serving is an opt-in mode with an enumerated list of tested models, DeepSeek and Qwen MoE families among them, which is why §30.4 describes hosted nondeterminism as a property of default serving rather than an impossibility in principle; the project also tracks configurations where the guarantee does not yet hold, so the mode is scoped rather than universal, and reproducibility additionally requires pinning kernels, sampling, hardware and versions. And vLLM exposes offline in-process inference through LLM.generate() and LLM.chat(), so it is not API-only (§37.4). Also cited, against the penalty implementation in vllm/model_executor/layers/utils.py, for the §31.2 distinction between the two count-based penalties: the frequency term multiplies an occurrence-count tensor and the presence term a binary occurrence mask, so only the former scales with repetition. An earlier revision of this document described both as scaled by how often the token occurred.docs.vllm.ai/en/latest/features/batch_invariance ↗
  4. 1Joshua Davis, The AI SDLC: An Operating Model, Control Framework, and Maturity Progression for Engineering Organizations Building With Agents, v1.0, September 2026 The companion framework covering governance, controls, and organizational absorption. Cited here for scope boundaries rather than for evidence.jdav.is/ai-sdlc ↗
PDF↓