The 3B Active Parameter Number Does Not Tell You What Fits

The 3B Active Parameter Number Does Not Tell You What Fits

A model can activate only 3 billion parameters per token and still need roughly 20 GB just to hold one quantized copy of its weights.

That is the uncomfortable detail hidden inside the current excitement around Holo4-35B-A3B, an alleged open-weight model for computer use. A September 30 post from The Broke Vibe Coder claims the model combines a 35 billion parameter mixture of experts design with about 3 billion active parameters, vision, GUI control, code execution, tool calls, MCP, and API access.

The post also claims Apache 2.0 licensing, open weights, GGUF availability, a 262K context window, and benchmark results of 80.8% on OSWorld, 30.9% on OSWorld 2.0, and 85.4% on a held-out MCP workflow benchmark.

Those claims are not independently confirmed here. The source account itself says the model had been out for roughly two days and that real hardware testing was only starting. Treat the specifications and scores as alleged until the model publisher provides a model card, weights, license text, evaluation harness, and reproducible results.

The post is still worth examining because it exposes a recurring mistake in local-agent discussions: confusing token-time compute with deployment requirements.

Active parameters are a compute number

In a mixture of experts model, the router selects a subset of experts for each token. The active parameter count describes the part of the network used for that token’s forward pass.

That can reduce inference compute compared with a dense model containing the same total number of parameters. It does not mean the unused experts disappear.

The claimed Holo4 Q4_K_M footprint is about 20 GB. That number is more operationally important than the 3B headline. The full set of quantized weights still needs to be available to the runtime, along with the vision projector, runtime buffers, and the key value cache.

A 24 GB GPU is therefore not a 24 GB model machine. It is a machine where a particular quantization might fit with a constrained context and careful memory accounting. The difference matters even more when the agent is expected to inspect screenshots, hold long tool traces, and switch between interfaces in one workflow.

The simple capacity model looks like this:

runtime memory = model weights + vision components + KV cache + workspace

The active parameter count mostly informs the compute term. It does not bound the other terms.

The context window is another trap

The same post claims a 262K context window. That is a useful upper bound only if the deployment can pay for the cache that stores attention history.

A computer-use agent does not consume context as a neat block of text. It accumulates screenshots, accessibility trees, action results, error messages, source files, tool schemas, and intermediate plans. A long context window can make a demo possible while making a sustained session expensive or unstable.

The practical question is not whether the model accepts 262K tokens. It is whether the full agent loop can keep enough of that history in memory while reserving capacity for the next action.

That is why a local operator should measure at least three separate things:

  1. Weight fit at the intended quantization.
  2. KV-cache growth at the intended context length.
  3. End-to-end action latency while vision and tools are active.

A model that looks small in a benchmark table can still be the wrong model for a 12 GB workstation or a long-running agent on a shared host.

One model for every interface is an architectural bet

The alleged feature list is more interesting than the parameter count. Holo4 is described as switching between screenshots, GUI actions, code, tools, MCP, and APIs inside one workflow.

If that description holds up, the model is making a bet against the usual collection of specialist components. Instead of routing visual understanding, computer control, code generation, and tool selection to separate models, the system tries to teach one action model the interfaces together.

That could simplify the control loop. There are fewer handoffs, fewer translation formats, and fewer model-specific adapters.

It could also make failures harder to diagnose. When a workflow fails, the operator needs to know whether the problem came from visual grounding, action selection, tool schema interpretation, permission handling, or the model’s plan. A single generalist can compress those failure modes into one opaque trace.

The right comparison is therefore not “35B versus 3B.” It is:

One multimodal action model

Advantage: fewer handoffs across GUI, code, and tools.

Operational cost: harder failure attribution and potentially larger memory footprint.

Routed specialist models

Advantage: clearer responsibilities and smaller per-task models.

Operational cost: more orchestration, serialization, and interface glue.

Dense smaller model

Advantage: simpler deployment and predictable memory.

Operational cost: less total capacity for unusual tasks.

The best design depends on the verifier around the model. A computer-use agent needs checks for the state it believes it sees, the action it is about to take, and the result it claims to have achieved.

The security boundary does not shrink with the model

The source post describes code execution, MCP, API access, and GUI control as capabilities. Those are not just features. They are trust boundaries.

A local model may keep source code and logs off a hosted inference API, which is a real privacy benefit if the runtime is configured correctly. It does not make arbitrary tool calls safe.

An agent that can click, execute code, call an API, and invoke an MCP server needs explicit separation between observation and action. It needs least-privilege credentials, isolated workspaces, approval gates for irreversible operations, and an audit trail that records the model decision, the tool payload, and the result.

The model’s open weights do not provide those controls. The harness does.

That is the part the active-parameter discussion routinely misses. A lower compute bill can make it easier to run a capable agent locally. It does not make the agent trustworthy by default.

What operators should verify

If the Holo4 claims survive publication of the weights and model card, the first test should not be a leaderboard reproduction. It should be a memory and control-loop test on the hardware where the agent will actually run.

Measure the quantized weight size, peak resident memory, cache growth, screenshot latency, tool-call latency, and failure recovery. Run the same task with a hard limit on filesystem access and network credentials. Record when the model asks for an action outside its permitted scope.

For the evaluation claims, require the task list, the exact model revision, the prompt and tool contract, the number of attempts, and the verifier. A score without those details is a marketing number, whether the model is closed or open.

The active number is not the deployment number

The Holo4 discussion may turn out to be accurate, exaggerated, or somewhere in between. The primary source available for this article is one early X post, and that post itself calls for skepticism while testing is still beginning.

The durable lesson does not depend on the release surviving scrutiny. Mixture of experts routing can reduce token-time compute while leaving weight memory, cache memory, interface complexity, and tool risk intact.

For local agents, the model you can run is not the model with the smallest active parameter count. It is the model whose complete control loop fits your memory budget, verifier, and trust boundary.

Sources

Keep reading