Update · local AI / 3 MIN READ

Budget a local AI model for the machine it needs.

NVIDIA’s Nemotron 3.5 Lightning has 30B total and 3B active parameters. The smaller active figure does not tell a buyer how much memory, hosting or support the deployment needs.

THE DECISION IN 30 SECONDS

WATCH until a local model passes the real workload on the proposed machine. ACT on total memory, language, licence and recurring cost. PASS on choosing a host from the “3B active” headline alone.

01 / EVIDENCE LEDGERHow I check ↗

What is checked

NVIDIA’s August 11 model card states 30B total and 3B active parameters and OpenMDW 1.1 terms. Ollama lists builds with different formats and sizes. We have not run this model; catalogue sizes are not runtime measurements.

What this cannot prove

No throughput, answer quality, minimum host size or cost saving is established for a VKV or client deployment. Declared language coverage is not a tested quality guarantee in every market.

01

The attractive part of the announcement

01 / ACTIVE / TOTALAn active-expert metaphor: the full set of weights still occupies storage.

NVIDIA describes a model using a subset of its total parameters for each generated token. That active subset concerns part of the computation; it does not remove the remaining weights from the deployment.

A buyer still pays for a usable host, storage, runtime memory and operations. Listed model-file sizes differ by format. Neither a parameter headline nor a download size measures the memory and response time of your full assistant.

02

Fit is more than memory

02 / THE FIT CHECKHardware capacity, language evaluation and licence are separate deployment checks.

For France, Sweden, Finland, the UK, the US and Slovenia, test the actual document and question languages. The model card names French and English among its declared use-case languages; omission of Swedish, Finnish or Slovenian from that list is not proof that they are unsupported.

Review the actual OpenMDW 1.1 licence and distribution plan separately from technical fit. Ask for the hardware tier, selected model format, support responsibility and recurring costs. Those are proposal inputs, not a measured quotation in this briefing.

03

What to test before switching

03 / THE TEST PLANCompare grounded answers, latency and tool permissions on the same workload.

ACT by comparing the proposed model with the existing approach on the same document questions. Check source-grounded answers, missing-evidence responses, your languages, memory and response time under the intended number of users. Name acceptance limits before the trial.

WATCH if a larger host or additional operational work is needed without a demonstrated benefit. For assistants using tools, also test missing or hostile tool results and enforce permissions outside the model. PASS on replacing a working system solely because a new model was announced.

Sources and artefacts

  1. NVIDIA: Nemotron 3.5 Lightning model card, August 11, 2026
  2. Ollama: NVIDIA Nemotron 3.5 Lightning
  3. Ollama: Nemotron 3.5 Lightning model tags and sizes

APPLY THE METHOD TO YOUR BUSINESS

A useful answer starts with your own evidence.

The on-prem AI scope is relevant when the task needs a private model and a defined operating boundary. Agree documents, languages, workload, hardware and acceptance cases before selecting a larger host.

Explore on-prem AI ↗Service details and current pricing are on VKVstudio.com. Project scope and agreements are handled by email.