Update · local AI / 3 MIN READ
Budget a local AI model for the machine it needs.
NVIDIA’s Nemotron 3.5 Lightning has 30B total and 3B active parameters. The smaller active figure does not tell a buyer how much memory, hosting or support the deployment needs.
WATCH until a local model passes the real workload on the proposed machine. ACT on total memory, language, licence and recurring cost. PASS on choosing a host from the “3B active” headline alone.
What is checked
NVIDIA’s August 11 model card states 30B total and 3B active parameters and OpenMDW 1.1 terms. Ollama lists builds with different formats and sizes. We have not run this model; catalogue sizes are not runtime measurements.
What this cannot prove
No throughput, answer quality, minimum host size or cost saving is established for a VKV or client deployment. Declared language coverage is not a tested quality guarantee in every market.
The attractive part of the announcement
NVIDIA describes a model using a subset of its total parameters for each generated token. That active subset concerns part of the computation; it does not remove the remaining weights from the deployment.
A buyer still pays for a usable host, storage, runtime memory and operations. Listed model-file sizes differ by format. Neither a parameter headline nor a download size measures the memory and response time of your full assistant.
Fit is more than memory
For France, Sweden, Finland, the UK, the US and Slovenia, test the actual document and question languages. The model card names French and English among its declared use-case languages; omission of Swedish, Finnish or Slovenian from that list is not proof that they are unsupported.
Review the actual OpenMDW 1.1 licence and distribution plan separately from technical fit. Ask for the hardware tier, selected model format, support responsibility and recurring costs. Those are proposal inputs, not a measured quotation in this briefing.
What to test before switching
ACT by comparing the proposed model with the existing approach on the same document questions. Check source-grounded answers, missing-evidence responses, your languages, memory and response time under the intended number of users. Name acceptance limits before the trial.
WATCH if a larger host or additional operational work is needed without a demonstrated benefit. For assistants using tools, also test missing or hostile tool results and enforce permissions outside the model. PASS on replacing a working system solely because a new model was announced.
Sources and artefacts
APPLY THE METHOD TO YOUR BUSINESS
A useful answer starts with your own evidence.
The on-prem AI scope is relevant when the task needs a private model and a defined operating boundary. Agree documents, languages, workload, hardware and acceptance cases before selecting a larger host.
Explore on-prem AI ↗Service details and current pricing are on VKVstudio.com. Project scope and agreements are handled by email.