Run an open-source model locally
Set up an open-weight model on your own machine, with the security trade-offs stated plainly.
Jamie Owen Updated 26 Sept 2026 fast-moving: check a current source
Open-weight releases move faster than anything else covered on this site.
Running a model locally means the text you send it never leaves your machine. For sensitive work, that is the whole point: no tenant boundary to reason about, because nothing crosses one. The trade-off is that you take on everything the cloud provider used to handle, from capability to updates to security.
What you need
A recent open-weight model, and a runner that hides the complexity. Tools like Ollama or LM Studio let you pull a model and chat with it in a few minutes. The practical limit is memory: a small quantised model runs on a modern laptop, while larger ones want a capable GPU and plenty of RAM.
The families worth starting with change every few months. As of September 2026 the compact end includes Google’s Gemma 4 and OpenAI’s gpt-oss-20b, which OpenAI says runs “within 16GB of memory” and which it suggests running through Ollama on ordinary hardware. Alibaba’s Qwen3.8-27B, released on 14 August, is a step up in size at 27 billion parameters, and wants memory to match. Check a current roundup rather than trusting a name you read last year.
Start small. A 7–8B parameter model in a quantised form is enough to judge whether local inference fits your workflow before you invest in hardware.
”Open” no longer means “runnable”
Worth setting expectations before you go shopping. Several 2026 open-weight releases are at frontier scale: Alibaba’s Qwen3.8 flagship, out on 12 August, has 2.4 trillion parameters, and Z.ai’s GLM-5.3 has 744 billion. The weights are open to download and modify, each under its own licence, but serving them needs a cluster, not a workstation.
So “open-source model” now spans two very different propositions: sovereignty over your infrastructure at datacentre scale, and a decent assistant on the machine in front of you. This page is about the second. Filter any release announcement by parameter count against your actual hardware before it earns your attention.
The security trade-offs, honestly
Local isn’t automatically secure. It’s differently secure.
You own patching now. A model or runner with a known issue stays vulnerable until you update it, and nobody does that for you. The supply chain matters for the same reason: pull weights and tooling from sources you trust, because a model file is data but the surrounding tooling is code that runs on your machine.
There’s a capability gap too. Local models are smaller than frontier hosted ones, and on hard reasoning they will underperform, so verify output as carefully as ever (see how these tools work). And the data still needs handling: “on my machine” isn’t “safe” if the machine is unencrypted, shared, or backed up to somewhere uncontrolled.
Local inference removes the cloud boundary problem and hands you the maintenance and hardening that came with it. That’s a fair trade for sensitive work, if you actually do the maintenance.
Where this fits
Local models are one answer to the shadow-AI problem: give people a sanctioned, private option so they stop pasting secrets into random web tools. Pair it with real data-loss-prevention controls rather than treating it as a substitute for them.
Sources
Everything above was checked against these on 26 Sept 2026. Providers change things without notice. If a detail matters to a decision, follow the link.