Advanced Open source News

Open-source model roundup: September 2026

Qwen and Z.ai both shipped at data-centre scale, and Z.ai says its new model's hacking skills grew faster than it expected.

Benchmark and capability claims are the labs' own, quoted as theirs. Verify on your own tasks.

The open-weight releases worth knowing about since July both landed in August, and both are far too large for a laptop. One of them also arrived with an admission from its makers that matters more than its size.

Notable since July

  • Qwen3.8 from Alibaba’s Qwen team: the flagship Qwen3.8-2.4T-A95B on 12 August and Qwen3.8-27B on 14 August. The names give the size. The flagship has 2.4 trillion parameters, the learned numbers that make up a model, of which about 95 billion do the work on each token. The team’s own summary: “For the first time, Qwen3.8 brings a Qwen-Max-class model to open release.”
  • GLM-5.3 from Z.ai, in late August: 744 billion parameters with 40 billion active. It uses the same base model as GLM-5.2, so “every gain comes from post-training”, the further training done after the main run. A smaller sibling, GLM-5.3-Flash (320 billion, 18 billion active), is built on a new design meant to make long inputs cheaper to serve. Z.ai calls GLM-5.3 “the most capable open-weights model for coding”, a claim that rests partly on its own in-house benchmark.

For a laptop, the practical choices are still the older compact models. OpenAI says gpt-oss-20b runs “within 16GB of memory”, and Google’s Gemma 4 is one command in Ollama, a tool for running models on your own machine. Qwen3.8-27B is the only release in this edition small enough to consider outside a data centre.

The capability nobody can recall

Z.ai’s release notes for GLM-5.3 include this sentence: “As we scaled post-training, cyber capability developed faster than we expected.” The lab says the model is “state of the art on CyberGym for vulnerability discovery”, a test of finding security flaws in software, and “more than doubles GLM-5.2 on exploitation benchmarks”, which measure turning a flaw into a working attack.

Anthropic’s answer to the same risk is a gate. Its standard Claude Fable 5.1 can be used to find software vulnerabilities, “though not to develop exploits for them”, and a version with looser safeguards goes only to vetted US organisations, as September’s agent security edition explains. An open-weight model has no gate. Once the weights are listed for download, anyone with the hardware can run the model, and nobody can take it back.

For most readers this changes nothing about which model to use. It changes the threat picture: skills of the kind Anthropic rations are now open to anyone with the hardware to serve a 744-billion-parameter model.

How to evaluate a release

Check the parameter count against your hardware before you check the benchmark scores, and read the licence before you build on anything. Open weights ship under different terms: gpt-oss uses Apache 2.0, a standard permissive licence, while Kimi K3 comes under its own Kimi K3 License. Then run a candidate on your own tasks. Running a model locally covers the set-up and the security trade-offs, and July’s roundup has the earlier picture.

Sources

Everything above was checked against these on 26 Sept 2026. Providers change things without notice. If a detail matters to a decision, follow the link.

  1. Qwen3.8Qwen team on GitHub
  2. GLM-5Z.ai on GitHub
  3. gpt-ossOpenAI on GitHub
  4. OllamaOllama on GitHub
  5. Kimi K3Moonshot AI on GitHub
  6. Introducing Claude Fable 5.1 and Claude Mythos 5.1Anthropic · 1 Sept 2026

Search