Eighteen months ago, "open-weight model" mostly meant Llama, with Mistral as the scrappy alternative. That's not the landscape anymore. The last few weeks of September 2026 alone brought a new DeepSeek release, and the months before it brought fresh drops from Qwen, Kimi, and GLM, each one landing within striking distance of closed frontier models on independent leaderboards, and all of them downloadable.
I'm not writing this to tell you open weights beat the closed APIs: for most teams building a product, they still don't, and the operational tax of self-hosting is real. This is a map of what's actually out there right now, what it costs to run, and the handful of situations where reaching for an open-weight model is the right call instead of a hobbyist detour.
Why open-weight at all
Three reasons teams actually make this switch, in order of how often I see them:
Data residency and compliance. If your data can't leave a jurisdiction, or can't touch a third-party API at all, healthcare records, government contracts, anything under strict data-sovereignty rules: self-hosting isn't a cost optimization, it's the only option. This is the single biggest driver I see in practice.
Cost at extreme volume. API pricing is per-token, and per-token costs stop looking cheap once you're running hundreds of millions of tokens a day. At that volume, amortized GPU cost can beat API pricing, but the crossover point is higher than people expect once you account for ops overhead, and it usually only pencils out for the largest, most predictable workloads.
Fine-tuning and customization. Closed APIs increasingly offer fine-tuning, but you don't get to touch the base weights, change the architecture, or run heavily customized inference stacks (speculative decoding tuned to your traffic pattern, custom quantization, etc.). If your product's edge depends on a model that's meaningfully different from the vendor's default, open weights are the only path there.
If none of those three apply to you, the closed APIs are probably still the better default, see prompting reasoning models if the real question is how to get more out of the model you're already using rather than which model to switch to.
The current lineup
DeepSeek V4.1-Flash
DeepSeek shipped V4.1-Flash on September 10, 2026: genuinely recent as of this post. It's a 552B-parameter mixture-of-experts model with a causal encoder-decoder architecture: roughly 8B parameters active per token on input and 16B active on output, which is what makes it runnable at a fraction of what a dense 552B model would cost in compute. Weights are on Hugging Face under the MIT license: about as permissive as it gets, full commercial use and redistribution rights.
What's notable isn't just the release: it's that DeepSeek said independent tests put V4.1-Flash ahead of their own larger V4-Pro flagship on performance, cost, and speed, and they've since routed V4-Pro API traffic to V4.1-Flash and billed it at the smaller model's rate. That's a company telling you directly which of its own models to use, which is a useful signal in itself.
Self-hosting reality: even with the sparse MoE activation, 552B total parameters means you're not running this on a single consumer GPU. You're in multi-GPU server territory, think 8x80GB-class cards for full-precision serving, less with aggressive quantization, but this is an infrastructure-team project, not a laptop experiment.
Qwen3.8-27B
Alibaba's Qwen team released Qwen3.8-27B in mid-August 2026, and it's architecturally the opposite of DeepSeek's approach: 27 billion parameters, fully dense (every parameter activates on every token, no MoE sparsity), licensed Apache 2.0. It's vision-capable with native image and video understanding and a context window past 262,000 tokens.
Dense models are simpler to reason about than MoE (no routing behavior to debug, more predictable latency) and 27B dense is squarely in the range people run locally. With 4-bit quantization this fits on a single high-end consumer GPU (a 24GB card gets you close, depending on context length used), which makes it the most realistic "run it on my own hardware" option of the four models here. If your actual goal is local inference rather than server-farm self-hosting, Qwen3.8-27B is the one to start with.
Kimi K3
Moonshot AI's Kimi K3, released in mid-July 2026, is the extreme end of the size spectrum: 2.8 trillion parameters, reportedly the largest open-weight model released to date. On release it landed at #3 on the Artificial Analysis leaderboard (behind the top closed frontier models, but ahead of most other open competitors) and it notably beat proprietary models on a front-end web development benchmark.
The license is worth reading before you plan around it: it's a custom license, not MIT or Apache. Companies above $20M in annual revenue need to negotiate a separate contract with Moonshot before offering Kimi K3 to customers as a service, and companies above $20M in monthly revenue or with more than 100 million monthly active users have to display attribution wherever the model is used. This isn't "fully open source" in the permissive sense: it's open weights with commercial strings attached once you're operating at scale. Read the actual license text before you build a paid product on top of it.
At 2.8T parameters, self-hosting K3 is not a small-team project even with MoE sparsity helping efficiency. This is the kind of model you'd realistically access through a hosted inference provider rather than running yourself, unless you already operate serious GPU infrastructure.
GLM-5.3
Zhipu AI (operating as Z.ai) released GLM-5.3 on August 14, 2026, initially API-only, with open weights following about two weeks later on Hugging Face under zai-org/GLM-5.3. It's roughly the same 744B-parameter MoE architecture as its predecessor GLM-5.2, with about 40B parameters active per token, Zhipu described the release as focused purely on scaling post-training rather than architecture changes.
Zhipu positions it specifically as a coding and agentic-workflow model, and pointed to a large jump on Terminal-Bench 3.0 as evidence. Worth noting: the company attributed part of the delay between the API release and the open-weights release to a safety review, specifically because of the model's strength on cybersecurity and vulnerability-finding tasks: an unusual and fairly candid thing for a lab to say about its own release timeline.
Self-hosting: similar tier to DeepSeek V4.1-Flash: 744B total parameters with MoE sparsity means multi-GPU server infrastructure, not consumer hardware.
The older guard: Llama and Mistral
Both families are still around, still used in production, and both now sit in what I'd call legacy territory relative to the four models above, they were the default open-weight choice for a couple of years and got outpaced on recent benchmarks by the newer MoE releases. We keep older Llama and Mistral guides live for reference since plenty of existing deployments still run on them, but if you're starting a new project today, benchmark against the current generation first.
Matching model to workload
A rough framework, ordered by what actually determines the choice:
Need it to run on hardware you already own, not a server farm? Qwen3.8-27B is the realistic option. Dense architecture, quantizes well, fits consumer-grade high-end GPUs.
Need the most permissive license for a commercial product, no negotiation required? DeepSeek V4.1-Flash (MIT) or Qwen3.8-27B (Apache 2.0). Both let you build and ship without talking to anyone.
Coding and agentic workflows specifically, and you have server-grade infrastructure? GLM-5.3 and Kimi K3 are both explicitly tuned for this, with GLM-5.3 leaning harder into the coding/terminal-agent use case specifically.
Chasing the top of the leaderboard regardless of infrastructure cost? Kimi K3 is closest to frontier-model performance among open-weight options right now, but budget for the licensing conversation if you're a mid-size or larger company, and budget for serious GPU spend regardless.
Not sure any of this beats just using an API? It's a fair question to ask before committing engineering time to a self-hosting project. Run the same eval set against both an API model and your open-weight candidate before deciding, the promptfoo testing guide covers building that harness, and it's cheaper to spend a day on evals than a month on infrastructure you didn't need.
The caveat that matters most
This space moves fast enough that "current" has a short shelf life: DeepSeek's release above landed less than two weeks before this post went up, and GLM-5.3's open weights followed its API launch by about two weeks too. Release dates, licenses, and parameter counts here are accurate as verified on September 23, 2026. If you're reading this months later, check each project's Hugging Face page or GitHub repo for the current version before you build a procurement decision on numbers from this post. For a wider view of how open-weight options stack up against closed frontier models generally, see the model comparison guide.



