July 2026 open-weight rush — before you trust the benchmarks
In short: In July 2026 open-weight models shipped in a two-week flood, but most headline benchmarks are vendor-reported numbers for models whose weights aren't even out yet — so if you run locally, look at license, memory fit, and reproducibility before the leaderboard's #1.
In July 2026 open-weight models shipped in a two-week flood, but most headline benchmarks are vendor-reported numbers for models whose weights aren't even out yet — so if you run locally, look at license, memory fit, and reproducibility before the leaderboard's #1.
In short: published benchmarks and 'what runs on your desktop' are different numbers — here is July's rush (Kimi K3, GLM-5.2, Inkling, Mistral, DeepSeek V4) from a local point of view.
What actually shipped in July 2026?#
Five moves overlapped in two weeks. Kimi K3 (Moonshot) was announced as a 2.8-trillion-parameter MoE (16 of 896 experts per token), 1M context, native vision — but API-only at launch, with weights due July 27 and no license named yet. GLM-5.2 (Z.ai) landed at 744B with MIT weights and a reported GPQA Diamond of 91.2%, credited with narrowing the gap to the Western closed frontier (about $1.40 input / $4.40 output per million tokens). Alongside them, Thinking Machines shipped Inkling under Apache 2.0, Mistral opened early access to a new sparse MoE family, and DeepSeek V4 was referenced on a mid-July schedule. Note that MiniMax's 2.7T open-weight is still a press-reported plan, not a release (the internal name may change).
Why not trust leaderboard numbers as-is?#
Three reasons. First, today's headline figures are mostly vendor-reported at max reasoning effort, and many models (like Kimi K3) don't have public weights yet — numbers nobody else can reproduce. Second, the same 'SWE-bench' is run with a custom coding agent by one vendor and a minimal shell loop by another, with different token budgets, so placing the scores side by side compares apples to oranges. Third, the aggregators changed — the HuggingFace OpenLLM Leaderboard retired both v1 and v2, and Vellum, Onyx, and Artificial Analysis now fill that gap. Tellingly, Moonshot itself stated that K3 trails Fable 5 and GPT-5.6 Sol overall — rare candor for a launch, and the right posture for reading any announcement.
Think of it this way: a leaderboard's #1 is a world record set under ideal conditions; running it 4-bit with a short token budget on your desktop is the local track — same athlete, different time on the clock.
| Model (as announced) | Architecture / license | Notable reported numbers | Local angle |
|---|---|---|---|
| Kimi K3 (Moonshot) | 2.8T MoE, 1M ctx / weights due 7/27, license unnamed | Program Bench 77.8, SWE Marathon 42.0 (vendor) | too big for a desktop — API / server-class |
| GLM-5.2 (Z.ai) | 744B, 1M ctx / MIT | GPQA Diamond 91.2% (announced) | multi-GPU server, free license |
| DeepSeek V4 (line) | MoE, long context / MIT (announced) | SWE-bench Verified 80.6% (announced) | price & license strengths |
| Gemma 4 12B | dense / open | beats last year's 27B at half the memory (announced/community) | practical for laptops & edge |
| Qwen3.6-27B | dense / open | common 24GB pick (community) | fits a 24GB single GPU |
If you're running locally, what should you look at?#
Not the leaderboard's #1, but three things first: license (are the weights actually out, is commercial use and redistribution allowed — K3's was undecided at launch), memory fit (do the parameters and quantization fit your VRAM — a 2.8T MoE is not a desktop model), and reproducibility (what precision, token budget, and agent produced the vendor number). Only when those line up does a published number mean anything in your setup. The conclusion is simple: July 2026's open weights are genuinely better, but for a local user, 'which one fits my GPU on license, memory, and reproducibility' is a far more practical question than 'which one is strongest.'
Note: the figures above are from each vendor's announcements, public leaderboards, and community aggregates (2026-07), not Hax's own measurements. Open weights update weekly, so rankings, weight availability, and license status may change after publication.
Reference links
Responses
No responses yet. Be the first to respond.