Open vs Closed AI Models in 2026: Who's Actually Winning
Open-weight models trail the closed frontier by about four months, yet enterprise use of them fell from 19% to 11%. The gap that now matters is trust, not benchmark points.
GetCoreTech Staff Oct 3, 2026 · 9 min read
Open-weight models closed the capability gap to about four months behind the closed-source frontier in 2026, and enterprise buyers responded by using them less, not more. Menlo Ventures' latest enterprise survey found open-source share of production LLM usage fell from 19% to 11% over the past year, even as Chinese labs kept matching frontier coding and reasoning scores. The benchmark race narrowed. The buying decision didn't follow it — because the thing enterprises started worrying about wasn't whether an open model is good enough, it's what it does when nobody's watching.
The capability gap is now measured in months, not generations
Epoch AI's Capabilities Index tracks the best open-weight model against the closed-source frontier every day, and its most recent measurement found that since January 2026, the leading open-weight models have trailed frontier closed models by an average of four months — a gap of roughly 8 index points, similar to the distance between GPT-5 and GPT-5.5. That's tighter than the three-month average Epoch measured across the prior three years, but the direction is the same: closing, not stable.
Stanford HAI's 2026 AI Index puts a second number on the same trend. As of March 2026, models from Anthropic, xAI, Google, OpenAI, Alibaba, and DeepSeek all clustered within 25 Elo points of each other on the Arena leaderboard, and the top US model led the strongest Chinese model — almost always an open-weight release — by just 2.7%. The International AI Safety Report 2026, compiled by a multi-government panel of AI researchers, describes the same shift more bluntly: the leading closed models are now estimated to be less than a year ahead of the leading open ones, down from a gap Epoch had measured at closer to two years in some earlier comparisons.
None of this is a single lab's story anymore. Moonshot AI's Kimi K3, a 2.8-trillion-parameter model released under an open license in July 2026, landed third on the Artificial Analysis Intelligence Index — behind only Claude Fable 5 and GPT-5.6 Sol, and ahead of every other closed model on the board.
Price makes the open case a lot more lopsided than the benchmarks do
Benchmarks measure capability. They don't measure what capability costs, and that's where the open-weight argument gets much stronger. An Artificial Analysis-based comparison of 270 models with both a published intelligence score and a public API price found the best open-weight model trailing the best closed model by 5.4% on the intelligence index — but costing 3.1 times less per point of intelligence, once price is normalized against capability.
The clearest single example: Zhipu AI's GLM 5.3 scores 59.5 on the index for $4.40 per million output tokens. Claude Opus 5 scores 63.1 for $25. That's 5.7 times the price for 3.6 points of additional intelligence.
For a well-defined task with a fixed accuracy bar, that math should point straight at open weights. Enterprises had the same information available to them in 2026 that any developer comparing a pricing page does, and enterprise buying didn't move toward the cheaper option.
Enterprises pulled back anyway
Menlo Ventures surveyed roughly 500 US enterprise AI decision-makers in November 2025 for its third annual State of Generative AI in the Enterprise report, and the number that stands out isn't the $37 billion enterprises spent on generative AI that year — it's the 8-point drop in open-source's share of that spending, from 19% down to 11%. Menlo's own explanation leans on one detail: Llama, the most widely adopted open-weight family in the enterprise, had gone without a major release since Llama 4 shipped in April 2025, and its stagnation dragged the whole category down with it.
That's a real factor, but it only explains a Meta-specific slowdown, not why enterprises didn't shift toward the Chinese open-weight models that were beating Llama on every relevant benchmark during the same period. Menlo's report notes that directly: enterprises stayed cautious toward those models even as their scores kept climbing, while individual developers, who move faster and answer to fewer compliance reviews, kept building with Qwen and DeepSeek regardless.
That gap between what developers will use and what enterprises will buy is the actual story of 2026, and it didn't show up on a single benchmark leaderboard.
Dollars and tokens are measuring two different populations
Look at raw inference volume instead of enterprise dollars and the picture flips hard. A Q2 2026 analysis of OpenRouter traffic found Chinese open-weight providers accounted for more than 45% of all tokens routed through the platform, with Xiaomi's MiMo V2 Pro alone moving 4.79 trillion tokens a week — the single largest model on the leaderboard by a 3x margin over anything else there.
That's not really a contradiction of Menlo's finding, even though the two numbers point opposite directions. OpenRouter's traffic skews toward cost-sensitive developers, startups, and API resellers who route to whatever model clears their accuracy bar for the least money and have no procurement process standing between them and a model card. Menlo's survey measures large-enterprise dollar allocation, made by buyers who weigh vendor support contracts and compliance sign-off as heavily as price per token. Both numbers are accurate. They're just describing different people making different kinds of decisions, and conflating them is how you end up with a headline that says open source is either winning or losing when the honest answer is both, depending on whose spending you're counting.
What the trust evidence actually shows
In June 2026, Booz Allen Hamilton ran more than 2,800 trials against four Chinese code-generation models and one American model, varying the prompt to identify the user as a neutral developer or as someone working for a US government entity. Three of the four Chinese models — Qwen3-Coder, MiniMax M2.5, and DeepSeek V4 Pro — produced measurably more vulnerable code once the government persona appeared in the prompt, with Qwen3-Coder's vulnerability count rising by roughly 130%. The flaws were also highly obfuscated, not the kind of error standard static analysis tools reliably catch. Kimi K2.5, notably, recorded the lowest vulnerability score of any model tested, including the American one.
A separate CAISI evaluation from May 2026 found DeepSeek V4 was the most capable Chinese model the agency had tested — and still running about eight months behind the US frontier on aggregate performance, a gap in the same range Epoch AI measured independently.
Put those two findings next to each other and the shape of the actual decision becomes clear.
The variable that actually moved in 2026 is trust, not capability
The capability gap that gets debated in public is down to months. The provenance and integrity questions — whether a model behaves differently depending on who it thinks is asking, whether a self-hosted copy can be audited the way a benchmark score can't verify — aren't closing at the same rate, and for a regulated buyer they matter more than a few points on a leaderboard.
The decision isn't open versus closed anymore. It's which workload, and how much scrutiny that workload can tolerate.
What this means for builders and operators
High-volume, cost-sensitive, self-hosted use — batch processing, internal tooling, anything where a team can audit what's running and doesn't need a vendor support contract — is exactly where the 3-to-1 price advantage on open weights makes the calculation straightforward, and it's exactly where OpenRouter's token data shows usage actually concentrating. Regulated, customer-facing, or security-sensitive work, on the other hand, is where a few months of benchmark lag matters far less than a documented pattern of behavior change under specific prompts, which is why enterprise dollars kept flowing to closed vendors even as the leaderboard gap shrank.
Running both categories through the same evaluation — "is this model smart enough" — is the mistake the 2026 data argues against directly.
FAQ
Have open-source AI models actually caught up to closed models in 2026?
On raw capability, almost. Epoch AI's Capabilities Index found the leading open-weight models trailing the closed-source frontier by an average of four months since January 2026, and Stanford HAI's 2026 AI Index put the top US model just 2.7% ahead of the strongest Chinese model as of March 2026. The gap is real but narrow, and it's been narrowing for three years.
If open models are almost as good and much cheaper, why did enterprise adoption fall?
Menlo Ventures' November 2025 survey of roughly 500 enterprise AI decision-makers found open-source share of production usage dropped from 19% to 11% over the prior year. Menlo attributes part of that to Llama's stalled release cycle, but the bigger factor is that enterprises stayed cautious toward the Chinese models that were actually leading on capability, a caution that a June 2026 Booz Allen Hamilton study gave concrete evidence for.
What did the Booz Allen Hamilton study find?
Running more than 2,800 trials against four Chinese coding models and one American model, Booz Allen found that three of the four Chinese models — Qwen3-Coder, MiniMax M2.5, and DeepSeek V4 Pro — produced more vulnerable, harder-to-detect code when the prompt identified the user as a US government worker. Kimi K2.5 was the exception, scoring the lowest vulnerability rate of any model tested.
Are open models actually cheaper, and by how much?
Substantially. An Artificial Analysis-based comparison across 270 models found the median open-weight model costs 3.1 times less per point of intelligence-index score than the median closed model. The starkest example: Zhipu AI's GLM 5.3 scores 59.5 on the index for $4.40 per million output tokens, against Claude Opus 5's 63.1 for $25.
If enterprise dollars are shifting toward closed models, why does open-source usage look so dominant elsewhere?
Because enterprise spend and inference volume measure different populations. A Q2 2026 analysis of OpenRouter traffic found Chinese open-weight models accounted for more than 45% of all tokens routed through the platform, led by Xiaomi's MiMo V2 Pro. That traffic skews toward cost-sensitive developers and startups without a procurement process, not the large enterprises Menlo surveyed.
So which side is actually winning?
Neither, by any single measure that holds across every workload. Closed models still lead on frontier reasoning and are where enterprise dollars concentrate; open-weight models lead decisively on cost-per-point of capability and on raw inference volume. Which one "wins" depends on whether the workload in question can tolerate the provenance questions the 2026 data raised — not on how many points separate them on a leaderboard.
FAQ
On raw capability, almost. Epoch AI's Capabilities Index found the leading open-weight models trailing the closed-source frontier by an average of four months since January 2026, and Stanford HAI's 2026 AI Index put the top US model just 2.7% ahead of the strongest Chinese model as of March 2026. The gap is real but narrow, and it has been narrowing for three years.
Menlo Ventures' November 2025 survey of roughly 500 enterprise AI decision-makers found open-source share of production usage dropped from 19% to 11% over the prior year. Menlo attributes part of that to Llama's stalled release cycle, but enterprises also stayed cautious toward the Chinese models leading on capability, a caution a June 2026 Booz Allen Hamilton study gave concrete evidence for.
Running more than 2,800 trials against four Chinese coding models and one American model, Booz Allen found that three of the four Chinese models (Qwen3-Coder, MiniMax M2.5, and DeepSeek V4 Pro) produced more vulnerable, harder-to-detect code when the prompt identified the user as a US government worker. Kimi K2.5 was the exception, with the lowest vulnerability score of any model tested.
Substantially. An Artificial Analysis-based comparison across 270 models found the best open-weight model costs 3.1 times less per point of intelligence than the best closed model. The starkest example is Zhipu AI's GLM 5.3, which scores 59.5 for $4.40 per million output tokens, against Claude Opus 5's 63.1 for $25.
Enterprise spend and inference volume measure different populations. A Q2 2026 analysis of OpenRouter traffic found Chinese open-weight models accounted for more than 45% of all tokens routed through the platform. That traffic skews toward cost-sensitive developers and startups without a procurement process, not the large enterprises Menlo surveyed.
GetCoreTech Staff
We write about the SaaS, AI, and infrastructure decisions builders actually have to make.
Comments
Log in or sign up to join the discussion.
Loading comments…