ai models13 min read

A 16.5 GB download now outscores February's best closed model

Three years ago the best model you could download lost to GPT-4 by 37 points on HumanEval and needed two graphics cards to run. Last week a 27B open-weight model scored 52.0 on Artificial Analysis's Intelligence Index against Claude Opus 4.6 Max's 44.9 — at one-ninth the price, from a 16.5 GB file. A before-and-after with the benchmark tables, the price curve, and what it means that the catch-up lag is down to 162 days.

K
Ken Jo
#open-llm#qwen#qwen3-8#open-weight#local-llm#ai-pricing#benchmarks#ai-models

Somebody on Hacker News is running a frontier-grade language model on a graphics card with 16 GB of memory, at 50 to 60 tokens per second, with a 128,000-token context window. The model is Qwen3.8-27B, released August 14, 2026. On Artificial Analysis's Intelligence Index it scores 52.0 — higher than Claude Opus 4.6 Max, the best proprietary model on the market as recently as February, which scores 44.9.

That sentence would have been science fiction in 2023, and not the fun kind.

This post is the before-and-after. Not the release coverage — we already wrote that one — but the comparison across three years: what the best downloadable model scored against the frontier then versus now, what it cost to buy a given level of intelligence then versus now, and what hardware you needed to own to run it yourself. Every figure below has a source and a date. The trend they describe is one of the more genuinely optimistic things happening in software right now.

TL;DR:

  • The gap closed, and we can date it. In July 2023 the best open model lost to GPT-4 by 17.5 points on MMLU and 37.1 points on HumanEval. In August 2026 an open 27B beats February's proprietary flagship on 15 of the 19 benchmarks they both report, and on an independent composite index.
  • The price fell about 208x in 23 months for a fixed capability threshold, per Epoch AI. Today's open 27B costs $1.09 per million tokens blended against Opus 4.6 Max's $10.00 — and $0 per token if you run it yourself.
  • It is not the frontier, and we say so. Claude Opus 5 Max scores 63.1 and GPT-5.6 Sol Max scores 60.9. The 27B loses to Opus 4.6 on broad-knowledge reasoning by a wide margin: 30.8 against 40.0 on HLE.

The interconnect panel of a Cray Y-MP supercomputer, dense rows of red and blue coaxial cables plugged into steel connectors

Image: Steve Jurvetson, Wikimedia Commons, CC BY 2.0. The interconnect of a Cray Y-MP, circa 1988. Serious compute used to be a room you walked into.

What "before" actually looked like

Set the clock to July 18, 2023, the day Meta published Llama 2. It was the best openly downloadable model in the world, and its own paper printed the scoreboard against the closed frontier without flinching.

Llama 2 70B scored 68.9 on MMLU against GPT-4's 86.4. On GSM8K it scored 56.8 against 92.0. On HumanEval, the coding benchmark, it scored 29.9 against 67.0 — less than half. The paper's own summary of the situation reads: "There is still a large gap in performance between Llama 2 70B and GPT-4 and PaLM-2-L." That is the open-weight flagship, in its own words, in 2023.

Now the hardware. TheBloke's 4-bit community build of Llama 2 70B — the standard way people ran it — was a 41.4 GB file, published September 4, 2023. To run that at a usable speed you needed two 24 GB graphics cards or a 64 GB Mac. The model that actually fit on a normal laptop was Llama 2 13B, a 7.9 GB file scoring 54.8 on MMLU: 31.6 points behind GPT-4, and useful for very little beyond demos.

So the 2023 deal was: pay a frontier lab, or run something visibly worse on hardware most people didn't own. That was the whole reason the "just use the API" argument won so decisively for two years. It was correct.

Now put the same question to August 2026

Alibaba published Qwen3.8-27B on August 14, 2026 under Apache 2.0, and printed its own head-to-head table against Claude Opus 4.6 Max on the model card. We counted every row where both models report a number.

Qwen3.8-27B wins 15 of those 19 rows and loses 4. The wins include SWE-bench Pro (61.7 against 53.4), instruction-following on IFBench (79.5 against 62.5), competitive coding on LiveCodeBench v6 (90.3 against 88.8), computer use on OSWorld-Verified (84.3 against 72.7), mobile agent tasks on AndroidWorld (81.9 against 62.0), and a long run of multimodal benchmarks where the margins are not close — 90.0 against 65.5 on MathVision, 83.7 against 66.0 on chart analysis in CharXiv, 65.5 against 40.8 on embodied reasoning in ERQA.

The four losses matter more than the wins, because they tell you what a 27-billion-parameter model still cannot do. It loses on Terminal Bench 2.1 (73.0 against 78.2), on repo-level code generation in NL2Repo-Bench (42.3 against 47.6), narrowly on GPQA Diamond (89.2 against 91.3), and badly on HLE, the broad multidisciplinary reasoning benchmark: 30.8 against 40.0. That last one is the shape of the whole limitation. You cannot compress the world's knowledge into a 16.5 GB file; you can compress an enormous amount of skill.

One row deserves an asterisk: the 79.0-against-63.8 result is on QwenSWEBench, a benchmark Qwen built. Vendor-run benchmarks on vendor-built benchmarks are the weakest evidence in any model card, and we are not counting on it.

The independent scoreboard says the same thing

Vendor tables are vendor tables, so here is a third party. Artificial Analysis, which runs its own evaluations rather than reprinting anyone's, has now benchmarked the 27B. As of August 21, 2026 it gives Qwen3.8-27B an Intelligence Index of 52.0, ranking it first of 135 models in the open-weight 4B–40B size class, against 44.9 for Claude Opus 4.6 Max. That is a 7.1-point gap in the open model's favour, measured by someone with no stake in the outcome, 190 days after Opus 4.6 shipped.

Track that comparison back through time and the trend is the story. In July 2023 the best open model under 40 billion parameters scored 2.6 on the same index. In July 2024 it was 7.4. In March 2025, 13.4. In February 2026, 34.6. Last week, 52.0.

Line chart of the Artificial Analysis Intelligence Index from 2023 to August 2026 comparing the best proprietary flagship with the best open-weight model of 4 to 40 billion parameters, with the open line reaching 52 in August 2026

Both lines climb; the interesting quantity is the horizontal distance between them. Artificial Analysis Intelligence Index, retrieved August 21, 2026.

Here is the number we think deserves a name. The frontier first crossed an Intelligence Index of 52 on March 5, 2026, when GPT-5.4 at extra-high reasoning effort scored 53.1. A model small enough to run on a personal machine got there on August 14 — 162 days later. Call it the catch-up lag: the time between a capability appearing at the frontier and appearing on your own hardware.

That lag used to be much longer. Llama 3.1 8B needed 260 days to reach a level GPT-4 Turbo had hit. GLM-4.7-Flash needed 410 days to match o1. The last three laptop-class releases have come in at 201, 155 and 162 days. The frontier is not slowing down — Opus 5 Max sits at 63.1 today — but the trailing edge is closing on it about twice as fast as it was two years ago.

The price didn't fall. It collapsed.

Capability is half the before-and-after. The other half is the invoice.

Epoch AI published the cleanest measurement of this in March 2025, and the method is worth understanding: hold a benchmark score fixed, then track the cheapest model that achieves it over time. Fix the threshold at GPT-4's 86 on MMLU and the ladder runs $37.50 per million tokens in March 2023, $15.00 by November 2023, $7.50 in May 2024, $2.19 later that month, and $0.18 by February 2025. Same score. 208 times cheaper in 23 months.

Epoch's headline finding is that the rate varies enormously by task — between 9x and 900x per year depending on which capability you fix — with the price of GPT-4-level performance on PhD-level science questions falling roughly 40x per year. The direction never varies.

Descending line on a logarithmic scale showing the lowest blended price per million tokens for a model scoring 86 or better on MMLU, falling from 37.50 dollars in March 2023 to 18 cents in February 2025

The benchmark score is held constant across every point on this line. Epoch AI, March 12, 2025.

Extend the series to today with live prices and it keeps going. On the OpenRouter catalogue as of August 21, 2026, Qwen3.8-27B costs $0.45 per million input tokens and $3.20 per million output tokens; Claude Opus 4.6 costs $5.00 and $25.00. Blended at Epoch's 3:1 input-to-output convention, that is $1.09 against $10.00 — the open model is 9.1 times cheaper and 7.1 index points better than the closed flagship of six months ago.

Scatter plot of Artificial Analysis Intelligence Index against blended price per million tokens on a logarithmic scale, with open-weight models in blue at low prices and high scores

Every model that matters is migrating toward the top-left corner. Scores from Artificial Analysis, prices from OpenRouter, both retrieved August 21, 2026.

And then there is the option that has no per-token price at all. Run the 27B on your own RTX 4090 at the roughly 48 tokens per second practitioners reported in the release thread, assume the card's 450-watt rating, and at the US average residential electricity price of 18.44 cents per kilowatt-hour (EIA, May 2026) the electricity works out to about $0.48 per million output tokens — before you count the card, and after you stop counting anyone's margin. That is a computed estimate with its assumptions on the table, not a vendor claim.

The before-and-after, on one line each

July 2023August 2026
Best downloadable modelLlama 2 70B (Meta, Jul 18, 2023)Qwen3.8-27B (Alibaba, Aug 14, 2026)
Parameters70B27B
Standard 4-bit community build41.4 GB16.5 GB
Smallest usable build7.9 GB (13B, MMLU 54.8)9.8 GB (2-bit, 128K context on a 16 GB card)
Realistic hardwaretwo 24 GB GPUs or a 64 GB Macone 16 GB consumer GPU
Frontier model of the dayGPT-4 (Mar 14, 2023)Claude Opus 4.6 Max (Feb 5, 2026)
Result against it−17.5 MMLU, −35.2 GSM8K, −37.1 HumanEvalwins 15 of 19 reported benchmarks; +7.1 Intelligence Index
Blended price of that frontier model$37.50 / M tokens$10.00 / M tokens
Blended price of the open modelnot offered at scale$1.09 / M tokens
LicenceLlama 2 Community LicenceApache 2.0

Sources for every cell are listed at the end. The row we keep coming back to is the second-to-last one: the frontier's own price dropped 3.75x across three years, while the open alternative arrived underneath it at roughly a tenth of that. Both things happened. The second is what changes who gets to build.

What is still not true

The enthusiasm around this release outruns the evidence in four specific ways, so let us mark them.

It is not the frontier. Claude Opus 5 Max scores 63.1 on the same index and GPT-5.6 Sol Max scores 60.9, both well clear of 52.0. If your work depends on the hardest available reasoning, the frontier is still a different product, and it is still closed.

It is not cheap for its size. Artificial Analysis's own summary calls Qwen3.8-27B "particularly expensive when comparing to other open weight models of similar size" — the median open model in its class charges $0.05 per million input tokens against Qwen's $0.43. You are paying for capability, not for parameters.

It is verbose, and verbosity is a bill. Running the Intelligence Index evaluation, the 27B generated 160 million output tokens against a median of 45 million. One practitioner prompted it with two words on a 64 GB Mac mini and watched it think for 17 minutes and 12 seconds. Reasoning models charge you in wall-clock time even when the tokens are free.

And the benchmark scores belong to the full-precision model. The 9.8 GB two-bit build that fits your graphics card is not the model that scored 52.0. Heavy quantization costs accuracy, especially over long contexts. Every "runs on a laptop" claim, including ours, carries that asterisk.

Why the lag should keep shrinking

Here is our take, and we will mark the pivot clearly: everything above is measurement, and what follows is inference from it.

The mechanism driving the catch-up lag down is not mysterious. Post-training techniques published at the frontier get reproduced in open labs within months. Reasoning-effort control, which was a proprietary trick eighteen months ago, ships as a documented parameter on this model card. Community quantization now finishes before the original lab's documentation stops changing. None of those three feedback loops has a reason to slow down, and two of them get faster as more people participate.

So the reasonable expectation for the next generation is not that open models overtake the frontier. It is that the 162-day lag keeps compressing while the absolute capability at both ends keeps rising — which means the interesting threshold is not "can open models win" but "how recent does the frontier have to be before the difference stops mattering for your task." For meeting transcription, translation, summarisation and most document work, that threshold was crossed some time ago. For frontier research and the hardest agentic engineering, it has not been.

That is a much better world than the 2023 one, and it is worth saying so plainly. The cost of giving a computer a competent understanding of language fell by two orders of magnitude in three years and is still falling. The people who benefit most are the ones who were priced out: solo developers, teams in countries where a $10-per-million-token bill is not a rounding error, researchers with a grant instead of a budget, anyone who needs the data to stay on their own machine for legal reasons.

The bottom line: the frontier stopped being a place and became a date

The old question was "which lab do you have access to." The new question is "how many months behind the frontier are you willing to be, given that it costs a ninth as much and runs on hardware you own."

For a growing number of jobs, the honest answer is: about five months, and falling.


Where Telli.sh fits: everything above is why we build on open models rather than around a single vendor. Telli.sh already runs an open Qwen model in its summarisation path, so each step down this curve arrives as better meeting notes at the same price, not as a pricing announcement. What we work on is the part no download gives you: capturing audio reliably for an hour, keeping speakers apart, translating live across 15 languages, and turning all of it into notes you can still search months later.

Start a live AI note

Sources


Back to Blog