JS Wei (Jack) Sun

Qwen 3.8 27B ties GPT-5.6 Luna on score, burns 2.3× the tokens

Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.

← Back to the issue

Sources

Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index simonwillison.net

Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index That’s the same score as GPT-5.6 Luna (max), and just one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) - that GLM is 753B and that DeepSeek is 1.7T parameters , and Luna is size unknown but presumably a whole lot bigger than 27B. Qwen 3.8 27B is a truly astonishing model . Via Hacker News Tags: ai , generative-ai , llms , qwen , ai-in-china , artificial-analysis

References

Simon Willison Substack — ‘Qwen 3.8 27B is excellent but it defaults to overthinking’ simonw.substack.com

The model consumed over 22,000 reasoning tokens to produce a simple SVG graphic; setting reasoning_effort to ‘medium’ or ‘low’ is essential for routine tasks.

r/LocalLLM — ‘I benchmarked every Qwen 3.8 27B quant that fits’ reddit.com

Q4_K_M retains 99.97% of original model quality at ~17.1 GB VRAM, effectively indistinguishable from Q8_0; the newer NVFP4 format underperforms standard K-quants.

Hacker News discussion (item 49304017) news.ycombinator.com

The 27B model ‘punches above its weight’ via RLVR, but generates upwards of 10,000 reasoning tokens on straightforward tasks — a ‘caveman’ note-form thinking style with an internal ‘desired oververbosity 9’ parameter.

Hacker News thread linked from Simon Willison (49334544) news.ycombinator.com

Parity on the Intelligence Index masks a ‘latency tax’ — Qwen 3.8 consumes up to 2.3x more tokens than GPT-5.6 Luna for the same task, making it slower in agentic workflows despite matching scores.

kie.ai — Qwen 3.8 27B model overview kie.ai

Vendor-reported 61.7 on SWE-bench Pro and 90.3 on LiveCodeBench v6; DeepSWE 1.1 jumped from 13.3 (Qwen 3.6) to 42.2 with day-zero llama.cpp support and native 262k context.

Simon Willison — earlier Qwen 3.8 27B writeup (Aug 16) simonwillison.net

Independent tests suggest a regression in specialized general knowledge — the model reportedly lost a majority of medical benchmark trials to its predecessor Qwen 3.6 27B.

Jack Sun

Jack Sun, writing.

Engineer · Bay Area

Hands-on with agentic AI all day — building frameworks, reading what industry ships, occasionally writing them down.

Digest
All · AI Tech · AI Research · AI News
Writing
Essays
Elsewhere
Subscribe
All · AI Tech · AI Research · AI News · Essays

© 2026 Wei (Jack) Sun · jacksunwei.me Built on Astro · hosted on Cloudflare