Qwen 3.8 27B ties GPT-5.6 Luna on score, burns 2.3× the tokens
Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.
Sources
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index simonwillison.net
Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index That’s the same score as GPT-5.6 Luna (max), and just one point behind GLM-5.2 (max) and DeepSeek V4 Pro 0813 (max) - that GLM is 753B and that DeepSeek is 1.7T parameters , and Luna is size unknown but presumably a whole lot bigger than 27B. Qwen 3.8 27B is a truly astonishing model . Via Hacker News Tags: ai , generative-ai , llms , qwen , ai-in-china , artificial-analysis
References
Simon Willison Substack — ‘Qwen 3.8 27B is excellent but it defaults to overthinking’ simonw.substack.com
The model consumed over 22,000 reasoning tokens to produce a simple SVG graphic; setting reasoning_effort to ‘medium’ or ‘low’ is essential for routine tasks.
r/LocalLLM — ‘I benchmarked every Qwen 3.8 27B quant that fits’ reddit.com
Q4_K_M retains 99.97% of original model quality at ~17.1 GB VRAM, effectively indistinguishable from Q8_0; the newer NVFP4 format underperforms standard K-quants.
Hacker News discussion (item 49304017) news.ycombinator.com
The 27B model ‘punches above its weight’ via RLVR, but generates upwards of 10,000 reasoning tokens on straightforward tasks — a ‘caveman’ note-form thinking style with an internal ‘desired oververbosity 9’ parameter.
Hacker News thread linked from Simon Willison (49334544) news.ycombinator.com
Parity on the Intelligence Index masks a ‘latency tax’ — Qwen 3.8 consumes up to 2.3x more tokens than GPT-5.6 Luna for the same task, making it slower in agentic workflows despite matching scores.
kie.ai — Qwen 3.8 27B model overview kie.ai
Vendor-reported 61.7 on SWE-bench Pro and 90.3 on LiveCodeBench v6; DeepSWE 1.1 jumped from 13.3 (Qwen 3.6) to 42.2 with day-zero llama.cpp support and native 262k context.
Simon Willison — earlier Qwen 3.8 27B writeup (Aug 16) simonwillison.net
Independent tests suggest a regression in specialized general knowledge — the model reportedly lost a majority of medical benchmark trials to its predecessor Qwen 3.6 27B.