JS Wei (Jack) Sun

Subquadratic's 12M-token breakthrough is a finetune with a self-run eval

Subquadratic's headline long-context numbers come from finetuning open weights and benchmarking against older FlashAttention, while investors price the pitch at $500M.

Subquadratic’s 12M-token breakthrough is a finetune with a self-run eval

TL;DR

  • Subquadratic’s 12M-token model is a sparse-attention finetune of open weights, per its own CTO.
  • 98% retrieval and 56× prefill come from a self-commissioned eval against older FlashAttention versions.
  • DeepSeek’s NSA and Moonshot’s MoBA shipped content-dependent sparse attention months earlier.
  • Investors priced the pitch at $500M post-money on a $29M seed led by JAM Fund.

Today’s lone AI-research read is a worked example of how far a long-context number can travel before anyone checks the substrate. Subquadratic pitched a 12M-token breakthrough — 98% retrieval, 56× prefill speedup — and raised a $29M seed at a $500M post-money on the strength of it. The CTO has since confirmed the model is a sparse-attention finetune of open-source weights, the eval was self-commissioned against older FlashAttention versions, and the core technique was shipped months earlier by DeepSeek’s NSA and Moonshot’s MoBA. None of that makes the engineering uninteresting; it does make the “solved a decade-old bottleneck” framing hard to defend. The interesting question isn’t whether the numbers are real — it’s why the diligence loop closed before anyone re-ran them.

Subquadratic’s 12M-context LLM is a finetune of open weights

Source: mit-tech-review-ai · published 2026-06-19

TL;DR

  • Subquadratic’s CTO confirmed the company’s “breakthrough” 12M-token model is a sparse-attention finetune of open-source weights, not trained from scratch.
  • The headline numbers — 98% retrieval at 12M tokens, 56× prefill speedup — come from a self-commissioned eval against older FlashAttention versions.
  • DeepSeek’s NSA and Moonshot’s MoBA shipped content-dependent sparse attention months earlier, undercutting the “solved a decade-old bottleneck” framing.
  • The company raised $29M seed at a $500M post-money valuation led by JAM Fund and ex-SoftBank’s Javier Villamizar.

The receipts have a footnote

MIT Tech Review’s framing is that Miami startup Subquadratic, after a thin stealth-exit in May, is “bringing the receipts” on a claim to have solved the quadratic-attention bottleneck. The receipts are real — but so is the asterisk the piece soft-pedals.

Within days of launch, former OpenAI engineer Will Depue argued on X that SubQ was “almost surely a sparse attention finetune of Kimi or DeepSeek.” CTO Alex Whedon later conceded the team had in fact started from open-source weights, defending it as a pragmatic move for a seed-stage company 1. That reframes what’s being sold: a novel attention mechanism grafted onto somebody else’s pretrained intelligence, not a ground-up frontier model. The distinction matters because “we solved the bottleneck holding back LLMs” reads very differently from “we swapped the attention layer on an existing checkpoint.”

What the benchmarks actually show

The technical report is specific and, on its face, impressive: 98% retrieval accuracy at a 12-million-token context window while attending to only 0.13% of token pairs, a 56× prefill speedup over FlashAttention-2 at 1M tokens, 99.12% on RULER at 128K, and 81.8% on SWE-Bench Verified 2. Those are the numbers the MIT TR piece leans on.

The numbers the piece omits: the speedup is measured against older FlashAttention versions rather than the current industry standard, and the “independent” Appen evaluation that produced the headline retrieval figure was commissioned and paid for by Subquadratic itself. Critics, including Stepan Goncharov, flagged the comparisons as cherry-picked, and noted a reported ~17-point gap between the research variant being benchmarked and the production variant being served 3.

The prior art the pitch elides

“Broke through a bottleneck that’s held back LLMs for almost a decade” is a strong claim in a field that has been shipping the answer for months. DeepSeek’s Native Sparse Attention (arXiv 2503.01868, February 2026) already demonstrated natively-trainable, content-dependent block sparsity with up to 11.6× decoding speedups on 64K sequences 4. Moonshot’s MoBA applies MoE-style gating to block selection in attention. Both predate SubQ’s launch and both are public.

More awkwardly, the rest of the industry has moved in a different direction entirely. The 2026 frontier conversation is dominated by hybrid stacks — Jamba, Zamba2, Nvidia’s Nemotron-3, IBM Bamba, Qwen3-Next with its 3:1 DeltaNet-to-attention ratio — precisely because pure subquadratic models hit a documented recall gap on multi-hop reasoning and needle-in-haystack retrieval, compressing context into a fixed-size state 5. SubQ’s pitch — pure subquadratic, at the frontier — is a bet against where practitioners have already converged.

What’s actually here

Strip the marketing and there’s a plausible engineering artifact: a sparse-attention layer that delivers genuine prefill speedups when bolted onto a capable open-weight base. That’s a useful contribution. It is not a decade-old bottleneck cracked open by a Miami seed-stage team — and the $500M valuation 6 is buying the second story, not the first.

Footnotes

  1. The Next Webhttps://thenextweb.com/news/subquadratic-subq-sparse-attention-llm-bottleneck

    Will Depue tweeted that SubQ is ‘almost surely a sparse attention finetune of Kimi or DeepSeek’… Subquadratic’s CTO Alex Whedon eventually confirmed on X that the company used open-source weights as a starting point due to limited initial funding.

  2. Subquadratic technical report (subq.ai)https://subq.ai/subq-1-1-small-technical-report

    SSA achieves 98% retrieval accuracy at a 12-million-token context window while attending to only 0.13% of token pairs, with a 56x prefill speedup over FlashAttention-2 at 1M tokens and 99.12% on RULER at 128K.

  3. AI Agents Directoryhttps://aiagentsdirectory.com/blog/subq-is-a-sub-quadratic-llm-built-for-12m-token-reasoning

    Critics like Stepan Goncharov labeled the initial results as ‘cherry-picked,’ pointing out that SubQ compared its performance against older versions of FlashAttention rather than current industry standards, and the Appen evaluation was commissioned by Subquadratic itself.

  4. DeepSeek NSA paper (arXiv 2503.01868)https://arxiv.org/abs/2503.01868

    Native Sparse Attention uses a hierarchical three-branch design (sliding window, compressed tokens, selective blocks) and is ‘natively trainable,’ reporting up to 11.6x decoding speedups on 64k sequences — establishing prior art for content-dependent sparse attention that predates SubQ’s launch.

  5. Galileo AI (Mamba scaling analysis)https://galileo.ai/blog/mamba-linear-scaling-transformers

    Pure subquadratic models often struggle with multi-hop reasoning and precise needle-in-a-haystack retrieval because they compress context into a fixed-size state bottleneck; the industry has converged on hybrid designs (Jamba, Zamba2, Nemotron-3, IBM Bamba) rather than pure subquadratic stacks.

  6. SiliconAnglehttps://siliconangle.com/2026/05/05/subquadratic-launches-29m-bring-12m-token-context-windows-ai/

    Subquadratic debuted with $29 million in seed funding at a reported $500 million post-money valuation, led by Tinder co-founder Justin Mateen (JAM Fund) and former SoftBank Vision Fund partner Javier Villamizar.

Jack Sun

Jack Sun, writing.

Engineer · Bay Area

Hands-on with agentic AI all day — building frameworks, reading what industry ships, occasionally writing them down.

Digest
All · AI Tech · AI Research · AI News
Writing
Essays
Elsewhere
Subscribe
All · AI Tech · AI Research · AI News · Essays

© 2026 Wei (Jack) Sun · jacksunwei.me Built on Astro · hosted on Cloudflare