JS Wei (Jack) Sun

Anthropic pays $1.5B for piracy, Fable 5 stalls at 8%, Ox Alpha is Zhipu

Anthropic books a $1.5B piracy settlement and a stalled Fable 5 launch, while tokenizer forensics unmask Ox Alpha as sanctioned Zhipu.

Anthropic pays $1.5B for piracy, Fable 5 stalls at 8%, Ox Alpha is Zhipu

TL;DR

  • Anthropic settled for $1.5B over ~500,000 pirated LibGen titles, roughly $3,000 per work.
  • Judge Alsup split training from sourcing, calling LLM training transformative but piracy ingestion actionable.
  • Fable 5 captured just 8% of Anthropic spend on Ramp in July, versus 28% for older Opus 4.8.
  • Anthropic ARR jumped $47B to $65B May-to-July, with 6,000 customers paying $100K+/year.
  • Tokenizer forensics match Ox Alpha to Zhipu at 99% confidence, exposing a US Entity List problem.

Two of today’s three frontier stories are Anthropic stories, and they cut in opposite directions. Revenue is up sharply — $47B to $65B ARR in two months — but the flagship Fable 5 captured only 8% of Ramp-tracked Anthropic spend in July, while the older Opus 4.8 took 28%. In parallel, Judge Alsup’s Bartz v. Anthropic settlement puts a $1.5B price tag on how the training data was sourced, even as the training itself was blessed as transformative. Adoption and legality are now being scored on axes the vendor doesn’t control.

The third story rhymes. An anonymous Ox Alpha on OpenRouter posted an 80% Pass@1 headline — until tokenizer forensics matched it to Zhipu’s unreleased GLM-5.3 at 99% confidence, dragging a stealth-launch playbook into US Entity List territory. Users, courts, and forensics researchers are all writing the meaning of these launches after the fact.

Anthropic pays $1.5B for pirated books, not for training on them

Source: techcrunch-ai · published 2026-08-23

TL;DR

  • Anthropic settled for $1.5B — ~$3,000 per work across ~500,000 pirated titles, 4× the statutory minimum.
  • Judge Alsup split training from sourcing in Bartz v. Anthropic: LLM training is “quintessentially transformative,” LibGen ingestion is not.
  • Meta’s Kadrey win is fragile — the judge called it a “technical” result and invited future plaintiffs to prove market dilution.
  • HarperCollins is pricing licenses at $2,500/title for three years, split 50/50 with publishers — a rate author Daniel Kibblesmith called “abominable.”

The doctrine has actually converged

TechCrunch frames the legal picture as “complicated,” but US courts have quietly aligned on one specific split: training a large language model on copyrighted books is transformative fair use, and how you obtained those books is a separate question with its own price tag. Judge Alsup’s June 2025 ruling in Bartz v. Anthropic put it plainly — training Claude on books was “quintessentially transformative,” while maintaining a permanent library of works pulled from shadow libraries like LibGen was not 1.

That distinction is what produced the largest copyright settlement in US history. In July 2026 Anthropic agreed to pay $1.5 billion to resolve claims over roughly 500,000 pirated works — about $3,000 per title, which the court itself noted was four times the statutory minimum for ordinary infringement, plus an order to destroy the pirated files 2. Every pending AI copyright suit now has a number to anchor to.

Kadrey is shakier than the headlines

Meta’s summary-judgment win in Kadrey got read as a second data point for “training is fair use.” The Authors Guild reads it differently: Judge Chhabria called it a “technical win” based on the specific record, noting that plaintiffs failed to develop sufficient evidence of market harm, and explicitly flagged that a better-developed record on market dilution could produce a different outcome 3.

That caveat is load-bearing because the ongoing OpenAI MDL is building exactly that record. In January 2026 Judge Stein compelled OpenAI to produce 20 million anonymized ChatGPT conversation logs so plaintiffs can quantify how often models reproduce or summarize copyrighted content in real-world use 4. Summary-judgment briefing runs through late 2026. The “training is fair use” consensus could narrow before it hardens.

The market isn’t waiting for the courts

While the litigation grinds on, publishers have started pricing the right directly. HarperCollins offered authors a flat $2,500 per title for a three-year training license, typically split 50/50 with the publisher — an offer author Daniel Kibblesmith publicly rejected as “abominable” 5. The Authors Guild’s April 2026 model contract pushes for 75–85% author share and separate line items for training, retrieval-augmented generation, and derivative summaries — a direct rebuke to the publisher-friendly split.

Europe answered the question and creators still object

The US fair-use frame obscures that the EU already resolved the doctrinal question by statute. The EU AI Act requires general-purpose model providers to comply with the 2019 DSM Directive, which permits text and data mining for commercial purposes unless rights holders opt out via machine-readable means, plus a “sufficiently detailed summary” transparency requirement 6. A coalition of 40 creative organizations formally protested in July 2024 that the implementation guidelines were too vague and machine-readable opt-outs are impractical for individual authors.

What’s actually at stake

The interesting questions are no longer whether training is legal. They are whether provenance liability alone makes large-scale training economically untenable at frontier scale, and whether the market-dilution evidence surfacing in OpenAI’s discovery retroactively narrows the Bartz and Kadrey wins. The $1.5B number already priced the first question. The second is still open.


Anthropic’s Fable 5 stalls at 8% as Opus 4.8 takes 28%

Source: simon-willison · published 2026-08-23

TL;DR

  • Fable 5 captured just 8% of Anthropic spend on Ramp in July, versus 28% for the older Opus 4.8 7.
  • Anthropic’s annualized revenue jumped from $47B to $65B May-to-July, with 6,000 customers paying $100K+/year.
  • Accel’s Miles Clements — an Anthropic investor — calls the flagship-by-default period “not a durable era” 7.
  • Opus 5’s new tokenizer emits ~30% more tokens per prompt, erasing most of the halved $5/$25 sticker cut 8.

The chart everyone is reading wrong

The headline number from the FT — Fable 5 stuck at 8% of Anthropic spend on Ramp’s index — is real, but it is not primarily a story about a model that launched July 24 not having time to ramp. Opus 5, released the same day, is at 3.5%. The model actually eating everyone’s lunch is Opus 4.8, at 28%. Enterprises are not waiting for the new thing; they are actively staying on the old thing.

Ramp’s own chief economist Ara Kharazian put a name on it: Fable 5 has established a “new upper bound” for what companies are willing to pay, and the performance gap versus cheaper tiers isn’t enough to justify the premium 9. Accel’s Miles Clements — pointedly, an Anthropic investor — told the FT that “most people don’t need to operate at the frontier” 7. When your own cap table is briefing that frontier pricing has hit a demand ceiling, the “give it time” reading gets harder to sustain.

Tokenizer economics and a stealth price hike

The developer-side signal matches. A widely-circulated PSA on r/Anthropic reports Opus 5 spending 43 minutes spinning up sandboxes and test suites for a config-file patch that Opus 4.8 finished in two, and identifies a new tokenizer that emits roughly 30% more tokens for the same input 8. At nominally halved $5/$25 rates, that erases most of the price cut in practice. It also explains the Opus 4.8 stickiness: buyers who’ve calibrated their monthly burn on 4.8 see the successor as more expensive per completed task, not less.

Fable 5 carries additional baggage the spend chart doesn’t show. The Trump administration forced a brief withdrawal in early June over national-security concerns; the July 1 re-release came bundled with 30-day data-retention requirements that spooked privacy-sensitive buyers 10. Anthropic sources are now briefing that the low share is “intentional” — Fable 5 is a Mythos-class showcase for extreme reasoning, while Opus 5 is meant to carry volume 11. That’s a reasonable retcon, but it’s still a retcon.

What the Ramp panel doesn’t see

Before declaring the frontier dead, note the methodology ceiling. Ramp’s 70,000-firm panel skews tech-forward and venture-backed, misses free-tier and “shadow AI” charged to personal cards, and undercounts AI embedded inside Salesforce and Microsoft contracts 12. The intra-Anthropic share ranking is credible; extrapolating it to “the market rejects flagship models” is not.

The real takeaway is narrower and sharper: the pricing power of a new flagship is now conditional on tokens-per-task, not headline $/Mtok. Anthropic’s revenue is still compounding fast — but on Opus 4.8, not on the model it wants customers to buy.


Ox Alpha fingerprinted as Zhipu’s unreleased GLM-5.3

Source: techcrunch-ai · published 2026-08-23

TL;DR

  • Tokenizer forensics match Ox Alpha to Zhipu’s GLM family at 99% confidence, with stack traces exposing the internal paas/v4/chat path.
  • 80% Pass@1 on a DeepSWE slice beat Fable (65%) and GPT-5.6 Sol (52%).
  • That headline rests on just 10 tasks, unaudited, with broader private evals reportedly weaker.
  • Zhipu has been on the US Entity List since January 2025, turning “free anonymous coding model” into a sanctions problem.
  • Stealth “Alpha” drops on OpenRouter are now a repeatable pre-launch playbook, not a mystery genre.

The whodunit is basically solved

TechCrunch framed Ox Alpha as an open mystery. The fingerprinting community has already closed it. A 95-of-95 tokenizer probe matches the GLM family with a constant 75-token offset attributable to a hidden system wrapper; malformed requests return a proprietary 1214 Incorrect role information error and Java stack traces containing paas/v4/chat, a path unique to Zhipu’s official API 13. Behavioral tells pile on: German-style LaTeX decimals like 0{,}375, a characteristic emoji density, and CCP-aligned answers — the model calls Taiwan an “inalienable part of China” when pushed 1314. The working hypothesis is Zhipu’s unreleased GLM-5.3, expected to drop open-weights on Aug 28.

The benchmark number is real but thin

The viral claim is an 80% Pass@1 on a 10-task DeepSWE subset, well ahead of Fable at 65% and GPT-5.6 Sol at 52% 15. That gap is eye-catching, but ten tasks is not a leaderboard, and private evals have reportedly placed Ox Alpha closer to two-generation-old systems on broader suites. The more defensible distinctions are structural: a 1,048,576-token context window, native video and image inputs, and a claimed 100T-token/day serving capacity — a footprint noticeably broader than the text-only GLM-5.3 Zhipu pre-announced on Aug 14.

Stealth drops are a playbook now, not an event

Ox Alpha is the fifth or sixth iteration of the same pattern. Anonymous “Alpha” endpoints appear on OpenRouter, coders benchmark them, then the lab claims them at launch:

Stealth handleActual modelRevealed
Quasar Alpha / Optimus AlphaOpenAI GPT-4.1Apr 2025
Sonoma Sky / Sonoma DuskxAI Grok 4 Fast2025
Hunter AlphaXiaomi MiMo-V2-Pro2026
Owl AlphaMeituan LongCat-2.02026
Ox AlphaZhipu GLM-5.3 (expected Aug 28)pending

Source: 16. Labs get brand-blind telemetry; OpenRouter gets traffic; power users get a free frontier-ish model for a week. Everyone in that loop knows the game.

The part TechCrunch underplayed

If the Zhipu attribution holds — and the forensics are hard to argue with — US developers piping proprietary code into a free anonymous endpoint are potentially routing it to an entity the Commerce Department added to the Entity List in January 2025 for advancing military modernization 17. OpenRouter’s own docs confirm the anonymous provider retains prompts and completions; SiliconAngle’s headline is blunt: “nobody knows… where the code goes” 18.

Anonymous provider, retained I/O, no DPA, and a likely sanctioned operator.

That’s not a privacy footnote. It’s a compliance tripwire, and it’s the reason the interesting question about Ox Alpha shifted this week from who built it? to who’s allowed to use it? The fingerprinting community treats the first question as settled. The security community treats the answer as the problem.

Round-ups

Flock CEO seeks compromise amid surveillance backlash

Source: techcrunch-ai

Flock Safety’s chief executive is pushing for a middle ground as public outcry grows over the company’s license-plate reader network and fears of misuse. The pitch lands as civil-liberties groups and local governments increasingly challenge deployments of the surveillance hardware.

Linkdaze’s smart calendar ships AI meal planner without paywall

Source: techcrunch-ai

Linkdaze positions its smart digital calendar as a household operating system rather than a scheduling tool, bundling an AI meal planner and other features with no subscription tier. The free-feature approach contrasts with rivals that gate AI extras behind monthly fees.

Footnotes

  1. StartupFortune on Judge Alsup’s Bartz v. Anthropic rulinghttps://startupfortune.com/how-one-judges-split-ruling-on-anthropic-became-ais-copyright-rulebook/

    using copyrighted books to train large language models like Claude was ‘quintessentially transformative’… However, the court drew a hard line at the ‘inputs’ side of the process… downloading and maintaining a permanent library of pirated works from ‘shadow libraries’ like Library Genesis (LibGen) was not fair use.

  2. Wolters Kluwer Copyright Blog on the $1.5B Anthropic settlementhttps://legalblogs.wolterskluwer.com/copyright-blog/the-bartz-v-anthropic-settlement-understanding-americas-largest-copyright-settlement/

    Anthropic agreed to a historic $1.5 billion settlement—the largest in U.S. copyright history—to resolve claims regarding the unauthorized acquisition of nearly 500,000 works… approximately $3,000 per work, which the court noted was four times the statutory minimum.

  3. Authors Guild statement on Kadrey v. Meta rulinghttps://authorsguild.org/news/meta-ai-ruling-meta-gets-technical-win-but-law-favors-authors/

    a ‘technical win’ based on the specific evidence presented rather than a broad legal precedent… the plaintiffs failed to provide sufficient evidence that Meta’s specific use had already caused such market harm… a better-developed record of ‘market dilution’ could lead to different outcomes.

  4. National Law Review on OpenAI discovery orderhttps://natlawreview.com/article/openai-loses-privacy-gambit-20-million-chatgpt-logs-likely-headed-copyright

    Judge Stein affirmed a discovery order compelling OpenAI to produce 20 million anonymized ChatGPT conversation logs… essential for authors to determine how often the models reproduce or summarize copyrighted content in the ‘real world’.

  5. StoryWriter.pro on 2025-26 publisher licensing dealshttps://storywriter.pro/knowledge/how_does_ai_licensing_for_book_authors_work_and_what_are_the_current_legal_standards_in_2026.php

    HarperCollins offered authors a flat fee of $2,500 per title for a three-year training license, typically split 50-50 with the publisher… author Daniel Kibblesmith publicly rejected the offers, labeling them ‘abominable’.

  6. Garrigues analysis of EU AI Act TDM opt-outhttps://www.garrigues.com/en_GB/garrigues-digital/ai-and-copyright-machine-readable-machine-actionable-opt-out-tdm-question

    The EU AI Act… mandates that providers of general-purpose AI models must comply with the 2019 DSM Directive, which permits TDM for commercial purposes unless rights holders explicitly ‘opt out’ using machine-readable means… a coalition of 40 creative organizations formally protested that the EU’s implementation guidelines were too vague.

  7. Financial Times (original article)https://www.ft.com/content/5ee49718-c258-4f01-aa32-7e5b76ae5245?syn-25a6b1a6=1

    Miles Clements, a partner at Accel [an Anthropic investor], said ‘most people don’t need to operate at the frontier’… describing the period where flagship models were the primary choice as ‘not a durable era.’

    2 3
  8. r/Anthropic developer PSAhttps://www.reddit.com/r/Anthropic/comments/1v5jw8x/psa_for_revs_be_careful_upgrading_to_claude_opus/

    Opus 5 spending 43 minutes spinning up sandboxes and testing suites for a simple config-file patch that its predecessor completed in two minutes… a new tokenizer that generates approximately 30% more tokens for the same text, effectively acting as a stealth price hike.

    2
  9. PYMNTS coverage of FT reporthttps://www.pymnts.com/news/artificial-intelligence/2026/anthropic-customers-switch-to-cheaper-models-ahead-of-ipo/

    Ara Kharazian, chief economist at Ramp, noted that Fable 5 has established a ‘new upper bound’ for what companies are willing to pay, with many finding the performance gap between it and cheaper models insufficient to justify the premium.

  10. Incrypted / analyst rounduphttps://incrypted.com/en/companies-ditch-claude-fable-5-in-favor-of-cheaper-ai-solutions/

    The Trump administration initially forced a withdrawal of the model in early June due to national security concerns… Although it was relaunched on July 1 with federal approval, analysts believe this regulatory uncertainty, combined with strict [30-day] data-retention requirements, further discouraged early enterprise adoption.

  11. Startup Fortunehttps://startupfortune.com/anthropic-cuts-claude-opus-prices-in-half-as-enterprises-balk-at-the-bill/

    Anthropic has since pivoted its commercial focus toward Claude Opus 5, which has already surpassed Fable 5 in total enterprise spending due to its superior performance-per-dollar ratio… sources framed the low adoption not as a failure, but as an ‘intentional’ strategy.

  12. ExplainX critique of Ramp AI Indexhttps://explainx.ai/blog/top-1-percent-ai-spend-per-employee-ramp-index-august-2026

    Ramp’s customer base disproportionately consists of tech-forward, venture-backed, and knowledge-work-intensive firms… this ‘transaction-first’ view creates a ‘blind spot’ for free tools and ‘shadow AI’ purchased via personal accounts.

  13. local-ai-zone technical analysishttps://local-ai-zone.github.io/blog/ox-alpha-stealth-model-comprehensive-analysis.html

    Ox Alpha’s tokenizer behavior matches the GLM family with 99% certainty, exhibiting a constant 75-token offset likely caused by a hidden system wrapper… malformed requests triggered a proprietary ‘1214 Incorrect role information’ error and Java stack traces containing the path ‘paas/v4/chat’, which is unique to Zhipu’s official API.

    2
  14. Business Insiderhttps://www.businessinsider.com/ox-alpha-ai-model-mystery-2026-8

    When questioned on sensitive topics such as the independence of Taiwan, the model provides responses aligned with Chinese state policy, stating it is an ‘inalienable part of China’.

  15. Pandailyhttps://pandaily.com/anonymous-ai-model-ox-alpha-crushes-coding-benchmarks-aug2026

    On a 10-task subset of the DeepSWE coding benchmark, Ox Alpha reportedly achieved an 80% Pass@1 rate, notably outperforming Fable (65%) and GPT-5.6 Sol (52%) in the same sample run.

  16. Adam Holter blog (historical context)https://adam.holter.com/sonoma-dusk-alpha-sonoma-sky-alpha-exploring-xais-stealth-llms-with-2m-token-context-windows/

    Quasar Alpha and Optimus Alpha were officially revealed on April 14, 2025 as early snapshots of OpenAI’s GPT-4.1; the Sonoma Sky/Dusk pair was later identified as xAI’s Grok 4 Fast — establishing a pattern where ‘Alpha’ stealth models are free previews that get retired once the official model launches.

  17. Digital Appliedhttps://www.digitalapplied.com/blog/stealth-ox-alpha-anonymous-frontier-model-appears

    In January 2025, the U.S. Commerce Department added Zhipu AI to the Entity List, citing its role in advancing military modernization… creates a legal tripwire for US firms sending proprietary code to an anonymous provider that may be a sanctioned entity.

  18. SiliconAnglehttps://siliconangle.com/2026/08/23/nobody-knows-who-built-ai-coding-model-ox-alpha-or-where-the-code-goes/

    Nobody knows who built AI coding model Ox Alpha or where the code goes — the anonymous provider retains all prompts and completions, prompting warnings against using the model with sensitive or proprietary code.

Jack Sun

Jack Sun, writing.

Engineer · Bay Area

Hands-on with agentic AI all day — building frameworks, reading what industry ships, occasionally writing them down.

Digest
All · AI Tech · AI Research · AI News
Writing
Essays
Elsewhere
Subscribe
All · AI Tech · AI Research · AI News · Essays

© 2026 Wei (Jack) Sun · jacksunwei.me Built on Astro · hosted on Cloudflare