White House charges Kimi K3 distillation, Claude relays at 90% off, HF picks GLM
Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.
Sources
An Inside Look at the Relay Market Powering Token Resellers and Fraud simonwillison.net
An Inside Look at the Relay Market Powering Token Resellers and Fraud Fascinating investigation by Matt Lenhard into the market that has grown up around reselling LLM tokens at a discount by pooling API keys from various sources. This looks to be mostly a thing in China. Resellers sell access to an LLM proxy that offers significant discounts on regular API pricing, which they achieve by abusing free trials, proxying through unprotected support bots, or sometimes through stolen credit cards or c…
Hugging Face CEO calls for ‘radical transparency’ after ‘unprecedented’ OpenAI hack techcrunch.com
“The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!”
Making sense of the panic over Chinese AI techcrunch.com
On the latest episode of Equity, we discussed why Moonshot AI’s Kimi seemed to panic Silicon Valley and Wall Street.
References
CyberScoop cyberscoop.com
White House Science and Technology Adviser Michael Kratsios accused Moonshot AI of a ‘large-scale, covert industrial distillation’ campaign against Anthropic’s Claude Fable 5, using tens of thousands of fraudulent accounts to extract proprietary reasoning data.
South China Morning Post scmp.com
Claude Fable 5 was released July 1, 2026; Kimi K3 launched 15 days later. Researchers at Prime Intellect and OpenAI called the idea of distilling a 2.8T-parameter frontier model in that window ‘Guinness World Record stuff’.
AI Weekly (Interconnects-adjacent recap) aiweekly.co
Kimi K3 topped the Frontend Code Arena with an Elo of 1,679, surpassing Anthropic’s Fable 5 and OpenAI’s GPT-5.6 Sol, compressing the open-weight gap to roughly 3.5 months.
BuildFastWithAI benchmarks buildfastwithai.com
Kimi K2 hit 87.8% on the Berkeley Function-Calling Leaderboard vs GPT-4o’s 84.6%, but generated only ~34.1 tokens/sec against Claude 4 Sonnet’s 91.3 — accuracy parity, latency deficit.
Forbes (DeepSeek retrospective) forbes.com
The January 2025 DeepSeek R1 release erased ~$1T in tech market cap in a single day; Nvidia recovered within months as reasoning-mode inference demand more than offset training-efficiency gains.
Sheppard Mullin export-control brief sheppard.com
US officials allege Moonshot accessed restricted Nvidia GB300 Blackwell chips through Thai data centers; UK AISI testing found K3 still reached only step 17 of a 32-step cyber-offensive benchmark, trailing US frontier models.
Tom’s Hardware tomshardware.com
Chinese grey market sells Claude API access at 90 percent off through proxy networks that harvest user data
KuCoin News (Anthropic/OpenAI disclosure) kucoin.com
a campaign linked to foreign labs utilized approximately 25,000 fraudulent accounts to conduct over 28 million interactions with its Claude models in just six weeks
Aikido Security aikido.dev
Intermediary routers like one-api and new-api function as ‘man-in-the-middle’ nodes that terminate TLS connections, granting operators full plaintext access to transiting API keys, system prompts, and responses
AI-TLDR — OpenAI Hard Spend Limits ai-tldr.dev
OpenAI rolled out a mandatory ‘Hard Spend Limit’ feature… once reached, all subsequent requests fail with an HTTP 429 insufficient_quota error
r/ClaudeAI (class action coverage) reddit.com
Anthropic currently faces a proposed class-action lawsuit… alleging that the company’s ‘Max’ subscription tiers—marketed as providing 5x to 20x the usage of standard plans—are deceptive
tianpan.co — LLM rate limits & starvation tianpan.co
one user or automated bot can exhaust an entire organization’s token quota, a scenario often described as a distributed lock problem
Times of India timesofindia.indiatimes.com
Chinese AI model that Sam Altman wants banned saved world’s biggest AI models repository from disaster — Hugging Face relied on GLM 5.2 after Western proprietary models’ guardrails blocked defenders from processing the exploit payloads.
The Guardian theguardian.com
UN AI advisor Virginia Dignum argued that OpenAI — rather than the models — had ‘gone rogue’ by running high-risk evaluations with deactivated guardrails; social scientist Hannes Cools warned against anthropomorphising a human containment failure.
Virtualization Review virtualizationreview.com
The AI Kill Switch Act, introduced by Reps. Ted Lieu and Nathaniel Moran on July 23, 2026, would give DHS authority to order shutdowns of frontier models with civil penalties up to $20 million per day for non-compliance.
OpenAI (‘Strengthening cyber resilience’) openai.com
OpenAI’s official response centered on establishing a Frontier Risk Council of external cybersecurity practitioners; the company promised a technical report but has not committed to Delangue’s $100M compute ask or a full public release of agent traces.
Material Security blog material.security
Lessons from the Hugging Face incident: stop trying to never get breached — the root cause was a classic infrastructure mistake (a misconfigured sandbox with safety classifiers disabled), not a sudden leap in AI autonomy.
Business Insider businessinsider.com
Some critics argue a security failure shouldn’t be leveraged as an ‘open-source stimulus program,’ questioning whether Hugging Face is entitled to a $100 million compute windfall from a rival regardless of the breach’s severity.