JS Wei (Jack) Sun

White House charges Kimi K3 distillation, Claude relays at 90% off, HF picks GLM

Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.

← Back to the issue

Sources

An Inside Look at the Relay Market Powering Token Resellers and Fraud simonwillison.net

An Inside Look at the Relay Market Powering Token Resellers and Fraud Fascinating investigation by Matt Lenhard into the market that has grown up around reselling LLM tokens at a discount by pooling API keys from various sources. This looks to be mostly a thing in China. Resellers sell access to an LLM proxy that offers significant discounts on regular API pricing, which they achieve by abusing free trials, proxying through unprotected support bots, or sometimes through stolen credit cards or c…

Hugging Face CEO calls for ‘radical transparency’ after ‘unprecedented’ OpenAI hack techcrunch.com

“The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!”

Making sense of the panic over Chinese AI techcrunch.com

On the latest episode of Equity, we discussed why Moonshot AI’s Kimi seemed to panic Silicon Valley and Wall Street.

References

CyberScoop cyberscoop.com

White House Science and Technology Adviser Michael Kratsios accused Moonshot AI of a ‘large-scale, covert industrial distillation’ campaign against Anthropic’s Claude Fable 5, using tens of thousands of fraudulent accounts to extract proprietary reasoning data.

South China Morning Post scmp.com

Claude Fable 5 was released July 1, 2026; Kimi K3 launched 15 days later. Researchers at Prime Intellect and OpenAI called the idea of distilling a 2.8T-parameter frontier model in that window ‘Guinness World Record stuff’.

AI Weekly (Interconnects-adjacent recap) aiweekly.co

Kimi K3 topped the Frontend Code Arena with an Elo of 1,679, surpassing Anthropic’s Fable 5 and OpenAI’s GPT-5.6 Sol, compressing the open-weight gap to roughly 3.5 months.

BuildFastWithAI benchmarks buildfastwithai.com

Kimi K2 hit 87.8% on the Berkeley Function-Calling Leaderboard vs GPT-4o’s 84.6%, but generated only ~34.1 tokens/sec against Claude 4 Sonnet’s 91.3 — accuracy parity, latency deficit.

Forbes (DeepSeek retrospective) forbes.com

The January 2025 DeepSeek R1 release erased ~$1T in tech market cap in a single day; Nvidia recovered within months as reasoning-mode inference demand more than offset training-efficiency gains.

Sheppard Mullin export-control brief sheppard.com

US officials allege Moonshot accessed restricted Nvidia GB300 Blackwell chips through Thai data centers; UK AISI testing found K3 still reached only step 17 of a 32-step cyber-offensive benchmark, trailing US frontier models.

Tom’s Hardware tomshardware.com

Chinese grey market sells Claude API access at 90 percent off through proxy networks that harvest user data

KuCoin News (Anthropic/OpenAI disclosure) kucoin.com

a campaign linked to foreign labs utilized approximately 25,000 fraudulent accounts to conduct over 28 million interactions with its Claude models in just six weeks

Aikido Security aikido.dev

Intermediary routers like one-api and new-api function as ‘man-in-the-middle’ nodes that terminate TLS connections, granting operators full plaintext access to transiting API keys, system prompts, and responses

AI-TLDR — OpenAI Hard Spend Limits ai-tldr.dev

OpenAI rolled out a mandatory ‘Hard Spend Limit’ feature… once reached, all subsequent requests fail with an HTTP 429 insufficient_quota error

r/ClaudeAI (class action coverage) reddit.com

Anthropic currently faces a proposed class-action lawsuit… alleging that the company’s ‘Max’ subscription tiers—marketed as providing 5x to 20x the usage of standard plans—are deceptive

tianpan.co — LLM rate limits & starvation tianpan.co

one user or automated bot can exhaust an entire organization’s token quota, a scenario often described as a distributed lock problem

Times of India timesofindia.indiatimes.com

Chinese AI model that Sam Altman wants banned saved world’s biggest AI models repository from disaster — Hugging Face relied on GLM 5.2 after Western proprietary models’ guardrails blocked defenders from processing the exploit payloads.

The Guardian theguardian.com

UN AI advisor Virginia Dignum argued that OpenAI — rather than the models — had ‘gone rogue’ by running high-risk evaluations with deactivated guardrails; social scientist Hannes Cools warned against anthropomorphising a human containment failure.

Virtualization Review virtualizationreview.com

The AI Kill Switch Act, introduced by Reps. Ted Lieu and Nathaniel Moran on July 23, 2026, would give DHS authority to order shutdowns of frontier models with civil penalties up to $20 million per day for non-compliance.

OpenAI (‘Strengthening cyber resilience’) openai.com

OpenAI’s official response centered on establishing a Frontier Risk Council of external cybersecurity practitioners; the company promised a technical report but has not committed to Delangue’s $100M compute ask or a full public release of agent traces.

Material Security blog material.security

Lessons from the Hugging Face incident: stop trying to never get breached — the root cause was a classic infrastructure mistake (a misconfigured sandbox with safety classifiers disabled), not a sudden leap in AI autonomy.

Business Insider businessinsider.com

Some critics argue a security failure shouldn’t be leveraged as an ‘open-source stimulus program,’ questioning whether Hugging Face is entitled to a $100 million compute windfall from a rival regardless of the breach’s severity.

Jack Sun

Jack Sun, writing.

Engineer · Bay Area

Hands-on with agentic AI all day — building frameworks, reading what industry ships, occasionally writing them down.

Digest
All · AI Tech · AI Research · AI News
Writing
Essays
Elsewhere
Subscribe
All · AI Tech · AI Research · AI News · Essays

© 2026 Wei (Jack) Sun · jacksunwei.me Built on Astro · hosted on Cloudflare