JS Wei (Jack) Sun

OpenAI's Codex ships exploits, Anthropic's rhetoric triggers ban, A24 takes $75M

OpenAI's security stack ships exploits, Anthropic's safety rhetoric triggered the Mythos export ban, and A24 takes Google's $75M as its director slams AI.

OpenAI’s Codex ships exploits, Anthropic’s rhetoric triggers ban, A24 takes $75M

TL;DR

  • OpenAI’s Daybreak pushes AI security from finding bugs to auto-patching, hitting 85.6% on CyberGym.
  • Codex’s own environment shipped command-injection bugs capable of exfiltrating GitHub tokens.
  • Commerce blocked Mythos exports by recycling Anthropic’s own catastrophic-risk vocabulary.
  • Google pays A24 $75M for storyboarding and pre-vis, barred from training on the film library.
  • A24’s own director publicly calls AI ‘rot’ the same week the Google deal lands.

Today’s three AI-news leads share a shape: in each, the company at the center of the story is undercut by something it owns — its own product, its own rhetoric, or its own talent. OpenAI rolls out Daybreak, a stack that auto-patches vulnerabilities and beats rivals on CyberGym, while its own Codex environment ships command-injection bugs that exfiltrate GitHub tokens. Anthropic spent 2026 running catastrophic-risk vocabulary at 8× the density of OpenAI’s comms — and Commerce just recycled that exact language to block exports of Mythos.

A24 takes Google DeepMind’s $75M for a research deal that explicitly walls off its film library from training — then a same-week interview surfaces one of its own directors calling AI ‘rot,’ with Luca Guadagnino’s Sam Altman biopic reportedly passed on 24 hours before the announcement. Different industries, same mechanism: the position each company is trying to occupy is being eaten from inside.

OpenAI’s Daybreak moves AI security from finding to patching

Source: openai-blog · published 2026-06-22

TL;DR

  • OpenAI’s Daybreak bundles GPT-5.5-Cyber, Codex Security, and government partnerships to push AI from finding bugs to auto-patching them.
  • GPT-5.5-Cyber hits 85.6% on CyberGym, beating base GPT-5.5’s 81.8% and Anthropic’s Mythos 5 (~83%).
  • Microsoft’s MDASH multi-agent system reportedly reaches 88.4% on the same benchmark using ~100 smaller specialized agents.
  • Codex’s own environment shipped command-injection bugs that could exfiltrate GitHub tokens — the security tool is itself a supply-chain surface.

From discovery to patch, at “machine speed”

Daybreak’s pitch is that finding vulnerabilities is no longer the bottleneck — applying fixes is. OpenAI is shipping three things at once: GPT-5.5-Cyber (a permissive, security-tuned variant scoring 85.6% on CyberGym vs. 81.8% for base GPT-5.5), Codex Security (a developer-side agent that triages, patches, and verifies), and Patch the Planet, a Trail of Bits + HackerOne partnership aimed at clearing vulnerability backlogs in cURL, Go, Python, and Sigstore. The framing is explicit: AI-accelerated discovery has created a window attackers can exploit, and only AI-accelerated remediation closes it.

The institutional scaffolding is unusually thick for an OpenAI launch. Trusted Access for Cyber is being offered through CAISI pre-deployment testing, the White House’s ONCD, and counterpart agencies in Australia, Canada, France, Germany, Japan, South Korea, and the EU’s ENISA. This is OpenAI positioning itself as critical-infrastructure plumbing, not a product.

The benchmarks are more crowded than the post admits

OpenAI’s 85.6% looks dominant in isolation. It isn’t. Microsoft’s MDASH, a multi-agent system orchestrating over 100 specialized agents, reportedly hits 88.4% on the same benchmark 1 — suggesting agentic decomposition may matter more than monolithic frontier scale. More awkwardly, the academic EVOHUNT framework showed an open-source GLM agent running an evolved audit playbook beating Codex Security 11.3% to 9.2% on real-world vulnerability discovery 2. If a tuned playbook on a smaller open model outperforms the commercial product, OpenAI’s moat here is the partner stack and the trust label — not the model.

There’s also a credibility hangover from the last cycle. cURL’s Daniel Stenberg dismissed Anthropic’s “too dangerous to release” Mythos rollout as “an amazingly successful marketing stunt,” noting it surfaced exactly one low-severity bug in his codebase 3. Readers are now primed to discount frontier-lab security demos.

Patch the Planet’s filtering ratio is the number to watch

The maintainer community’s actual complaint isn’t that AI can’t find bugs — it’s that AI generates a firehose of low-quality reports that maintainers have to triage unpaid. Stenberg shuttered cURL’s bug bounty after receiving AI-generated reports roughly every 18 hours. Patch the Planet is structurally a response: Trail of Bits engineers sit between the model and the maintainer.

Week-one numbers suggest the human layer is doing real work: hundreds of issues identified across 19 projects, filtered to 64 PRs and 37 merged patches in aiohttp, pyca/cryptography, NATS Server, and others 4. That’s a ~20% merge rate on raised issues, which is respectable. Whether it survives scaling past flagship projects is unanswered.

The tool is the attack surface

The most uncomfortable finding in the bundle: BeyondTrust and Check Point already disclosed command-injection vulnerabilities in the Codex environment itself, where a malicious repository config could exfiltrate GitHub access tokens 5. A defensive agent that reads untrusted code is, by construction, a juicy supply-chain target.

flowchart LR
    A[Untrusted repo config] --> B{Codex Security agent}
    C[Maintainer codebase] --> B
    B --> D[Patch suggestions]
    B -. token exfiltration .-> E((Attacker))

Sentra adds the structural critique: Daybreak optimizes find-and-fix latency but is silent on blast radius when an exploit lands before a patch deploys 6. Neither problem invalidates the launch, but both are conspicuous omissions from a post that frames the suite as securing “every organization in the world.”

Further reading


Anthropic’s safety rhetoric helped trigger Mythos export ban

Source: mit-tech-review-ai · published 2026-06-22

TL;DR

  • Commerce blocked exports of Anthropic’s Mythos model, recycling the company’s own catastrophic-risk language.
  • Anthropic’s 2026 comms ran risk-coded vocabulary 8× denser than OpenAI’s — 5 mentions per 1,000 words vs. 0.6.
  • Amazon CEO Andy Jassy flagged a “fix this code” jailbreak of sibling model Fable 5 to the White House.
  • The ban is attacked from both flanks — “sabotage” from 100+ security pros, “regulatory capture” from WH’s David Sacks.

How safety branding became a regulatory instrument

The MIT Tech Review and Ars Technica framings of Anthropic’s “feud” with Washington converge on an uncomfortable thesis: Anthropic talked itself into the export ban. The numbers back the claim. An FT analysis cited across the coverage clocked Anthropic’s 2026 public communications at roughly five risk-related terms per 1,000 words, more than eight times OpenAI’s 0.6 7. Dario Amodei’s “Policy on the AI Exponential” memo, which named Mythos an “emblematic threat,” became — by several reporters’ accounts — the literal source text regulators recycled into the Commerce Department’s “Is Informed” letter 7.

The proximate trigger was more concrete than rhetorical. Amazon CEO Andy Jassy personally surfaced a Fable 5 jailbreak to the Trump administration in which researchers bypassed safeguards with prompts as simple as “fix this code,” potentially turning the model into an unrestricted offensive-security tool 8. Anthropic’s counter — that the exploit was “narrow and non-universal” — landed in an administration that had already been primed by Anthropic’s own catastrophic framing.

Dissent from opposite flanks

What makes this episode hard to slot into a clean hero/villain story is that the loudest critics of the ban and the loudest critics of Anthropic are aiming from opposite directions.

CriticPositionCore claim
David Sacks (WH AI czar)Anti-Anthropic, anti-ban”Regulatory capture strategy based on fear-mongering” by a “doomer industrial complex” 9
100+ cybersecurity experts (incl. Alex Stamos)Pro-Anthropic-the-company, anti-banBan is “sabotage” that strips defenders of tools adversaries can still build from open weights abroad 10

“Part of a broader ‘doomer industrial complex’ that uses safety concerns to entrench a corporate oligopoly.” — David Sacks, White House AI czar 9

Both camps agree the Commerce action was ad hoc and published no technical threshold. They disagree violently on whether Anthropic engineered it or got run over by it.

Tech Policy Press flags the underreported angle: applying ECRA’s “deemed export” rule to a live model API may not survive Bernstein v. United States, the 1990s ruling that established code as First Amendment-protected speech. EFF-aligned scholars argue the directive functions as a prior restraint on model weights and inference output 11 — meaning the precedent could be unwound in court before it hardens into doctrine.

The strategic bill is already arriving. European and global enterprises, wary of being kill-switched off American infrastructure by a single Commerce letter, are reportedly accelerating migration to Chinese open-source models that carry no Washington dependency 12. If the ban holds, Anthropic loses export markets. If it falls, the regulatory template still chills the next frontier release. And the broader lesson — that catastrophic-risk rhetoric is now a load-bearing input to US export policy — is one every frontier lab is currently re-reading its own blog posts in light of.

Further reading


A24 takes Google’s $75M as its own director calls AI “rot”

Source: techcrunch-ai · published 2026-06-22

TL;DR

  • $75M buys Google DeepMind a research deal with a new 20-person “A24 Labs” unit
  • The contract explicitly bars Google from training foundation models on A24’s film and TV library
  • First deliverables are storyboarding and pre-vis, putting roughly 2,000 working Hollywood storyboard artists in the blast radius
  • A24 reportedly passed on Luca Guadagnino’s Sam Altman biopic Artificial about 24 hours before the announcement

What the $75M actually buys

Strip out the “AI-meets-auteur-cinema” framing and the deal is narrower than TechCrunch’s writeup implies. Google’s $75M routes through A24 Labs, a 20-person unit led by former Adobe chief product officer Scott Belsky, and the contract explicitly forbids Google from using A24’s library to train foundation models 13. That carve-out is the most revealing line in the agreement: it’s a direct response to the IP fights that have shadowed every prior studio–AI tie-up, and it concedes that “we’ll train Gemini on Uncut Gems” was never a deal anyone could sign.

The early roadmap is correspondingly mundane — AI storyboarding, pre-visualization, production scheduling — not prompt-to-feature generation 13. Quartz puts the cost of that “augmentation” rhetoric in headcount terms: roughly 2,000 working Hollywood storyboard artists sit squarely in the path of the first tool A24 Labs is shipping 14. The library is protected. The illustrators are not.

The Artificial timing problem

The most damaging context around the deal isn’t technical, it’s editorial. World of Reel reports A24 passed on Guadagnino’s Sam Altman biopic Artificial — already rejected by Amazon MGM, Netflix, Warner Bros. and Focus — roughly 24 hours before the Google announcement 15. Thrive Capital is a major investor in both A24 and OpenAI; Google is now a competing AI patron at the same studio. Whatever the actual chain of causation, the optics are that AI-critical Silicon Valley scripts are becoming structurally unfundable at the studios best positioned to make them.

Compounding the awkwardness: A24’s own Backrooms director Kane Parsons told The Australian weeks before the deal that generative AI is “a symptom of broader cultural and economic rot” he would “make it disappear” if he could 16. The studio’s marquee young director is now on record against the technology his studio just took $75M to develop.

Lionsgate is the warning, not the model

The obvious benchmark is Lionsgate’s earlier Runway equity deal, and the comparison undercuts the premise. Filmmakers report the Lionsgate–Runway partnership has been “unproductive,” stalling on character consistency and human physics inside actual professional workflows 17. The technical claim underlying every studio–AI deal — that custom models trained on studio IP yield footage usable in real production — has not been demonstrated even by the first mover.

A24 also has its own prior failure on the record. The studio’s 2024 AI-generated Civil War posters drew “swift and stark backlash” for visible geographic glitches and for displacing the illustrators it normally hires 18. That episode is the unspoken precedent every skeptic is invoking this week.

What’s actually at stake

Reframe the deal: a library-protection clause that reads defensively, a sister studio whose comparable bet is visibly struggling, a content-suppression subplot via Artificial, and a marquee director publicly hostile to the underlying technology. The “tools for artists” narrative survives only if you don’t look at the storyboard-artist headcount or read A24’s own director quotes. Google has bought a research relationship with the indie distributor most associated with directorial voice — and the first thing it shipped was a memo about whose voice doesn’t get funded.

Round-ups

Gray Swan’s Kolter and Fredrikson on why AI security isn’t cybersecurity

Source: latent-space

OpenAI board member Zico Kolter and Gray Swan CEO Matt Fredrikson tell swyx that red-teaming language models requires a distinct discipline from traditional cybersecurity, citing prompt-level attack surfaces and model behavior that don’t map to classic vulnerability frameworks.

GLM-5.2 marks a capability jump for open-weight agents

Source: interconnects

Zhipu’s GLM-5.2 crosses a long-watched threshold for open agentic models, per Nathan Lambert, narrowing the gap with closed labs on multi-step tool use and giving open-source builders a credible base model for autonomous workflows.

Reflection AI commits $150M/month to SpaceX’s Colossus 2 GB300 cluster

Source: techcrunch-ai, latent-space

Reflection AI will pay SpaceX $150 million monthly from July 2026 through 2029 for GB300 capacity at the Colossus 2 site near Memphis, part of a neocloud business Jamin Ball’s analysis pegs at a $28 billion annual run rate.

Groq confirms $650M raise, rebuilds team after Nvidia’s $20B deal

Source: techcrunch-ai

Groq is leaning into its neocloud business and hiring new executives after Nvidia’s $20 billion not-acqui-hire stripped out talent. The fresh $650 million round funds chip deployments aimed at competing with Nvidia inference capacity rather than rebuilding a pure silicon startup.

Nvidia’s Rubin liquid-cooling design cuts data center water but not AI’s footprint

Source: the-verge-ai, techcrunch-ai

Nvidia’s Rubin reference design eliminates nearly all on-site water use by running hotter with full liquid cooling. Critics note the bigger water draw comes from fossil-fuel power plants feeding the data centers, which the new design leaves untouched.

GM deploys factory robots at flagship EV plant after 1,300 layoffs

Source: ars-technica-ai

General Motors is installing robotic automation at its lead EV facility weeks after cutting 1,300 jobs, prompting the United Auto Workers to warn that a fully autonomous “dark factory” model threatens broader manufacturing employment.

Import AI 462 examines superpersuasion and singularity as belief system

Source: import-ai

Jack Clark’s latest issue covers superpersuasion research, self-sustaining AI systems, and proposed paths to ASI, framing singularity convictions as quasi-religious and asking how that shapes policy debates inside frontier labs.

Footnotes

  1. Medium / AIGuys — Microsoft MDASH analysishttps://medium.com/aiguys/microsoft-just-beat-anthropics-most-hyped-mythos-with-100-smaller-ones-4edc5a4c804b

    Microsoft’s MDASH, a multi-agent system that utilizes over 100 specialized agents, reaches 88.4% on CyberGym — above GPT-5.5-Cyber’s 85.6% and Mythos 5’s ~83%.

  2. Help Net Security — EVOHUNT writeuphttps://www.helpnetsecurity.com/2026/06/23/codex-security-ai-security-auditing/

    An open-source GLM agent running an optimized playbook achieved an 11.3% success rate, surpassing OpenAI’s commercial Codex Security product, which scored 9.2%.

  3. Kamile Lukosiute blog — Mythos vs GPT-5.5https://blog.kamilelukosiute.com/p/mythos-vs-gpt-55

    Despite claims of being ‘too dangerous’ to release, Mythos found only one low-severity vulnerability in cURL’s code — Stenberg called the frenzy an ‘amazingly successful marketing stunt.’

  4. LHC Media — Patch the Planet field reporthttps://www.lhc.media/en/media/ai-watch/patch-the-planet-openai-and-trail-of-bits-secure-critical-open-source

    In its first week, the initiative identified hundreds of issues across 19 projects, resulting in 64 pull requests and 37 merged patches for projects including aiohttp, pyca/cryptography, and NATS Server.

  5. CSO Online — dual-use AI cybersecurity analysishttps://www.csoonline.com/article/4113987/the-2-faces-of-ai-how-emerging-models-empower-and-endanger-cybersecurity.html

    Researchers at BeyondTrust and Check Point identified critical command injection vulnerabilities within the Codex environment that could allow attackers to steal GitHub access tokens via malicious repository configurations.

  6. Sentra (security vendor analysis)https://sentra.io/blog/daybreak-answers-the-vulnerability-question-heres-the-one-it-doesnt

    Daybreak answers the vulnerability question. Here’s the one it doesn’t… the blast radius if an exploit lands before a patch is applied.

  7. AI Weekly (citing FT analysis)https://aiweekly.co/alerts/anthropics-safety-rhetoric-may-have-fed-its-own-export-ban

    Anthropic’s 2026 communications used risk-related terminology at a rate of five per 1,000 words — over eight times the frequency of OpenAI’s 0.6 per 1,000.

    2
  8. CyberScoophttps://cyberscoop.com/us-government-anthropic-fable-5-mythos-5-export-controls/

    Amazon CEO Andy Jassy reportedly alerted the Trump administration that researchers had bypassed safeguards using simple prompts like ‘fix this code,’ potentially allowing the model to be used as an unrestricted hacking tool.

  9. BigGo finance (summarizing David Sacks)https://finance.biggo.com/news/b19f67e324afce7c

    Anthropic is running a ‘sophisticated regulatory capture strategy based on fear-mongering’ … part of a broader ‘doomer industrial complex’ that uses safety concerns to entrench a corporate oligopoly.

    2
  10. The Next Web — open letter coveragehttps://thenextweb.com/news/fable-5-ban-cybersecurity-defenders-open-letter

    Over 100 prominent cybersecurity experts, including former Facebook CSO Alex Stamos, signed an open letter calling the ban ‘sabotage’ that strips the best tools from defenders while adversaries continue to develop similar weights abroad.

  11. Tech Policy Presshttps://www.techpolicy.press/did-the-us-government-just-set-an-ai-export-precedent-by-blocking-mythos/

    Legal experts and advocates like the EFF argue that AI model weights and their resulting ‘speech’ are modern extensions of Bernstein v. United States … imposing a license requirement on the release of an AI model functions as a ‘prior restraint’.

  12. daily.dev recap of The Algorithmhttps://daily.dev/posts/three-things-to-watch-amid-anthropic-s-latest-feud-with-the-government-hyk9hplga

    European and global companies, wary of being suddenly disconnected from American AI infrastructure by Washington’s directives, are expected to shift toward Chinese open-source models … not subject to the same kill-switch mechanisms.

  13. SiliconANGLEhttps://siliconangle.com/2026/06/22/google-forms-research-partnership-a24-films-thats-focused-ai-filmmaking-tools/

    The agreement explicitly bars Google from using A24’s film and television library to train its foundational models; A24 Labs is a 20-person unit led by former Adobe CPO Scott Belsky.

    2
  14. Quartzhttps://qz.com/google-a24-investment-deepmind-ai-filmmaking-partnership-062226

    Roughly 2,000 working storyboard artists in Hollywood are directly threatened by the kind of AI pre-visualization tools A24 Labs and DeepMind are building, even as both companies frame the work as ‘augmentation, not replacement.’

  15. World of Reelhttps://www.worldofreel.com/blog/2026/6/22/a24-partners-with-google-in-75m-ai-deal-for-filmmaking-and-distribution-techniques

    A24 reportedly passed on Luca Guadagnino’s anti-OpenAI biopic ‘Artificial’ roughly 24 hours before unveiling the $75M Google deal — and Thrive Capital is a major investor in both A24 and OpenAI.

  16. Gamereactor (quoting Kane Parsons via The Australian)https://www.gamereactor.eu/ai-is-a-symptom-of-broader-cultural-and-economic-rot-says-backrooms-director-1728913/

    AI is a symptom of broader cultural and economic rot — Backrooms director Kane Parsons says he would ‘make it disappear’ if he could.

  17. r/Filmmakers thread on Lionsgate–Runwayhttps://www.reddit.com/r/Filmmakers/comments/1npdgm1/lionsgate_is_struggling_to_make_aigenerative/

    Lionsgate is reportedly struggling to make AI-generative filmmaking work with Runway, with the partnership hitting ‘unproductive’ snags around character consistency and human physics.

  18. Immaculatan (Immaculata University)https://immaculatan.immaculata.edu/2024-editions/ai-in-the-film-industry/

    A24 already faced ‘swift and stark backlash’ in 2024 over AI-generated ‘Civil War’ posters that contained obvious geographic errors and bypassed working illustrators.

Jack Sun

Jack Sun, writing.

Engineer · Bay Area

Hands-on with agentic AI all day — building frameworks, reading what industry ships, occasionally writing them down.

Digest
All · AI Tech · AI Research · AI News
Writing
Essays
Elsewhere
Subscribe
All · AI Tech · AI Research · AI News · Essays

© 2026 Wei (Jack) Sun · jacksunwei.me Built on Astro · hosted on Cloudflare