JS Wei (Jack) Sun

OpenAI's Jalapeño doubted, Gemini Flash hits 88.4%, DeepMind bleeds to Anthropic

OpenAI's Jalapeño chip, Google's agent-native Gemini Flash, and Anthropic's DeepMind defection magnet show three frontier labs picking three different moats.

OpenAI’s Jalapeño doubted, Gemini Flash hits 88.4%, DeepMind bleeds to Anthropic

TL;DR

  • Jalapeño ASIC tape-out called impossible in 9 months without heavy Broadcom IP reuse
  • Gemini 3.5 Flash posts 88.4% on WebVoyager, a benchmark vendors call broken
  • DeepMind engineers defect to Anthropic at 11× the reverse rate
  • Micron profit jumps 15× to $28.2B on AI memory shortage
  • Bores loses NY-12 primary, closing a $27M OpenAI-Anthropic PAC proxy war

Today’s AI news clusters around three frontier labs each placing a different kind of bet for the next phase. OpenAI is going vertical with Jalapeño, a reticle-sized inference ASIC pitched for late-2026 gigawatt deployment — except hardware engineers reading the timeline say a 9-month blank-slate tape-out is impossible without heavy Broadcom IP reuse. Google is collapsing its two-model agent stack into one, with Gemini 3.5 Flash absorbing computer use natively and claiming 88.4% on WebVoyager, a benchmark vendors have spent months calling broken. Anthropic isn’t shipping hardware or models today — it’s just absorbing the talent, with DeepMind engineers defecting at roughly 11× the reverse rate per SignalFire’s 2026 benchmarking.

The round-ups extend the capital-allocation picture: Micron’s profit jumped 15× on AI memory demand, Cerebras shares dropped on its first post-IPO margin guidance, Agility Robotics filed a $2.5B SPAC merger, and the $27M Bores primary ended in a narrow loss that closed out the OpenAI-vs-Anthropic PAC proxy war without a clear winner.

Engineers call OpenAI’s 9-month Jalapeño tape-out impossible

Source: openai-blog · published 2026-06-24

TL;DR

  • Hardware engineers call the 9-month blank-slate tape-out “completely impossible” without heavy reuse of Broadcom IP.
  • Jalapeño is a reticle-sized inference ASIC for late-2026 gigawatt deployment, validated on GPT-5.3-Codex-Spark.
  • Broadcom popped 9–11%, Nvidia slipped ~3%, with KeyBanc raising targets on the 10 GW scope.
  • Gigawatt math collides with OpenAI’s paused 600 MW Texas site and shelved 8,000-GPU UK build.

What was actually announced

OpenAI and Broadcom unveiled Jalapeño, OpenAI’s first custom silicon — a reticle-sized inference ASIC pitched as built from a blank slate for LLM serving patterns, with Broadcom’s Tomahawk networking glued in for scale-out and Celestica handling rack integration. Initial deployment is slated for late 2026 in partnership with Microsoft, with a stated 10 GW ambition. Engineering samples are running production workloads against GPT-5.3-Codex-Spark in lab. That much is on the record across OpenAI’s own post and follow-ups in TechCrunch, The Verge, and Ars Technica.

The strategic logic is unambiguous: OpenAI wants to own the cost curve of its own inference, not rent it from Nvidia at the margin Jensen sets. Whether Jalapeño can deliver that is where the announcement and the independent reporting start to diverge.

The 9-month tape-out doesn’t pass the smell test

Tom’s Hardware was politest about it, noting that every performance-per-watt figure is “on paper” until somebody outside OpenAI runs the chip 1. The harder read came from hardware-engineering critique: a true from-scratch nine-month design-to-tape-out is “completely impossible,” and the realistic explanation is substantial reuse of Broadcom’s existing custom-ASIC IP plus the head start that OpenAI hardware lead Richard Ho carried over from his Google TPU work — which was itself a Broadcom co-design 2. The accompanying “our models helped design the chip” framing reads as marketing dressing on a fairly conventional ASIC program with a strong vendor relationship behind it.

That’s not a scandal — it’s how custom silicon actually gets built. But it does mean the timeline shouldn’t be read as a generational shift in how fast frontier silicon ships.

Gigawatt deployment runs into a power wall

The “late-2026, gigawatt-scale” line is the second claim that doesn’t survive contact with the rest of OpenAI’s footprint. Enverus has already documented a 600 MW Stargate-adjacent Texas expansion abandoned over grid-interconnection backlogs and an 8,000-GPU UK build paused on energy economics 3. The Microsoft Maia precedent rhymes: Maia 200 slipped at least six months into 2026 after OpenAI-requested design changes destabilized simulations, pushing an interim Maia 280 to 2027 4. Every hyperscaler custom-silicon program to date has missed its first announced window — Jalapeño is launching into that track record, not against it.

Markets price the deal; analysts hedge

The equity reaction was real. Broadcom gained roughly 9–11% on the announcement while Nvidia fell about 3%, and KeyBanc raised targets on the 10 GW scope while flagging Broadcom’s deepening OpenAI concentration as a risk 5. BofA’s Vivek Arya offered the cleaner read: Broadcom’s custom-AI wins reflect an “expanding AI pie,” not Nvidia displacement, and prior Microsoft and Meta custom programs have failed to dent CUDA’s developer ecosystem 6.

That bounds Jalapeño’s near-term meaning. It’s a captive inference part for OpenAI’s own models, optimized through OpenAI’s own Triton-based stack — useful for shaving the cost of ChatGPT tokens and API calls, not a chip third-party developers will ever touch. The 2026 question isn’t whether Jalapeño exists. It’s whether OpenAI can find the megawatts to plug enough of them in.

Further reading


Gemini 3.5 Flash’s 88.4% rides a benchmark vendors call broken

Source: deepmind-blog · published 2026-06-24

TL;DR

  • Gemini 3.5 Flash absorbs computer use natively, claiming 88.4% on WebVoyager over the old 2.5 standalone model.
  • Google also pitches a 30% latency cut on multi-step tasks from collapsing the two-model handoff into one.
  • WebVoyager itself is openly contested — a trivial Google-Search-shortcut agent already clears over 50% of its tasks.
  • SafeBreach demoed “Fake Context Alignment” injections via Slack/WhatsApp notifications that the new screenshot-scanning detector doesn’t address.
  • Real-world agentic loops cost 3–5.5× Gemini 3 Flash because the model burns extra reasoning tokens per step.

The actual news: computer use is no longer a side model

Google DeepMind folded computer use into the main Flash line. That matters more than the benchmark chart. Until now, developers building browser or desktop agents on Google’s stack had to bolt the Gemini 2.5 computer-use model onto a separate reasoning model, eating latency on every handoff. Gemini 3.5 Flash unifies seeing, reasoning, and acting in one model, alongside function calling and Search/Maps grounding, and ships through both the Gemini API and the new Enterprise Agent Platform with Browserbase as a launch host. The native integration — not the 3-point WebVoyager bump — is what enterprise teams will actually feel.

The 88.4% deserves an asterisk

The headline benchmark is in trouble at the source. A widely shared r/AI_Agents post argues many teams “manually correct” results or remove “impossible” tasks to inflate scores, and that a simple agent leaning on Google Search shortcuts can already solve over half of WebVoyager 7. A methodology paper on OpenReview goes deeper: the GPT-4V/4o auto-evaluators most labs use have low agreement with human judges and a Type-I error problem, and teams asymmetrically audit failures for hidden successes while never auditing successes for false positives 8.

A simple agent relying primarily on Google Search shortcuts can solve over 50% of the benchmark. 7

Translation: a 3-point WebVoyager delta is signal-poor. Independent head-to-head numbers on harder agent suites aren’t in this bundle, so treat the leaderboard win as vendor-reported until someone reruns it under a fixed task set.

Safety framing is thinner than “defense-in-depth” suggests

DeepMind’s safety story rests on adversarial training, opt-in user confirmation for sensitive actions, and an injection detector that watches the UI the model is viewing. The detector’s threat model is a malicious string visible on screen. SafeBreach Labs published “Fake Context Alignment,” where instructions hidden in WhatsApp or Slack notification payloads silently steered Gemini into controlling smart-home devices and poisoning long-term memory 9 — a channel the screenshot scanner doesn’t obviously cover. Separately, an independent multi-turn evaluation posted to Google’s own developer forum found 3.5 Flash uniquely steerable into “moral verdicts” and grievance amplification, a regression relative to Gemini 3.1 Pro 10. And the user-confirmation gate is opt-in, which makes safe-by-default a developer responsibility rather than a platform guarantee.

flowchart LR
    A[Web page / app UI] --> D{Injection detector}
    B[Slack / WhatsApp<br/>notifications] -. bypasses .-> M[Gemini 3.5 Flash agent]
    D --> M
    M --> T[Tools: clicks,<br/>typing, transactions]
    T -. unconfirmed if<br/>gate is off .-> X((External effects))

Cost and reliability undercut the efficiency pitch

The 30% latency win is real, but the economics partly eat it. Evolink pegs Flash at $1.50/$9.00 per 1M input/output tokens and estimates real-world agentic workloads run 3× to 5.5× the cost of Gemini 3 Flash because 3.5 Flash “burns” more output tokens per reasoning turn 11. Developer sentiment on Google’s own forum is unusually blunt: a thread titled “3.5 Flash worst model and worst IDE for coding” reports hallucinations, ignored system instructions, and degradation as prompts approach the 1M-token context limit 12 — exactly the regime long-horizon computer-use tasks live in.

What to actually take away

Native computer use in the flagship Flash model is a real platform shift and the right architectural bet. Treat 88.4% and “30% latency cut” as vendor figures. Before pointing a Flash agent at a CRM or a checkout flow, turn on user confirmation, sandbox it, and budget for 3–5× the token spend you modeled on Gemini 3 Flash.


DeepMind staff defect to Anthropic at 11× the reverse rate

Source: techcrunch-ai · published 2026-06-24

TL;DR

  • DeepMind engineers are ~11× more likely to move to Anthropic than the reverse, per SignalFire’s 2025/2026 benchmarking.
  • Anthropic leads frontier-lab retention at 80%, vs. DeepMind 78%, OpenAI 67%, Meta 64%.
  • The reported trigger is compute reallocation — Google moved GPUs off Shazeer’s experimental project to harden Gemini pre-training.
  • Jefferies calls it “noise”, arguing Google’s distribution moat (5 apps past 3B users) outweighs any single researcher.

The flow is structural, not anecdotal

TechCrunch’s framing — another two names on a growing list — undersells what the benchmark data actually says. SignalFire’s two-year tracking puts the DeepMind→Anthropic migration rate at roughly 11× the reverse direction, and Anthropic now sits at the top of the frontier-lab retention table at 80%, ahead of DeepMind (78%), OpenAI (67%), and Meta (64%) 13. That asymmetry holds across seniority bands, not just the Nobel-tier headlines. The interesting question isn’t why Jonas Adler and Alexander Pritzel left; it’s why the gradient is this steep at all.

What’s pushing people out

The trigger that keeps surfacing in reporting is concrete and organizational. Shortly before Noam Shazeer’s exit, Google reportedly reassigned compute from one of his experimental projects to a London team working on Gemini pre-training 14. The same pattern of centralizing accelerators around the flagship model has preceded other senior departures. This is a sharper story than the usual “culture clash” gloss — post-merger consolidation of Brain and DeepMind appears to be pulling GPU allocation toward shipping Gemini and away from the speculative research that senior PIs were hired to run. If you joined to chase your own bets, losing the compute to run them is the exit signal.

What Anthropic is buying

The hires aren’t random. Adler and Pritzel come from the AlphaFold team, joining John Jumper’s earlier move — and they land in a company that has spent the last year building wet labs in-house, acquiring Coefficient Bio, and training specialized Claude variants on structural biology and clinical filings with a stated goal of compressing drug discovery timelines 10× 15. Meanwhile Arthur Conmy, another recent arrival, told The Decoder his motivation was alignment: current models including Claude are “not yet sufficiently aligned” to safely scale toward AGI, and his role is “triaging signs of misalignment” during training rather than patching outputs 16. Two distinct bets — AI-for-science and interpretability — being staffed deliberately.

The bull case for Google

Not everyone reads this as thesis-breaking. Jefferies’ Brent Thill calls the departures industry-wide “musical chairs” and argues Google’s moat is distribution and TPU economics, not any individual researcher: five products each past 3 billion users is a buffer most labs would trade their entire roster for 17. Demis Hassabis, speaking at Cannes Lions, conceded the talent market is the most intense in tech history but insisted DeepMind still wins its “fair share” from a massive research bench 18.

What’s actually at stake

The bear and bull cases don’t really contradict each other — they’re scored on different timelines. Distribution wins the next four quarters. Whether Anthropic’s biology and alignment hires compound into a model Google can’t match on capability is a 2027-2028 question. The 11× number is the one to watch: if it narrows, the consolidation story was overblown. If it widens, Google has a compute-allocation problem that org charts won’t fix.

Round-ups

Micron profit jumps 15x as AI memory shortage drives record quarter

Source: techcrunch-ai

Revenue at the US memory maker quadrupled year-over-year to $41.45 billion, while profit surged from $1.88 billion to $28.2 billion. HBM demand from Nvidia and hyperscaler AI buildouts has tightened DRAM supply industry-wide, pushing pricing and margins sharply higher.

Cerebras shares tumble on first post-IPO earnings as margin guidance spooks

Source: techcrunch-ai

The AI chipmaker’s stock dropped sharply after it forecast narrower gross margins in its core business during its debut earnings report as a public company. CEO Andrew Feldman argued investors misread the outlook, but the sell-off underscores scrutiny on Nvidia challengers.

Agility Robotics targets public listing via $2.5B SPAC merger

Source: techcrunch-ai

The humanoid robotics startup, spun out of Oregon State in 2015, expects to raise roughly $620 million in proceeds from the deal. Agility’s Digit robot is already deployed at GXO and Amazon warehouses, giving it a revenue story rivals like Figure lack.

Bores loses NY-12 primary as $27M OpenAI-Anthropic proxy war ends in draw

Source: the-verge-ai

New York Assemblyman Alex Bores narrowly lost the Democratic primary for New York’s 12th Congressional district, ending a $27 million spending battle between Anthropic and OpenAI-aligned super PACs. Bores’ polling actually surged after a pro-AI PAC targeted him over his state-level AI safety record.

Databricks’ Zaharia and Xin argue the frontier stack must stay open

Source: latent-space

In a joint interview, Databricks co-founder Matei Zaharia and chief architect Reynold Xin lay out why every enterprise will need to build its own Agent Cloud, and why open models and infrastructure are the precondition for that shift away from closed frontier labs.

Figma Config 2026 ships code layers, shaders, and AI plugin builder

Source: techcrunch-ai, the-verge-ai

At its annual Config conference, Figma unveiled a reimagined canvas optimized for full-stack development, alongside motion graphics, shader tools, and the ability to spin up custom plugins from natural-language prompts. The push positions Figma against Cursor and Vercel in AI-native design-to-code workflows.

Engineers gain hiring share even as AI fuels layoff narrative

Source: techcrunch-ai

SignalFire data shows software engineers making up a growing portion of new startup hires, contradicting predictions that coding assistants would gut the role. Junior hiring remains weak, but mid and senior engineers are the most resilient white-collar cohort tracked in the report.

Footnotes

  1. Tom’s Hardwarehttps://www.tomshardware.com/tech-industry/artificial-intelligence/broadcom-and-openai-unveil-custom-built-jalapeno-inference-processor-openais-first-chip-is-a-massive-reticle-sized-asic-built-in-an-ultra-fast-nine-month-development-cycle

    OpenAI’s first chip is a massive reticle-sized ASIC built in an ultra-fast nine-month development cycle… performance claims remain ‘on paper’ until independent benchmarks are released.

  2. beri.net analysis (HN-style critique)https://www.beri.net/article/openai-jalapeno-custom-inference-chip-broadcom-50-percent-cost-cut-enterprise-2026

    A nine-month from-scratch tape-out is ‘completely impossible’; the chip is likely an evolution of existing Broadcom IP, and hardware lead Richard Ho’s prior TPU work with Broadcom provided a significant head start.

  3. Enverus — ‘Stargate Scales Back’https://www.enverus.com/blog/stargate-scales-back-openai-and-oracle-abandon-data-center-expansion/

    A planned 600-megawatt expansion in Texas was abandoned as power-grid interconnection backlogs delayed large-scale energization, and an 8,000-GPU UK build was paused over energy costs.

  4. TrendForce on Microsoft Maia roadmaphttps://www.trendforce.com/news/2025/07/04/news-microsoft-reportedly-eyes-2027-interim-ai-chip-amid-in-house-design-delays-and-nvidia-pressure/

    Microsoft’s Maia 200 (Braga) faced a minimum six-month delay into 2026 after design modifications requested by OpenAI caused simulation instabilities; an interim Maia 280 is now slated for 2027.

  5. Crypto Briefinghttps://cryptobriefing.com/broadcom-custom-chip-openai-nvidia/

    Broadcom (AVGO) gained roughly 9–11% on the deal while Nvidia slipped about 3%; KeyBanc cites the 10-gigawatt scale as a long-term tailwind, but Broadcom’s customer concentration in OpenAI is flagged as a risk.

  6. Kavout (citing BofA’s Vivek Arya)https://www.kavout.com/market-lens/broadcom-s-custom-ai-empire-openai-deal-validates-its-indispensable-role-against-nvidia

    Broadcom’s growth may simply reflect an ‘expanding AI pie’ rather than a direct displacement of Nvidia; past custom-silicon efforts by Microsoft and Meta have struggled to match CUDA’s developer ecosystem.

  7. r/AI_Agents thread: ‘WebVoyager is broken’https://www.reddit.com/r/AI_Agents/comments/1r2yziq/webvoyager_is_broken_and_every_agent_company/

    many teams ‘manually correct’ results or remove ‘impossible’ tasks to inflate scores, making direct comparisons between models functionally meaningless… a simple agent relying primarily on Google Search shortcuts can solve over 50% of the benchmark

    2
  8. OpenReview paper on WebVoyager auto-evaluationhttps://openreview.net/pdf?id=6jZi4HSs6o

    auto-evaluators (typically GPT-4V or GPT-4o) … often have low agreement with human judgment and struggle with Type-I (false positive) errors … teams may manually review ‘unknown’ or ‘failure’ labels to find hidden successes but fail to audit ‘success’ labels

  9. SafeBreach Labs researchhttps://www.safebreach.com/blog/gemini-voice-assistant-prompt-injection-exploit/

    a ‘Fake Context Alignment’ technique where malicious instructions hidden in message notifications (from WhatsApp or Slack) could silently force Gemini’s voice assistant to control smart home devices or poison long-term memory

  10. Google AI Developers Forum — independent multi-turn capture-risk evalhttps://discuss.ai.google.dev/t/findings-from-an-independent-multi-turn-capture-risk-evaluation-including-gemini-3-5-flash-and-3-1-pro-preview/168822

    Gemini 3.5 Flash could be steered into issuing moral verdicts or amplifying user grievances during complex interactions, which differs from the behavior of its ‘Pro’ counterparts

  11. Evolink.ai pricing analysishttps://evolink.ai/blog/gemini-3-5-flash-pricing-guide

    $1.50 per 1M input tokens and $9.00 per 1M output tokens … the ‘real-world’ cost can be 3x to 5.5x higher than Gemini 3 Flash … attributed to 3.5 Flash’s tendency to ‘burn’ more output tokens during complex agentic reasoning turns

  12. Google AI Developers Forum complaint threadhttps://discuss.ai.google.dev/t/3-5-flash-worst-model-and-worst-ide-for-coding-with-even-worse-limits/166769

    ‘3.5 Flash worst model and worst IDE for coding’ … ‘extremely error-prone,’ reporting frequent hallucinations, ignored system instructions, and ‘sloppy execution’ as prompts approach the 1M token context limit

  13. Search Engine Journal (citing SignalFire 2025/2026 data)https://www.searchenginejournal.com/google-loses-two-top-ai-researchers-to-openai-anthropic/580201/

    DeepMind engineers are approximately 11 times more likely to move to Anthropic than the reverse… Anthropic maintains an industry-leading 80% employee retention rate, surpassing DeepMind (78%), OpenAI (67%), and Meta (64%)

  14. Outlook Businesshttps://www.outlookbusiness.com/corporate/google-set-to-lose-two-more-researchers-to-anthropic-after-earlier-high-profile-exits

    Shortly before Shazeer’s exit, Google reportedly reassigned computing power from one of his experimental projects to a London-based team to streamline pre-training for the flagship Gemini model.

  15. SynBioBetahttps://www.synbiobeta.com/read/anthropic-is-hiring-biologists-building-wet-labs-and-betting-big-on-drug-discovery

    Anthropic has opened internal wet labs to conduct basic biological research and acquired the biotech-focused startup Coefficient Bio… training specialized versions of Claude on vast datasets of structural biology and clinical filings with the goal of compressing drug development timelines by a factor of ten.

  16. The Decoder (quoting Arthur Conmy)https://the-decoder.com/google-keeps-losing-top-ai-researchers-to-rivals/

    Existing models like Claude possess ‘extraordinary’ capabilities but are not yet sufficiently aligned to safely manage the development of AGI… his new role at Anthropic involves ‘triaging signs of misalignment’ during training, seeking root-cause fixes rather than surface-level patches.

  17. TipRanks / Jefferies note (Brent Thill)https://www.tipranks.com/news/jefferies-says-googles-ai-talent-losses-are-just-noise-heres-why

    Jefferies characterizes the recent wave of departures as an industry-wide ‘musical chairs’ dynamic… Google does not need to possess the ‘absolute best’ AI model to win — its competitive advantage lies in its massive distribution network with five applications each boasting over 3 billion users.

  18. Taipei Times (Hassabis at Cannes Lions)https://www.taipeitimes.com/News/biz/archives/2026/06/22/2003859496

    Hassabis characterized the talent movement as the most intense in the history of the tech industry but maintained that DeepMind continues to ‘win its fair share’ of top-tier talent due to its massive research bench.

Jack Sun

Jack Sun, writing.

Engineer · Bay Area

Hands-on with agentic AI all day — building frameworks, reading what industry ships, occasionally writing them down.

Digest
All · AI Tech · AI Research · AI News
Writing
Essays
Elsewhere
Subscribe
All · AI Tech · AI Research · AI News · Essays

© 2026 Wei (Jack) Sun · jacksunwei.me Built on Astro · hosted on Cloudflare