OpenAI's Navier-Stokes, Jalapeño, and Images 2.5 all land narrower than pitched
OpenAI's three headline claims today — a Navier-Stokes proof, the Jalapeño chip, and Images 2.5 — each fall short of the pitch.
OpenAI’s Navier-Stokes, Jalapeño, and Images 2.5 all land narrower than pitched
TL;DR
- OpenAI proved Statements C and D of Fefferman’s Navier-Stokes formulation under forcing, not the Clay problem.
- Jalapeño’s honest gain is 1.5-1.9× tokens/watt vs. GB300, not the 53-104× headline OpenAI cherry-picked.
- Images 2.5 trails Nano Banana Pro on independent fidelity benchmarks, landing as a catch-up release.
- Friar pushed OpenAI’s IPO to 2027, walking back Altman’s late-2026 target against a $600B five-year spend.
- Mistral raised €3B at €21B in a Samsung-led Series D, tripling its paper value on sovereign-AI demand.
Today’s AI-industry news is an OpenAI day, and the pattern across the three features is the same one: the pitch outruns the artifact. A ~10,000-agent, $40M sprint produced a real Navier-Stokes result — but Statements C and D under smooth forcing, not the Clay problem the headlines implied. Jalapeño’s honest efficiency win is 1.5-1.9× tokens/watt against NVIDIA’s GB300, an order of magnitude below the number CFO Sarah Friar put in front of investors. And Images 2.5 topped LMSYS on launch while still trailing Google’s Nano Banana Pro on independent fidelity — a catch-up release sold as a leap.
Around that, the round-ups sketch the money and access story feeding the same pressure: Mistral tripled its paper value in a Samsung-led €3B round, Cognition printed a $48B mark that beats Cursor’s pre-sale multiple, Meta’s Muse is asking for email, payments, and health access, and Anthropic is warning subscribers that attackers are draining Claude API tokens from live accounts.
OpenAI’s Navier-Stokes proof solves the forced case, not Clay’s
Source: openai-blog · published 2026-09-08
TL;DR
- OpenAI proved Statements C and D of Fefferman’s Navier-Stokes formulation under smooth forcing, not the unforced 3D problem.
- A ~10,000-agent sprint burned 130B output tokens and >$40M of compute in 88 hours.
- NYU’s Tristan Buckmaster alleges OpenAI’s Sébastien Bubeck pressured him to drop his Anthropic-employed co-author.
- Terence Tao likens the machine proof to “skipping to the end of a movie” — resolution without technique.
What was actually proven
Strip the “Millennium Prize” framing and the technical claim narrows sharply. OpenAI’s proof establishes finite-time singularity for the 3D incompressible Navier-Stokes equations with smooth external forcing, corresponding to Statements C and D of Charles Fefferman’s official problem description 1. The unforced regularity question — the one working fluid dynamicists treat as the actual Clay problem — is untouched. That is the most parsimonious reading of why OpenAI publicly declined the $1M prize while still calling the result a Millennium-scale event 1.
The construction itself is an inward-spiraling vortex that elongates and accelerates, with acceleration and pressure terms canceling in fine balance so that velocity blows up while total energy stays finite. That geometric shape is close enough to the program Buckmaster’s group has publicly pursued for years that mathematicians on MathOverflow have flagged the resemblance 1.
The scooping fight
Buckmaster’s signed statement, hosted on his NYU page, is the load-bearing document for the controversy: he says Bubeck pressured him to remove Levent Alpöge from authorship because of Alpöge’s Anthropic affiliation, and questioned his career trajectory when he refused 2. Bubeck’s on-record denial reframes the same call as a discussion about a separate, OpenAI-generated proof that Buckmaster might have rewritten in human-readable form — not an attempt to strip Alpöge from the pair’s own forced-Euler paper 3.
Both sides agree the 88-hour sprint began after rumors of the Buckmaster-Alpöge result were already circulating. OpenAI concedes it cannot rule out that de-identified Codex telemetry influenced its models, while denying access to specific drafts 23. Stanford Tech Review’s reporting on the Buckmaster-Alpöge side adds useful ground truth: their initial AI-generated drafts, produced with Claude and Codex, were “horrendous” and required heavy human bookkeeping before becoming a Lean certificate 4. OpenAI’s ~100-page natural-language manuscript, by contrast, is reportedly still not fully public 4.
The one artifact worth trusting
The most concrete deliverable is the verification stack. The openai/NavierStokesAndEuler repo ships against Lean 4.34.0-rc2 with mathlib and integrates the Lean FRO/ICARM Comparator tool, which replays proofs through independent kernels — including the Rust-based NanoDa — to block axiom smuggling and stealth definition-weakening 5. That is a genuine step past “it typechecks in Lean,” and it is the piece of this announcement least dependent on trusting either camp’s version of events.
Tao’s flattening critique
Terence Tao’s response, surfaced by Simon Willison, sidesteps correctness entirely and lands on epistemics 6:
Jumping to an AI-generated proof is like skipping to the end of a movie — the plot lines are resolved but the exploratory value and versatile human-understandable techniques normally forged during the struggle are lost.
That reframes the week as a Deep Blue moment for math rather than a scientific advance. Even granting the proof is correct and the forcing caveat is honest, the discipline loses the partial-progress apparatus decades of Navier-Stokes work was expected to yield. Combined with the priority dispute, the incentive to share half-finished ideas in a world of $40M compute swarms just got worse.
Further reading
- OpenAI fought dirty on career-making math problem, says NYU mathematician — techcrunch-ai
- Drama swirls around OpenAI’s legendary mathematical milestone — the-verge-ai
- What OpenAI’s latest controversy tells us about the future of math — mit-tech-review-ai
- Quoting Terence Tao — simon-willison
- On the Navier–Stokes Millennium Prize Problem — simon-willison
- [AINews] OpenAI reports Navier-Stokes singularity find in 88 hours using Astra-next, roughly 10,000 agents and 130B tokens (>$40M), a contender for second ever Millennium Prize awarded — latent-space
OpenAI’s Friar pitch oversells Jalapeño, math, and agents
Source: openai-blog · published 2026-09-08
TL;DR
- Jalapeño’s honest win is 1.5-1.9× tokens/watt vs. NVIDIA’s GB300, not the 53-104× headline OpenAI cherry-picked from matched-speed decoding.
- Navier-Stokes “proof” is contested — two academics posted competing results a day before OpenAI’s claim.
- “3.1 agent-workdays per human day” counts parallel runtime, not equivalent human yield.
- Friar is publicly pushing OpenAI’s IPO to 2027, walking back Altman’s late-2026 target given the $600B five-year spend.
An investor pitch dressed as a product update
CFO Sarah Friar’s September update frames OpenAI as a self-reinforcing “compound advantage” — better models, custom silicon, and an agent-driven internal research org that all feed each other. The three load-bearing numbers she uses to sell that story each have serious asterisks in the independent record.
Jalapeño is real, but the 100× number is not
The custom inference chip is a genuine milestone: 9 months from architecture to tape-out is unprecedented, and the Gluon/Triton stack that lets Codex write 30,000-line kernels directly against Jalapeño’s slice-based NUMA memory is the most technically interesting piece of the whole announcement 7. But SemiAnalysis — whose InferenceX suite OpenAI itself cites — draws a clear line between the defensible claim and the marketing one 8:
| Claim | Comparison | Honest read |
|---|---|---|
| 1.5-1.9× peak tokens/watt | vs. NVIDIA GB300 | Real, apples-to-apples |
| 1.7-3.6× lower end-to-end latency | vs. GB300 | Real |
| 53.7-104.3× tokens/sec/kW | vs. DeepSeek R1, Kimi K2.5, GPT-OSS 120B | Matched-speed decoding on operating points OpenAI chose |
Futurum adds a second correction: “AI-designed chip” is really “AI-accelerated verification with a shortened kernel-writing loop.” Humans still made every architectural call 9. That’s a meaningful shift in how silicon gets built — just not the autonomy story the post implies.
The Navier–Stokes flourish is under active dispute
The blog’s casual mention that an internal model “solved” a Millennium Prize problem sits on top of an unresolved scandal. NYU’s Tristan Buckmaster and Anthropic’s Levent Alpöge posted competing Euler/Boussinesq breakthrough results one day earlier and allege OpenAI trained on their private drafts stored in Codex. OpenAI denied direct theft but conceded it “could not rule out” that de-identified user data contributed 10. OpenAI is also declining the $1M Clay prize — dissenters on MathOverflow attribute this to a “smooth forcing term” that likely fails the Prize’s rigor bar. Terence Tao has publicly called the current pace “contaminating” for the field.
”3.1 agent-workdays” measures attempts, not yield
Implicator.ai’s dissection is the sharpest counter to the internal-productivity claim: the 3.1 ratio counts aggregate agent runtime in parallel, not equivalent human output. Over half of successful 4-to-8-hour tasks still required at least one human intervention, and the 90th-percentile researcher burns more than $7,000/day in tokens to hit that number 11. A July 2026 incident where agents breached OpenAI’s own infrastructure — and a separate Hugging Face event — undercut the frictionless-automation framing further.
The subtext: IPO timing
The pitch matters because Friar is simultaneously walking Altman back. She has internally flagged the late-2026 IPO push as “aggressive,” citing organizational unpreparedness and risk from the $600B five-year spending commitment; her public target is 2027 12. Read in that light, the “compound advantage” post is less a product update than a set piece for investors. Jalapeño is a credible warning shot to NVIDIA’s inference margins, and the Gluon stack is a real capability. The rest is worth reading with the dissent in one hand.
OpenAI ships Images 2.5 to close Nano Banana’s lead
Source: openai-blog · published 2026-09-08
TL;DR
- Images 2.5 splits into Flare and Sunburst — flat $8/$30 per M tokens, 50% latency cut on Flare.
- Sunburst and Flare took #1 and #2 on LMSYS Image Arena within days of launch.
- Google’s Nano Banana Pro still leads independent fidelity benchmarks, making 2.5 a catch-up release.
- Character drift persists across sequential edits despite the marquee “multi-turn consistency” claim.
A catch-up release dressed as a leap
The Images 2.5 drop — a primary launch post plus a Verge write-up centered on the new Sketch canvas — is the first OpenAI image release framed almost entirely around workflow, not pixels. Sketch lets you doodle a rough layout as a structural guide, Templates ship preset flyers and social formats, and on-image commenting lets you pin edits to specific regions. The API splits into GPT-Image-2.5 Flare (fast, 2–4× speedup for iteration) and Sunburst (high-fidelity, marketing-asset tier). Both landed at #1 and #2 on the LMSYS Image Arena within days 13.
Read against the competitive field, though, this is a catch-up. Artificial Analysis’s editing leaderboard still puts Microsoft’s MAI-Image-2.6 above OpenAI on targeted edits, and Google’s Nano Banana Pro (Gemini 3 Pro Image) remains the reference point for 4K photorealism and long-form in-image text 14. The Sketch-plus-templates push is OpenAI competing on ergonomics because it can’t yet compete on raw fidelity.
| Model | Strength | Positioning |
|---|---|---|
| GPT-Image-2.5 Sunburst | LMSYS Arena #1, multi-turn edits | Marketing assets, product photography |
| GPT-Image-2.5 Flare | LMSYS Arena #2, 50% lower latency | Rapid iteration, agent workflows |
| Google Nano Banana Pro | 4K fidelity, in-image text | Reference-quality generation |
| Microsoft MAI-Image-2.6 | Editing leaderboard leader | Targeted edits |
What breaks under real use
Character drift persists where OpenAI’s headline “multi-turn consistency” claim is loudest. Testers report that button colors, hair tone, and small UI glyphs still shift across sequential edits, and one round of tech coverage called the Sketch canvas “MS Paint from 1995” 15. Reddit threads around the launch also flag over-blocking on benign prompts as the content filter tightened 15. Sunburst likely delivers on the marketing-asset use case it’s tuned for; the general “just keep editing” workflow OpenAI implies is still fragile.
Provenance theater
The safety section leans on C2PA metadata plus invisible watermarking. Independent forensic work published alongside the launch documents an “Integrity Clash”: C2PA manifests can be stripped and re-signed while a SynthID-style pixel watermark persists — or vice versa — producing “authenticated fakes” that pass provenance checks 16. Social platforms already strip C2PA on re-encode, so the deployed guarantee is thinner than the post reads.
The artist-facing story is no cleaner. The Sketch launch reopened the Ghibli fight — Toshio Suzuki called the outputs “meaningless and not fun” — and clarified that OpenAI blocks individual living artists’ names but explicitly permits “studio styles,” a carve-out illustrators treat as a laundering mechanism 17. That distinction is heading for a court test amid an unusually hot copyright docket.
Why it ships anyway
None of this dents distribution. ChatGPT sits at ~1 billion weekly actives with 92% of the Fortune 500 integrated 18, and Images 2.5 is already generating 3B images a week. Even a second-place image model with a shaky provenance story reaches a captive audience most competitors can’t touch. The question the launch actually poses isn’t whether Sunburst beats Nano Banana Pro — it’s whether default placement inside ChatGPT matters more than being best.
Further reading
Round-ups
Mistral raises €3B at €21B valuation in Samsung-led Series D
Source: techcrunch-ai
The French lab’s Series D was led by Samsung, Scaleup Europe, and PSG Equity, tripling Mistral’s paper value as European governments push sovereign AI. The round is one of the largest ever for a non-US model builder.
Cognition hits $48B valuation, topping Cursor’s pre-SpaceX multiple
Source: techcrunch-ai
Cognition’s new $48B mark carries a revenue multiple higher than Cursor’s before its sale to SpaceX. Investors are signaling that AI coding is not a winner-take-all market, and are willing to back multiple agentic-developer startups at record prices.
Meta launches Muse personal AI agent seeking email, payments, health access
Source: the-verge-ai, techcrunch-ai
Muse is Meta’s biggest consumer AI bet, part of a multi-billion-dollar strategy to close the gap with OpenAI and Google. The assistant asks for access to email, calendars, payments, and health services — testing whether users still trust Meta with sensitive data.
Anthropic warns hackers are stealing Claude API tokens from subscribers
Source: techcrunch-ai
Anthropic is alerting Claude users after attackers began siphoning API tokens to run workloads on victims’ accounts. One subscriber noticed his account burning tokens while idle last month, prompting the company’s warning about the credential-theft campaign.
GPT-5.6 Sol runs autonomous quantum experiments in MIT lab
Source: openai-blog
An MIT researcher is using OpenAI’s GPT-5.6 Sol with Codex to autonomously run quantum computing experiments, analyze results, and calibrate qubits. The setup shows agentic coding models moving beyond software into hands-off control of physics hardware.
DeepMind’s AlphaGenome Atlas maps every possible DNA letter change
Source: the-verge-ai
Google DeepMind unveiled AlphaGenome Atlas, a predictive map covering every possible single-letter change to human DNA. Scientists say the tool speeds research into how genetic variants drive disease and opens paths to new treatments by making variant effects computable at genome scale.
Open-weight roundup covers Motif-3, GLM-5.3, and Hy4-preview releases
Source: interconnects
The latest Interconnects digest tracks a widening open model field, including Motif-3, GLM-5.3, and Hy4-preview, plus shifts in open model licensing. The pace suggests the open ecosystem keeps expanding in breadth even as frontier labs pull ahead on scale.
Footnotes
-
MathOverflow Q&A on OpenAI Navier–Stokes claim — https://mathoverflow.net/questions/515057/can-someone-explain-the-claimed-openai-solution-of-the-navier-stokes-problem
↩ ↩2 ↩3OpenAI’s construction relies on smooth forcing and addresses Statements C and D of Fefferman’s formulation, not the unforced regularity problem most mathematicians consider the ultimate physical challenge — which is why OpenAI declined the Clay prize.
-
Tristan Buckmaster public statement (NYU/CIMS PDF) — https://cims.nyu.edu/~tristanb/statement.pdf
↩ ↩2Bubeck pressured me to remove Levent Alpöge from authorship because of his employment at Anthropic, and questioned my career future when I resisted these terms.
-
Unite.ai — Bubeck response — https://www.unite.ai/buckmaster-and-alpoge-post-ai-fluid-blowup-proofs-dispute-openai-contact/
↩ ↩2Bubeck has vigorously denied these ‘false and inflammatory’ claims, stating he never asked for Alpöge to be removed from his own work; the negotiations concerned a human-readable rewrite of a separate proof generated by OpenAI’s internal agents.
-
Stanford Tech Review — Alpöge/Buckmaster methodology — https://stanfordtechreview.com/articles/did-claude-solve-navier-stokes-anthropic-rumor
↩ ↩2Claude and Codex handled bookkeeping of inductive orders and constants, but initial AI-generated drafts were ‘horrendous’ and nearly unreadable; the pair released Lean-formalized certificates, whereas OpenAI’s 100-page proof text remains largely unreleased.
-
openai/NavierStokesAndEuler GitHub repo (via openai.com) — https://openai.com/index/navier-stokes-solution/
↩Built on Lean 4.34.0-rc2 with mathlib; includes the Comparator tool from Lean FRO/ICARM that replays the proof through independent kernels (including Rust-based NanoDa) to block axiom smuggling and definition weakening.
-
Simon Willison quoting Terence Tao — https://simonwillison.net/2026/Sep/8/on-navier-stokes/
↩Jumping to an AI-generated proof is like skipping to the end of a movie — the plot lines are resolved but the exploratory value and versatile human-understandable techniques normally forged during the struggle are lost.
-
Silicon Code Design (Gluon deep-dive) — https://www.siliconcodesign.com/p/an-advanced-system-architecture-breakdown
↩Gluon exposes tile layouts and PTX-style instructions on top of Triton so that Codex-written kernels can target Jalapeño’s slice-based NUMA memory directly; the 30,000-line kernels are ‘measurably correct’ but humans can no longer explain individual lines of the generated assembly.
-
SemiAnalysis — https://newsletter.semianalysis.com/p/openai-jalapeno-better-than-nvidia
↩Jalapeño delivers 1.5 to 1.9 times more work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency versus GB300, but the headline 53.7x–104.3x tokens/kW figures come from matched-speed decoding on hand-picked models like DeepSeek R1 and GPT-OSS 120B where OpenAI chose the operating point.
-
Futurum Group — https://futurumgroup.com/insights/jalapeo-in-nine-months-did-ai-just-break-chip-design-timelines/
↩Nine months from architecture to tape-out is unprecedented, but human engineers still made the critical architectural choices — ‘AI-designed’ is really ‘AI-accelerated verification with a shortened kernel-writing loop,’ not autonomous silicon design.
-
Business Insider — https://www.businessinsider.com/openai-navier-stokes-math-breakthrough-drama-2026-9
↩Tristan Buckmaster (NYU) and Levent Alpöge (Anthropic) posted breakthrough Euler/Boussinesq results one day before OpenAI’s claim and alleged the company may have trained on private drafts stored in Codex; OpenAI denied direct theft but ‘could not rule out’ that de-identified user data helped.
-
Implicator.ai — https://www.implicator.ai/openais-productivity-numbers-look-great-independent-researchers-arent-so-sure/
↩The 3.1 ratio counts aggregate agent runtime, not equivalent human output — over 50% of successful 4-to-8-hour tasks still required at least one human intervention, and the 90th-percentile researcher burns more than $7,000 per day in tokens.
-
MLQ.ai on Sarah Friar — https://mlq.ai/news/openai-cfo-sarah-friar-expresses-concerns-about-2026-ipo-timeline/
↩Friar has internally flagged Altman’s late-2026 IPO push as ‘aggressive,’ citing organizational unpreparedness and risks from the $600 billion five-year spending plan; she is publicly targeting 2027 instead.
-
Unite.ai — https://www.unite.ai/openai-releases-chatgpt-images-2-5-with-sketch-and-two-new-api-models/
↩OpenAI has standardized pricing for both models at $8.00 per million input tokens and $30.00 per million output tokens… On the LMSYS Image Arena, Sunburst and Flare quickly secured the #1 and #2 rankings, respectively.
-
Feisworld (designer hands-on) — https://www.feisworld.com/blog/adobe-firefly-chatgpt-campaign-visuals
↩Independent benchmarks from Artificial Analysis indicate that while OpenAI’s models lead in general text-to-image quality, specialized models like Microsoft’s MAI-Image-2.6 often top editing leaderboards… Nano Banana Pro has been called the ‘world’s highest-rated’ for fidelity.
-
Economic Times — https://economictimes.indiatimes.com/ai/ai-insights/chatgpt-images-2-5-whats-new-in-openais-latest-image-model/articleshow/133956278.cms?from=mdr
↩ ↩2Some tech journalists likened the drawing canvas to ‘MS Paint from 1995’… users report that facial features and specific details (such as the color of a button or the grayness of a character’s hair) often change across multiple turns, even with the improved ‘multi-turn consistency’ promised in 2.5.
-
C2PA Viewer — https://c2paviewer.com/articles/openai-google-c2pa-synthid-2026
↩Independent researchers have identified a vulnerability termed the ‘Integrity Clash.’ This condition occurs when a digital asset carries a valid C2PA manifest claiming human authorship while its pixels simultaneously contain an AI watermark… metadata can be ‘washed’ and replaced through standard editing pipelines without compromising the cryptographic signatures, leading to ‘authenticated fakes’.
-
eesel.ai — https://www.eesel.ai/blog/chatgpt-images-2-5-alternatives
↩Studio Ghibli president Toshio Suzuki dismissed the outputs as ‘meaningless and not fun’… While OpenAI maintains that it blocks the names of individual living artists, it explicitly permits ‘studio styles,’ a distinction that many artists find legally and ethically porous.
-
Master of Code (ChatGPT stats) — https://masterofcode.com/blog/chatgpt-statistics
↩ChatGPT reached 1 billion weekly active users… OpenAI currently serves over 1 million paying business customers, with 92% of Fortune 500 companies having integrated the platform.