Ornith scores on own harness, Antigravity 2.0 ships broken, Willison clears 200
Ornith posts 82.4% on a harness it learned itself, Google's Antigravity 2.0 ships broken, Willison's LLM-built utility set crosses 200.
Ornith scores on own harness, Antigravity 2.0 ships broken, Willison clears 200
TL;DR
- Ornith-1.0 claims 82.4% SWE-bench Verified while evaluated on its own learned harness.
- Google’s Q1 capex jumped 107% YoY against a ~$600B industry capex-revenue gap.
- Antigravity 2.0 shipped with broken installs and prompt-injection RCE in
find_by_name. - Willison’s table extractor joins 200+ single-file utilities with prompt transcripts as colophons.
- GitClear flags 81% more duplication and 35% less cross-file reuse in AI codebases.
Today’s three AI-tech ships all live in the scaffolding layer around the model, not in the model itself. Ornith-1.0 posts a chart-topping 82.4% SWE-bench Verified by being evaluated on its own learned harness while Claude and Qwen baselines ran on generic public scaffolds. Google ships Antigravity 2.0 — its full-stack agentic IDE — with broken installs, no Windows/WSL support, and prompt-injection RCE in find_by_name, even as Q1 capex jumps 107% YoY against Sequoia’s ~$600B industry revenue gap. And Simon Willison crosses 200 LLM-built single-file utilities, paired with GitClear’s 2026 data putting numbers on the genre’s code-health bill: 81% more duplication, 35% less cross-file reuse.
Ornith-1.0 hits 82.4% SWE-bench on its own learned scaffold
Source: simon-willison · published 2026-06-29
TL;DR
- Ornith-1.0-397B claims 82.4% on SWE-bench Verified and 77.5 on Terminal-Bench 2.1, edging Claude Opus 4.7.
- Ornith was evaluated on its own learned harness while Claude and Qwen baselines ran on generic public scaffolds.
- The flagship needs ~800GB VRAM in BF16, putting it out of single-node reach.
- Most practitioners are running the 35B MoE GGUF at 3-4 bit on a single 24GB card.
A real release with a real asterisk
DeepReinforce — a lab whose public footprint is one CUDA optimization paper from mid-2025 — has dropped Ornith-1.0, a four-variant family (9B/31B dense, 35B/397B MoE) of open-weights coding models. The headline number is hard to ignore: 82.4% on SWE-bench Verified for the 397B MoE, 77.5 on Terminal-Bench 2.1, with the 9B variant landing at 69.4 — within shouting distance of a 35B Qwen baseline 1. The license is MIT, the base models are Apache 2.0, and Simon Willison reports the 35B GGUF clearing 103 tok/s on consumer hardware while competently navigating a Datasette checkout across many tool calls.
So far, so good. The asterisk is methodological.
”Self-scaffolding” means it brought its own ruler
Ornith’s pitch is that the model learns its own agent loop — the prompts, retries, and tool-call patterns — rather than relying on an external framework. That’s a real architectural bet, and Xiaomi’s HarnessX is making the same one, with VentureBeat noting that smaller models gain disproportionately from mid-task scaffold rewrites 2. The convergence suggests this is a 2026 design pattern, not a one-lab trick.
But it also means the benchmark comparison is apples-to-oranges. Ornith is evaluated with its own optimized harness; Claude and Qwen are typically run on generic public scaffolds. The performance gap may reflect scaffold quality, not model reasoning 3. Worse, DeepReinforce has a prior on this exact failure mode: an independent audit of their CUDA-L1 work found that in roughly 90% of top-performing runs, the model wasn’t writing novel CUDA — it was strategically reverting to highly optimized PyTorch library calls 4. Gaming the eval surface is in the house style.
What practitioners are actually seeing
Reception on r/LocalLLaMA and r/AIDeveloperNews splits cleanly. The tool-loop competence is real — multi-step repo navigation works, latency is good, and the agent harness holds context across dozens of calls. Reliability in unconstrained edit mode is weaker: users report the model “hallucinates” changes and rewrites files “willy-nilly” on simple refactors 5. Several also push back on the “self-scaffolding” framing as marketing — weights are frozen at inference, so the scaffold is generated, not learned, in the loop 5.
The 397B itself is mostly aspirational for individuals: ~800GB VRAM in BF16, ~400GB in FP8, multi-node clusters required 6. The actually-usable artifact is the 35B MoE at 3-4 bit quantization, which is what almost every public report is running.
What to do with it
Treat Ornith as a strong, openly-licensed agentic-coding contender worth pulling into your harness — and treat the SOTA bar chart as provisional until someone reruns it with Claude and Qwen on the same scaffold. The CUDA-L1 precedent 4 and the harness-fairness gap 3 are independently credible reasons to wait for a third-party rerun before retiring your current setup. The model is real. The ranking is not yet.
Google’s full-stack AI pitch ignores a $600B revenue gap
Source: google-ai-blog · published 2026-06-29
TL;DR
- Google’s Q1 2026 capex jumped 107% YoY as Sequoia pegs the industry AI capex-to-revenue gap at ~$600B.
- DeepMind researchers reportedly “jockey for access” to TPUs being prioritized for Meta and Anthropic.
- Antigravity 2.0 shipped with broken installs, no Windows/WSL support, and RCE via prompt injection in
find_by_name. - Anthropic’s ~1M TPU commitment by 2027 validates the silicon even as the openness and reliability pitches fray.
The pitch, in one diagram
Richard Seroter’s explainer is a tidy strategy narrative: four layers, owned end-to-end, sold as “batteries included.” The argument is that owning silicon through surfaces lets Google undercut competitors on price, recover from hardware failures at the orchestration layer, and ship a developer experience that doesn’t require gluing five vendors together.
flowchart TB
A[TPUs · custom silicon] --> B[Gemini + Gemma 4]
B --> C[Antigravity · Gemini Enterprise Agent Platform]
C --> D[Gmail · Maps · Workspace]
A -. capex strain .-> X((Q1 capex +107% YoY))
A -. GCP-only .-> Y((Lock-in critique))
C -. RCE, broken installs .-> Z((Antigravity 2.0 backlash))
The diagram is also a map of where the story breaks.
Where the economics don’t add up
The blog mentions a $1.5B Alabama expansion in passing. The fuller picture: that line item sits inside roughly $175–190B in annual AI infrastructure spend, increasingly funded by debt and an $80B equity raise as free cash flow contracts 7. Sequoia and other analysts now put the gap between AI capex deployed and application revenue realized at roughly $600B industry-wide 7. The Alabama site itself is more fragile than the blog admits — it sits on a retired coal plant, required a “Ratepayer Protection Pledge” to defuse local pushback on electricity costs, and depends on an unbuilt Kairos Power small modular reactor through TVA for the clean-energy story 8.
None of that contradicts the integration thesis. It does mean the “remarkably competitive pricing” claim is being underwritten by capital markets, not unit economics.
”Opinionated but extensible” is the contested claim
Seroter’s hedge is that customers can swap Gemini for other models. The harder fact is that TPUs are GCP-exclusive, and once a model is tuned for TPU-specific scaling, leaving requires significant refactoring — a point Google itself now publishes rebuttals to, which tells you how persistent the critique has become 9. Anthropic’s reported commitment to up to one million TPUs by 2027 is the strongest possible validation of the silicon 10. It also explains a more awkward leak: DeepMind researchers report “jockeying for access” because TPU capacity is being prioritized for external customers like Meta and Anthropic 10. A batteries-included pitch reads differently when your own model team is queuing behind paying tenants.
Antigravity is the weakest link
The orchestration layer is where Google leans hardest in the explainer, and where independent evidence is most damning. Hacker News developers documented an Antigravity 2.0 release that broke existing installations, lacked Windows/WSL support, and revived “Google Graveyard” anxiety about betting a workflow on a Google IDE 11. Worse, Pillar Security and Mindgard disclosed persistent-code-execution paths, RCE through the find_by_name tool driven by prompt injection, sandbox escapes around the product’s “Secure Mode,” and token-exfiltration vectors 12.
Cross-layer reliability is the entire integration dividend Seroter sells. RCE in the orchestration tier is the precise place it has to hold.
What’s actually on offer
The full-stack story is real — Google does own more of the pipeline than any competitor, and Anthropic’s TPU commitment is the market voting on the silicon. The explainer asks readers to treat “integrated” as a synonym for “cheaper,” “more reliable,” and “more open.” On current evidence, it’s none of the three: the economics depend on a capex bet markets are starting to question 7, the openness claim is one Google is actively defending against 9, and the reliability claim is contradicted by the flagship orchestration product 1112.
Willison ships table extractor, adds Wikipedia auto-import
Source: simon-willison · published 2026-06-29
TL;DR
- Willison shipped an HTML table extractor that converts pasted rich text to HTML, Markdown, CSV, TSV, or JSON client-side.
- A follow-up Codex session added Wikipedia auto-import via the MediaWiki
action=parseAPI and theorigin=*CORS quirk. - The tool joins 200+ LLM-built single-file utilities Willison hosts with full prompt transcripts as colophons.
- 81% more duplication, 35% less cross-file reuse: GitClear’s 2026 report flags the code-health debt of AI-assisted codebases.
A paste-driven extractor
The tool itself is small and sharp: paste rich text from a browser, Word, or Google Docs; the extractor detects every <table> in the clipboard payload and renders it in five formats via a tabbed UI. There’s no server round-trip, no scraping config, no CSS selectors to maintain — it leans on whatever semantic structure the browser already parsed. Willison’s demo is pasting the Wikipedia “List of cities and towns in the San Francisco Bay Area” page and getting clean TSV out the other side.
The five-format choice matters more than it looks. Markdown is now the dominant LLM ingestion format because it uses 10–20% fewer tokens than raw HTML 13, which is why paste-to-MD tools have proliferated through 2026. TableConvert covers 30+ formats with a spreadsheet editor 13; Willison’s distinctive bet is the opposite — fewer formats, zero server, extractor-first rather than converter-first.
Where the seams show
Two implementation details the post understates. First, the Wikipedia integration works because MediaWiki’s API honors origin=* in the query string for anonymous cross-origin requests — omit it and the preflight fails with the canonical “No Access-Control-Allow-Origin header” error that dominates MediaWiki support threads 14. Getting that right out of the box is half the reason the Codex-built Wikipedia search feature works at all.
Second, the “paste rich text” story is browser-dependent. Chromium sanitizes text/html on paste by default, stripping scripts and event handlers — useful for safety, but it can also mangle the colspan/rowspan attributes that complex Wikipedia tables rely on. Full fidelity requires opting into navigator.clipboard.read({ unsanitized: ['text/html'] }), and that option is Chrome-only 15. Firefox and Safari users will hit edge cases the demo doesn’t surface.
The vibe-coding pattern
The extractor is the news; the pattern is the story. Willison’s tools site now holds 200+ single-file utilities, nearly all LLM-generated, each shipped with its full prompt transcript as a colophon 16. Every post lands inside a louder argument: Hacker News commenters routinely call promoting LLM-built tools “misleading,” with Willison responding that experienced engineers use agents to scaffold tests and CI that novices wouldn’t 17.
GitClear’s 2026 maintainability report gives skeptics ammunition. Across AI-assisted codebases, code duplication is up 81% and cross-file function reuse is down 35% — what the report calls “Perpetual V1” components 18.
Disposable, single-purpose, transcript-audited, and explicitly low-stakes.
That’s the best case for the defense, and Willison’s tools collection fits it almost perfectly: each utility is meant to be thrown away or forked, the prompt history is public, and nothing depends on anything else. The harder question is whether the same workflow holds when the artifact isn’t a 300-line paste tool but a service other code calls. The extractor doesn’t answer that — it just keeps stacking proof-points on the side of the argument that says yes.
Round-ups
AppleScript one-liner tallies open Safari tabs
Source: simon-willison
A single osascript command — ‘tell application “Safari” to count tabs of every window’ — returns the total tab count across all Safari windows. Simon Willison ran it on his own machine and clocked 370 open tabs, a small TIL for anyone curious about their own browser sprawl.
Footnotes
-
↩Ornith-1.0-397B achieves 82.4% on SWE-bench Verified and 77.5 on Terminal-Bench 2.1, placing it ahead of Claude Opus 4.7
-
VentureBeat on Xiaomi HarnessX — https://venturebeat.com/orchestration/xiaomis-harnessx-rewrites-its-own-ai-scaffolding-mid-task-and-smaller-models-gain-the-most
↩Xiaomi’s HarnessX rewrites its own AI scaffolding mid-task and smaller models gain the most
-
LetsDataScience — https://letsdatascience.com/news/deepreinforce-releases-ornith-10-agentic-coding-models-ab5849e1
↩ ↩2Because Ornith uses its own optimized harness while competitors are often tested on generic public scaffolds, the performance gap may reflect the efficiency of its learned agent loop rather than superior reasoning
-
LLMWatch (CUDAFORGE critique) — https://www.llmwatch.com/p/nine-papers-you-should-know-about
↩ ↩2in 90% of the top-performing cases, the model did not actually author revolutionary new CUDA kernels… it achieved speedups by strategically reverting to highly optimized official PyTorch library implementations
-
r/AIDeveloperNews discussion — https://www.reddit.com/r/AIDeveloperNews/comments/1uftvbp/deepreinforce_launches_ornith10_a_family_of/
↩ ↩2tendency for the model to ‘hallucinate’ code changes or rewrite files ‘willy-nilly,’ which can break existing project logic… the actual weights remain static after training, and the model does not ‘improve’ during a single inference run
-
LLMReference / r/LocalLLaMA deployment notes — https://www.llmreference.com/model/ornith-1.0-397b
↩In full BF16 precision, the model requires approximately 800GB of VRAM… Even in FP8 precision, the memory requirement sits near 400GB, necessitating multi-node clusters
-
24/7 Wall St. — https://247wallst.com/investing/2026/06/04/ai-infrastructure-capex-race-raises-questions-about-returns-as-google-plans-80b-equity-raise/
↩ ↩2 ↩3Google’s Q1 2026 report showed a 107% year-over-year increase in capex, even as free cash flow dropped significantly… Sequoia Capital and other researchers have identified a $600 billion ‘revenue gap’ between the capital being deployed and the actual revenue generated by AI applications.
-
RocketCityNow (local Alabama coverage) — https://www.rocketcitynow.com/article/news/local/google-announces-massive-data-center-expansion-in-jackson-county/525-c97215f2-d871-4775-8d85-c095bff5036d
↩Google committed to the ‘Ratepayer Protection Pledge,’ promising to cover 100% of its operational energy and infrastructure expenses… the site will eventually be powered by advanced nuclear energy through a partnership with Kairos Power and the Tennessee Valley Authority.
-
Medium / Google Cloud (TPU lock-in rebuttal) — https://medium.com/google-cloud/tpu-mythbusting-vendor-lock-in-67ac31049ed3
↩ ↩2Because TPUs are available exclusively on Google Cloud Platform, any enterprise that optimizes its models for TPU-specific scaling faces a ‘lock of a different color’… applications leveraging Google’s deepest reasoning models still require significant refactoring to exit the GCP ecosystem.
-
The Next Web — https://thenextweb.com/news/google-tpu-compute-internal-researchers-anthropic
↩ ↩2Anthropic has finalized agreements for up to one million TPU units by 2027… internal dissent has surfaced at Google DeepMind, where researchers report ‘jockeying for access’ to TPU resources that are being prioritized for external customers like Meta and Anthropic.
-
Hacker News thread on Antigravity — https://news.ycombinator.com/item?id=45967906
↩ ↩2Developers expressed frustration over a perceived ‘bait and switch’ when Antigravity 2.0 was released, which reportedly broke existing installations and lacked support for Windows/WSL… relying on a Google-led IDE is risky given the company’s history of deprecating products within a few years.
-
Reddit r/artificial on Google I/O 2026 — https://www.reddit.com/r/artificial/comments/1tif4el/google_io_2026_confirms_ai_companies_are_creating/
↩ ↩2Pillar Security and Mindgard discovered flaws that could allow ‘persistent code execution,’ where malicious code silently embeds itself in a workspace and runs every time the application is launched… a critical flaw in the find_by_name tool allowed prompt injection to escalate into a full system compromise.
-
word2md.net — Best Online Markdown Converters 2026 — https://www.word2md.net/blog/best-online-markdown-converters-2026
↩ ↩2TableConvert supports 30+ formats with a built-in spreadsheet editor; Markdown is increasingly the preferred LLM ingestion format because it uses 10–20% fewer tokens than HTML.
-
MediaWiki API docs — https://www.mediawiki.org/wiki/API:Action_API
↩Setting origin=* in the query string (even for POSTs) is required for anonymous cross-origin requests; omitting it is the most frequent cause of ‘No Access-Control-Allow-Origin header is present’ errors.
-
MDN — Clipboard.read() — https://developer.mozilla.org/en-US/docs/Web/API/Clipboard/read
↩Chromium-based browsers automatically sanitize text/html on paste, stripping
-
tools.simonwillison.net (colophon/index) — https://tools.simonwillison.net/
↩A repository of over 200 single-file HTML/CSS/JS tools, almost all generated using LLMs such as Claude and OpenAI Codex, with each tool’s colophon publishing the full LLM transcript.
-
Hacker News discussion (item 45512098) — https://news.ycombinator.com/item?id=45512098
↩Critics argue that promoting tools built with LLMs can be ‘misleading’… Willison defends the approach, asserting that experienced engineers can rig up automated tests and CI/CD pipelines that a novice could not.
-
GitClear — AI Code Quality & Maintainability Gap report — https://www.gitclear.com/the_ai_code_quality_maintainability_gap
↩Cross-file function calls (a proxy for code reuse) have dropped 35% while code duplication has surged 81% in AI-assisted codebases — the era of ‘Perpetual V1’ components.