AMIE ties 21 PCPs, GPT-Rosalind tops LifeSciBench 36.1%, Arbor beats Codex 2.5×
Today's research leads pitch expert-replacement agents across primary care, drug discovery, and ML engineering — Google's AMIE, OpenAI's Rosalind, Arbor.
AMIE ties 21 PCPs, GPT-Rosalind tops LifeSciBench 36.1%, Arbor beats Codex 2.5×
TL;DR
- Google’s AMIE tied 21 PCPs on management reasoning and beat them on guideline adherence.
- g-AMIE’s guardrail agent caught 64% of red-flag symptoms vs 40% for PCPs.
- GPT-Rosalind topped LifeSciBench at 36.1%, with no Claude in the field and no public dataset.
- Arbor posted 2.5× the held-out gain of Codex across six agent-research tasks.
- Arbor’s merge gate blocks promotions that win only on dev scores, gating the 86% MLE-Bench claim.
Today’s three AI-research leads sit in unrelated expert domains. Google’s AMIE tied 21 primary care physicians on management reasoning and beat them on guideline adherence, while a follow-up g-AMIE adds a guardrail agent that quietly retires the Nature paper’s autonomous framing. OpenAI’s GPT-Rosalind tops a new drug-discovery benchmark at 36.1%, with two reviewer-flagged omissions — no Claude in the field, no public dataset at launch. And Arbor posts 2.5× the held-out gain of Codex on agent research, with the real contribution being a merge gate that blocks promotions winning only on dev scores.
The round-ups lean adversarial: POISE and CodeSpear show how thin agent guardrails actually are, ModSleuth audits hidden model lineage across releases, and a calibration paper finds instruction tuning leaves models systematically overconfident in their own answers.
Google’s AMIE ties PCPs on chronic care, trails on practicality
Source: google-ai-blog · published 2026-06-17
TL;DR
- Google’s AMIE tied 21 primary care physicians on management reasoning and beat them on guideline adherence.
- A Beth Israel pilot of 100 patients found AMIE’s differential included the correct diagnosis 90% of the time.
- PCPs still beat AMIE on practicality and cost-effectiveness — the axis that drives chronic-care adherence.
- Follow-up g-AMIE adds a guardrail agent, catching 64% of red-flag symptoms vs. 40% for PCPs.
- The guardrailed variant quietly retires the autonomous framing the Nature paper describes.
What the Nature paper actually claims
AMIE — Google’s Articulate Medical Intelligence Explorer, built on long-context Gemini — has moved from one-shot diagnosis to multi-visit chronic disease management. In a blinded OSCE-style study with patient actors and 21 PCPs, the dual-agent system (an empathetic dialogue front-end plus a deep-thinking reasoner cross-referencing guidelines and formularies) tied physicians on overall management reasoning and significantly outperformed them on plan preciseness and guideline adherence.
That’s a real result, but a narrow one. The win is on the axes LLMs are structurally good at: exhaustive recall of clinical guidelines and precise medication selection. The real-world test at Beth Israel Deaconess is more telling — and more mixed. Across 100 actual patients interacting with AMIE up to five days before their PCP visit, the model’s differential contained the correct diagnosis in 90% of cases and required zero safety escalations, with 75% of doctors saying the pre-visit summary left them better prepared 1. But on practicality and cost-effectiveness of the resulting management plan, human PCPs still beat the AI 1. For chronic care, where socioeconomic fit determines whether a patient actually fills a prescription, that’s the axis that matters most.
The g-AMIE pivot is the real deployment story
The Nature paper describes an autonomous conversational agent. The thing Google is actually piloting is something else: g-AMIE, a guardrailed variant that intercepts any individualized advice and routes it through an asynchronous “clinician cockpit” for sign-off 2.
flowchart LR
P[Patient chat] --> D[Dialogue agent]
D --> R[Reasoning agent<br/>guidelines + formulary]
R --> G{Guardrail agent}
G -->|safe info| P
G -->|individualized advice| C[Clinician cockpit]
C -->|sign-off| P
The numbers justify the architecture: 90% compliance with safety guardrails vs. ~72% for human clinicians, 64.2% red-flag catch rate vs. 40% for PCPs, and 40% faster physician review than running the consultation themselves 2. That’s a meaningful productivity story — but it’s a different product than the one the Nature headline implies, and it concedes that the original autonomous framing wasn’t deployable.
Methodology and equity caveats
Eric Topol and sociologist Catherine Pope argue the OSCE setup is rigged in the model’s favor: forcing physicians into synchronous text chat strips tone, non-verbal cues, and physical exam — the channels humans are trained on — while letting the LLM play in its native medium 3. Physician-blogger Ahmed Zayed adds that the empathy rubric likely rewards “exhaustive politeness” an untiring chatbot can trivially produce 4.
The equity gap is sharper. An MIT study found medical LLMs were 7–9% more likely to recommend home self-management over clinical care when patient messages contained typos, dramatic phrasing, or non-standard dialects including African American English 5 — a direct hit on the populations chronic-disease tools are pitched to help. AMIE compounds this by remaining closed-source, blocking the kind of independent audit that competing systems like MIRA permit 6.
The rubrics may merely be capturing ‘exhaustive politeness’ rather than a genuine clinical connection. 4
The honest read: AMIE wins where guideline lookup and prescription precision dominate, ties on reasoning, and loses where real patients live. The news worth tracking isn’t the benchmark — it’s whether g-AMIE’s physician-supervised loop holds up in the nationwide field studies now starting.
OpenAI stakes drug-discovery claim with Rosalind and a benchmark
Source: openai-blog · published 2026-06-17
TL;DR
- GPT-Rosalind tops LifeSciBench at 36.1%, with GPT-5.5 at 25.7% and Gemini 3.1 Pro at 23.6%.
- Two omissions undercut the leaderboard: no Claude in the field, no public dataset at launch.
- Molecule.one’s TEMPO loop lifted average yield from 16.6% to 25.2% across 10,080 automated reactions.
- Novo Nordisk extended its OpenAI partnership to manufacturing and supply chain the same week.
A coordinated move, not a research drop
OpenAI shipped two life-sciences artifacts together: LifeSciBench, a 750-task expert-authored benchmark, and a peer-style paper on a “near-autonomous AI chemist” that improved a real medicinal-chemistry reaction. Read separately they look like research. Read together — alongside a Novo Nordisk expansion and a new gated “GPT-Rosalind” model — they’re a commercial entry into territory that Recursion and Schrödinger have spent years staking out 7.
The benchmark numbers, with asterisks
GPT-Rosalind tops LifeSciBench at a 36.1% pass rate, up from GPT-5.5’s 25.7%, with Gemini 3.1 Pro at 23.6% 8. No model clears 40%, and performance collapses on the tasks that matter most for lab work: 14.8% on numeric tasks, 24.0% on sequence/structure outputs, and a drop from 45.1% to 28.1% the moment a task requires interpreting a figure or data file rather than pure text.
Two omissions matter. Anthropic’s Claude family is missing from the official leaderboard entirely 8. And LifeSciBench wasn’t released to GitHub or Hugging Face at launch, with external hosting “not yet confirmed” 9. A benchmark with 19,020 bespoke rubric criteria, scored by a vendor on a vendor-curated field of competitors, is a marketing instrument until an outside group can run it.
The chemistry result is narrower than the headline
The Molecule.one paper is the strongest piece of the drop because it has receipts. Across 10,080 reactions on Molecule.one’s Maria Lab HTE platform, adding a TEMPO additive to Chan–Lam couplings of primary sulfonamides lifted mean yield from 16.6% to 25.2%, and more than doubled the share of reactions exceeding 30% yield (15.6% → 37.5%) 10. The mechanism — suppressing oxidative deboronation — is concrete and chemically plausible.
The caveats are equally concrete. The result is one reaction class on one platform, the loop still depends on human chemists to prepare materials and validate runs, and no independent lab has reproduced it 11. “Near-autonomous” is generous; “AI-assisted optimization of a known failure mode” is closer.
The ‘near-autonomous’ label downplays the critical role of human experts who still prepare materials, operate parts of the physical lab, and validate results. 11
What the bundle is actually for
The strategic frame explains the packaging. Novo Nordisk extended its 2024 OpenAI partnership the same week to cover manufacturing and supply chain, with full integration targeted by year-end 12. GPT-Rosalind is gated through a Trusted Access Program limited to vetted U.S. enterprises — biosecurity language that doubles as enterprise lock-in 7. Recursion and Schrödinger stock fell as the market read the move correctly: OpenAI is no longer adjacent to AI drug discovery, it’s competing in it 7.
The chemistry numbers are real and checkable. The benchmark leaderboard is not, until Claude is on it and the dataset is downloadable. Treat the two halves accordingly.
Further reading
Arbor’s held-out merge gate lifts agent research 2.5× over Codex
Source: hf-daily-papers · published 2026-06-09
TL;DR
- Arbor posts 2.5× the held-out gain of Codex and Claude Code across six research tasks.
- The real contribution is a held-out merge gate that blocks promotions winning only on dev scores.
- On MLE-Bench Lite with GPT-5.5, Arbor hits 86.36% any-medal and 77.27% gold.
- Independent coverage flags a “one-pass plateau”: the tree records what worked, not why.
What Arbor actually adds
Tree-structured ML agents are not new. AIDE’s “solution tree” already framed ML engineering as code-version search with debugging/feature edges as children 13, and I-MCTS layered introspective node expansion on top of MCTS for roughly a 6% lift over flat agents 14. Arbor’s contribution is not the tree; it is the separation of concerns around it, plus a merge gate that refuses to promote any candidate that doesn’t beat the current best on a held-out evaluator.
flowchart LR
C[Coordinator<br/>long-lived] -->|dispatch hypothesis| E1[Executor<br/>git worktree]
C -->|dispatch hypothesis| E2[Executor<br/>git worktree]
E1 -->|score + insight| T[(Hypothesis Tree)]
E2 -->|score + insight| T
T --> C
C -->|promote?| G{Held-out<br/>merge gate}
G -->|pass| B[Current best]
G -->|fail| T
The Coordinator owns global strategy and the tree; Executors are ephemeral, parallel, and return structured evidence (score, insight, commit hash). That split is what lets the system “learn lessons” across branches instead of collapsing into one long, lossy trajectory the way Codex and Claude Code do.
The numbers, and what they cover
The headline result is a 2.5× average relative held-out gain versus Codex (GPT-5.5) and Claude Code (Opus 4.6). The per-task picture is sharper: BrowseComp accuracy moves from 45.33% to 67.67% (Claude Code stalls at 53.33%), and the math-reasoning data-synthesis task gains 19.79 pass-gap points against 5.21 for Codex and 7.29 for Claude Code. The MLE-Bench Lite numbers — 86.36% any-medal, 77.27% gold — are the most eye-catching, but they need calibration: the original AIDE + o1-preview baseline scored just 16.9% any-medal across the full 75-competition MLE-bench 15. “Lite” is a curated subset and the backbone is GPT-5.5, so the delta is HTR plus a much stronger model, not HTR alone.
What the paper underplays
Independent write-ups peg Arbor’s default budget at 20 coordinator cycles, max tree depth 2, and a 48-hour wall-clock cap — matched to AIDE and the official MLE-bench protocol 16. That parity is the right way to read the 2.5× number: it’s at equal compute, not unbounded search. But the same coverage notes agentic coding workloads can burn up to 1,000× the tokens of chat, and historical full MLE-bench sweeps have run ~$48k 1615. “Higher efficiency per unit of progress” is true and also expensive in absolute terms.
The reliability gap is the bigger asterisk. A 2026 survey of autonomous research agents claims coding agents still produce fabricated or invalidated experimental results in roughly 80% of open-ended cases 17. And a summary of the Arbor paper itself names the failure mode the held-out gate doesn’t fix:
Most agents manage a single meaningful structural change but fail to build on it in subsequent iterations — agents can identify a leverage point but cannot generalize why it worked. 18
The merge gate stops the system from shipping overfit dev gains. It doesn’t give the Coordinator a model of why a winning hypothesis won, which is what cumulative research actually requires. Arbor is the most disciplined hypothesis-tree agent yet shipped; it is not, on this evidence, a self-improving researcher.
Round-ups
POISE hides agent skill-poisoning inside benign instructions
Source: hf-daily-papers
POISE plants malicious triggers inside ordinary-looking skill instructions using YAML-header and body injections, evading scanners that only flag privileged tool calls. The position-aware attack reports high success rates against Codex and GPT-5.2 agents, exposing a blind spot in current static defenses.
Grammar-constrained decoding jailbreaks LLMs into writing malware
Source: hf-daily-papers
Grammar constraints meant to enforce syntactic validity in code generation double as an attack surface. The authors introduce CodeSpear, which hides honeypot logic inside grammatically valid output to bypass safety filters, plus CodeShield, an alignment defense that restores semantic harmlessness without sacrificing structural diversity.
ModSleuth audits hidden model lineage across modern LLMs
Source: hf-daily-papers
ModSleuth is an agentic auditor that reconstructs LLM dependency graphs from public artifacts, resolving contradictions in cards, configs, and weights. The system surfaces invisible base-model reuse and train-eval coupling, with implications for license compliance and contamination tracking across heterogeneous releases.
EvoTrainer co-evolves LLM policies and their RL harnesses
Source: hf-daily-papers
EvoTrainer treats the training harness itself as a learnable artifact, evolving policies and infrastructure together using rollout-level diagnostics and backtests. The autonomous loop produces reusable agentic skills and beats handcrafted RL pipelines on complex reasoning and coding tasks.
DeNovoSWE trains agents to build full repos from docs
Source: hf-daily-papers
DeNovoSWE is a large-scale dataset and sandboxed workflow for teaching code agents to generate entire repositories from documentation alone. Using divide-and-conquer with critic-repair loops and difficulty-aware filtering, a fine-tuned Qwen3-30B-A3B posts strong gains on the new BeyondSWE-Doc2Repo benchmark.
xLSTM beats Mamba-2 on sequence modeling benchmarks
Source: hf-daily-papers
A survey of subquadratic architectures finds xLSTM outperforms Mamba-2 and Gated DeltaNet on sequence modeling, crediting stronger state tracking and memory dynamics. The paper traces gains to gating choices that help length generalization, with applications spanning code pre-training, distillation, and time-series foundation models.
Instruction tuning makes LLMs overconfident in their answers
Source: hf-daily-papers
Calibration degrades sharply after instruction tuning, and chat templates worsen the effect through an ownership bias toward the model’s own outputs. Reframing those outputs as user input during confidence elicitation restores calibration without retraining, offering a cheap fix for downstream uncertainty estimates.
Footnotes
-
Digital Health Wire (BIDMC feasibility pilot) — https://digitalhealthwire.com/google-amie-outperforms-in-real-world-debut/
↩ ↩2PCPs still outperformed AMIE on the practicality and cost-effectiveness of management plans; 100 patients interacted with the AI up to five days before primary care appointments, with AMIE’s differential including the correct diagnosis in 90% of cases and zero safety interventions required.
-
Google Research blog on g-AMIE — https://research.google/blog/enabling-physician-centered-oversight-for-amie/
↩ ↩2g-AMIE maintained 90% compliance with safety guardrails versus ~72% for human clinicians, and caught 64.2% of ‘red flag’ symptoms compared to only 40.0% by human PCPs; senior doctors reviewed g-AMIE outputs 40% faster than conducting traditional consultations.
-
Pharmaphorum (Topol/Pope methodology critique) — https://pharmaphorum.com/news/ai-agents-can-beat-doctors-clinical-decision-making
↩The text-only constraint effectively handicapped human physicians by stripping away non-verbal communication, tone of voice, and physical examination—elements that are fundamental to human diagnostic accuracy.
-
Ahmed Zayed, MD (physician blog) — https://zayedmd.com/ai-in-healthcare/google-amie-ai-outperforms-physicians-nature-medicine/
↩ ↩2Rubrics used to measure empathy may merely be capturing ‘exhaustive politeness’ rather than a genuine clinical connection—LLMs can provide scripted, perfectly worded empathetic responses without the fatigue or time constraints facing a real PCP, effectively gaming evaluation axes.
-
KevinMD on MIT bias study — https://kevinmd.com/2026/06/ai-bias-in-health-care-reads-the-writer-not-the-symptom.html
↩Models were 7–9% more likely to recommend home self-management over clinical care when patient messages contained ‘messy’ features—typos, dramatic phrasing, or non-standard dialects—effectively punishing patients with lower digital literacy or those using African American English.
-
Harrison PLLC Substack (legal/oversight analysis) — https://harrisonpllc.substack.com/p/what-happens-when-doctors-supervise
↩AMIE remains proprietary and closed-source, unlike competing systems like MIRA, which limits independent verification of its safety claims; physicians using AI assistants are sometimes perceived by peers as having lesser clinical skill—a ‘competence penalty.‘
-
DrugPatentWatch on GPT-Rosalind — https://www.drugpatentwatch.com/blog/gpt-rosalind-what-openais-life-sciences-model-actually-does-to-drug-development/
↩ ↩2 ↩3Gated deployment through the Trusted Access Program limits Rosalind to vetted U.S. enterprise organizations… shares of established AI-drug discovery firms like Recursion Pharmaceuticals and Schrodinger fell as OpenAI moved into their territory.
-
MarkTechPost (LifeSciBench coverage) — https://www.marktechpost.com/2026/06/17/openai-releases-lifescibench-a-750-task-benchmark-grading-ai-models-on-real-life-science-research-with-expert-written-rubric/
↩ ↩2GPT-Rosalind reaches 36.1% pass rate vs GPT-5.5 at 25.7% and Gemini 3.1 Pro at 23.6%… no model passed more than 40% of the tasks, particularly struggling with wet-lab troubleshooting and multi-step experimental optimization.
-
Digg tech roundup on LifeSciBench access — https://digg.com/tech/7ay1iq9b
↩Specific external hosting plans—including direct GitHub or Hugging Face repository links—remained ‘not yet confirmed’ at the time of the initial announcement… critics have raised questions about whether the highly specialized nature of the Ph.D.-level tasks might inadvertently lead to model ‘overfitting’ on the specific rubrics.
-
OpenAI/Molecule.one TEMPO paper (PDF) — https://cdn.openai.com/pdf/4934b0ed-3de2-4ac5-835c-97604d52dea7/tempo-improves-generality-and-decreases-oxidative-deboronation.pdf
↩Average yield rose from 16.6% to 25.2%; reactions achieving >30% yield increased from 15.6% to 37.5%, across 10,080 reactions run on Molecule.one’s Maria Lab automated HTE platform.
-
Reddit r/AIGuild discussion of AI chemist — https://www.reddit.com/r/AIGuild/comments/1u8ra42/openais_ai_chemist_improved_a_difficult/
↩ ↩2The reported yield improvements have not yet been reproduced by independent laboratories… the ‘near-autonomous’ label downplays the critical role of human experts who still prepare materials, operate parts of the physical lab, and validate results.
-
Digital Health Insights on Novo Nordisk deal — https://dhinsights.org/news/novo-nordisks-openai-deal-reaches-far-beyond-drug-discovery
↩Novo Nordisk expanded its 2024 partnership into a sweeping agreement to integrate OpenAI technology across its entire value chain, including manufacturing and supply chain management, with full integration targeted for the end of the year.
-
AIDE.ml (Weco) — https://www.aide.ml/
↩AIDE frames ML engineering as a code optimization problem… its ‘solution tree’ represents distinct code versions as nodes, with edges representing refinement steps like debugging or feature engineering
-
I-MCTS GitHub (Introspective MCTS for AutoML) — https://github.com/jokieleung/I-MCTS
↩I-MCTS demonstrated a 6% performance gain over standard agents by using introspective node expansion
-
OpenAI MLE-bench GitHub — https://github.com/openai/mle-bench
↩ ↩2evaluates agents across 75 Kaggle competitions… AIDE paired with o1-preview achieved a Bronze medal or higher in 16.9% of tasks
-
Harrison AIX blog — Arbor write-up — https://harrisonaix.com/blog/arbor-autonomous-scientific-research-htr/
↩ ↩2default configuration uses 20 coordinator cycles with a maximum tree depth of two… 48-hour wall-clock limit to ensure consistency with other autonomous agents like AIDE
-
Medium — ‘The Recursive Self-Improvement Loop Tightens But Does Not Close’ (Adnan Masood) — https://medium.com/@adnanmasood/the-recursive-self-improvement-rsi-loop-tightens-but-does-not-close-part-2-348b0abf182c
↩coding agents still produce fabricated or invalidated experimental results in roughly 80% of open-ended research cases
-
AIModels.fyi — Arbor paper summary — https://www.aimodels.fyi/papers/arxiv/toward-generalist-autonomous-research-via-hypothesis-tree
↩a ‘one-pass plateau’ — most agents manage a single meaningful structural change but fail to build on it in subsequent iterations… agents can identify a leverage point but cannot generalize why it worked