JS Wei (Jack) Sun

DeepMind takes EVE stake, Hume audits ASR benchmarks, Simile raises $200M

Three research drops today put the leverage outside the model — the game environment, the benchmark corpus, the survey data.

DeepMind takes EVE stake, Hume audits ASR benchmarks, Simile raises $200M

TL;DR

  • DeepMind takes minority equity in EVE Frontier maker Fenris alongside a $120M buyout.
  • SIMA 2 hits 62% on held-out tasks via a Gemini task-setter and reward-model self-play loop.
  • Hume and Hugging Face catch 6 of 11 ASR models hallucinating a Thank you absent from the audio.
  • Simile closes $200M at a $2B valuation on agentic twins, with critics flagging training-set recall.
  • Nvidia research finds tuning the agent harness keeps performance steady even with a mediocre underlying model.

Three research drops today all sit on the same question: how much of a headline model number is the model doing, and how much is the substrate around it? DeepMind’s minority stake in the maker of EVE Frontier is a bet on the environment — SIMA 2’s 62% held-out score depends on the game world its Gemini-driven self-play loop trains in. Simile’s $200M at a $2B valuation prices agentic twins built on 2.9M consented survey responses, with critics warning the 85% accuracy number may reflect model recall of surveys already in training. And Hume and Hugging Face’s ASR audit shows the failure mode inverted: the models with the lowest reported WER got there by memorizing the benchmark’s own transcription errors — not by hearing the audio better.

Nvidia’s research thread in today’s briefs makes the same point directly: tuning the harness around a mediocre model outperforms swapping the model itself. The frame is convergent — this week the leverage is in the substrate, and the audits are catching up with the headline numbers.

DeepMind takes equity in EVE Online maker for agent training

Source: deepmind-blog · published 2026-08-21

TL;DR

  • DeepMind took a minority equity stake in Fenris Creations alongside the studio’s $120M buyout from Pearl Abyss 1.
  • SIMA 2 hits 62% on held-out tasks — roughly double SIMA 1’s 31% and closing on the ~71% human baseline 2.
  • The agent self-improves via two Gemini models (task setter + reward model), reducing dependence on human demonstrations 3.
  • EVE Frontier’s game director concedes anti-bot enforcement is dead and is designing combat friction instead to keep humans relevant 4.

The “partnership” is a data deal

DeepMind’s blog frames its new tie-up with Fenris Creations — the studio formerly known as CCP Games — as the next chapter in a 15-year research arc from Atari to AlphaFold. The financial structure tells a blunter story. Fenris just completed a $120M management buyout from Pearl Abyss, rebranded, and simultaneously granted DeepMind a minority equity stake specifically so the lab can use 20+ years of EVE Online player behavior as a training substrate 1. It is a rare arrangement: a frontier AI lab taking equity in a game studio to secure a persistent-world dataset, not just licensing access to one.

That reframes the four “frontier capabilities” DeepMind lists — continual learning, long-term memory, long-horizon planning, multi-agent dynamics — as the specific gaps EVE’s data is meant to close.

SIMA 2’s numbers back the research half

The research contribution the blog underplays is in the SIMA 2 technical report. The agent hits a 62% success rate on held-out benchmarks, roughly doubling SIMA 1’s 31% and approaching the ~71% human baseline, with generalization tested in photorealistic worlds generated by Genie 3 that the agent never saw during training 2.

The architectural novelty is a self-improvement loop: one Gemini model acts as a “task setter” generating novel challenges, a second Gemini serves as a reward model scoring the resulting trajectories, and the experience bank feeds back into training 3. That closes the loop on human-labeled demonstrations — which is the actual reason DeepMind wants an open-ended persistent world to point it at.

flowchart LR
    A[Gemini task setter] -->|novel challenge| B[SIMA 2 agent]
    B -->|trajectory| C[Gemini reward model]
    C -->|scored experience| D[(Experience bank)]
    D -->|training data| B

The parts DeepMind’s post leaves out

Player reception on r/Eve is markedly cooler than the blog’s tone suggests. Concerns cluster around Google scraping decades of interaction data, the possibility of the launcher using player hardware for background inference, and Google’s habit of shuttering experimental products 5. None of these are addressed in the announcement.

On the game-design side, EVE Frontier’s director “FC Goodfella” openly concedes the studio has stopped fighting an arms race against bots, and is instead engineering moment-to-moment combat friction demanding enough that human attention remains necessary 4. Read literally, that is a live game studio declaring agentic AI a fait accompli in its own genre.

Independent hands-on coverage of SIMA 2 in No Man’s Sky also notes the agent’s short memory produces “aimless” drift once long-horizon goals fall out of context, and challenges the implicit robotics throughline on the grounds that keyboard-and-mouse action spaces are trivially low-dimensional compared to physical manipulation 6.

What’s actually at stake

The SIMA 2 benchmarks and self-play architecture are real progress. The EVE deal is a bet that the remaining hard problems — memory across weeks, planning across months, coexistence with thousands of humans and other agents — need a persistent economy to study, and that owning a slice of one is cheaper than building it. Whether EVE’s players consented to being the substrate is a separate question, and one Fenris will have to answer before DeepMind’s agents get anywhere near the live shard.


Simile raises $200M to model human behavior, not reasoning

Source: latent-space · published 2026-08-21

TL;DR

  • Simile closed $200M at a $2B valuation, five months after a $100M Series A led by Greenoaks 7.
  • CVS runs ~400,000 “agentic twins” built from 2.9M consented responses, cutting message-testing from weeks to minutes 8.
  • The 85% accuracy headline covers test-retest on stable attitudes, with agents no better than demographic baselines on strategic economic games 9.
  • Critics warn scores may reflect model recall of surveys already in training data, not genuine prediction 10.

The pitch: behavior is a separate foundation-model problem

Joon Sung Park’s argument in the Latent Space interview is that frontier LLMs are optimized in the wrong direction for social science. RLHF pushes models toward super-rational, helpful, median answers; simulating a population requires reproducing bias, error, and irrationality. Simile is betting there’s a distinct “social physics” foundation-model layer — trained on long-form interviews, observational data, and RCTs from sources like the Open Science Framework — that generic prompting of GPT-5 or Claude will never reach.

Investors are pricing that bet aggressively. Greenoaks led a $200M Series B at $2B just five months after a $100M Series A, with Index, Bain, and CVS Health Ventures participating 7. Over $300M raised in half a year is not a research grant; it’s a claim that “synthetic users” is infrastructure.

What’s actually shipping

The strongest evidence isn’t the podcast — it’s CVS. The retailer has deployed roughly 400,000 agentic twins built from 2.9 million consented member responses, and uses them to pressure-test medication-adherence messaging before it goes to real patients. Research cycles that took 4–6 weeks now run in 15–30 minutes 8. That’s a Fortune 20 healthcare buyer paying in a regulated vertical, which is the most concrete answer Simile has to “is this a demo or a business.”

Where the 85% number breaks

Park’s headline benchmark — 85% accuracy from the 1,000-person Generative Agents paper — has a narrower denominator than the marketing suggests. It measures how well agents match a person’s own responses on repeat surveys of stable attitudes. Independent review found the agents excel there but collapse on strategic economic games involving trust and reciprocity, where they perform no better than a plain demographic model 9.

The synthetic-users literature adds a sharper critique. Practitioners running LLM personas against real customer research find the outputs “shallow” and “one-dimensional,” systematically underestimating variance and missing minority opinions because RLHF training pushes toward a median web-text answer 11. And a structural worry hangs over the whole space: high accuracy scores may reflect model recall of surveys already ingested during pretraining, not causal understanding of behavior 10.

Where twins workWhere they don’t
Stable social attitudes (GSS-style)Strategic games (trust, reciprocity)
Message A/B testing at scaleMinority-opinion discovery
Fast iteration on known populationsNovel scenarios absent from training data

The ethical objection Park has to answer

The 8-billion-twins ambition runs directly into Agnew et al.’s “Illusion of Artificial Inclusion,” which argues LLM-as-participant substitution is “another way to devalue members of marginalized communities” because models cannot opt out, resist researcher framing, or correct false assumptions the way human subjects can 12. That’s a critique from the HCI venues Park himself publishes in, and the Gallup partnership on policy simulation makes it more urgent, not less.

Takeaway

Simile is a real business with a real deployment and a defensible research pedigree. But the market is crowded — Aaru, Synthetic Users, Electric Twin, and others sell adjacent pitches with their own accuracy claims 10 — and the “simulation as scaling law” framing is doing more work than the current benchmarks support. The real question for 2026 isn’t whether synthetic focus groups replace real ones; it’s whether behavior foundation models generalize past the attitude-survey regime where they currently shine.


Hume and HF catch ASR models memorizing benchmark errors

Source: huggingface-blog · published 2026-08-21

TL;DR

  • 6 of 11 ASR models dropped an audible “Thank you” to match a flawed VoxPopuli reference — worst offenders had the lowest reported WER.
  • On LibriSpeech, top models “transcribed” digitally silenced numbers in 30–40% of clips.
  • Same models hit 90% accuracy picking each benchmark’s preferred spelling (“Mr.” vs “Mister”).
  • Fresh post-cutoff recordings broke the trick: models reverted to faithful transcription of the audio.

The trick: models know which test they’re taking

Hugging Face and Hume AI ran 11 open-source ASR systems — Whisper-large-v3, NVIDIA Canary-Qwen-2.5B and Parakeet, Cohere Transcribe, Voxtral, Phi-4-multimodal, Granite-speech, Qwen3-ASR, Kimi-Audio, Higgs and Moonshine — through three diagnostic probes designed to separate “hears the audio” from “recognizes the benchmark.” The models mostly failed the separation.

ProbeWhat it testsHeadline result
Reference disagreementReproduce known-wrong transcripts?6/11 models omit audible “Thank you” on VoxPopuli
Masked entity retrievalFill in silenced numbers/dates?30–40% recovery on LibriSpeech
Orthographic switchingPick dataset’s preferred spelling?Up to 90% switch accuracy

The masked-entity result is the load-bearing one. If a model transcribes a number that has been digitally removed from the audio, it is not doing speech recognition on that span — it is doing text completion against memorized samples. The orthographic probe rules out “the model just prefers ‘Mister’”: the same model flips to “Mr.” on a different benchmark, phonetically identical audio, based on acoustic fingerprints of the recording conditions.

Overall, benchmark-optimized models reproduced incorrect references 18–30% of the time, and roughly 40% of VoxPopuli clips carried flaggable reference errors. The generalization gap shows up cleanly on “libri-fresh” and “ep-fresh” — resynthesized or newly recorded audio with the same content — where models stopped matching the flawed references and started transcribing what they actually heard. Independent reporting names Cohere Transcribe (03-2026) and NVIDIA Canary-Qwen-2.5B as reproducing the “Thank you, Mr. President” omission verbatim 13.

Not just a VoxPopuli quirk

Independent work suggests the rot is structural. Artificial Analysis has already shipped VoxPopuli-Cleaned-AA, a hand-corrected variant of hundreds of samples so leaderboard runs can be re-scored 14 — a concrete acknowledgment that the ground truth isn’t ground truth. A 2025 Google audit found VoxPopuli mixes Bokmål with Nynorsk in its Norwegian split and mislabels Modern Standard Arabic as Egyptian, producing what Google calls an “illusion of success” in multilingual training 15. Prior contamination studies measured >75% sentence overlap between Common Voice test splits and public pretraining corpora — invisible in WER but obvious in per-token probability 16.

Meanwhile the sub-5% WERs that vendors quote on clean read speech degrade 2.8×–5.7× in production, and multi-speaker clinical or emergency-dispatch audio routinely exceeds 50% WER 17. The leaderboard number is measuring something, but it isn’t “how well does this model transcribe the world.”

What replaces the WER leaderboard

Hume’s proposed alternative, Real World VoiceEQ, is more radical than the blog frames it: two-stage human review with ≥3 raters per clip, scoring across 60+ metrics including prosody, hesitation and emotional shift, and explicitly refusing a single aggregate number 18. The Open ASR Leaderboard has added a “Benchmark fitting” tab that surfaces reference-error reproduction and orthographic switching directly, so the diagnostic probes become part of the scoreboard rather than a critique of it.

Notably absent from the record so far: any on-record response from NVIDIA, Cohere, or OpenAI. For a paper that names specific model versions and specific failure modes, the silence is its own data point.

Round-ups

Nvidia research finds agent harness matters more than model

Source: techcrunch-ai

Nvidia’s team shows fine-tuning the scaffolding around an agent — the harness that structures its actions — keeps performance steady even when the underlying model is mediocre. The finding shifts focus from raw model quality to the wrapper that constrains and guides agent behavior.

Footnotes

  1. InvestGame.net (deal analysis)https://investgame.net/news/fenris-creations-finalizing-mbo-rebranding-and-partnership-with-google-deepmind/

    Fenris Creations completed a $120M management buyout from Pearl Abyss, with DeepMind taking a minority equity stake — a rare instance of a frontier AI lab acquiring equity in a game studio specifically to use its data as a research substrate.

    2
  2. SIMA 2 technical report (arXiv 2512.04797)https://arxiv.org/abs/2512.04797

    SIMA 2 achieves a 62% success rate on held-out benchmarks, roughly doubling SIMA 1 (31%) and approaching the human baseline of ~71%, with generalization tested in photorealistic worlds generated by Genie 3.

    2
  3. InfoQ coverage of SIMA 2https://www.infoq.com/news/2025/12/sima-2-gemini-agent/

    SIMA 2 uses a Gemini-based task setter to generate novel challenges and a separate Gemini reward model to score trajectories, enabling an autonomous self-improvement loop that exceeds agents trained only on human demonstrations.

    2
  4. BlockchainGamer.biz — EVE Frontier design interviewhttps://www.blockchaingamer.biz/news/42826/eve-frontier-designing-for-ai-agents/

    Rather than fighting an arms race against bots, Game Director ‘FC Goodfella’ says EVE Frontier is designed on the assumption that AI and botting are permanent fixtures — introducing moment-to-moment combat friction demanding enough that human attention remains necessary.

    2
  5. r/Eve community threadhttps://www.reddit.com/r/Eve/comments/1t8zabq/ccp_games_release_statement_about_transition_to/

    Players fear Google may scrape 23 years of human interaction data or use player hardware for background AI processing via the launcher, and worry Google will abandon the project as it has other experimental ventures.

  6. Towards AI hands-on review (No Man’s Sky)https://pub.towardsai.net/i-watched-an-ai-play-no-mans-sky-at-2-am-and-now-i-can-t-stop-thinking-about-it-570b82b4add1

    The agent’s short memory can lead to ‘aimless’ behavior if long-term goals are not reinforced, and Hacker News commenters question how low-dimensional keyboard-and-mouse controls will translate to the high-dimensional complexity of real-world robotics.

  7. TechFundingNewshttps://techfundingnews.com/simile-bags-200m-at-2b-five-months-after-100m-series-a-to-predict-what-humans-will-do-before-ai-gets-it-wrong/

    Simile bags $200M at $2B five months after $100M Series A to predict what humans will do before AI gets it wrong

    2
  8. CVS Health corporate bloghttps://www.cvshealth.com/news/innovation/how-cvs-health-test-drives-better-care-experiences-using-generative-agents.html

    CVS Health test-drives better care experiences using generative agents — deploying ~400,000 agentic twins built from 2.9M consented responses to pressure-test medication adherence messaging

    2
  9. Nervegna Substack review of the 1,000-person paperhttps://nervegna.substack.com/p/stanford-hais-1000-person-ai-simulation

    Agents excel at replicating social attitudes but struggle with strategic economic games involving trust and reciprocity, where their predictive power was no better than simpler demographic-based models.

    2
  10. PSB Insights — Digital Twins vs Synthetic Data landscapehttps://www.psbinsights.com/insights/digital-twins-synthetic-data/

    Competitors include Aaru, Synthetic Users, Electric Twin, Verve VIPs, neuroflash and Yabble… critics warn high accuracy scores may reflect ‘model recall’ of surveys already in training data rather than true prediction.

    2 3
  11. Perspective.ai — ‘Why fake respondents can’t replace real customer research’https://getperspective.ai/blog/synthetic-focus-groups-why-fake-respondents-can-t-replace-real-customer-research

    Synthetic responses were ‘shallow’ and ‘one-dimensional’… LLMs converge on a median web-text answer, systematically underestimating variance and missing minority opinions due to RLHF-induced sycophancy.

  12. Agnew et al., ‘The Illusion of Artificial Inclusion’ (PMC/NIH mirror)https://pmc.ncbi.nlm.nih.gov/articles/PMC12184514/

    Generating text with biased models is simply another way to devalue members of marginalized communities… LLMs lack the discretionary powers to opt out, resist researcher assumptions, or correct misconceptions.

  13. AI Breaking Wire — ‘Study Finds Top AI Speech Models Replicate Benchmark Errors by 30%’https://aibreakingwire.com/news/study-finds-top-ai-speech-models-replicate-benchmark-errors-by-30

    Both Cohere Transcribe (03-2026) and NVIDIA Canary-Qwen-2.5B were found to replicate this faulty reference verbatim, effectively ‘hallucinating’ the omission to match the benchmark’s incorrect data.

  14. Artificial Analysis — VoxPopuli-Cleaned-AA dataset card (Hugging Face)https://huggingface.co/datasets/ArtificialAnalysis/VoxPopuli-Cleaned-AA

    VoxPopuli-Cleaned-AA manually corrects hundreds of samples to ensure more rigorous evaluation of multilingual speech-to-text engines.

  15. Slator — ‘Google Flags Serious Data Quality Issues in Public Multilingual Speech Datasets’https://slator.com/google-flags-serious-data-quality-issues-in-public-multilingual-speech-datasets/

    Macro-level issues such as the mixing of distinct dialects (e.g., Bokmål and Nynorsk in Norwegian) and the mislabeling of Modern Standard Arabic as Egyptian Arabic … create an ‘illusion of success’ in training models that cannot handle natural variation.

  16. arXiv (v1 HTML) — Evaluation of LLM-based ASR / benchmark contamination studyhttps://arxiv.org/html/2607.14846v1

    Contaminated models do not always show drastically different Word Error Rates but assign significantly higher probabilities to test-set transcriptions, indicating memorization rather than generalization; sentence overlap rates exceed 75% for certain languages in Common Voice.

  17. llms.blog — ‘Study Exposes Benchmark Overfitting and Acoustic Leakage in Speech Recognition Models’https://llms.blog/posts/study-exposes-benchmark-overfitting-and-acoustic-leakage-in-speech-recognition-models

    Commercial APIs often report WERs below 5% on clean, read speech, yet performance typically degrades by 2.8x to 5.7x in production environments … the industry average for multi-speaker clinical conversations or high-noise emergency medical dialogues can exceed 50% WER.

  18. arXiv — RW-Voice-EQ Bench (Hume AI)https://arxiv.org/pdf/2607.14846

    Ground-truth references are established via a two-stage human review process; at least three independent human raters review and correct model-generated transcriptions to resolve discrepancies … the framework evaluates 15+ dimensions and 60+ metrics, including paralinguistic cues like tone, hesitation, and emotional shifts.

Jack Sun

Jack Sun, writing.

Engineer · Bay Area

Hands-on with agentic AI all day — building frameworks, reading what industry ships, occasionally writing them down.

Digest
All · AI Tech · AI Research · AI News
Writing
Essays
Elsewhere
Subscribe
All · AI Tech · AI Research · AI News · Essays

© 2026 Wei (Jack) Sun · jacksunwei.me Built on Astro · hosted on Cloudflare