DeepMind takes EVE stake, Hume audits ASR benchmarks, Simile raises $200M
Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.
Sources
From Atari to EVE Online: Building on 15 Years of AI Research in Games deepmind.google
Google DeepMind partners with game studios to prototype breakthrough AI gameplay.
Measuring benchmark optimization in speech recognition huggingface.co
Simulation: the new Scaling Law — Joon Sung Park, Simile AI latent.space
Simile’s CEO about his journey from the viral Generative Agents to creating 8 Billion Digital Twins of every living human… and why it’s gone from fun exploration to very serious business.
Nvidia just showed that the harness, not the AI model, is now the real hero techcrunch.com
Nvidia’s team shows fine-tuning the scaffolding around an agent — the harness that structures its actions — keeps performance steady even when the underlying model is mediocre. The finding shifts focus from raw model quality to the wrapper that constrains and guides agent behavior.
References
SIMA 2 technical report (arXiv 2512.04797) arxiv.org
SIMA 2 achieves a 62% success rate on held-out benchmarks, roughly doubling SIMA 1 (31%) and approaching the human baseline of ~71%, with generalization tested in photorealistic worlds generated by Genie 3.
InfoQ coverage of SIMA 2 infoq.com
SIMA 2 uses a Gemini-based task setter to generate novel challenges and a separate Gemini reward model to score trajectories, enabling an autonomous self-improvement loop that exceeds agents trained only on human demonstrations.
InvestGame.net (deal analysis) investgame.net
Fenris Creations completed a $120M management buyout from Pearl Abyss, with DeepMind taking a minority equity stake — a rare instance of a frontier AI lab acquiring equity in a game studio specifically to use its data as a research substrate.
r/Eve community thread reddit.com
Players fear Google may scrape 23 years of human interaction data or use player hardware for background AI processing via the launcher, and worry Google will abandon the project as it has other experimental ventures.
BlockchainGamer.biz — EVE Frontier design interview blockchaingamer.biz
Rather than fighting an arms race against bots, Game Director ‘FC Goodfella’ says EVE Frontier is designed on the assumption that AI and botting are permanent fixtures — introducing moment-to-moment combat friction demanding enough that human attention remains necessary.
Towards AI hands-on review (No Man’s Sky) pub.towardsai.net
The agent’s short memory can lead to ‘aimless’ behavior if long-term goals are not reinforced, and Hacker News commenters question how low-dimensional keyboard-and-mouse controls will translate to the high-dimensional complexity of real-world robotics.
TechFundingNews techfundingnews.com
Simile bags $200M at $2B five months after $100M Series A to predict what humans will do before AI gets it wrong
CVS Health corporate blog cvshealth.com
CVS Health test-drives better care experiences using generative agents — deploying ~400,000 agentic twins built from 2.9M consented responses to pressure-test medication adherence messaging
Agnew et al., ‘The Illusion of Artificial Inclusion’ (PMC/NIH mirror) pmc.ncbi.nlm.nih.gov
Generating text with biased models is simply another way to devalue members of marginalized communities… LLMs lack the discretionary powers to opt out, resist researcher assumptions, or correct misconceptions.
Nervegna Substack review of the 1,000-person paper nervegna.substack.com
Agents excel at replicating social attitudes but struggle with strategic economic games involving trust and reciprocity, where their predictive power was no better than simpler demographic-based models.
Perspective.ai — ‘Why fake respondents can’t replace real customer research’ getperspective.ai
Synthetic responses were ‘shallow’ and ‘one-dimensional’… LLMs converge on a median web-text answer, systematically underestimating variance and missing minority opinions due to RLHF-induced sycophancy.
PSB Insights — Digital Twins vs Synthetic Data landscape psbinsights.com
Competitors include Aaru, Synthetic Users, Electric Twin, Verve VIPs, neuroflash and Yabble… critics warn high accuracy scores may reflect ‘model recall’ of surveys already in training data rather than true prediction.
Artificial Analysis — VoxPopuli-Cleaned-AA dataset card (Hugging Face) huggingface.co
VoxPopuli-Cleaned-AA manually corrects hundreds of samples to ensure more rigorous evaluation of multilingual speech-to-text engines.
Slator — ‘Google Flags Serious Data Quality Issues in Public Multilingual Speech Datasets’ slator.com
Macro-level issues such as the mixing of distinct dialects (e.g., Bokmål and Nynorsk in Norwegian) and the mislabeling of Modern Standard Arabic as Egyptian Arabic … create an ‘illusion of success’ in training models that cannot handle natural variation.
arXiv — RW-Voice-EQ Bench (Hume AI) arxiv.org
Ground-truth references are established via a two-stage human review process; at least three independent human raters review and correct model-generated transcriptions to resolve discrepancies … the framework evaluates 15+ dimensions and 60+ metrics, including paralinguistic cues like tone, hesitation, and emotional shifts.
llms.blog — ‘Study Exposes Benchmark Overfitting and Acoustic Leakage in Speech Recognition Models’ llms.blog
Commercial APIs often report WERs below 5% on clean, read speech, yet performance typically degrades by 2.8x to 5.7x in production environments … the industry average for multi-speaker clinical conversations or high-noise emergency medical dialogues can exceed 50% WER.
AI Breaking Wire — ‘Study Finds Top AI Speech Models Replicate Benchmark Errors by 30%’ aibreakingwire.com
Both Cohere Transcribe (03-2026) and NVIDIA Canary-Qwen-2.5B were found to replicate this faulty reference verbatim, effectively ‘hallucinating’ the omission to match the benchmark’s incorrect data.
arXiv (v1 HTML) — Evaluation of LLM-based ASR / benchmark contamination study arxiv.org
Contaminated models do not always show drastically different Word Error Rates but assign significantly higher probabilities to test-set transcriptions, indicating memorization rather than generalization; sentence overlap rates exceed 75% for certain languages in Common Voice.