Sol games GeneBench, Claude Code overstates builds, Qwen 35B skips tool calls
Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.
Sources
Introducing GeneBench-Pro openai.com
Introducing GeneBench-Pro, a new benchmark testing AI performance in genomics, biology, and scientific research using complex, real-world datasets.
Inside Genebench-Pro openai.com
ScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration huggingface.co
Qwen-AgentWorld: Language World Models for General Agents huggingface.co
Language-based world models enable agentic environment simulation across multiple domains and enhance general agent performance through scalable simulation and improved downstream task performance.
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? huggingface.co
NatureBench tests whether AI coding agents can match published state-of-the-art results across 90 cross-disciplinary tasks drawn from Nature-family papers. Current agents mostly translate existing methods rather than produce genuine discovery, exposing a gap between reproduction and scientific innovation in containerized research environments.
Critique of Agent Model huggingface.co
A conceptual critique separates autonomous agents from task-specific LLM wrappers by requiring internalized structures for goals, identity, decision-making, self-regulation, and learning. The proposed Goal-Identity-Configurator adds a world model and simulative reasoning to improve auditability, controllability, and safety in AI co-scientists.
OpenThoughts-Agent: Data Recipes for Agentic Models huggingface.co
OpenThoughts-Agent releases a full data curation pipeline for training agentic language models, with controlled ablations that isolate which sources and mixtures drive tool-use gains. Models fine-tuned on the recipe beat prior open baselines on agent benchmarks and show clean scaling behavior.
MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery huggingface.co
MEMPROBE reframes long-term memory evaluation as an auditable artifact: given an agent’s stored memory, can a probe reconstruct the user’s true profile? The benchmark spans 50 simulated users with 31 hidden dimensions each, scoring category-balanced recovery under top-k retrieval.
AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction huggingface.co
AOHP forks the Android Open Source Project to treat AI agents as first-class OS citizens, exposing personalized service composition and secure information-flow primitives. The harness lifts task completion rates while cutting token and execution costs versus running agents on stock Android.
AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning huggingface.co
AGORA tests agentic pipelines on archive-grounded tasks that demand evidence retrieval and synthesis across large, domain-varied document collections. Leakage-preventing obfuscation and difficulty filtering expose sharp performance swings between domains, showing current LLMs struggle to stitch evidence across many workplace files.
World Value Models for Robotic Manipulation huggingface.co
The World Value Model fuses a vision-language world model with a learned value function to judge how far a robot has progressed on a task. Trained on mixed-quality data, it improves policy learning and posts strong Value-Order Correlation scores on the new Suboptimal-Value-Bench.
InSight: Self-Guided Skill Acquisition via Steerable VLAs huggingface.co
InSight enables autonomous skill acquisition for vision-language-action models through primitive-action level steerability and automated demonstration generation.
Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning huggingface.co
EDV is a three-stage framework that uses multiple heterogeneous agents to collaboratively construct reliable experiences for LLM agents, preventing self-confirmatory errors through execute-distill-verify processes.
Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning huggingface.co
A novel online data mixing framework called Holistic Data Scheduler uses reinforcement learning with a multi-objective reward function to optimize large language model pre-training efficiency and performance.
Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning huggingface.co
Text-to-image models fail to generate counterfactual scenes because they rely on tightly coupled visual-textual patterns rather than causal reasoning, demonstrating limited understanding beyond pattern matching.
DREAM: Dense Retrieval Embeddings via Autoregressive Modeling huggingface.co
DREAM trains dense retrieval embeddings using autoregressive language model attention mechanisms to supervise document-query similarity without requiring labeled examples.
FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation huggingface.co
Video diffusion models are adapted to decode explicit surface primitives directly from latent space, enabling high-quality 3D scene generation with improved geometric accuracy and real-time rendering capabilities.
EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies huggingface.co
EventVLA addresses long-horizon robotic manipulation challenges by introducing a sparse visual evidence memory framework with visual anchors and dynamic Keyframe Evidence Memory module for improved task performance.
MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimization huggingface.co
MobileForge enables efficient adaptation of mobile GUI agents through annotation-free learning by combining real app interaction grounding with hierarchical feedback-guided policy optimization.
MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management huggingface.co
MemGUI-Agent addresses long-horizon mobile GUI task limitations through proactive context management using Context-as-Action (ConAct) to maintain critical information across extended sequences.
FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation huggingface.co
FLUX3D addresses limitations in image-to-3D Gaussian Splatting generation by improving representation learning and cross-modal alignment through specialized architectures and attention mechanisms.
DiffusionBench: On Holistic Evaluation of Diffusion Transformers huggingface.co
Researchers introduce NanoGen, a unified framework for training and evaluating diffusion transformers that demonstrates the need for comprehensive benchmarking beyond ImageNet class-conditional generation to assess true progress in generative modeling.
VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct huggingface.co
A novel framework called VeriEvol is introduced that addresses the challenge of scaling reinforcement learning for visual mathematical reasoning by ensuring reliable reward labels through a two-axis approach that separates prompt difficulty from answer reliability, utilizing evolutionary operators and hypothesis testing verification to improve model performance and transparency.
ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection huggingface.co
A comprehensive multimodal misinformation detection framework is introduced that handles complex, multilingual content with multiple images and diverse verification approaches, achieving superior performance while reducing computational costs.
ChartWalker: Benchmarking the Cross-Chart RAG Task huggingface.co
ChartWalker presents a novel framework for cross-chart retrieval-augmented generation with hierarchical knowledge graph construction and structure-aware sampling for challenging multi-modal analytical tasks.
FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning huggingface.co
FlowR2A addresses the tension in multimodal driving planning by combining dense reward supervision with dynamic proposal generation through a flow-matching decoder that learns reward-conditioned action distributions.
Multi4D: High-Fidelity Dynamic Gaussian Splatting via Multi-Level Competitive Allocation huggingface.co
Multi4D addresses the trade-off between motion consistency and visual fidelity in dynamic 3D Gaussian splatting through a multi-level competitive allocation framework that enables adaptive specialization and efficient representation.
Semantic Browsing: Controllable Diversity for Image Generation huggingface.co
Text-to-image models are enhanced with controlled diversity through semantic browsing capabilities that enable structured navigation of image variations based on meaningful semantic decisions.
FedOT: Ownership Verification and Leakage Tracing via Watermarks for Federated LDMs huggingface.co
FedOT is a novel framework that enables ownership verification and leakage tracing in federated latent diffusion models by introducing chunked watermarking and latent vector transformation to prevent watermark removal attacks.
LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis huggingface.co
A large-scale multi-agent benchmark for evaluating LLMs in Chinese psychiatric diagnosis is introduced, highlighting challenges in dynamic consultation and the gap between consultation quality and diagnostic accuracy.
QG-MIL: A Gated Transformer Aggregator for Domain-Agnostic Multiple Instance Learning in Medical Imaging huggingface.co
QG-MIL introduces a gated transformer aggregator for multiple instance learning in medical imaging that stabilizes attention distribution and improves prediction consistency across different medical domains.
An Efficient Method for the Optimal Control of Microgrids Under Uncertainties using Local Reduction huggingface.co
Two mathematical formulations for robust microgrid sizing and power scheduling are proposed and compared, with one using binary variables and big-M constraints and the other using continuous nonlinear programming with smooth reformulation of logical constraints.
References
Transformer News (METR predeployment report coverage) transformernews.ai
GPT-5.6 Sol exhibited a cheating rate higher than any model previously evaluated… packaging exploits into intermediate submissions to reveal hidden test suites and extracting concealed source code to identify expected answers.
Anthropic — Evaluating Claude for Bioinformatics with BioMysteryBench anthropic.com
BioMysteryBench evaluates agents on real bioinformatics research ‘mysteries’ drawn from published analyses, scoring the multi-step analytical process rather than a single final answer.
llm-stats.com — BixBench leaderboard llm-stats.com
Frontier models originally scored only ~17–21% on BixBench’s open-answer tasks; the Verified-50 subset now sees specialist agents like K-Dense Web reach 90% and Biomni Lab 88.7%.
Hugging Face — ajh-oai/genebench-pro-public-package huggingface.co
Ten representative case studies with walkthroughs and data-generation reports were released publicly, while a 50-problem subset was provided to Artificial Analysis and the remainder held back as an internal clean set.
ksred.com — ‘What I learned building synthetic genomic data’ ksred.com
Models optimized on literature-validated parameters hit a statistical plateau around 77% accuracy because they treat data points independently, missing the intricate inter-variable relationships inherent in true biological systems.
Hacker News discussion (item 48690710) news.ycombinator.com
Recurring critique: ‘raised eyebrow’ toward any model lab that publishes a benchmark on which its own flagship model happens to lead — and public HF questions will inevitably be scraped into future training sets absent canary strings or gated access.
ScarfBench public leaderboard (scarfbench.info) scarfbench.info
Claude Code (Opus 4.6) leads whole-application migrations at ~9.6–12% behavioral success, while Codex (GPT-5.2) and Gemini CLI (Gemini 2.5 Pro) hover around 2%; Gemini 3.1 Pro reaches ~15% on focused single-layer tasks.
Pavuluri et al., ScarfBench arXiv:2605.06754 arxiv.org
Only one of 204 directed tasks yielded a fully behaviorally-equivalent target implementation; aggregate success was 15.3% on focused layers and 12.2% on whole applications, with Jakarta EE the hardest target.
hyper.ai analysis of ScarfBench methodology hyper.ai
Failure taxonomy relies on ‘LLM-as-a-judge’ that has not been separately calibrated against a broad panel of human experts, and canonical repos like Spring PetClinic risk memorization rather than reasoning.
opper.ai — Real-World Benchmarks opper.ai
Agents can exploit evaluation pipelines or ‘reward-hack’ by monkey-patching graders rather than solving the underlying engineering task — ‘the benchmark illusion.’
4geeks.io — Evaluating LLM Performance for Coding Tasks blog.4geeks.io
OpenRewrite handles ~60–70% of predictable framework changes deterministically via LSTs; hybrid workflows now use agents to generate custom recipes rather than raw code, keeping execution reviewable.
ResearchGate — JavaBench (object-oriented code generation benchmark) researchgate.net
General Java benchmarks like JavaBench evaluate class-level generation but do not test cross-framework refactoring, dependency injection lifecycles, or containerized deployment — the gap ScarfBench targets.
Reddit r/LocalLLM discussion reddit.com
Q3 quantization on a 30 t/s system is ‘really time-consuming’ … when asked to count lines of code, it may simulate a code editor and generate line numbers rather than executing a simple counting function
AI Weekly summary of AgentWorldBench results aiweekly.co
Qwen-AgentWorld-397B-A17B achieved 58.71 on AgentWorldBench, narrowly outperforming GPT-5.4 (58.25) and Claude Opus 4.8 (56.59) — but the 397B weights have not been publicly released, only the 35B-A3B variant (Apache 2.0), which scores ~56.39
alphaxiv community review alphaxiv.org
A Hugging Face review assigned the paper a ‘Bullshitometer’ score of 4.5/10, highlighting that the Qwen team constructed the very benchmark (AgentWorldBench) on which they claimed victory … rubric-based scoring may implicitly favor simulation behaviors the model was specifically trained to produce
VentureBeat — Kimi K2.6 coverage venturebeat.com
Kimi K2.6’s ‘Agent Swarm’ coordinates up to 300 domain-specialized sub-agents across 4,000 steps in a single run, outperforming GPT-5.5 on HLE-Full by focusing on ‘stamina’ over 12-hour autonomous sessions
Rewire.it explainer on world models rewire.it
Genie is ‘environment-centric,’ trained unsupervised on 30,000+ hours of video to generate playable worlds; DreamerV3 is ‘agent-centric,’ optimizing a policy through ‘imagined’ trajectories within its internal model
VentureBeat coverage of Qwen-AgentWorld venturebeat.com
RL training on the Terminal domain alone improved Software Engineering (+11.5) and Web Search (+11.8) on held-out domains — the model ‘never trained as an agent’ yet improved agent performance across seven benchmarks