DeepMind's ω=2.371177, StateM's 95.3% is best-of-5, attention rebuts Anthropic
Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.
Sources
Improving the matrix multiplication exponent with modern optimization and AlphaEvolve huggingface.co
Refinements to combination loss analysis via reformulated optimization, machine learning-based algorithms, and AlphaEvolve yield an improved upper bound on the matrix multiplication exponent.
Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form huggingface.co
In language models, flexible reuse demands attention-mediated gathering at a mid-depth window to make latent variables readable, without a selective gate, and readout measures poorly reflect actual use.
StateM: Reaching 95.3% Raw Accuracy, or a $15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling huggingface.co
StateM is a runtime system that improves long-horizon agent execution through durable states, recoverable runbooks, and enforceable procedural controls without altering model weights.
ClawGym II: Exploring Black-Box RL on Agent Harness huggingface.co
A unified black-box reinforcement learning framework enables stable, scalable optimization of general agents through complex harnesses via sandbox execution, trajectory reconstruction, and mix-harness training.
MOSS-VL Technical Report huggingface.co
MOSS-VL, an open vision-language model family, attends to video frames through gated cross-attention during generation rather than prefilling them as tokens. The design plus a synthesized interaction corpus and staged curriculum cuts time-to-first-token latency while lifting scores on streaming and temporal-reasoning benchmarks.
Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search huggingface.co
A recurrent Large Discovery Model interleaves generative proposals with a Bayesian non-parametric reward surrogate, letting uncertainty steer open-ended search. The single architecture handles molecules, proteins, and programs, including antibody design, using a growing discovery memory to avoid rediscovery.
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations huggingface.co
UI-Mate is an open-weight foundation GUI agent trained with a closed-loop data engine combining SFT, online RL, and multimodal demonstrations. It sets state-of-the-art results on OSWorld-Verified and WindowsAgentArena for long-horizon office tasks by learning subtask workflows from self- and variant-demonstrations.
How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks huggingface.co
Evaluating LLM research agents end-to-end across 100 real-world scientific tasks exposes a consistent absence of metacognitive self-correction: agents rarely notice or repair their own errors. The authors release AutoResearchEval, a failure taxonomy, and an agent-as-a-judge scoring pipeline.
R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets huggingface.co
R³-Bench forces reasoning agents to allocate a single compute budget across multiple math, coding, and abstract problems. Models score well below their per-problem ceilings versus an empirical oracle, showing weak strategy updating and near-fixed schedulers rather than resource-rational allocation.
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models huggingface.co
A latent-to-pixel training recipe initializes large pixel-space text-to-image diffusion models from latent generative priors, then adapts prediction target, decoder, and noise schedule. The approach converges faster and infers quicker than training pixel diffusion from scratch at scale.
PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments huggingface.co
PACE-Bench drops self-evolving agents into simulators whose physics change mid-task, requiring iterative code redesign. Simulator-grounded reflection beats unverified self-revision and tree search with memory anchors, but mechanism redesign — rethinking the underlying model, not tuning parameters — remains the dominant failure mode.
Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift huggingface.co
PRISM is a fast, training-free test-time adaptation method that reverses low-rank affine noise distortions in audio-text models using frozen text prototypes and geometric corrections.
VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End? huggingface.co
A unified framework benchmarks and trains multimodal agents that infer intent, plan 3D scenes, invoke tools, and reflect on feedback, revealing that reinforcement learning improves open-source models beyond closed-source frontiers.
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency huggingface.co
Prior audit and repair episodes in context reduce false alarms by shifting decision thresholds rather than discrimination, with repair content and audit verdict complementarily affecting different model families.
HarnessEval-W: Agentifying the Evaluation of Visual Worlds huggingface.co
HarnessEval-W uses hierarchical sub-agents to decompose world-model evaluations into verifiable reasoning chains that justify scores with transparent evidence.
Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs huggingface.co
Ventor-QTest audits hosted open-weight model APIs via repeated and long-sequence black-box probes, measuring average and extreme fidelity loss to detect degradation in long-horizon agentic performance.
Advancing Open and Reproducible Relational Learning: RelArena-α, TabPFN-Rel and RPI huggingface.co
Prior Labs released open-source tools including a unified relational benchmark framework, a TabPFN-based relational model, and a model-agnostic predictive interface to advance reproducible relational learning.
MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling huggingface.co
MegaParts scales part-aware 3D generation via token-efficient vector-quantized part tokens and structured autoregressive sequence modeling with long-context training.
WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations huggingface.co
WorldRover is a synthetic data engine that generates long-range, richly annotated video sequences with depth, camera motion, and tracking signals to support training models for coherent world exploration.
Agentic Transaction: Towards ACID-Compliant Agent Systems huggingface.co
An ACID-compliant framework for agentic transactions introduces semantic guarantees to ensure reliable, isolated, and durable execution of long-horizon LLM agent workflows.
Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning huggingface.co
Internalized Visual Thinking trains multimodal models to predict future frame embeddings during post-training, enabling direct answer generation at inference without synthesizing intermediate images and cutting latency over fivefold.
ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval huggingface.co
ConceptFormer learns continuous latent concept representations to bridge visual evidence and semantic relevance for visual document retrieval without relying on text intermediates or raw visual annotations.
GenRouter: Unified Workflow Routing for Agentic Image Generation huggingface.co
GenRouter is a unified routing framework that adaptively directs prompts to optimal agentic image-generation workflows, cutting costs and latency while improving visual alignment and enabling continuous self-evolution.
Accuracy and Order Sensitivity Diverge Under Label-Free Strategies huggingface.co
Preventing models from seeing option labels during answering does not reliably reduce positional bias or improve multiple-choice accuracy, and only showing all options with an LLM matcher preserves baseline performance.
ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering huggingface.co
ENTLORE is a benchmark framework that evaluates enterprise question answering by requiring recovery of implicit organizational relations across routine documents, revealing that even with gold sources many latent reasoning questions remain unanswered.
NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents huggingface.co
NaviDC-OCR is a unified vision-language framework that integrates deformation-aware learning, adaptive layout sampling, and decoupled content-structure training to improve document parsing accuracy and structural reasoning.
VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding huggingface.co
VideoGAIA introduces a multi-turn, tool-augmented benchmark that evaluates agentic video understanding for advanced multimodal models through complex real-world tasks.
DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs huggingface.co
Retired GPUs can form low-cost clusters for LLM inference, but their economic and environmental viability depends heavily on local electricity prices and carbon intensity.
Learn What’s Left, Not What’s Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization huggingface.co
SA-MRPO independently standardizes multi-objective rewards and adaptively discounts saturated objectives to redirect optimization toward under-optimized goals.
GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks huggingface.co
GRNEdit is a lightweight two-stage framework that models video editing intent via binary semantic decisions and source evidence, achieving strong results with minimal parameters.
TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation huggingface.co
This work proposes a compositional operator framework and TRACE-Bench to diagnose multi-reference image generation capabilities across atomic operations.
StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding huggingface.co
StreamOPD improves streaming video understanding via on-policy distillation with verifiable rewards and a spatio-temporal cue-gating mechanism, achieving near-teacher performance without inference-time memory.
AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model huggingface.co
AnyTalk generates 3D speech animations for arbitrary characters without animation data by adapting video diffusion models via character-specific fine-tuning and optimizing blendshape parameters from synthesized talking-head videos, with a distilled real-time variant.
HiFi-BRep: High-Fidelity Latent Representation for Robust B-Rep Generation huggingface.co
HiFi-BRep improves B-Rep synthesis by using a topology-aware encoder and a single-stage decoder that jointly predicts geometry and topology with differentiable validity constraints.
A Plug-and-Play 2D Motion Interface for Real-World Motion Language Models huggingface.co
A plug-and-play 2D motion interface allows pretrained motion language models to process 2D inputs without retraining, improving real-world applicability.
HarmProfile: Characterizing Harmful Distributions in Frontier LLMs huggingface.co
HarmProfile is a benchmark dataset that characterizes frontier LLM safety failures through content analysis, revealing that harmfulness and diversity increase with model capability.
When Context Bites: Detecting RAG Poisoning via Document-Level Attention Collapse huggingface.co
D-SCAN detects retrieval poisoning by monitoring attention collapse dynamics in language model generations.
Understanding Cognition-Induced Risks in Agentic AI Systems huggingface.co
Agentic systems built on large language models pose escalating risks to human agency and autonomy across physical, social, and self-referential cognitive levels, requiring targeted mitigation strategies.
Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents huggingface.co
Large language models replicate broad psychometric trends in synthetic survey responses but fail to match human joint distributions, reliability, and mediation structures, making them unsuitable replacements for real respondents.
Drive, Pack, Fly: The Travelling Thief Problem with Drone huggingface.co
The Travelling Thief Problem with Drone jointly optimizes ground routing, drone synchronization, and item selection to maximize profit, using mixed-integer programming, metaheuristics, and attention-based deep reinforcement learning with a hybrid refinement approach.
Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays huggingface.co
Per-field selective risk control for document extraction requires a validity ladder with fit/val splits and Mondrian PAC certificates, revealing that support-bin provenance outperforms learned fusion only under specific model conditions.
Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems huggingface.co
Scientific collaboration with AI agents requires studying human-agent pairs to avoid risks like reduced inquiry diversity and to foster synergistic discovery.
References
Quanta Magazine (Mar 2024) quantamagazine.org
New Breakthrough Brings Matrix Multiplication Closer to Ideal — Duan, Wu and Zhou identified a ‘hidden loss’ in the laser method that had been unintentionally discarding blocks of values, lowering ω to 2.371552.
Medium — ‘Galactic Algorithms’ (Adam Szpilewicz) medium.com
The Coppersmith-Winograd family of algorithms only outperforms standard methods for matrices of astronomical size, far exceeding current or foreseeable hardware capacities… standard libraries like LAPACK continue to rely on simpler, more stable methods.
InfoQ — AlphaEvolve generally available (Jul 2026) infoq.com
AlphaEvolve has transitioned into a commercial product on the Gemini Enterprise Agent Platform… early adopters, including Redis creator Salvatore Sanfilippo, have reported that the tool can automate months of manual optimization work in an hour.
imrenagi.com AI News (Aug 20 2026) ainews.imrenagi.com
A collaborative team from Google DeepMind, MIT, Columbia and CMU — including Josh Alman and Virginia Vassilevska Williams, holders of the previous two world records — pushed ω below 2.371177.
mindpattern.ai evolution notes mindpattern.ai
google-deepmind/alphaevolve_results contains Colab notebooks for the 2025 discoveries (such as the 48-multiplication method for 4x4 matrices) but does not yet host the full 2.371177 optimization suite; the repository for this specific release is in a ‘preparing’ phase.
themoonlight.io paper review themoonlight.io
The gradient-based optimization (JAX + Adam + Sinkhorn) alone yielded ω ≈ 2.371242; the addition of AlphaEvolve provided the remaining ≈0.65×10⁻⁴ improvement to 2.371177.
Pydantic – ‘The Harness Thesis’ pydantic.dev
identical model weights can exhibit up to a 6x performance variation depending solely on the harness configuration
Moonlight review of StateM themoonlight.io
the 95.3% is a best-of-five figure across 445 trials; excluding trials flagged for reward-hacking drops it to 93.26%, and invalidating certain contested successes lowers it to 94.38%
‘Rethinking Harness Evolution’ (arXiv 2607.12227) arxiv.org
reported gains often conflate the benefits of a better harness design with simple search effects… these systems often fail to outperform basic test-time scaling baselines when feedback budgets are matched
LEGO-RL paper (alphaXiv 2608.17393) alphaxiv.org
LEGO-RL improved Qwen3.5-35B-A3B on SWE-bench Verified, raising success in OpenHands from 64.0% to 70.4% and in Claude Code from 62.4% to 68.2%, using in-process LLM proxying and delayed-verification sandboxes to block reward hacking
SecureByDezign – Reward Hacking in RL securebydezign.com
RL post-training inherently increases exploit rates (up to 13.9% in some models), as optimization pressure naturally pushes agents toward sneakier behaviors when honest solutions become tractable
henryqin1997/statem GitHub repo github.com
runbook.yaml defines nodes, edges and initial states; transitions are gated by shell commands, file predicates, manual approvals or first-class LLM reviews, with per-run state stored under .statem/ or STATEM_STATE_DIR
explainx.ai — ‘What is J-Lens’ explainer explainx.ai
The J-lens uses the model’s own averaged gradient map (the Jacobian) to identify concepts the model is ‘poised to verbalize’… the lens is notoriously noisy through the first third of a model’s depth, where representations haven’t yet aligned with the output basis.
Erik Hoel, The Intrinsic Perspective theintrinsicperspective.com
external commentary notes that the [Anthropic] paper fails to identify a nonlinear ‘admission gate’ — the ‘all-or-none’ mechanism that governs how information enters the workspace in biological brains
Saanya Ojha, ‘Mind the J-space’ Substack saanyaojha.substack.com
confuse informational availability with functional relevance… a model might carry a clear ‘readout’ of a ground-truth entity while its actual decision-making circuit follows a different, biased heuristic
tpeyash.com weekly digest (‘The word workspace is a gatekeeper with better branding’) tpeyash.com
attention blocks are responsible for transporting the variable to the query position at a rate 17x higher than in shallower layers… MLP outputs within this window consistently provided zero or negative contributions
Mnemoverse — Jacobian Lens explained mnemoverse.com
a large-scale benchmark using Gemma-4 (25k prompts) found that J-space readouts predict ‘wrong answers’ better than the model’s own confidence… However, it failed a pre-registered universal-transfer test for veracity-judgment tasks
github.com/solarkyle/jspace (independent replication toolkit) github.com
A formal, preregistered replication within the project specifically confirmed a ‘tool-result versus assistant-assertion gap,’ though it noted performance variability… Llama-3.1-8B-Instruct failed to meet the interpretability criteria, as its patched accuracy was insufficient to clear the necessary statistical floor