Prime Agent hits 95.5% on ARC-AGI-3, AX-RAY catches 192/192, FORGE fools 27%
Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.
Sources
Prime Agent: A Self-Improving RLM Harness huggingface.co
Prime Agent is an open-source harness that uses recursive subagents, persistent computation, and agent-to-agent coordination to extend language models’ long-horizon capabilities across coding and reasoning tasks.
The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models huggingface.co
The study proposes a lightweight audit to detect causality violations in sequence models by verifying prefix invariance, revealing that attention-mask checks miss leaks from scans or normalization.
One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders huggingface.co
Search-augmented LLM recommenders are highly vulnerable to web content polluted by generative engine optimization, frequently promoting fake products despite reasoning and defenses.
What AstroPT knows about galaxies, and what that can teach us about LLMs huggingface.co
AstroPT, a transformer trained on galaxy images, encodes real astrophysical properties recoverable by linear probes on frozen checkpoints. Concepts such as redshift, band magnitude, and specific star formation rate emerge in a consistent difficulty-based sequence across training runs, offering a testbed for mechanistic interpretability lessons that transfer to LLMs.
Apodex 1.1: Scaling Agentic Intelligence for Complex Work huggingface.co
Apodex 1.1 targets sustained progress on complex real-world work by expanding executable training environments and adding an AgentOS execution harness. Agents learn long-horizon task decomposition, asynchronous integration, and state recovery through trajectory training, aiming for verifiable completion rather than plausible one-shot responses.
From Generation to Simulation: How Far Are World Models from Being True Simulators? huggingface.co
Generative world models are measured against traditional physics engines across eight capabilities, exposing persistent gaps in physical guarantees, state feedback, and long-horizon stability. Diffusion and joint-embedding approaches show gains in controllability and interaction, but the study argues cross-route hybridization is needed to reach true simulator fidelity.
ReWorld: An Interactive World Model with Long-Horizon Memory huggingface.co
ReWorld trains short-horizon control and long-horizon memory separately, then bounds compute at inference using mixed per-head attention windows and a pose-indexed landmark KV bank. Distribution-matching LoRA distillation delivers real-time playback with strong action fidelity and long-range recall across palindrome trajectories.
Tomatoes, Potatoes, and Onions: Questioning the Need for Faces in Face Presentation Attack Detection huggingface.co
Training presentation attack detectors on images of tomatoes, potatoes, and onions rather than faces yields transferable representations that improve cross-dataset AUC. The result suggests foundation-model PAD relies on generic presentation cues like frequency artifacts, not facial content, easing privacy and licensing constraints on training data.
One Success Isn’t Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows huggingface.co
Thinkingbox is a sandbox and benchmark from Microsoft that scores agents on consequential, multi-turn business tasks requiring policy adherence, dependent tool coordination, and correct persistent state transitions. The framing rejects single-run success as a reliability signal, pushing evaluation past code repair and simple tool-call benchmarks.
The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search huggingface.co
A leave-one-out causal probe isolates evidence-utilization bottlenecks in retrieval-augmented generation, quantifying structural dilution across a deconfounded factorial grid. An iterative submodular scheduler and attribution-steered contrastive decoder then orchestrate context in a closed loop, lifting portfolio recall during sequential generation.
LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures huggingface.co
LongRCA Bench evaluates failure diagnosis across lengthy agent trajectories, and the training-free RCTA method improves attribution of responsible roles and root-cause steps.
GameXpert-Bench: How Far Are Coding Agents from Expert Game Development? huggingface.co
GameXpert-Bench evaluates coding agents across three game development stages—generation, repair, and optimization—using interactive and behavioral tests to reveal strengths in building playable foundations and weaknesses in defect discovery and regression preservation.
Same Agent, Different Answers: A Repeat-Aware Audit of Corpus-Induced Answer Churn in Retrieval-Augmented QA huggingface.co
Retrieval-augmented QA systems can exhibit hidden answer churn during index updates without noticeable accuracy changes, motivating compatibility audits alongside utility evaluations.
TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration huggingface.co
TileMix routes attention score tiles to mixed FP16 or INT8 precision within fused dense attention, recovering long-context accuracy while improving prefill throughput without retraining.
Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection huggingface.co
Task-CoEvolve improves LLM harness optimization by adaptively selecting validation tasks and estimating full-set performance from partial evaluations, cutting evaluation costs by 80%.
AutoResearch: Insight In, Hallucination Out huggingface.co
AutoResearch is a two-stage autonomous system that grounds research ideas through integrated generation and evidence-based execution to improve experimental reliability and measurable outcomes.
Better Retrieval, Worse Robustness:How Multi-hop RAG Amplifies Upstream ASR Errors huggingface.co
Retrieval-augmented generation extensions amplify automatic speech recognition errors in spoken multi-hop question answering, primarily through corrupted query entities.
MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks huggingface.co
MobilePA-Bench is an interactive sandbox benchmark that evaluates mobile planning agents on tool-calling, sub-agent collaboration, memory usage, and composite skill invocation under real runtime constraints.
Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization huggingface.co
ERPO replaces action-side policy regularization with input-side query distribution control to stabilize reinforcement learning for language models while preserving response exploration.
LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks huggingface.co
EvoMap externalizes verified execution experience into reusable structured Gene, improving long-workflow task completion and reducing token costs across diverse models.
ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts huggingface.co
ClawProBench evaluates agent configurations via execution traces across live and frozen tracks, revealing that final-answer rankings obscure native-runtime failures and process-quality differences.
Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress huggingface.co
R2-OPD improves on-policy distillation by filtering teacher rewards that conflict with reasoning progress via within-trajectory ranking comparisons.
ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction huggingface.co
ARC improves fairness in group-based reinforcement learning for open-ended agents by conditioning rollout comparisons on strategy, enabling more context-appropriate behavior in responsive user-agent interaction.
WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning huggingface.co
WorldToken fuses heterogeneous robot observations into per-timestep world tokens processed by a causal Transformer and diffusion action head, with scaling and temporal-context analyses on RoboCasa and RMBench.
EchoWM: Open and Enterable Omnimodal World Models huggingface.co
EchoWM is an omnimodal world model that generates synchronized high-resolution video, sound, music, and speech while following continuous 6-DoF navigation trajectories across first- and third-person views.
RISE: Adaptive Imagination for World Action Models huggingface.co
RISE adaptively decides when to continue or stop imagination rollouts for planning by weighing expected benefit against cost, supported by a counterfactual driving dataset with expert annotations.
Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision huggingface.co
A hierarchical taxonomy and dense supervision strategy improve diffusion-based image editing through fine-grained concepts, large-scale paired data, and granular evaluation.
RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA Modeling huggingface.co
RIBOSPAN is a large bidirectional RNA foundation model pretrained on up to 10,240 nucleotides that enables high-resolution full-transcript modeling, strong long-context representations, and discrete-diffusion-based mRNA generation and redesign.
Towards a Densing Law for User Representation Learning at Billion-Scale Capacity huggingface.co
The study proposes a scaling law linking data size to tokenization capacity for user behavior modeling and introduces an adaptive tokenization method to improve efficiency.
Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion huggingface.co
Block3D accelerates text-to-3D generation by using block-wise diffusion with confidence-guided correction to reduce inference time while preserving geometric fidelity.
EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment huggingface.co
EXPL-FR explains face recognition similarity scores by aligning vision-language embeddings to the recognition space, enabling label-free auditing of semantic attributes and model comparison.
TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming huggingface.co
TLive-Omni is an omni-modal model for live-commerce that unifies image, video, audio, and text via timestamped token grouping, staged supervised training, and reinforcement fine-tuning with verifiable feedback to enable accurate real-time understanding.
Industrial technical reports contain high-value knowledge for maintenance, troubleshooting, and product engineering, but their heterogeneous structure (dense prose, specifications, tables) makes them difficult to index and reason over with standard retrieval and QA pipelines, and no public instruction-tuning or benchmark datasets are built from such documents. We address this gap with Industrial-Instruction, contributing (i) two open QA datasets built from real industrial technical reports and (
Hybrid Quantum-inspired Kolmogorov-Arnold Networks for Privacy-Aware Federated Biosignal Learning huggingface.co
A hybrid quantum-inspired Kolmogorov-Arnold network improves federated ECG classification accuracy with fewer parameters and lower communication costs than standard multilayer perceptrons.
Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs huggingface.co
Quantization-aware healing recovers compressed 4-bit language models faster and more stably than quantization-aware training by distilling directly from the original uncompressed model.
References
NVIDIA Developer Blog — AVO on ARC-AGI-3 developer.nvidia.com
NVIDIA AVO reaches 100% on ARC-AGI-3, demonstrating a frontier-level general-purpose architecture for long-horizon autonomous agents.
AI Weekly — Prime Agent v0.8.1 stability complaints aiweekly.co
Users reported the harness frequently crashes within minutes; the IPython kernel fails and the TUI produces ‘empty sessions’ that cannot be deleted, prompting a proposed redesign of the supervisor from a message broker to a registry system.
rlm.md — Alex Zhang (MIT CSAIL) rlm.md
Instead of cramming millions of tokens into a fixed context, the root model treats the prompt as an external variable in a Python REPL and decomposes it via grep/partition/map/reduce, calling subagents whose full histories never flood the parent’s context.
Towards AI — ‘AGI is not a compute problem’ pub.towardsai.net
OpenAI showed GPT-5.6 Sol jumping from 13.3% to 38.3% with no weight changes, purely by retaining private reasoning and swapping truncation for context compaction — the harness, not the model, is doing the work.
ThursdAI — ARC Prize eligibility notes thursdai.news
The 95.5% was achieved on the public set with Claude Opus 5 via API; official ARC Prize rules require a no-internet Kaggle environment, so Prime Agent sits on the ‘Community’ leaderboard rather than as a verified prize contender.
dev.to — Codex vs Claude Code 2026 benchmark dev.to
On SWE-bench Verified, Claude Code (Opus 5) reaches ~97.0% while Codex (GPT-5.5) trails at ~88.7%, but Codex is 3–4x more token-efficient per task ($8.39 vs $11.84 on DeepSWE).
dev.to (VIDRAFT team writeup) dev.to
AX-RAY validates this by performing two forward passes on inputs that are identical except for the very last token… the tool can identify the exact point where causal leakage occurs — meaning information from the ‘future’ has incorrectly influenced earlier computations.
AI Weekly — ‘Zamba2 and Nemotron-H fail new hybrid-model causality audit’ aiweekly.co
The defect specifically affects the PyTorch-based ‘slow path,’ which is triggered whenever optimized fused kernels are unavailable, such as in CPU environments, Continuous Integration (CI) pipelines, or stock library installations.
OpenReview — Mamba2 causality bug (Issue #700) openreview.net
Mamba2 outputs for the same prefix differed when the full sequence was processed versus a truncated version… suggesting future leakage, where information from subsequent tokens erroneously influences the current state.
AI Weekly — editors’ blog on the two-pass audit aiweekly.co
In the same 192 trials, the industry-standard check (inspecting the attention mask) detected zero of the faults (0/192)… While gradient-based methods matched the localization accuracy, they were significantly more expensive in terms of memory and compute.
Daum News (Korean coverage) v.daum.net
AX-RAY is now being integrated into government-level cybersecurity initiatives to provide a more robust ‘weights-level’ audit for foundational models.
ResearchGate — FINAL-Bench: Measuring Functional Metacognitive Reasoning researchgate.net
FINAL-Bench uses a 5-axis rubric to evaluate procedural error recovery, identifying it as a critical bottleneck in modern LLMs.
arXiv — PoisonedRAG lineage paper arxiv.org
By surgically crafting just five malicious documents, researchers achieved a 90% attack success rate against databases containing millions of texts.
Level Up Coding — ‘Your RAG is not the same as my RAG’ levelup.gitconnected.com
In controlled N=1 poisoning tests, success rates plummeted from 81.9% in vanilla RAG to 24.4% in Recursive Language Models (RLM).
Business Insider businessinsider.com
AI Overviews incorrectly attributed negative competitor reviews to nascent legitimate businesses, such as ‘The Plastics Shed,’ causing significant reputational damage while the business was concurrently paying for Google Ads.
ALM Corp — Lily Ray study summary almcorp.com
Google might cite a brand’s self-promotional listicle as a source, [but] it refuses to recommend that brand 69% of the time, opting instead for verified third-party competitors.
Zenodo replication — skepticism prompting evaluation zenodo.org
Instructing a model to ‘distrust unfamiliar brands’ or ‘be skeptical of potential pollution’ increased attack success rates by an average of 10.5 percentage points across several models.
Fast Company fastcompany.com
One fake webpage can be enough to trick AI shopping recommendations.