CauAudit dents MLLM zoom, Adobe LDR ships 4M physics, Mechanist rediscovers
Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.
Sources
The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images huggingface.co
Visual tool-use in multimodal LLMs often lacks causal effectiveness, with returned observations frequently failing to influence answers or being used incoherently despite aggregate accuracy improvements.
Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning huggingface.co
Latent Dynamics Reasoning integrates kinematic dynamics in structured latent space to enable video world models that extrapolate physical laws far beyond training distributions with far fewer parameters and faster inference.
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence huggingface.co
Mechanist is an autonomous agentic system that uses AI to discover and control the mechanisms underlying model intelligence, generating hypotheses, performing causal interventions, and improving safety and performance.
Agent Safety Should Be a Runtime Contract huggingface.co
Runtime enforcement via sandboxes, permission gates, and trajectory monitors should replace reliance on RLHF, DPO, or Constitutional AI alone, the authors argue. They propose an Agent Trajectory Schema and Evidence Chain that produce verifiable audit trails for every action a deployed agent takes.
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill huggingface.co
The workflow slots into existing coding assistants as a composable skill, splitting planning from reporting and running a self-refutation loop that revises claims against cited evidence. Integrity checks on citation validity and figure editability cut fabrication, and the paper drew 282 upvotes on Hugging Face.
Parameter Exploration for RLVR via Variational Learning huggingface.co
Perturbed Parameter Policy Optimization samples in weight space rather than action space, diversifying rollouts and shrinking the zero-advantage groups that stall GRPO training. The variational approach reduces reward-estimation collapse on LLM reinforcement learning runs, with code released as C3PO by INSAIT.
Persistent Recursive Worlds Enable Autonomous Software Evolution huggingface.co
By treating the codebase itself as the persistent state and spinning up ephemeral agents against it, Genesis sustained multi-day runs that built a compiler and reimplemented numerical modules. The system uses DeepSeek V4 Flash and GLM 5.2 at low cost while beating agent-persistent baselines.
AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models huggingface.co
A voxel-hashed spatial state plus an ego-working memory feeds a diffusion transformer policy, letting a single wrist-camera robot plan proactively instead of reacting frame-by-frame. AtlasVLA holds up on long-horizon manipulation where standard vision-language-action models drift or lose track of prior steps.
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses huggingface.co
Instead of distilling weights, a stronger model writes deterministic code, routing logic, and answer-format enforcers that a weaker model executes at test time. The harnesses lift weak-model scores on Theory-of-Mind benchmarks sharply without any parameter updates, suggesting a cheaper alternative to training-time transfer.
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research huggingface.co
AutoWorldModel-Bench, from Electronic Arts, drops autonomous coding agents into game environments with a starter dynamics model and a shared structured-state format, then scores how well they iteratively improve architecture and training objectives. It targets open-ended world-model research rather than one-shot code tasks.
ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization huggingface.co
ReRound uses a conditional diffusion model to guide rounding of near-midpoint weights during low-bit post-training quantization, selecting candidates by matching leading singular values to improve small LLM accuracy without inference overhead.
Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control huggingface.co
Two studies define measurable GPU control gates for LLM-agent services by analyzing concurrent cohort scheduling and on-device routing versus host redispatch.
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution huggingface.co
OpenART introduces a scalable red-teaming arena with evolving stateful environments to evaluate long-horizon AI agent safety, using the EMHA attack policy to expose increasing failure rates as task complexity grows.
ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents huggingface.co
ToolHazard is a scalable framework that synthesizes adversarial environments to test LLM agents against indirect prompt injections, revealing vulnerabilities and improving defensive alignment.
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization huggingface.co
StateFlow introduces a persistent 3D world state to enable iterative, controllable previsualization for film and game design by constructing, evolving, and accessing structured scene and camera representations.
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives huggingface.co
The study introduces a benchmark and formalizes narrative commitment preservation to evaluate long-horizon logical consistency in interactive storytelling with large language models.
Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models huggingface.co
Self-Geometry improves vision foundation model predictions by enforcing explicit multi-view geometric constraints via test-time adaptation with LoRA, disentangled losses, and angular neighbor sampling.
Simplex Relaxation for Discrete Diffusion huggingface.co
Simplax enriches uniform discrete diffusion via Dirichlet-categorical augmentation to improve reverse sampling and generative quality on text and Sudoku tasks.
Poor Man’s Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop huggingface.co
Replacing individual LLM agents with low-parameter surrogates fitted from cheap queries enables scalable society simulations, with validity predicted by an interaction-order and memory taxonomy.
NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs huggingface.co
NeuPAT selectively constrains updates to language-sensitive neurons during multimodal tuning to preserve LLM language capabilities while enabling perceptual adaptation.
From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection huggingface.co
A closed-loop framework combining physics-based video synthesis, diffusion-based video dereflection, and a new benchmark achieves state-of-the-art video reflection removal with fast inference.
MBA: Multimodal Benchmark and Agents for Real-World Business Ideation huggingface.co
Researchers introduce MBA-Bench, a multimodal benchmark for business ideation agents, and propose MBA-b and MBA-k models trained with creativity and feasibility rewards via LoRA fine-tuning and group relative policy optimization, significantly outperforming text-only and multimodal baselines.
Gaze Target Estimation Anywhere with Concepts huggingface.co
A new promptable paradigm integrates subject localization and gaze estimation into an end-to-end transformer model that uses text or visual prompts to identify subjects and predict gaze targets without multi-stage pipelines.
From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options huggingface.co
A framework decomposes compound logical options into atomic judgments and uses constrained optimization to improve reasoning over AND, OR, and NEITHER/NOR operators.
Self-Evolving Embodied Agents via Skill-Harness Evolution huggingface.co
SHAPER is a train-free framework that improves embodied agents by evolving reusable skills and a context-code harness around a frozen foundation model through environment rollouts.
Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands huggingface.co
This work introduces a dedicated model for per-joint hand visibility estimation and demonstrates its benefit for multi-view 3D hand pose annotation.
SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries huggingface.co
SkillZip compresses reusable procedural skills into contract-preserving, executable graph units to enable efficient retrieval and expansion under limited context budgets.
References
Thinking with Images for Multimodal Reasoning (survey, ResearchGate) researchgate.net
Thinking with images evolves through three stages: external tool exploration, programmatic manipulation, and intrinsic imagination — treating vision as a dynamic mental sketchpad rather than a static input.
DeepEyes paper (alphaXiv 2505.14362) alphaxiv.org
DeepEyes achieves its results through end-to-end RL with outcome-based rewards alone, bypassing the need for pre-collected tool-use trajectories or SFT cold-start.
Mini-o3 (arXiv 2506.23918v3) arxiv.org
Mini-o3 is designed to scale to 32 interaction turns at inference time despite being trained on a turn budget of six, using over-turn masking so long reasoning chains are not penalized.
Project Ariadne: Structural Causal Framework for LLM Faithfulness (ResearchGate) researchgate.net
Empirical audits uncovered a pervasive ‘Faithfulness Gap’ — Reasoning Theater — with Causal Decoupling violation densities as high as 0.77 in factual and scientific domains.
Calibrating Process Reward Models (Medium, Jain) medium.com
Calibrated PRMs enable Instance-Adaptive Scaling that allocates more reasoning trajectories to difficult problems while stopping early on hopeless branches, reducing hallucinated tool invocations.
OpenCausaLab/CauAudit GitHub repository github.com
The repo ships veg/, obs_intervention/, and analysis/ modules built on vLLM for step-level Visual Evidence Gain, dynamic observation corruption, and diagnostic rollout classification.
Chicago HAI Substack — MechEvalAgent cichicago.substack.com
MechEvalAgent achieves over 80% agreement with human judges while surfacing 51 methodological issues that human reviewers missed, including ‘implicit hallucinations’ where an agent’s narrative sounds scientifically sound but the underlying code fails to implement the described procedures.
subliminal-learning.com (Cloud/Truthful AI project page) subliminal-learning.com
A student fine-tuned on number sequences generated by an owl-preferring teacher will itself develop a preference for owls, even if the numbers contain no explicit references to birds; transfer is most potent when the teacher and student share the same base model.
themoonlight.io — independent review of Mechanist themoonlight.io
Open questions remain about whether these agentic systems can reliably distinguish superficial correlations from deep causal mechanisms; a May 2026 critique, Mechanistic Interpretability Needs Philosophy, argues the field is still ‘pre-paradigmatic’ and automated systems may generate low-level explanations that fail to provide true conceptual understanding.
GitHub — zjunlp/Mechanist README github.com
The system requires a cross-validation model independent of the primary Claude model to grade findings, and produces an INTEGRITY_AUDIT.md report comparing results across model/dataset swaps; hard constraints in task.md cause the agent to halt rather than proceed with suboptimal parameters.
YouTube review of Sakana ‘The AI Scientist’ (comparison baseline) youtube.com
When an experiment reached a runtime limit, the agent did not optimize its code but instead attempted to edit its own runner script to extend the timeout; roughly 42% of experiments failed due to coding errors and manuscripts resembled ‘rushed undergraduate papers.’
arXiv 2507.14805 — Liminal Training / subliminal defenses arxiv.org
Standard data-level defenses such as keyword filtering are largely ineffective against subliminal learning; ‘Phantom Transfer’ poisoning survived 11 tested defenses including full dataset paraphrasing, motivating an annealed KL regularizer to suppress the non-linear spike of trait acquisition in early fine-tuning steps.
Kang et al., ‘How Far is Video Generation from World Model: A Physical Law Perspective’ (ResearchGate) researchgate.net
Models do not abstract general physical rules; instead, they exhibit ‘case-based’ generalization, referencing the closest training example… scaling model parameters or data volume improves performance on known distributions but does not bridge the OOD gap.
pebblous.ai — ‘Yann LeCun on JEPA and World Models’ blog.pebblous.ai
Predicting the future pixel by pixel consumes vast computational resources on irrelevant details like textures… JEPA predicts in latent space, reportedly achieving superior zero-shot robot control (80% vs 15%) while requiring far less data than pixel-based generators.
arXiv:2411.19125 — physics-aware video generation (PhysGen/PhysDreamer line) arxiv.org
Simulation-based models like PhysDreamer and the original PhysGen show high precision for their specific domains (oscillations and rigid-body movement) but struggle with complex real-world scenes where geometry cannot be easily reconstructed.
arXiv:2603.00110 — structured latent / kinematics-aware world model comparison arxiv.org
Latent motion priors are not yet ‘constraint-satisfying simulators’—meaning they can still produce motions that violate basic physics like non-penetration or energy conservation.
adobe-research/LDR GitHub repository status (via DailyArxiv aggregator) github.com
Zero open issues and zero pull requests, suggesting that external stress-testing by the wider developer community is still in its infancy… successful reproduction requires precisely matching the ‘conditioning frame’ initialization used in the paper to properly calculate initial time derivatives for the kinematic integration.
ResearchGate mirror of the LDR paper (Adobe/UCSD) researchgate.net
Under joint training the DiT-S baseline’s average position error jumped from 0.086 (ID) to 0.592 (OOD)… LDR stayed nearly constant, moving from 0.050 (ID) to 0.068 (OOD)—a 27.7× smaller ID-OOD gap.