MIT's modules vanish in GPT-2, Dion3 needs Hopper, RLVR trades tasks
Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.
Sources
Modular Cognitive Architecture Emerges in Large Language Models huggingface.co
Large language models develop modular neural architectures that mirror human brain specialization across language, reasoning, and physical cognition, suggesting modularity is a fundamental property of intelligent systems.
Dion3: Full-Stack Orthogonal Updates huggingface.co
Dion3 accelerates the Muon optimizer by reducing orthogonalization and communication overhead through algorithmic, kernel-level, and update-rule improvements.
Verifier-Induced Support Reshaping in On-Policy Optimization huggingface.co
On-policy reinforcement learning with verifiable rewards can improve immediate task performance while reducing the diversity of successful responses needed for future training, a phenomenon called verifier-induced support reshaping.
Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models huggingface.co
Deliberative behaviors like self-correction and hypothesis testing get boosted most by reasoning training, yet correctness-linked traits such as confidence calibration barely improve. The Behavioral Lift analysis exposes a gap between what thinking models perform and what actually predicts right answers across vision-language tasks.
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning huggingface.co
Mobius-v0 routes global knowledge into FFN memory modules and iterative reasoning into self-attention, letting Intern-S2-Mobius match baseline performance with less training data and quicker inference. The decoupled design targets compositional reasoning without paying the usual dense-model compute tax.
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems huggingface.co
Committees of language-model agents making clinical decisions fall for shortcuts that sound socially reasonable rather than isolated spurious cues. Only an independent referee agent reliably catches the cascade, suggesting benchmark-gaming defenses need external oversight instead of peer deliberation among the deciding agents.
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development huggingface.co
Systematic evaluation across long-horizon AI R&D tasks finds autonomous agents strong at engineering optimization yet unstable in solution framing, weak on methodological novelty, and inconsistent at reusing prior experience. AutoResearchEval scores execution and feedback control separately from final task metrics.
Scaling Domain Data Repetition in LLM Pretraining huggingface.co
When model size and token budgets scale proportionally, the best number of passes over high-quality domain data rises only slowly with scale. Repetition count tracks domain validation loss more tightly than the volume of unique data, reshaping pretraining mix decisions.
Marionette: Predicting World States, Rendering Geometry, Painting Appearance huggingface.co
Marionette factors interactive game generation into an autoregressive 3D articulated world-state predictor, a zero-parameter geometry renderer, and a video-diffusion appearance model. The split enables direct state-level control and long-horizon consistency repair, improving FVD over end-to-end video models.
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe huggingface.co
During the gaps where ReAct agents wait for tool observations, Second Thought spawns auxiliary reasoning branches in parallel with the main decode. The training-free scheme cuts sequential steps and turn counts on agentic benchmarks while keeping Pass@1 accuracy intact.
Claim-Level Reliability Assessment for Efficient Test-Time Reasoning huggingface.co
Claim-Level Reliability Assessment improves reasoning accuracy by verifying critical claims instead of sampling more solutions, reducing token use while boosting performance.
Multimodal Model Diffing for Feature Discovery and Control huggingface.co
MMDiff uses multimodal sparse autoencoders to isolate, detect, and control specific features in multimodal language models, improving interpretability and targeted steering of visual and safety behaviors.
Latent On-Policy Self-Distillation huggingface.co
Latent On-Policy Self-Distillation learns privileged teaching context end-to-end from experience to provide dense token-level supervision, improving agent performance and efficiency.
UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers huggingface.co
UNMASK automatically discovers and mitigates spurious correlations in text classifiers via causal verification and group-based reweighting without manual annotations.
UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations huggingface.co
UniProbe is a lightweight learnable detector that uses a directed graph and alternating GNN, ViT, and GRU modules to identify hallucinated tokens in frozen large vision-language models, enabling real-time resampling during generation.
Is this Citation on Point? huggingface.co
Large language models reliably detect incorrect case citations but frequently miss wrong pinpoint page references, conflating topical relevance with precise legal support.
Self-Supervised Visual On-Policy Distillation huggingface.co
Self-supervised visual on-policy distillation improves small vision-language models by distilling from original images into strongly augmented student views without privileged annotations or larger teachers.
LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure huggingface.co
A curated elementary-grade pretraining corpus and 5B-parameter model create a controlled sandbox for studying knowledge acquisition, representation, and bounded capability growth via post-training and in-context learning.
DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data huggingface.co
Mimir v1 is a 1-billion-parameter Hierarchical Reasoning Model trained solely on permissible data that achieves competitive English results and state-of-the-art Danish performance across multiple benchmarks.
SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning huggingface.co
On-policy distillation from a long-context reasoning teacher to short-context students improves mathematical proof reasoning and generalizes to science benchmarks by aligning token spans, constraining length growth, and stabilizing training.
SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation huggingface.co
SPARGen unifies 3D reconstruction, dense correspondence, and spatial reasoning into a single instruction-conditioned multimodal generative model that jointly learns shared spatial representations.
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence huggingface.co
Apodex Discovery introduces a framework for verifiable, extended AI investigations using a heavy-duty solver and structured evaluation across real-world scientific tasks.
Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction huggingface.co
GAS improves multimodal understanding by using generation as auxiliary supervision via next embedding prediction and a decoupled mixture-of-transformers architecture, with no inference overhead.
PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment huggingface.co
PRM-as-a-Judge 1.5 provides fine-grained process metrics and reliability tools to evaluate embodied robotic models beyond binary success rates.
Forecast Collapse in Time-Series Foundation Models huggingface.co
Forecast collapse in hourly equity return prediction stems from low predictability and per-series objectives, and the proposed CalibRank objective balances calibration and ranking to restore cross-sectional structure.
MobileMem: Learning from a Year of Mobile Experiences huggingface.co
MobileMem is a benchmark and framework for evaluating on-device long-term memory through year-scale, multimodal mobile experience trajectories that require temporal reasoning, knowledge updating, and preference inference.
HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark huggingface.co
HumanTracker introduces a large-scale benchmark and preference-aligned metric to evaluate humanoid motion tracking based on perceptual quality and physical contact stability.
Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings huggingface.co
ParliamentRAG is a retrieval-augmented generation system for Italian parliamentary records that uses topic-dependent speaker authority to retrieve expert perspectives and generate faithful, multi-perspective summaries.
CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing huggingface.co
CPI-Bench is a comprehensive benchmark for real-world image editing that evaluates multi-image tasks, practical applications, and reasoning-based editing to better differentiate model performance.
A new benchmark for AI-generated video detection reveals that current detectors fail to generalize across realistic crisis-related videos and become less reliable as content spreads socially.
A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images huggingface.co
The ALD/E-ImageMiner benchmark and ICDAR 2026 competition advance machine interpretation of scientific figures through tasks spanning visual reading, domain reasoning, and evidential justification, proposing long-term goals for verifiable multimodal scientific AI.
Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead huggingface.co
A 3B-parameter agentic model using a Looped Transformer requires bug fixes and a chunked-prefill strategy to run reliably on Apple Silicon and complete agentic tasks.
References
AlKhamissi & Tuckute — ‘The LLM Language Network’ project page bkhmsi.github.io
Applying functional localizers to 18 LLMs identified units selective for language over math or code; ablating these units causes a dramatic collapse in language modeling while leaving reasoning intact.
Semantic Scholar — AlKhamissi, Tuckute et al., ‘The LLM Language Network’ (NAACL 2025 precursor) semanticscholar.org
Language-selective units emerge in LLMs and are causally necessary for linguistic behavior, but analogous ‘Theory of Mind’ and ‘Multiple Demand’ networks are far less cleanly separable than the language module.
ResearchGate — ‘Empirical Limits of Neuron-Level Ablation for AI Safety’ researchgate.net
Ablation studies can ‘lie’ because of the Hydra Effect / self-repair: downstream components compensate for silenced neurons, and polysemantic MLP units encode multiple unrelated features, making ‘surgical’ single-neuron ablation an interpretability illusion.
NYU News — ‘Does the brain work like an LLM in predicting words?’ (2026) nyu.edu
LLMs predict words from immediate local context while the human brain predicts by grouping words into hierarchical grammatical constituents — a structural mismatch even where surface behavior converges.
AI Weekly — ‘LLMs mirror the brain’s modular neuron layout, study finds’ aiweekly.co
The paper’s framing is ‘unusually direct’ in asking whether modularity is a biological accident or a convergent requirement of intelligence; modularity was absent in GPT-2, appearing only once models reach reasoning competence.
alphaXiv discussion of Han et al. (2608.13567) alphaxiv.org
The study establishes that modularity emerges, but does not demonstrate that modular organization improves raw task accuracy compared to monolithic processing — leaving the functional payoff of specialization open.
Tri Dao blog — ‘Gram Newton-Schulz’ tridao.me
Gram Newton-Schulz provides a 40–50% reduction in runtime for the orthogonalization step compared to the standard Newton-Schulz routine, yielding up to a 2x speedup on the total optimizer step while staying within 0.01 validation perplexity of the baseline.
Microsoft Research — original Dion paper page microsoft.com
Dion uses amortized power iteration to orthonormalize only a low-rank subspace of the momentum matrix, computing updates without ever reconstructing a full parameter matrix on a single device — retaining synchronous semantics under FSDP/TP.
Hugging Face blog — ‘Scaling is not plug-and-play’ huggingface.co
Naively applying Muon to multi-billion-parameter models often results in ‘paralysis’ (vanishing updates) or ‘explosions’ (loss spikes and NaNs); stability requires weight decay and per-parameter update-scale corrections that early Muon variants lacked.
NorMuon paper (arXiv 2504.05295) arxiv.org
Original Muon effectively reduces condition numbers but leads to non-uniform neuron norms; NorMuon adds neuron-wise second-moment normalization and outperforms Muon by 11.31 percentage points on 1.1B-parameter pretraining.
Emergent Mind — Dion3 explainer emergentmind.com
The complexity tax is real: Dion3’s headline speedups depend on custom CuTeDSL kernels and ‘megabatching’ communication patterns, and some practitioners note this makes it harder to integrate into legacy pipelines than plain AdamW.
alphaXiv — Dion3 discussion page alphaxiv.org
For square matrices (α=1) Gram Newton-Schulz is FLOP-identical to standard Newton-Schulz and can be slightly slower in wall-clock due to extra kernel launches; the library falls back to standard NS with symmetric kernels in this regime.
Fu et al., ‘Scaling Reasoning, Losing Control’ (MathIF, arXiv 2505.14810) arxiv.org
as models are scaled or fine-tuned specifically for reasoning — such as through distilled long chains-of-thought — their ability to follow user-specified constraints often degrades
ReasonIF GitHub (Kwon et al.) github.com
many ‘state-of-the-art’ models fail to follow reasoning instructions more than 75% of the time… highest Instruction Following Scores (IFS) often remain below 0.25
Yue et al., ‘Does RLVR Really Incentivize Reasoning Capacity?’ (arXiv 2606.15455) arxiv.org
base models often maintain a higher pass@256 or pass@1024… for the hardest problems requiring novel strategies, the RL-tuned model becomes less likely to stumble upon the correct answer than its original base version
‘Battle for Entropy: RL Algorithms & LLMs’ (gopubby practitioner blog) ai.gopubby.com
collapse is not uniform but is triggered by ‘premature overconfidence’ at a small subset (~5%) of structurally critical decision points
Robust Policy Optimization (FRPO) — ResearchGate 400622099 researchgate.net
optimizing rewards across a ‘KL-bounded neighborhood’… ensures that a model’s math accuracy remains stable even when it is subsequently adapted for new downstream instruction sets
Wilson Wu, ‘PPO vs GRPO’ (practitioner blog) wilsonwu.me
GLM-5.2 switched from GRPO back to PPO to achieve ‘qualitative improvements’ in training controllability and generalization