Intern-S2 leads science evals, AICode refactors 189 files, Microsoft loops state
Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.
Sources
Intern-S2-Preview: Scientific Agentic Foundation Model huggingface.co
Intern-S2-Preview is a scientific agentic foundation model series that integrates multimodal pre-training, multi-task reinforcement learning, and memory-augmented extensions to support long-horizon scientific reasoning and forecasting.
This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated code and no pre-existing oracle to validate the target behaviour. The task, dismantling a central invariant across a large interdependent codebase, was assessed by the author as effectively infeasible through incremental refactoring, the kind of change that conventionally calls for a rewrite instead
Full-bandwidth transformer huggingface.co
Full-bandwidth transformers use latent feedback of top-layer hidden states to improve reasoning and efficiency without altering the core architecture.
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus huggingface.co
Massive activations in layer-interleaved hybrid linear-attention LLMs form pre-attention spikes with inter-spike plateaus, driven by cancellation timing in gated deltanet output gates. The morphology recovers as layers approach the full-attention limit, giving a systematic outlier analysis for hybrid architectures.
DarwinX: Evolving Agent Harnesses Through Natural Selection huggingface.co
DarwinX treats agent scaffolding as an evolving population, recombining harnesses around frozen base models and keeping winners in an archive. The approach lifts verified benchmark scores without benchmark-specific patches, showing self-improvement can happen outside the weights.
Thought-Level Beam Search for Reasoning huggingface.co
Gambit prunes and expands partial reasoning trajectories at the thought level, steering fixed hardware budgets toward promising traces. The scheme improves large reasoning models’ test-time compute scaling by keeping GPUs saturated with parallel samples instead of long single chains.
Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning huggingface.co
CaRL applies reinforcement learning with refusal incentives and hindsight augmentation so LLMs quit reasoning when a problem exceeds their capability. The recipe cuts specious chain-of-thought and miscalibration without hurting scores on tasks the model can actually solve.
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design huggingface.co
AutoDesign wraps a code-writing agent in a meta-harness optimizer that recursively rewrites its own DesignHarness from rollout feedback. Applied to structured media generation, it sets state of the art on paper-to-poster synthesis, a long-horizon design task.
LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time huggingface.co
LiveAnimate drives pose-conditioned human video from a 14B diffusion transformer in real time, using teacher-forcing adaptation, self-forcing distillation and pose-retrieval sink attention. Ulysses sequence parallelism and operator fusion keep long-form streams stable without drift.
Maglev: Sliding Recurrent Memory huggingface.co
Maglev pairs sliding-window attention with a fixed-size recurrent K/V memory, jointly trained via a prefiller-decoder setup with parameter sharing and a memory consistency loss. Long-context quality improves while inference cost drops and training stays parallel.
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review huggingface.co
Rhetorical framing significantly biases AI scientific review scores in structured ways, with effects shaped by reviewer identity, score range, and evaluation strictness rather than rewriting complexity.
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers huggingface.co
LLM routing is formalized as a sequential decision process with a unified benchmark and modular infrastructure to compare and improve cost-effective model selection.
From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs huggingface.co
Researchers propose a black-box red-teaming method using inaudible low-frequency waveforms to expose vulnerabilities in audio-language models, alongside a defense that detects distribution shifts and requests a second recording to recover accuracy.
Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing huggingface.co
HPSE improves unstructured knowledge editing by distilling from hybrid rollouts that insert missing facts into the model’s reasoning paths, enabling composable multi-hop reasoning.
Are You Sure You’re Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity huggingface.co
Instruction tuning changes model confidence and reduces rationale diversity without improving calibration, indicating distinct effects on reasoning and certainty.
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence huggingface.co
A frozen vision-language model improves spatial reasoning by self-evolving through verified experience, reflection, and reusable memory retrieval without parameter updates or external tools.
LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation huggingface.co
LycheeMemory V2 improves long-term agent memory by batching interactions into semantic segments for efficient consolidation and structured retrieval, reducing construction costs while maintaining high accuracy.
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World huggingface.co
Evoke is an interactive world model that uses external persistent memory and a redesigned long-horizon teacher to enable responsive, open-ended video generation with bounded context and low latency.
SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models huggingface.co
SKILLER is a reinforcement learning framework that automatically generates tailored skills for small open-source models to reduce inference costs while maintaining high task performance.
An AI4AI Framework for Visual Token Pruning huggingface.co
AutoPrune uses large language models to automatically design visual-token pruning policies for multimodal models via a domain-specific language and residual search formulation, achieving high efficiency with minimal performance loss.
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation huggingface.co
DreamX-Phi 1.0 is an action-conditioned video world model for robotic manipulation that uses geometric attention encoding, depth estimation, object masks with a frozen teacher, and distillation to generate faithful future observations.
Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation huggingface.co
Context-Matched Distillation aligns teacher supervision with causal generation context for few-step autoregressive video models, improving control adherence and long-video quality.
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos huggingface.co
UniSwap enables synchronized appearance and voice replacement in talking videos through a unified streaming audio-visual diffusion transformer with specialized training and inference adaptations.
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives huggingface.co
PlayWorld benchmarks interactive video world models by using multi-modal agents to pursue long-horizon objectives, evaluating geometry consistency, interaction fidelity, and state evolution.
H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models huggingface.co
H2R-Bench evaluates video generation models on transforming human manipulation videos into robot-centric demonstrations across embodiment constraints and interaction fidelity.
PixSDS: Why Latent SDS Makes Noisy Pixels huggingface.co
PixSDS fixes VAE-induced pixel drift in latent score distillation sampling by guiding optimization with decoded image directions, reducing artifacts in text-to-3D generation.
RibAssist 3D: Biplanar Rib-Fracture Detection, Addressing, and Selective 3D Localization from CT-Derived Projections huggingface.co
Cross-view pairing of rib fractures in orthogonal CT projections enables accurate 3D localization, but correspondence confidence remains the primary bottleneck.
Mitigating Gender Bias in English to Romanian Machine Translation huggingface.co
A hybrid pipeline combining LLM gender classification with tag-aware neural translation improves gender accuracy in English-to-Romanian machine translation.
CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers huggingface.co
CW-BASS v2 selects pseudo-labels by measuring teacher reliability on held-out data and applying either strict filtering or an adaptive floor to avoid confirmation bias under saturated confidence.
TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement huggingface.co
TailBooster uses a dual-layer generative framework with statistical tail extraction and deep autoencoder cleaning to synthesize operationally valid extreme air-transport events, substantially improving extreme-event prediction accuracy.
References
interfaze.ai model card interfaze.ai
Mixture-of-Experts (MoE) architecture with a total of 397 billion parameters, of which approximately 17 billion are active per token, mirroring the efficiency of its Qwen3.5 heritage
Hugging Face — Intern-MemDec-4B model card huggingface.co
Memory Decoder is not a standalone model and cannot be used for independent chat tasks; it requires a compatible inference runtime and a specific router configuration tailored to the backbone-memory pair
arXiv — GEPO paper (2607.16850) arxiv.org
GEPO attenuates positive advantages in low-entropy groups to prevent the model from over-exploiting narrow, safe reasoning paths [and] reduces the weight of negative advantages in high-entropy groups, which preserves exploration
r/LocalLLaMA discussion reddit.com
self-hosting the full-precision model … requires approximately 8x H100 GPUs (794GB VRAM), whereas quantized versions can fit on 2-3x H100s … unusually robust to aggressive low-bit quantization, remaining functional at 4-bit or even ternary precisions
arXiv — Intern-BioBreaker red-team study (2510.03255) arxiv.org
widespread bio-risk jailbreak vulnerabilities in frontier models of this class … some attacks reaching a 100% success rate in bypassing standard text-level safeguards
StartupFortune analysis startupfortune.com
‘task scaling’ philosophy, where Shanghai AI Lab emphasizes the diversity and difficulty of training data over raw parameter counts … the 35B version achieves performance comparable to the trillion-scale Intern-S1-Pro
bizstack.tech bizstack.tech
34,770 insertions and 16,422 deletions landing in just two commits … total inference cost of $2,430 over three days
Dark Factory Dev darkfactory.dev
reporting omits certain platform-side migration issues that forced the author to estimate rather than precisely count specification corrections
Matsuoka HyperDev blog hyperdev.matsuoka.com
models remain ‘goal-seeking’ rather than ‘problem-solving’, occasionally attempting to bypass constraints with TypeScript
anytypes or by commenting out failing tests to reach a ‘finished’ state
codemyspec.com (Spec Kit vs Kiro) codemyspec.com
AWS Kiro enforces a strict EARS notation with a built-in SMT solver that mathematically proves requirements are free of contradictions or gaps before any code is generated
Snorkel AI — Senior SWE-bench snorkel.ai
replaces static oracles with a validation agent and a taste judge … scores code on minimality, hygiene, and fluency relative to repository standards
Codacy blog blog.codacy.com
up to 40% of AI-generated code contains security vulnerabilities … agents that write both refactor and tests often ‘grade their own homework’
TheMoonlight review of Full-Bandwidth Transformer themoonlight.io
the model cannot be retrofitted onto existing ‘vanilla’ checkpoints … it requires a specialized ‘multi-pass’ training objective—where a small fraction of training batches (e.g., 3%) use deeper feedback passes to ensure numerical stability over long sequences
AI Weekly alert on FBT efficiency claims aiweekly.co
if the baseline standard transformer were trained with the same total FLOP budget as the full-bandwidth version, it might achieve similar or superior results, potentially making the ‘efficiency’ claim a measurement artifact
labml.ai Feedback Transformer implementation notes nn.labml.ai
a Feedback Transformer with 126M parameters matched the perplexity (18.3) of a 257M parameter Transformer-XL … this typically results in training speeds 5 to 10 times slower than standard models
Meta Coconut paper (arXiv 2412.06769) arxiv.org
continuous thoughts can encode multiple alternative next steps simultaneously, [enabling] the model to explore several paths without prematurely committing to a single deterministic word, outperforming traditional Chain-of-Thought in tasks requiring complex planning and backtracking
EmergentMind analysis of FBT (2608.08888) emergentmind.com
instruction-tuning datasets are ‘off-policy’ relative to latent-feedback decoding … the model is forced to re-imitate this wordy style, negating the inherent efficiency of the widened latent channel
AIcerts news commentary aicerts.ai
there is no formal theoretical proof that the learned mapping is ‘globally contractive’ … leaves open the possibility of local instability, limit cycles, or divergence during long-horizon generation where small errors might compound