JS Wei (Jack) Sun

Claude clears 67% zeta, DiffusionGemma slips 19pts, Qwen3 leaks evicted tokens

Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.

← Back to the issue

Sources

Learning more about Claude’s mathematical capabilities anthropic.com

DiffusionGemma Technical Report huggingface.co

DiffusionGemma is a fine-tuned mixture-of-experts language model that uses discrete diffusion to generate text blocks in parallel, achieving high speed while preserving capabilities like multimodal inputs and reasoning.

Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV huggingface.co

Retained KV-cache entries can independently preserve specific past observations for long-horizon agents, defining a bounded memory contract for sparse serving.

To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing huggingface.co

Code-repair models prefer adding guard patterns over removing unnecessary lines, a bias the authors call deletion avoidance. Measured on SWE-bench Verified and a new CanItDelete benchmark, targeted post-training reduces the behavior without hurting overall repair accuracy.

GPTQ-2D: Cubic-Time Two-Sided Adaptive Rounding huggingface.co

Extending GPTQ’s adaptive rounding to two-sided matrix quantization, the method processes anti-diagonals in parallel to reach cubic runtime while producing the same output as the quartic vectorized baseline, making Kronecker-structured quantization tractable at scale.

Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures huggingface.co

The framework classifies agent failures by which component interaction broke — model, tools, or harness scaffolding — rather than lumping errors together. Annotators reach strong agreement (Cohen’s κ), letting teams target repairs across reasoning and multi-agent architectures instead of blaming the base model.

LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks huggingface.co

A manage-execute-audit loop stores verified subtask states externally, letting agents survive beyond a single context window. The AgentAdapter harness posts gains on WeaveBench, Terminal-Bench, and OSWorld, and drew 166 upvotes on Hugging Face’s daily papers.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? huggingface.co

Benchmarking self-evolving LLM agents across three streaming regimes shows no method dominates universally. Reliability of self-evolution depends on base model capability and how tasks are composed in the stream, undercutting claims that continual self-improvement generalizes across settings.

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code huggingface.co

The benchmark injects Counter-Edits from a User Patch Generator into SWE-bench Verified and SWE-Bench Pro tasks, simulating shared workspaces. Agents including DeepSWE fail to notice conflicting human changes, exposing weak workspace awareness and verification in current coding harnesses.

ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step huggingface.co

By hiding tool behaviors behind trial-and-error interaction, the benchmark isolates behavioral reasoning from documentation lookup. Agents discover mappings but then ignore them under structural changes, exhibiting belief inertia and exhaustive re-search instead of applying deductive strategies from persistent memory.

GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning huggingface.co

GradCuit improves reasoning by directly optimizing internal latent states via gradient paths through transformer layers, yielding higher accuracy and interpretability than token-based methods.

Zero-Mem: Zero-Token Memory Operations for LLM Agents huggingface.co

Zero-Mem enables structured agent memory without intermediate LLM generation by organizing interaction traces into complementary graph and temporal views for deterministic retrieval.

MemSFT: Mitigating Alignment Tax with an External Parametric Memory huggingface.co

MemSFT reduces alignment tax by using a reusable parametric memory and dynamic router to add domain expertise to LLMs without degrading general capabilities.

Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models huggingface.co

Multimodal language models retain coarse visual evidence but struggle to reliably control reliance on it versus language priors, with controllability improvable via fine-tuning and steering vectors.

SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation huggingface.co

Verified data synthesis improves language-model agents’ ability to identify, apply, and coordinate reusable skills through high-quality supervised training.

Progressive Agent Skill Generation via Reinforcement Learning huggingface.co

Skill-α uses reinforcement learning with rollback rewards to progressively generate agent skills by evaluating sequential edits against downstream task performance.

Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts huggingface.co

ReBA improves routing balance in vision-language mixture-of-experts by separately balancing image and text loads and enforcing per-image routing equality, reducing load imbalance across resolutions and tiling without harming accuracy.

DAPD: Dual-Anchored Policy Distillation huggingface.co

Dual-Anchored Policy Distillation resolves privilege illusion in on-policy self-distillation by aligning matched-information paths and reducing reliance on privileged guidance, improving performance across model scales.

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation huggingface.co

Visual Attribution Distillation uses counterfactual reconstruction to isolate visually supported corrections from mixed teacher signals, improving multimodal on-policy distillation.

WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning huggingface.co

A world critic model that jointly predicts future latent states and estimates values improves temporal reasoning and state-of-the-art performance in vision-language-action reinforcement learning.

WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity huggingface.co

WorldExam is a hierarchical benchmark that evaluates controllable video generation models on visual quality, control adherence, spatial consistency, and inherent world reactivity, revealing that current models lack consistent reactive capabilities despite high visual fidelity.

RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems huggingface.co

RecHarness automates recommender optimization by combining a bandit router for selecting modification directions with an LLM that generates executable edits, using a jump-basin mechanism to escape local stagnation.

Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge huggingface.co

Wnuan is a three-stage pipeline that adapts large language models to enterprise knowledge via supervised fine-tuning with replay and reinforcement learning on residual errors, improving domain accuracy while reducing general capabilities.

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs huggingface.co

Proposed multiple-choice planning and deferred trajectory verification reduce reasoning bias and improve verifiable autonomous driving decisions.

DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents huggingface.co

DeepVoyager-VL enables long-horizon multimodal search by integrating visual evidence into intermediate reasoning through event graphs, active visual acquisition, and supervised fine-tuning.

UEmbed: Unified Sparse and Dense Multimodal Embeddings huggingface.co

UEmbed is a decoder-only multimodal model that jointly produces dense and sparse embeddings in a single forward pass, extending sparse retrieval to unified text and multimodal inputs.

Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI huggingface.co

A model-agnostic framework evaluates per-example and per-modality failure modes when clinical modalities are dropped, revealing silent versus loud errors and complementarity across embeddings.

3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering huggingface.co

3DZip compresses 3D vision-language tokens via voxelization, diversity-based selection, and spatial merging to retain performance with far fewer tokens.

CADENA: Stepwise CAD Reverse Engineering huggingface.co

CADENA reconstructs 3D meshes into parametric CAD programs step-by-step with intermediate geometry checks and introduces a benchmark for mechanical part reverse engineering.

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space huggingface.co

SG-WAM learns geometry-aware, action-conditioned dynamics within a policy representation space using self-guided latent prediction and geometric supervision, achieving strong robot manipulation results without large-scale pretraining.

DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents huggingface.co

DreamTraj predicts 6-DoF object trajectories from a single image and instruction by decoding motion directly from intermediate representations of a frozen image-to-video diffusion model, achieving state-of-the-art results without privileged inputs.

A Frozen Pixel-Space Diffusion Model Can Guide Itself with Its Own Samples huggingface.co

Freezing a pretrained pixel diffusion model and training a lightweight head on synthetic samples enables self-guidance that improves image generation with minimal compute.

InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis huggingface.co

InfiniSplat improves single-image 3D Gaussian Splatting by aligning Gaussian primitives to scene surfaces via geometry-guided sampling and implicit decoding, yielding more coherent renderings under large viewpoint changes.

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation huggingface.co

LeapTalk enables real-time, long-form talking-head generation via single-step bridge distillation with Brownian-bridge transport, heterogeneous SNR-aligned distillation, and audio-driven guidance.

GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation huggingface.co

GEOID-Flood is a large multi-modal benchmark for flood segmentation that evaluates geospatial foundation models across SAR, optical, and elevation data, showing modest gains and improved transfer to unseen events.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks huggingface.co

SwanTale is a multi-speaker expressive speech and audio generation model that supports both instruction-based and zero-shot synthesis with high-quality multi-modality outputs.

Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis huggingface.co

Roomer is a reflective repair framework that identifies object-level layout violations and proposes validated local edits to fix them while preserving valid regions.

Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations huggingface.co

A two-stage framework transfers video motion across objects with different shapes by learning abstract dynamics and enabling direct reference-conditioned generation.

ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures huggingface.co

The Sci-ImageMiner benchmark evaluates multimodal AI on scientific figure tasks, revealing strong classification and summarization but weak data extraction and reasoning performance.

StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field huggingface.co

StyleForge selects coherent indoor furniture via a dynamic hypergraph style field and counterfactual preference learning to resolve cross-slot style conflicts without altering fixed layouts.

Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis huggingface.co

Poplar is a reproducible pipeline that synthesizes and curates human-centric image-text datasets through structured specification, realism-adapted rendering, and vision-language inspection.

References

GitHub — anthropics/zeta-23-lean github.com

The formalization is ‘sorry-free’ and rests only on Lean’s three standard axioms (propositional extensionality, classical choice, quotient soundness); it proves Theorem A (≥2/3 on the critical line), B (≥1/2 simple and on-line), C (≥3/4 distinct), D (Montgomery–Taylor constant c₁* ≈ 0.75329), and E (extension to primitive Dirichlet L-functions).

vibemathed.com — problem writeup vibemathed.com

The 41.66% (5/12) record set by Pratt, Robles, Zaharescu and Zeindler in 2020 had been a stagnation point; BGSTB’s 2024–25 work removed the RH assumption from Montgomery’s 1973 pair-correlation method, and Claude combined it with Bombieri (2000) via a ‘linear-algebraic reading’ using Sylvester’s law of inertia.

resultsense.com resultsense.com

Anthropic thanks Conrey and Goldston for ‘generously examining the paper on short notice’ — a rapid endorsement, but distinct from a formal months-long journal refereeing process.

Hacker News discussion (item 49247070) news.ycombinator.com

Commenters called the announcement ‘copium’, argued that burning 31M tokens is ‘an improper metric of intelligence’, and noted that because the checkpoint and full agent transcripts are withheld, the capability claim remains ‘anecdotal marketing’ even though the Lean proof itself is rigorous.

seanbreeden.com — 2026 AI-math survey seanbreeden.com

DeepMind’s AlphaProof Nexus solved nine Erdős problems (two open >50 years) via a Lean-scaffolded AlphaZero-style loop, and Gemini Deep Think earned a certified IMO gold — framing Claude’s zeta result as one entry in a broader 2026 wave rather than a singular event.

Reddit r/accelerate thread reddit.com

Skeptics argued Claude ‘recombined’ existing Bombieri and BGSTB machinery rather than inventing new methods, and that Sumner’s contribution was largely sending ‘keep going’ and ‘believe in yourself’ messages through 650 failed attempts.

r/LocalLLM independent benchmark reddit.com

6.7x throughput advantage for DiffusionGemma over its autoregressive counterpart, averaging 1,062 tokens per second in a local Docker-based vLLM setup on an NVIDIA RTX 6000 Blackwell

r/LocalLLaMA — ‘Diffusion Gemma is 4x faster but makes 6x more mistakes’ reddit.com

DiffusionGemma was found to make roughly six times more mistakes than its autoregressive counterpart, Gemma 4, often hallucinating smooth-sounding but incorrect names and dates

DIJA paper (arXiv 2507.11097) arxiv.org

DIJA has achieved a keyword-based ASR of 100% on benchmarks like Dream-Instruct, outperforming strong autoregressive baselines such as ReNeLLM by over 78%

‘Theoretical Benefit and Limitation of Diffusion Language Model’ (arXiv 2502.09992) arxiv.org

the number of steps required to ensure a low sequence error rate — crucial for logical chains and mathematical proofs — scales linearly with length, often negating the model’s speed advantages in reasoning-heavy contexts

Medium — ‘Diffusion LLMs are not better chatbots, they may be better text editors’ medium.com

for non-linear tasks like code refactoring, document infilling, and bidirectional constraint satisfaction, diffusion models excel because they do not suffer from the ‘early token bias’ inherent in left-to-right prediction

google/hackable_diffusion GitHub / vLLM issue reports github.com

popular local inference tools like Ollama and LM Studio do not support the specific diffusion-gemma architecture, leading to crashes… implementing dynamic per-sequence causal attention within batching frameworks like vLLM has been cited as a major engineering challenge

Xiao et al., ‘Efficient Streaming Language Models with Attention Sinks’ (arXiv:2309.17453) arxiv.org

we introduce StreamingLLM … keeping the KV of attention sink tokens (with just 4 initial tokens sufficing) together with the sliding window’s KV

LMCache blog — ‘Stop calling it KV cache, it’s something much bigger’ blog.lmcache.ai

Calling it a ‘cache’ is misleading — it behaves more like a persistent data pipeline with intrinsic semantic value that outlives the tokens that produced it.

EpiCache paper (huggingface.co/papers/2604.10098) huggingface.co

clusters conversation history into semantically coherent episodes and applies episode-specific KV cache eviction, achieving near-lossless performance at 4–6x compression

oklen/Compute-Globally-Materialize-Locally GitHub repo github.com

donor pair = two agent histories byte-identical in every served token and position but differing strictly in the value written by an earlier donor event that is excluded from the final prompt

OpenTrain.ai paper summary opentrain.ai

passive natural mentions in real dialogs (LoCoMo) yielded no detectable benefit over just re-encoding the text; the primitive is only dependable when developers deliberately phrase carrier events

Maharana et al., LoCoMo benchmark (ResearchGate) researchgate.net

evaluates AI performance over extended multi-session dialogues averaging 300 turns and 9,000 tokens … single-hop, multi-hop, and temporal reasoning

Jack Sun

Jack Sun, writing.

Engineer · Bay Area

Hands-on with agentic AI all day — building frameworks, reading what industry ships, occasionally writing them down.

Digest
All · AI Tech · AI Research · AI News
Writing
Essays
Elsewhere
Subscribe
All · AI Tech · AI Research · AI News · Essays

© 2026 Wei (Jack) Sun · jacksunwei.me Built on Astro · hosted on Cloudflare