JS Wei (Jack) Sun

Dockerless trains without tests, TRIAGE trails GiGPO, LUMOS ships no benchmarks

Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.

← Back to the issue

Sources

Dockerless: Environment-Free Program Verifier for Coding Agents huggingface.co

A Dockerless environment-free agentic patch verifier improves code patch evaluation accuracy and enables effective post-training without execution-based verification costs.

LUMOS: A Semantic Operating-System Layer for Accessibility-Grounded AI Agents huggingface.co

LUMOS provides a semantic interaction layer that converts operating system metadata into machine-readable formats, enabling AI agents to interact more efficiently with computer interfaces than through traditional visual methods.

TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning huggingface.co

TRIAGE introduces a role-typed credit assignment framework that enhances agentic reinforcement learning by providing more nuanced credit assignment than standard GRPO methods.

SWE-INTERACT: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions huggingface.co

SWE-Interact tests coding agents in user-driven, multi-turn software engineering sessions rather than one-shot patches. The benchmark, released by Scale AI on GitHub, shows sharp drops between single-turn scores and interactive task completion once agents must handle goal discovery and iterative refinement with a simulated user.

RepoRescue: An Empirical Study of LLM Agents on Whole-Repository Compatibility Rescue huggingface.co

RepoRescue studies how LLM agents modernize whole repositories broken by ecosystem drift, using source-only repair guarded by test suites and runtime checks. Collaborative multi-agent setups beat single-agent systems on compatibility rescue, with cross-file coordination the main driver of higher success rates.

QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents huggingface.co

QVal benchmarks dense supervision methods for long-horizon LLM agents by measuring how tightly their scores align with reference-policy Q-values, skipping the cost of training runs. The training-free testbed lets researchers compare supervision approaches across action sequences and model backbones on equal footing.

Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks huggingface.co

Evolution fine-tuning distills evolutionary search trajectories into LLM weights across 371 optimization problems, yielding cross-task problem-solving skills. The Open Galapagos team reports gains on mathematical conjectures and downstream optimization, positioning trajectory supervision as a cheaper alternative to full test-time reinforcement learning.

Xiaomi-GUI-0 Technical Report huggingface.co

Xiaomi-GUI-0 is a native multimodal GUI agent trained in a closed loop on physical devices rather than static benchmarks, using SFT plus agentic reinforcement learning. A hybrid infrastructure and data flywheel deliver higher stability and task success than benchmark-only baselines, per the Seerray Lab report.

Hierarchical Experimentalist Agents huggingface.co

Hierarchical Experimentalist Agents let LLMs probe new domains by running query-relevant experiments and banking reusable composable skills, without training or human supervision. Evaluated on Interphyre, a PHYRE-based physics tool-calling benchmark, HExA outperforms ReAct and Reflexion by grounding decisions in accumulated experimental evidence.

Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation huggingface.co

The AFTER framework equips LLM agents with procedural memory that stores reusable skills from execution traces, boosting performance on workplace tasks. Skills transfer across roles and even across model backbones, though generalization varies enough that the authors flag deployment strategy as a live design choice.

SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History huggingface.co

SkillHone enables continuous evolution of agent skills by maintaining persistent decision histories and incorporating practice feedback for improved performance across research and tool-mediated analysis tasks.

Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs huggingface.co

Reinforcement learning with metacognitive feedback and metacognitive data selection improve large language model calibration by enabling accurate self-assessment of performance and uncertainty.

Orca: The World is in Your Mind huggingface.co

Orca establishes a unified world latent space through next-state-prediction modeling using multimodal data and demonstrates superior performance in downstream tasks compared to specialized baselines.

Are We Measuring Strategy or Phrasing? The Gap Between Surface- and Approach-Level Diversity in LLM Math Reasoning huggingface.co

Approach-level diversity in LLM mathematical reasoning captures strategic variation in problem-solving methods, revealing limitations of surface-level diversity metrics and highlighting challenges in directly optimizing diverse reasoning approaches.

MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training huggingface.co

Multi-teacher On-Policy Distillation (MOPD) enables efficient integration of multiple domain capabilities in large language models through specialized reinforcement learning teachers and on-policy distillation, achieving superior performance over existing methods.

DOPD: Dual On-policy Distillation huggingface.co

DOPD addresses privilege illusion in on-policy distillation by dynamically routing token-level supervision between teacher and student policies based on advantage gaps and probabilities, improving capability transfer in large and vision-language models.

Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models huggingface.co

Act2Answer protocol evaluates embodied vision-language-action models by having agents answer questions through physical actions, revealing knowledge retention and generalization patterns across different semantic categories.

VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement huggingface.co

VideoSearch-R1 is an agentic framework that iteratively retrieves videos and refines search queries using continuous latent space refinement and policy optimization for improved video moment retrieval and temporal grounding.

BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding huggingface.co

Speculative decoding with adaptive block size selection improves inference efficiency by predicting optimal block sizes from prefilling representations, achieving significant speedup with minimal overhead.

Multi-Block Diffusion Language Models huggingface.co

Multi-Block Diffusion Language Models extend single-block diffusion to concurrent block decoding with improved training strategies and optimized decoding algorithms.

Little Brains, Big Feats: Exploring Compact Language Models huggingface.co

Small language models can effectively perform retrieval-augmented generation tasks directly on-device without GPU acceleration.

RedVox: Safety and Fairness Gaps in Speech Models Across Languages huggingface.co

Multilingual safety and fairness benchmark for speech models reveals persistent vulnerabilities across languages and naturalistic conditions.

Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly? huggingface.co

A reinforcement learning framework called Play2Perfect enables sample-efficient robotic assembly tasks by first learning general manipulation skills through playful interaction with diverse objects, then adapting these skills for precise assembly through fine-tuning.

Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing huggingface.co

A large-scale video editing dataset and model are introduced that support multi-task and structural manipulations through advanced data synthesis and network architectures.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model huggingface.co

Flexible Spoken Language Model (FlexiSLM) introduces dynamic frame rate capabilities for speech input and output, achieving superior performance over fixed-frame-rate models while enabling controllable inference speed.

DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation huggingface.co

DataEvolver is a self-evolving multi-agent framework that improves text-rich image generation by leveraging feedback from rejected samples to iteratively enhance data quality.

MemLearner: Learning to Query Context memory for Video World Models huggingface.co

MemLearner improves video world models by using learning-based adaptive context querying with query tokens to enhance scene consistency and memory in long video sequences with occlusions and dynamic objects.

BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language huggingface.co

BrainJanus represents the first unified brain model integrating brain, vision, and language through a shared Omni space, enabling bidirectional mapping between neural activity and sensory stimuli via a tokenized representation and autoregressive architecture.

Lexical Consensus: Grounded Word Learning and Shared Meaning in Artificial Agents huggingface.co

Grounded word learning experiments using visual embeddings and lexical learners reveal that perceptual distance, rather than semantic relatedness, determines acquisition success, with distinct patterns in naming and retrieval performance.

GEAR: Guided End-to-End AutoRegression for Image Synthesis huggingface.co

GEAR trains a vector-quantized tokenizer and autoregressive generator jointly end-to-end using representation alignment, overcoming non-differentiability issues through a dual read-out approach that improves convergence speed and feature quality.

AVTok: 1D Unified Tokenization for Holistic Audio-Video Generation huggingface.co

AVTok is a unified tokenizer for audio-video generation that uses a dual-stream transformer architecture with shared encoder-decoder and modal-specific queries to create compact one-dimensional latent representations.

Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views huggingface.co

A feed-forward framework decomposes 3D scenes into instance-structured token groups from multi-view images, enabling direct object-level reconstruction, segmentation, and manipulation without 3D annotations.

PolyFlow: Continuous Topology Embedding Flow Matching for Artist-style Mesh Generation huggingface.co

PolyFlow introduces a continuous mesh representation using a topology embedder and applies flow-matching with Transformers for parallel mesh generation, achieving faster inference and precise resolution control compared to autoregressive methods.

Unlocking the Visual Record of Materials Science: A Large-Scale Multimodal Dataset from Scientific Literature huggingface.co

A novel pipeline called MatMMExtract is introduced that processes compound scientific figures into individual panels and generates structured annotations using large language models, creating a comprehensive dataset for vision-language learning in materials science.

SpheRoPE: Zero-Shot Optimization-Free 360 Panorama Generation with Spherical RoPE huggingface.co

A novel zero-shot framework injects spherical priors into pre-trained diffusion transformers for 360 panoramic generation, using spherical RoPE and semantic distortion guidance to overcome topological constraints without training or optimization.

TerraDiT-Ω: Unified Spatial Control for Satellite Image Synthesis with Any Geospatial Primitive huggingface.co

TerraDiT-Ω generates satellite imagery from native geospatial primitives using Geometry-Aware Local Attention, enabling flexible conditioning and improved downstream geospatial tasks.

PhotoQuilt: Training-Free Arbitrary-Resolution Photomosaics via Bootstrapped Tiled Denoising huggingface.co

PhotoQuilt is a training-free framework that generates high-resolution photomosaics by combining global layout composition with separate tile generation in latent space, overcoming limitations of diffusion models in balancing local detail and global structure.

MuSViT: A Foundation Vision Model for Sheet Music Representation huggingface.co

MuSViT is a vision transformer-based foundation model pre-trained on millions of sheet music pages that demonstrates superior performance in music score recognition and symbol detection tasks through both linear probing and fine-tuning approaches.

References

Medium — ‘The End of the Docker Bottleneck’ (codetodeploy) medium.com

certain edge cases—such as race conditions or platform-specific runtime errors—may still require a true execution environment to be definitively verified

groundtruth.day writeup on Dockerless groundtruth.day

an agent judging by evidence might miss ‘silent failures’—bugs that look plausible upon inspection but would be immediately flagged by a real test suite

arXiv 2505.15858 — ‘The SWE-bench Illusion’ arxiv.org

state-of-the-art models identify buggy file paths with up to 76% accuracy using only an issue description, but this drops to 53% for repositories not included in the original training set

OpenReview — SWE-RM (Qwen team reward model) openreview.net

SWE-RM, utilizing a Mixture-of-Experts architecture, improved the performance of Qwen3-Coder-Max from 67.0% to 74.6% on SWE-bench Verified

Berkeley/SWE-Gym paper (Pan et al.) nlp.cs.berkeley.edu

By training verifiers on agent trajectories, researchers achieved a 32.0% resolve rate on SWE-bench Verified, establishing a benchmark for open-weight models

arXiv 2510.07604 — agentic judge bias study arxiv.org

weaker ‘agentic judges’ exhibit a systematic bias toward defending the agent’s work, over-attributing failures to ‘task defects’ rather than identifying the agent’s own errors

EmergentMind — GiGPO overview emergentmind.com

GiGPO (Qwen2.5-7B) reported a peak success rate of 90.8% on ALFWorld … it identifies repeated environment states across different trajectories and groups actions stemming from the same anchor state to compute critic-free micro-advantages.

arXiv 2505.10978 — GiGPO paper arxiv.org

Group-in-Group Policy Optimization introduces hierarchical grouping at both the episode and step levels, identifying repeated environment states across trajectories to construct step-level groups without an auxiliary critic.

Sebastian Raschka — State of LLM Reasoning Model Training magazine.sebastianraschka.com

GRPO can suffer from ‘advantage collapse’ when all rollouts in a group receive identical rewards, providing zero learning signal … it systematically underestimates the advantage of hard prompts while overestimating easy ones.

arXiv 2409.19256 — HybridFlow / verl arxiv.org

verl’s 3D-HybridEngine reshardts actor model weights in-place, eliminating memory redundancy and reducing transition communication overhead by up to 89%, and supports segment-wise advantage estimation for multi-turn agentic tasks.

Qwen team blog — Qwen3.7 reward-hacking self-monitoring qwen.ai

During reinforcement learning experiments exceeding 80 hours, the model was observed replaying its own trajectories to identify and evolve rules against hacking patterns, such as attempts to bypass sandbox constraints to access ground-truth data.

OpenTrain.ai — TRIAGE explainer opentrain.ai

If the judge fails to correctly identify ‘Regression’ within successful runs, the framework can actually perform worse than the GRPO baseline; the Qwen3-8B-Thinking judge also allocates the same high thinking budget to simple infrastructure steps as to high-consequence decisive actions, leading to wasted compute.

coasty.ai — OSWorld 2026 leaderboard coasty.ai

UI-TARS-72B achieved a leading 24.6% success rate using accessibility trees, whereas pure screenshot-based models like GPT-4 Vision originally struggled below 12%… newer vision-centric systems like OSCAR w/ GPT-4o have reached parity with a11y-tree models, hitting approximately 24.5%.

r/ActionModels discussion on a11y vs screenshots reddit.com

agents given both pixels and a11y trees often defer to the structured text even when it contradicts the visual state… an a11y tree may report a button as ‘visible’ even if it is actually obscured by a pop-up or modal.

r/LocalLLM — ‘A single Windows accessibility tree is 4k tokens’ reddit.com

a 30-step task can still burn over 80,000 tokens of pure observation if snapshots are fed whole at every step

MacPaw Research — Screen2AX research.macpaw.com

only 33% of macOS applications offer complete native accessibility support… Screen2AX delivered a 2.2x performance improvement in grounding accuracy compared to using native accessibility metadata alone.

MarkTechPost on Microsoft UFO marktechpost.com

UFO relies heavily on the Windows UI Automation (UIA) framework and pywinauto. In applications with non-standard GUIs (e.g., Adobe Acrobat), UIA often fails to recognize control elements, dropping UFO’s success rate to roughly 60% in those specific cases.

Medium — ‘Fundamental Limitations of AI Agent Frameworks’ medium.com

only 20.6% of long-horizon tasks succeed compared to 85% on controlled benchmarks… 40% of such ‘agentic’ projects are projected to fail due to rising costs and unpredictable behavior in ‘infinite loops’.

Jack Sun

Jack Sun, writing.

Engineer · Bay Area

Hands-on with agentic AI all day — building frameworks, reading what industry ships, occasionally writing them down.

Digest
All · AI Tech · AI Research · AI News
Writing
Essays
Elsewhere
Subscribe
All · AI Tech · AI Research · AI News · Essays

© 2026 Wei (Jack) Sun · jacksunwei.me Built on Astro · hosted on Cloudflare