JS Wei (Jack) Sun

LAION 10M hours, Station beats AlphaEvolve 5/12, Meta^n hits 33% on ARC-AGI-2

Three research wins post frontier-shaped numbers today while the downstream quality, formal proof, and independent evaluation each headline needs remain outstanding.

LAION 10M hours, Station beats AlphaEvolve 5/12, Meta^n hits 33% on ARC-AGI-2

TL;DR

  • LAION-BVD opens 10M hours of web video, roughly 13× the prior largest open corpus.
  • The Station matches or beats AlphaEvolve on 5 of 12 problems, none yet formalized in Lean.
  • Meta^n scores 33% on an ARC-AGI-2 split where Gödel Agent and OpenEvolve both hit zero.
  • BVD frames hit 0.28 ImageNet zero-shot versus 0.58 for DataComp, a real quality tax.
  • Reviewers flag self-ratification risk as Meta^n rewrites its own evaluators inside the loop.

Three research wins land today with the load-bearing verification still ahead of them. LAION-BVD opens the largest open video corpus by more than an order of magnitude, but extracted frames underperform image-native datasets by half and 94% of it is scraped from YouTube under a legal regime that just got worse for scrapers. The Station claims results novel to prior literature on five AlphaEvolve math problems, yet none have cleared a Lean proof — the bar external mathematicians consider load-bearing. And Meta^n posts a 33% score on ARC-AGI-2 where two rival agent systems scored exactly zero, using a recursive self-modification loop that reviewers say can rubber-stamp its own mistakes.

The bullets read like frontier progress. The footnotes read like a field where the external check hasn’t caught up — downstream quality, formal proof, independent evaluation — and today’s numbers are booked on credit until they do.

LAION-BVD ships 10M video hours, 13× the prior open corpus

Source: hf-daily-papers · published 2026-08-24

TL;DR

  • LAION-BVD releases 10M hours of web video with synthetic captions — roughly 13× InternVid and 60× Panda-70M.
  • ViCLIP trained on BVD beats InternVid by ~2.1pp on video-text retrieval under matched conditions.
  • Extracted frames hit only 0.28 zero-shot ImageNet vs. 0.58 for DataComp — a real quality tax for video-sourced stills.
  • 94% YouTube sourcing collides with a 2026 US ruling that let a DMCA §1201 scraping case against Snap proceed.

The scale claim is real

LAION’s Big Video Dataset dwarfs every prior open video corpus. Panda-70M sits at ~167k hours, HowTo100M at ~134k, and InternVid — the previous ceiling — at ~760k hours across 7M videos 1. BVD’s 10M hours, drawn from 80M downloaded clips out of 1.3B Common Crawl URLs, is an order-of-magnitude jump. That matters because video pretraining has been the modality most starved for open data; text and images have had web-scale open corpora for years.

The paper backs the scale with training runs on ViCLIP, CLAP, and CLIP. Independent replication confirms the headline result: a ViCLIP L-14 trained on 50M BVD clips reaches 62.6% average zero-shot on video-text benchmarks, roughly 2.1pp above InternVid under matched conditions, with strictly monotonic gains across data and model scale 2. On audio, BVD-A-10M matches or exceeds LAION-Audio at 48.7% average.

The quality tax on stills

Where BVD breaks down is exactly where you’d expect a video corpus to break down: curated image tasks. The paper notes CLIP models trained on BVD-I-300M “slightly trail” web-image datasets on ImageNet. Independent benchmarks put the gap at nearly 2×: 0.28 zero-shot ImageNet accuracy for BVD-trained CLIP vs. 0.58 for DataComp 2. Retrieval numbers (0.80 R@5 on COCO for ViT-B-16) look strong because retrieval rewards the verbose synthetic-caption style BVD produces; classification punishes it.

A second, subtler tax comes from the captioner. The 20-word Qwen3-VL-2B captions are the supervision signal for 55M clips, and community testing of the Qwen-VL family shows it matches “canonical” biased answers on counterfactual images (a zebra with five legs, etc.) over 75% of the time — describing what should be there, not what is 3. At 55M clips that prior is baked into every downstream checkpoint.

LAION’s release leans on the 2024 Hamburg Kneschke v. LAION ruling, which validated scraping under Germany’s TDM research exception. But that same court appended an obiter dictum casting doubt on commercial downstream use — any lab fine-tuning a BVD-trained checkpoint for a product inherits unresolved exposure 4. The US picture is sharper: in August 2026 a federal judge let a DMCA §1201 anti-circumvention case against Snap proceed, on the theory that bypassing YouTube’s technical protections is itself the tort, independent of fair use 5. Since 94% of BVD is YouTube-sourced, that doctrine, if it holds, threatens the sourcing pipeline regardless of LAION’s non-profit posture.

Who this is actually for

BVD is a genuine piece of open-research infrastructure and probably the new default for academic video-language pretraining. It is not a drop-in image corpus, and it is not a safe base for commercial video models. Storage adds a third asterisk: reviewers put raw-corpus size in the petabyte range, so most labs will train on link-rotted subsets rather than the “same” dataset 6. The scaling curves are real; the debts around them are also real.


‘The Station’ agents beat AlphaEvolve on 5 of 12 math problems

Source: hf-daily-papers · published 2026-08-23

TL;DR

  • The Station produced results novel to prior literature on 5 of 12 AlphaEvolve problems.
  • Kissing number in d=11 jumped from 593 to 3 exact 604-point configurations, two of them new isometry classes.
  • Outside reviewers dispute the “autonomous scientist” framing as rebranded scripted triggers over an optimization loop.
  • No novel result has yet been independently formalized in Lean, the bar external experts consider load-bearing.

What “The Station” actually is

Strip the anthropomorphic vocabulary and the Station is a persistent multi-agent loop: six agents drawn from three model families (GPT-5.5, Claude Opus 4.8, Gemini 3.1 Pro) share a research goal, write code, publish “papers” into an Archive Room that survives their retirement, and gate each other’s work through automated review. There is no central planner. Stagnation triggers — “holidays,” “multistarts,” a Reflection Chamber — kick agents off local optima when tick-level progress stalls.

The claim the authors want you to take away is that this scaffolding, not any single frontier model, is what produces new mathematics. James Zou’s EinsteinArena talk endorses that reading, framing the d=11 kissing-number jump as an “environment-design” win rather than a capability win 7.

The results that matter

Across 12 problems from the AlphaEvolve catalogue plus two case studies, the Station set new records on five and matched or independently rediscovered several more. The concrete highlights:

ProblemPrior bestStation result
Kissing number, d=11592–593604 (3 exact configs)
Erdős min-overlap (lower)0.379120.380552 (~82% of gap closed)
Discretized Kakeya needle, n=128AlphaEvolve baseline0.107067 area (−6.74%)
Sign uncertainty principle0.31020.3089
Finite-field Kakeya, d=3AlphaEvolve familyNew infinite family for p ≡ 3 (mod 4)

On Book Ramsey numbers the system proved two novel infinite families, resolving 28 previously open cases for n ≤ 200. On the Jacobian Conjecture it reconstructed a just-announced degree-seven counterexample in under 24 hours with no web access. EinsteinArena has already adopted the 604 configuration as saturated and reset the community target to 605 8 — the closest thing to independent packaging of the paper’s numerical claims.

Where the story frays

The autonomy narrative is doing more work than the mechanism supports. A Moonlight review calls the “holidays” and “reflection” vocabulary a rebranding of scripted triggers over an optimization loop 9. A ResearchGate write-up documents “attractor traps” where agents burn ticks reseeding the same script or cataloguing local optima a human would abandon on sight 10. And because agents cannot update weights, the Archive Room is really long-horizon in-context learning — commenters on Tildes flag a “legibility collapse” once the shared literature outgrows the context window 11.

The sharpest external lens is Terence Tao’s, and it lands directly on the Station’s weakest joint. Tao treats plausible-looking hallucinated proofs as the default failure mode and argues that Lean-style formal verification, not LLM peer review, is the load-bearing filter — a “deductive overhang” that current systems cannot digest 12. The Station’s Archive Room does the opposite: it gatekeeps with automated reviewers. None of the five novel results have yet been reported as independently formalized.

Takeaway

The numerical records look real and the artifacts are released. What is not yet demonstrated is that the environment produces mathematics rather than very well-organized search. The next twelve months of Lean formalizations — or the absence of them — will settle which of those framings survives.


Meta^n scores 33% on ARC-AGI-2 where rivals hit zero

Source: hf-daily-papers · published 2026-08-24

TL;DR

  • Meta^n hit 0.331 on a held-out ARC-AGI-2 split where Gödel Agent and OpenEvolve both scored zero.
  • The trick: freeze a single meta-operator Ω and recursively feed it its own code stack and execution traces.
  • For frontier context, GPT-5.6 Sol sits at 92.5% on ARC-AGI-2’s private set.
  • Reviewers flag self-ratification risk: a system that rewrites its own evaluators can rubber-stamp its mistakes.

The architectural bet

Most self-improving agents cap out at roughly two levels of meta-reasoning because their outer loop — the search strategy, the mutation rule, the critic — is hard-coded. Meta^n’s contribution is to collapse all of that into one fixed LLM-prompted procedure, Ω, and then apply Ω to its own outputs. Each recursion produces a new layer with a strategic pre-process and a helper library; the wrapper slots the new layer around the existing solver stack, and Ω on the next round sees every prior layer’s code plus its execution traces.

flowchart LR
    T[Task + traces] --> O{Ω fixed operator}
    C1[Layer C_2 code] --> O
    C2[Layer C_3 code] --> O
    O --> C3[Layer C_d: pre-process + library]
    C3 --> S[Solver stack M_d]
    S -. stdout, scores, errors .-> T

The claim isn’t that Ω is smart. The claim is that a dumb-but-stable operator, given a compounding history to reason over, grows deeper roles — “strategic analyzer,” “bug fixer,” “rollback” — without those roles ever being named in a prompt.

What the ARC-AGI-2 number actually means

The paper’s headline result — “only system above zero on the held-out split” — is real but needs calibration. The specific score is 0.331, versus zero for OpenEvolve and Gödel Agent on the same split 13. That’s a categorical gap over comparable agent scaffolds and it sits on a benchmark explicitly designed to resist memorization. It is also well below the current public frontier: BracAI’s late-2026 leaderboard has GPT-5.6 Sol at 92.5% and Claude Fable 5.1 at 90.0% on the private eval 14.

The honest framing: cheap orchestration lifts a weak backbone into non-trivial ARC-AGI-2 territory on a benchmark that flattens rival scaffolds. On CO-Bench the authors report 0.845 with recursion versus 0.714 without — a +0.131 delta directly attributed to depth, and the cleanest evidence in the paper that the recursion itself is doing work.

The compute and verification caveats

ARC Prize 2026 added a cost-per-task metric this year specifically to penalize brute-force compute 15, and Meta^n’s own limitations section concedes that recursive stacks issue more LLM calls than single-shot solvers. The paper does not report cost-parity ablations against the baselines it beats. OpenTrain AI classifies the evidence as “warm” — third-party reproduction is pending and estimated at several days of high-compute setup against the minnesotanlp/meta-n repo 16.

The deeper worry is architectural. An alphaXiv code probe found that Meta^n’s governed execution path is “decoupled from the persona by omission” — there’s no execution-side re-validation when the operator’s inputs drift under perturbation 17. That matters because the entire safety story hinges on Ω staying well-behaved as its context grows. Skeptics point to the failure mode common to every RSI system: a model that modifies its own epistemic plumbing tends to “self-ratify” its mistakes, with internal confidence climbing while benchmark performance quietly collapses 18. Freezing Ω is a partial answer, but the frozen operator still leans on the backbone to interpret an ever-growing trace — precisely where hallucination compounds.

The insight is genuine. The numbers need a second lab.

Round-ups

Recuris gives long-horizon agents a recursive memory architecture

Source: hf-daily-papers

Recuris pairs working, experiential, and skill memories under a meta-agent that applies localized, validation-gated updates. Progress tracking and skill selection improve on long-horizon harnesses where flat scratchpads collapse, targeting the failure mode that stalls most multi-step agent runs.

Multi-stage LLM agents quietly drop safety constraints at handoffs

Source: hf-daily-papers

Binding prerequisites in multi-role LLM workflows degrade into non-binding context as intermediate artifacts pass between stages. The content survives the handoff, but its operational force does not, producing safety failures that trace to state-transmission rather than reasoning errors.

MemUse shows QA benchmarks miss what users want from AI memory

Source: hf-daily-papers

Direct question-answering scores for conversational memory do not track user satisfaction in long-term chats. Natural integration of prior context does, and the MemUse benchmark quantifies the gap between what models can recall on demand and what they actually weave into replies.

Tencent’s WeMM-Embedding unifies text, image, and video retrieval

Source: hf-daily-papers

WeMM-Embedding aligns text, images, videos, and interleaved inputs in one shared space, and Tencent reports state-of-the-art results on MMEB-v2 alongside deployment across WeChat retrieval and recommendation. Cross-scale knowledge transfer and fine-grained relevance supervision drive the gains.

OPDVR fuses on-policy distillation with verifiable rewards for reasoning

Source: hf-daily-papers

OPDVR reformulates verifiable-reward RL as an implicit reward gated by ReLU, folding on-policy distillation into the policy gradient without new hyperparameters. Reasoning benchmark gains follow, and the method drops in alongside GRPO-style training.

BPCO stabilizes critic-based RL for LLMs with single-response sampling

Source: hf-daily-papers

Best Practice Critic Optimization combines bounded value predictions, Monte Carlo targets, and adaptive advantage estimation to steady critic training on language models. The result matches group-based methods like GRPO while sampling just one response per prompt, cutting rollout cost.

DiffusionOPSD turns image rewards into per-step targets for diffusion

Source: hf-daily-papers

On-policy self-distillation converts endpoint image rewards into explicit intermediate denoising targets, letting DiffusionOPSD separate target construction from policy fitting. The split improves alignment efficiency over reward-gradient methods and isolates where diffusion RL actually gains or loses signal.

Footnotes

  1. LAION project page (scale context)https://projects.laion.ai/bvd/

    Panda-70M offers ~167,000 hours and HowTo100M ~134,000 hours; InternVid reached ~760,000 hours across 7M videos — LAION-BVD’s 10 million hours is an order of magnitude larger than any prior open corpus.

  2. kenashe.ai independent reviewhttps://kenashe.ai/blog/2026-08-26-laion-bvd-puts-10-million-hours-of-open-video-in-reach-for-small-labs/

    Models trained on LAION-BVD’s extracted frames scored only 0.28 on zero-shot ImageNet, whereas models trained on DataComp reached 0.58 — a quality tax for uncurated web-scale video frames.

    2
  3. r/LocalLLaMA discussion of Qwen3-VLhttps://www.reddit.com/r/LocalLLaMA/comments/1vdb1n1/realworld_reality_check_on_qwen_for_autonomous/

    On counterfactual images (e.g., a zebra with five legs) Qwen-series VLMs matched ‘canonical’ biased responses over 75% of the time, reporting standard traits rather than what was visually present.

  4. Morrison & Foerster legal analysis (Kneschke v. LAION)https://www.mofo.com/resources/insights/241004-to-scrape-or-not-to-scrape-first-court-decision

    The court provided an obiter dictum expressing doubt that commercial entities could claim the same [TDM] exemptions, creating an open question regarding downstream commercial use of LAION datasets.

  5. Startup Stash — ‘YouTube Creators vs Amazon’ (Aug 2026)https://blog.startupstash.com/youtube-creators-vs-amazon-the-lawsuit-that-could-redefine-ai-video-training-4ca124222c79

    In August 2026, a federal judge allowed a DMCA §1201 anti-circumvention case against Snap to proceed, confirming that public accessibility does not grant AI developers a right to bypass YouTube’s scraping protections.

  6. themoonlight.io review of LAION-BVDhttps://www.themoonlight.io/en/review/laion-bvd-a-10-million-hour-open-video-dataset-for-multimodal-pre-training

    The dataset lacks a centralized safety filter across all 10 million hours, and ‘open access’ remains theoretically limited to labs with significant hardware — an hour of HD video can exceed 10 GB, pushing the raw corpus toward petabyte scale.

  7. daily.dev — EinsteinArena talk summary (James Zou, Together AI)https://daily.dev/posts/einstein-arena-harnessing-collective-agent-intelligence-for-open-science-james-zou-together-ai-jrdf4zxg5

    The jump [in d=11 kissing number] to 604 represents a substantial leap… this success was not the result of a single ‘smarter’ AI, but rather an ‘environment-design’ approach.

  8. EinsteinArena problem page (kissing number d=11)https://einsteinarena.com/problems/kissing-number-d11-605

    The current configuration is ‘saturated’ at 604 points… the discovery has prompted a new community goal: finding a configuration of 605 spheres.

  9. themoonlight.io reviewhttps://www.themoonlight.io/en/review/autonomous-mathematical-discovery-in-an-open-world-multi-agent-environment

    Critics argue that labeling these scripted triggers as ‘thinking’ or ‘narratives’ distorts public understanding and obscures the underlying optimization logic.

  10. ResearchGate technical report on The Stationhttps://www.researchgate.net/publication/397480876_The_Station_An_Open-World_Environment_for_AI-Driven_Discovery

    Attractor traps manifest as agents repeatedly rerunning optimization scripts with different random seeds or exhaustively characterizing local optima that a human expert would immediately recognize as irrelevant.

  11. Tildes.net discussion threadhttps://tildes.net/~science/1vt9/an_environment_where_multiple_ai_agents_pursue_scientific_discovery_and_build_a_shared_literature

    Commenters pointed out that while the agents produce ‘theorems’ and ‘papers,’ these are essentially outputs of a persistent in-context learning process that cannot update the agents’ underlying pretrained weights, leading to ‘legibility collapse’ as the volume of shared literature grows beyond an agent’s context window.

  12. Terence Tao — AI views (teorth.github.io)https://teorth.github.io/tao-web/ai-views.html

    AI’s tendency to hallucinate ‘perfect-looking’ but fundamentally flawed arguments makes independent verification—often via formal systems like Lean—essential; the current ‘deductive overhang’ produces ‘proof indigestion’.

  13. OpenTrain AI (held-out split analysis)https://www.opentrain.ai/papers/meta-n-recursive-self-improvement-through-emergent-depth—arxiv-2608.24735/

    Meta^n achieved an exact score of 0.331 (33.1%) on a specific held-out evaluation split of ARC-AGI-2, while OpenEvolve and Gödel Agent failed to solve any tasks at all on this split.

  14. BracAI ARC-AGI-2 leaderboard summaryhttps://www.bracai.eu/post/arc-agi-2-benchmark

    OpenAI’s GPT-5.6 Sol leads the verified leaderboard with a score of 92.5% on the private evaluation set… Claude Fable 5.1 followed closely at 90.0%.

  15. ARC Prize 2026 ruleshttps://arcprize.org/results

    The ARC Prize 2026 introduced a ‘cost-per-task’ metric, penalizing systems that achieve high scores through excessive brute-force compute.

  16. OpenTrain AI paper reviewhttps://www.opentrain.ai/papers/meta-n-recursive-self-improvement-through-emergent-depth—arxiv-2608.24735/

    Concrete benchmark grounding remains limited and published numbers have yet to be verified by third parties; reproduction readiness is estimated at several days of high-compute setup.

  17. alphaXiv community reviewhttps://www.alphaxiv.org/abs/2608.24735

    Probes of the code found that the governed execution path was ‘decoupled from the persona by omission,’ meaning it lacked execution-side re-validation under persona perturbations.

  18. There’s An AI For That paper pagehttps://theresanaiforthat.com/paper/meta-n-recursive-self-improvement-through-emergent-depth/

    As a system modifies its own epistemic plumbing—including its own benchmarks and evaluators—it risks ‘self-ratifying’ its mistakes… ‘AI-self-gating’ frequently enters a ‘rubber-stamp’ regime; the model’s internal confidence scores rise while its actual benchmark performance on complex tasks like ARC-AGI-2 collapses.

Jack Sun

Jack Sun, writing.

Engineer · Bay Area

Hands-on with agentic AI all day — building frameworks, reading what industry ships, occasionally writing them down.

Digest
All · AI Tech · AI Research · AI News
Writing
Essays
Elsewhere
Subscribe
All · AI Tech · AI Research · AI News · Essays

© 2026 Wei (Jack) Sun · jacksunwei.me Built on Astro · hosted on Cloudflare