JS Wei (Jack) Sun

Mollick's 2026 guide reorganizes around agent harnesses, not model picks

Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.

← Back to the issue

Sources

An opinionated guide to which AI to use to do stuff oneusefulthing.org

The Summer 2026 Edition

References

ExplainX — on Wharton Generative AI Labs prompting research explainx.ai

Magic words—being polite, offering tips, or using Chain-of-Thought formulas—have become largely valueless in the agentic era; the unlock is a real spec that clearly defines goals and success criteria.

Kili Technology — AI Benchmarks Guide 2026 kili-technology.com

Standard benchmarks like MMLU, GSM8K, and HumanEval are saturated; frontier evaluation has moved to ARC-AGI-2 (top systems ~54%), SWE-bench Pro on private repos, and Humanity’s Last Exam where AI still only hits ~35%.

Digg tech coverage of Mollick’s guide digg.com

Mollick argues public leaderboards like LM Arena are ‘saturated’ and failing to capture the true capability gaps of frontier models, advocating ‘stress-testing’ against ‘impossible’ real-world problems instead.

One Useful Thing comments thread oneusefulthing.org

Community feedback highlights frustration with Microsoft’s branding of ‘Cowork’ and the complexity of Claude’s ‘Max effort’ and ‘Extended’ modes — the lack of clear documentation and ‘badly named’ interfaces remain a barrier to entry for non-experts.

Jack Sun

Jack Sun, writing.

Engineer · Bay Area

Hands-on with agentic AI all day — building frameworks, reading what industry ships, occasionally writing them down.

Digest
All · AI Tech · AI Research · AI News
Writing
Essays
Elsewhere
Subscribe
All · AI Tech · AI Research · AI News · Essays

© 2026 Wei (Jack) Sun · jacksunwei.me Built on Astro · hosted on Cloudflare