Mollick's 2026 guide reorganizes around agent harnesses, not model picks
Every URL the pipeline pulled into ranking for this issue — primary sources plus the supporting and contradicting findings each Researcher returned. Inline citations in the issue point back here.
Sources
An opinionated guide to which AI to use to do stuff oneusefulthing.org
The Summer 2026 Edition
References
ExplainX — on Wharton Generative AI Labs prompting research explainx.ai
Magic words—being polite, offering tips, or using Chain-of-Thought formulas—have become largely valueless in the agentic era; the unlock is a real spec that clearly defines goals and success criteria.
Kili Technology — AI Benchmarks Guide 2026 kili-technology.com
Standard benchmarks like MMLU, GSM8K, and HumanEval are saturated; frontier evaluation has moved to ARC-AGI-2 (top systems ~54%), SWE-bench Pro on private repos, and Humanity’s Last Exam where AI still only hits ~35%.
Digg tech coverage of Mollick’s guide digg.com
Mollick argues public leaderboards like LM Arena are ‘saturated’ and failing to capture the true capability gaps of frontier models, advocating ‘stress-testing’ against ‘impossible’ real-world problems instead.
One Useful Thing comments thread oneusefulthing.org
Community feedback highlights frustration with Microsoft’s branding of ‘Cowork’ and the complexity of Claude’s ‘Max effort’ and ‘Extended’ modes — the lack of clear documentation and ‘badly named’ interfaces remain a barrier to entry for non-experts.