JS Wei (Jack) Sun

Mollick's 2026 guide reorganizes around agent harnesses, not model picks

Mollick's Summer 2026 practitioner guide demotes model choice and elevates agent harnesses, even as enterprise pilots keep failing at 95%.

Mollick’s 2026 guide reorganizes around agent harnesses, not model picks

TL;DR

  • Mollick’s Summer 2026 guide leads with harnesses — Claude Code, Codex, Antigravity — over model picks.
  • Consequential work gets Opus, Fable, or GPT-5.6 Sol at High thinking level.
  • Prompt engineering is dead per Wharton: magic words no longer move frontier agents.
  • Independent data cuts against the frame: 95% of enterprise AI pilots failed in early 2026.

Ethan Mollick’s Summer 2026 practitioner guide is the single AI-tech lead today, and it’s a useful marker of where the frontier-user consensus has drifted. The top-line knob is no longer which model — it’s which agent harness. Claude Code, Codex, and Antigravity get the first-order attention; model choice (Opus, Fable, GPT-5.6 Sol) is a secondary dial you set once you’ve picked the scaffolding. Prompt engineering, in Wharton’s phrasing, is dead.

The counter-signal sits inside the same piece: a reported 95% failure rate for enterprise AI pilots in early 2026. The harness-first frame is what works for a power user with taste and patience; it isn’t what’s landing in the org charts paying for the seats. Read the guide as advice from the front of the adoption curve, not evidence the middle has caught up.

Mollick’s 2026 guide picks agent harnesses over chatbots

Source: one-useful-thing · published 2026-07-23

TL;DR

  • Mollick’s Summer 2026 guide reorganizes around agentic harnesses — Claude Code, Codex, Antigravity — demoting model choice to a secondary knob.
  • For consequential work he names Opus, Fable, or GPT-5.6 Sol on at least the “High” thinking level.
  • Prompt engineering is dead — Wharton’s line: “magic words” no longer move frontier agents.
  • Independent data undercuts the optimism — a reported 95% failure rate for enterprise AI pilots in early 2026.

The guide isn’t about models anymore

The Summer 2026 edition of Ethan Mollick’s “which AI to use” post drops the chatbot-era shape entirely. Instead of ranking Claude vs. ChatGPT vs. Gemini as conversational partners, it organizes recommendations around agentic harnesses — Claude Code, Codex, Antigravity — and treats the underlying model as a parameter you set inside them. For anything consequential, Mollick’s prescription is unambiguous: reach for Claude’s Opus or Fable, or GPT-5.6 Sol on at least the “High” thinking level 1.

That’s a genuine break from the 2024–25 guides, which treated model selection and prompt phrasing as the two levers users needed to master. Both levers just got demoted.

Specs, not spells

The prompting-as-craft consensus is dissolving in parallel. Wharton’s Generative AI Lab, whose research Mollick leans on, now argues that classical tricks — politeness, tip offers, Chain-of-Thought incantations — have “become largely valueless in the agentic era.” The unlock instead is “a real spec that clearly defines goals and success criteria” 2. If you’re writing a paragraph of instructions to Claude Code, the pay-off is in acceptance criteria and constraints, not in “you are an expert” preambles.

The benchmark story tracks. Mollick argues public leaderboards like LM Arena have “saturated” and stopped surfacing the capability gaps that matter, advocating stress-tests against “impossible” real-world problems 3. Independent benchmark surveys agree — the evals that still discriminate between frontier systems have moved elsewhere:

EvalFrontier scoreStatus
MMLU / GSM8K / HumanEval~saturatedRetired as signal 4
ARC-AGI-2~54%Still discriminating 4
SWE-bench Pro (private repos)contestedHarder to game 4
Humanity’s Last Exam~35%Wide headroom 4

The counterweights Mollick underplays

Two things puncture the power-user tone. First, naming. Substack commenters on the guide itself flag Microsoft “Cowork,” Claude’s overlapping “Max effort” vs. “Extended” modes, and undocumented harness switches as real barriers the guide waves past 5. The tools Mollick recommends are genuinely hard to find the right button in, and the guide reads as if that friction doesn’t exist.

Second, and more damagingly: production. Kili Technology’s 2026 evaluation guide cites reports of a 95% failure rate for enterprise AI pilots in early 2026, attributing the collapses to “vibe coding,” hallucinated planning, infinite loops, and memory poisoning — pathologies specific to the agentic deployments Mollick centers 6. The frontier-model + harness stack he’s recommending is exactly where the field’s production track record is weakest.

What to take from it

Directionally, Mollick is right and the independent literature backs him: agents over chat, specs over prompts, real-world stress-tests over MMLU. But the guide is a power-user manual dressed as a general one. If you’re not already comfortable orchestrating Claude Code against a spec you wrote, the honest read of the 2026 evidence is that you’re closer to the 95% than to the demo.

Footnotes

  1. Mollick, One Useful Thing (verbatim)https://www.oneusefulthing.org/p/an-opinionated-guide-to-which-ai-b22

    For these issues, you will want to use the most advanced models you can get access to, which is either Claude’s most powerful models, Opus and Fable, or ChatGPT’s GPT-5.6 Sol, set to at least the ‘High’ thinking levels.

  2. ExplainX — on Wharton Generative AI Labs prompting researchhttps://explainx.ai/blog/ethan-mollick-wharton-prompting-science-specs-not-tricks-2026

    Magic words—being polite, offering tips, or using Chain-of-Thought formulas—have become largely valueless in the agentic era; the unlock is a real spec that clearly defines goals and success criteria.

  3. Digg tech coverage of Mollick’s guidehttps://digg.com/tech/mjasfep2

    Mollick argues public leaderboards like LM Arena are ‘saturated’ and failing to capture the true capability gaps of frontier models, advocating ‘stress-testing’ against ‘impossible’ real-world problems instead.

  4. Kili Technology — AI Benchmarks Guide 2026https://kili-technology.com/blog/ai-benchmarks-guide-the-top-evaluations-in-2026-and-why-theyre-not-enough

    Standard benchmarks like MMLU, GSM8K, and HumanEval are saturated; frontier evaluation has moved to ARC-AGI-2 (top systems ~54%), SWE-bench Pro on private repos, and Humanity’s Last Exam where AI still only hits ~35%.

    2 3 4
  5. One Useful Thing comments threadhttps://www.oneusefulthing.org/p/an-opinionated-guide-to-which-ai-b22/comments?utm_source=post&utm_medium=web&triedRedirect=true

    Community feedback highlights frustration with Microsoft’s branding of ‘Cowork’ and the complexity of Claude’s ‘Max effort’ and ‘Extended’ modes — the lack of clear documentation and ‘badly named’ interfaces remain a barrier to entry for non-experts.

  6. Kili Technology — production reality sectionhttps://kili-technology.com/blog/ai-benchmarks-guide-the-top-evaluations-in-2026-and-why-theyre-not-enough

    Early 2026 reports show a ‘95% failure rate’ for AI pilots in production, driven by ‘vibe coding,’ hallucinated planning, infinite loops, and memory poisoning in agentic deployments.

Jack Sun

Jack Sun, writing.

Engineer · Bay Area

Hands-on with agentic AI all day — building frameworks, reading what industry ships, occasionally writing them down.

Digest
All · AI Tech · AI Research · AI News
Writing
Essays
Elsewhere
Subscribe
All · AI Tech · AI Research · AI News · Essays

© 2026 Wei (Jack) Sun · jacksunwei.me Built on Astro · hosted on Cloudflare