28% of execs skip AI, Medicare WISeR stretches waits, Lila chases $8.5B
Executive adoption data, a Medicare AI pilot, and a biotech mega-round each expose the gap between AI pitch and measured result.
28% of execs skip AI, Medicare WISeR stretches waits, Lila chases $8.5B
TL;DR
- 28% of senior executives use no AI at all, with 69% under an hour weekly.
- 74% of executives admit overstating confidence in their AI strategies.
- Washington hospitals report WISeR AI review stretching procedure waits from 2 to 4-8 weeks.
- Lila Sciences is raising $2B at an $8.5B valuation on a flagship result peers call a reach.
- Senate 46-50 vote failed to kill WISeR, locking the pilot in through December 2031.
Today’s AI-news leads sit in three different worlds — enterprise adoption, Medicare prior-authorization, and biotech venture funding — but each one is a measurement story. A fresh adoption survey finds 28% of senior executives use no AI at all and 95% of GenAI pilots produce no P&L impact, undercutting the Claudeonomics-style dashboards vendors use to sell the story. In Washington state, hospitals report CMS’s WISeR pilot turning 2-week procedure windows into 4-to-8-week ones, with vendor Virtix Health already on a corrective action plan — and a Senate resolution to kill the program failed 46-50, locking it in through 2031. And Lila Sciences is chasing a $2B round at an $8.5B valuation, up 6.5× in under a year, on an mRNA result an outside biologist called “a reach.” Three different pitches, three different measurements, one recurring pattern: the number the vendor leads with isn’t the one that survives contact with the data.
Data vindicates Suresh: 28% of execs use no AI at all
Source: simon-willison · published 2026-07-19
TL;DR
- 28% of senior executives use no AI at all, with 69% logging under an hour a week
- 74% of executives admit overstating confidence in their own AI strategies
- 95% of GenAI pilots produce no measurable P&L impact, with budgets skewed to sales and marketing
- Meta’s “Claudeonomics” dashboard ranked 85,000 employees by token consumption in a single month
The executive usage gap is measurable
Suresh’s headline anecdote — an executive shipping a $2B+ AI strategy after admitting they’d never opened ChatGPT — reads like a bit. The Stanford/Charter numbers say it isn’t. 28% of C-suite leaders don’t use AI at all in their professional lives, and 69% use it less than an hour per week 1. A KUNGFU.AI survey adds the second half of the pincer: 74% of executives concede they’ve overstated confidence in their own AI strategy 2. That’s not “vendor sales fluff meeting excitable buyers.” That’s board-level mandates outpacing personal literacy, on the record, by admission.
MIT’s NANDA report puts the money at stake into focus: 95% of GenAI pilots fail to produce a measurable P&L impact, and more than half of AI budgets flow into sales and marketing — exactly the domains with the weakest returns 3. Spending is decoupled from utility in a way that matches Suresh’s diagnosis line-for-line.
The credibility trap has a shape
Suresh’s most interesting claim isn’t about incompetent buyers — it’s about why honest vendors stay quiet. A vendor rep who tells a customer executive that 100x productivity is fantasy has attacked that executive’s public credibility, and enterprise contracts die that way. So the vendor echoes the number. The customer executive hears their own claim reflected back and treats it as validation.
flowchart LR
A[Customer exec claims<br/>100x productivity] --> B[Vendor hears claim<br/>in review]
B --> C{Correct the number?}
C -->|Yes| D[Attacks exec credibility<br/>→ contract at risk]
C -->|No| E[Echoes the claim]
E --> F[Exec treats echo<br/>as validation]
F --> A
METR’s randomized trial is what breaks this loop empirically: experienced developers using Cursor+Claude were 19% slower, yet self-reported being 20% faster — a ~40-point perception gap 4. If practitioners in a controlled trial can’t accurately judge their own throughput, executives two org-levels removed have no chance of correcting the story.
Tokenmaxxing is worse than the joke suggests
The “rewrite our Go repo in Zig to keep my job” anecdote turns out to be the tame version. InfoWorld reports Meta’s internal “Claudeonomics” dashboard ranked 85,000 employees by token consumption with titles like “Token Legend,” and Amazon shuttered its KiroRank leaderboard in May 2026 after leadership discovered systemic cheating and vanity usage 5. Costs land on the operational side too: an HN commenter describes Copilot hallucinating a downed service, then the team burning hours convincing ops that the service didn’t exist 6.
Where Suresh overreaches
The weaker leg of the essay is the implicit “and therefore the tools don’t work.” HN pushback flags the tone as “extreme grandiosity,” and a follow-up METR analysis with late-2025 tooling flipped the 19% slowdown into an 18% speedup — though the researchers themselves flagged selection bias, since developers who considered AI essential refused the no-AI arm 4.
The distinction matters. Executive detachment, suppressed skepticism, and metric-gaming are independently documented corporate pathologies. Whether the underlying models are useful is a separate empirical question that the mania itself is making harder to answer honestly.
Medicare’s WISeR AI pilot stretches waits to 8 weeks
Source: ars-technica-ai · published 2026-07-18
TL;DR
- Washington hospitals report procedure completion times jumping from 2 weeks to 4–8 weeks under WISeR’s AI review layer.
- CMS put vendor Virtix Health on a corrective action plan for missing the mandated 72-hour clinical determination window.
- WISeR vendors earn 10–20% of “averted expenditures” — a bounty structure across six state contractors.
- A Senate resolution to kill the pilot failed 46–50 in July 2026, locking WISeR in through December 2031.
The pilot is no longer hypothetical
The Ars piece frames WISeR — the Wasteful and Inappropriate Service Reduction Model — as a coming test of whether AI helps or hurts prior authorization. Six months in, the answer is on the record. Washington State Hospital Association surveys show procedure completion times for covered services stretching from a historical two-week average to between four and eight weeks under the new machine-learning review layer 7. CMS has already been forced to place one of the six state vendors, Virtix Health, on a corrective action plan for repeatedly missing the mandated 72-hour clinical determination window 8 — the exact turnaround the same administration is holding up as its consumer-protection win against private insurers.
Those are not the physician-sentiment numbers the AMA survey captures. They are the first operational data points from a live CMS demonstration, and they cut the wrong way.
The 10–20% bounty
Ars mentions vendors are paid a share of “averted expenditures” without naming the figure. Independent advisory analysis pegs the cut at 10–20% of averted spend across the six contractors — Cohere, Innovaccer, Humata, Zyter, Genzeon, and Virtix 9. That is what fuels former CMS Administrator Donald Berwick’s “bounty” language and what makes the denial-for-profit critique more than rhetorical: the vendor’s revenue line is a direct function of how much care it stops.
Compounding the opacity, the Electronic Frontier Foundation has sued CMS under FOIA seeking training data, bias assessments, and the actual decision logic 10. CMS has disclosed none of it. Physicians appealing a denial against traditional Medicare are now arguing against a black box the public agency itself will not describe.
Medicare Advantage is the tell
The strongest independent evidence that AI prior auth tends toward wrongful denial sits in the Medicare Advantage litigation. Court filings in Lokken v. UnitedHealth allege the nH Predict tool carries a roughly 90% error rate — measured by the share of denials overturned when patients actually appeal — while internal targets pressured clinicians to stay within 1% of the algorithm’s projected length of stay 11. Cigna’s PXDX system faces parallel allegations of 1.2-second-per-claim reviews.
That is the base rate the Ars article’s 81%-overturned-on-appeal figure sits on top of. When automated denials are challenged, they collapse. Most are never challenged.
Providers in Oklahoma describe the WISeR clinical review process as “horrendous.” 12
Political durability
Opposition is loud but not yet decisive. A Senate Congressional Review Act resolution to kill WISeR failed 46–50 in July 2026, and a GAO finding that CMS bypassed required congressional review has not produced a rollback 12. The Gold Card exemption CMS floated is the administration’s pressure valve, not a retreat.
The framing question in the Ars headline — will AI fix prior auth or make it worse — is already answered for the six states inside the pilot. The live question is whether the vendor payment model survives contact with the first wrongful-death suit.
Lila Sciences chases $8.5B valuation on unproven science
Source: latent-space · published 2026-07-16
TL;DR
- Lila Sciences is reportedly raising $2B at an $8.5B valuation, up from $1.3B under a year ago.
- Johns Hopkins RNA biologist Jeff Coller called Lila’s flagship “Move 37” mRNA result “a reach.”
- Berkeley’s A-Lab precedent — 41 “novel” materials that turned out to be misidentified known phases.
- The “lab as data center” metaphor hides how much of the workflow still runs on humans-below-the-API.
The pitch
On the latest Latent Space, Lila’s Andy Beam and Rafa Gómez-Bombarelli argue that science, not the internet, is the last untapped training-data reservoir — and that the way to unlock it is a lab that feels like a data center: instrumented, closed-loop, AI-designed experiments running against AI-verified results. It’s a compelling frame, and investors are buying it. Lila is reportedly in talks for a $2B Series B at an $8.5B pre-money, anchored by CalPERS and Nvidia 13. That’s up from a $1.3B mark set less than a year ago 14.
The problem is that the capital arc has decoupled from the evidence arc. Lila has not named a paying customer or published a peer-reviewed result. Independent commentary reads the round as “pricing an option on a hypothetical outcome rather than a proven business model” 14.
The “Move 37” claim doesn’t survive contact with domain experts
The hyperstable-mRNA result the founders lean on as their AlphaGo moment has drawn direct pushback. Jeff Coller, an RNA biologist at Johns Hopkins, called the Move 37 branding “a reach” and characterized Lila’s output as “meaningful incremental gains on well-understood biology” 15.
The methodological objection is sharper than the rhetorical one. Lila benchmarked against an open academic baseline. State-of-the-art mRNA constructs at Moderna- and BioNTech-class shops are engineered behind closed doors and never published. Beating the public number does not establish frontier parity — it establishes that a well-funded lab beat a paper 15. The podcast does not engage this critique.
Autonomous labs have a validation-crisis track record
Lila’s “AI proposes, AI verifies” architecture inherits a specific failure mode from prior self-driving labs. Berkeley’s A-Lab announced 41 novel materials in 17 days in late 2023; a subsequent audit by Palgrave and Schoop showed the automated XRD pipeline had mistaken disordered known phases and mixtures for new compounds — a “beginner-level” error in solid-state chemistry 16. When the verifier is itself a learned model, reproducibility gets worse, not better: identical code across different GPUs, library versions, or RNG seeds can produce divergent scientific conclusions 17.
Beating a public baseline does not necessarily prove that Lila has outperformed the commercial-grade technology used by established pharmaceutical companies. 15
The data-center metaphor leaks
The “lab feels like a data center” framing invites listeners to picture racks of instruments and an API. In practice, the closest analog — Emerald Cloud Lab — still recruits Amazon-warehouse-style human “pickers” to move containers between workcells 18. That’s humans-below-the-API, not humans-out-of-the-loop. Lila’s own leadership concedes a similar pattern on the podcast, but the metaphor does the work of hiding it.
What’s actually at stake
The platform bet is real: nobody else has stacked this much capital, robotics, and generative modeling into one wet-lab campus. But the specific scientific claims Lila is using to justify an $8.5B mark are unvalidated by peers, benchmarked against weak baselines by the founders’ own domain, and shadowed by a recent autonomous-lab retraction the founders don’t address. Until a paying pharma customer or a peer-reviewed frontier result lands, this is a valuation story, not a science story.
Round-ups
Moonshot’s new Kimi release sparks US ‘AI communism’ backlash
Source: techcrunch-ai
Chinese lab Moonshot AI shipped an updated Kimi model this week, drawing alarmed commentary from figures including David Sacks, Dean Ball, and Travis Kalanick over open Chinese frontier models pressuring US labs on price and access.
Dave Eggers tells OpenAI staff ChatGPT is silencing a generation
Source: the-verge-ai
Novelist Dave Eggers used a Sam Altman-arranged talk to roughly 200 OpenAI employees to argue ChatGPT is muting young writers’ voices. The McSweeney’s founder and school-builder turned the invited lecture into a critique of the product his hosts ship.
Simon Willison pitches swapping golf courses for datacenter water offsets
Source: simon-willison
Google used 10.9 billion gallons of water in 2025, or 30 million gallons a day. Coachella Valley’s 120 golf courses each burn about 750,000 gallons daily, so buying 40 and rewilding them as birdwatching parks would cover Google’s footprint.
Latent Space calls July 18 a quiet day in AI
Source: latent-space
The daily AINews digest marks an unusually slow news cycle, with no major model launches, papers, or company announcements worth leading on. A useful signal for readers deciding whether to skip the day’s feeds.
Ben’s Bites rounds up weekend AI reading list
Source: bens-bites
The newsletter’s weekend edition collects longer-form AI links for readers catching up outside the news cycle. No single headline anchors the roundup — it’s a curated skim rather than a breaking story.
The Verge’s Installer #136 picks apps and gadgets for readers
Source: the-verge-ai
Installer No. 136 focuses on tools for people who read — apps, devices, and accessories — alongside host David Pierce’s note that he’s recording the next season of the Version History podcast.
Footnotes
-
Time/Charter — ‘CEOs Are Betting Big on AI While Barely Using It’ — https://time.com/partner-content/charter/7382209/ceos-are-betting-big-on-ai-while-barely-using-it/
↩Approximately 28% of CEOs, CFOs, and other senior executives report that they do not use AI at all in their professional lives, while 69% engage with it for less than one hour per week.
-
Inc. — KUNGFU.AI executive survey — https://www.inc.com/joe-galvin/91360446/91360446
↩74% of executives surveyed… confessed to overstating their confidence in their current AI strategies, suggesting that board-level mandates often outpace actual executive literacy.
-
Forbes — Jason Snyder on MIT NANDA GenAI Divide report — https://www.forbes.com/sites/jasonsnyder/2025/08/26/mit-finds-95-of-genai-pilots-fail-because-companies-avoid-friction/
↩95% of AI pilots fail to achieve a measurable return on investment… more than half of AI budgets are typically funneled into sales and marketing pilots, despite these areas showing some of the lowest measurable returns.
-
InfoWorld on METR randomized trial — https://www.infoworld.com/article/4020931/ai-coding-tools-can-slow-down-seasoned-developers-by-19.html
↩ ↩2Developers were 19% slower when using AI… before the study, developers predicted AI would make them 24% faster; even after experiencing the measured slowdown, they still self-reported a 20% productivity gain.
-
InfoWorld — ‘From story points to tokenmaxxing’ — https://www.infoworld.com/article/4195756/from-story-points-to-tokenmaxxing-why-engineering-keeps-measuring-the-wrong-things.html
↩Meta’s internal ‘Claudeonomics’ dashboard reportedly logged 60.2 trillion tokens in a single month, ranking 85,000 employees with titles like ‘Token Legend’… Amazon ultimately shuttered its ‘KiroRank’ leaderboard in May 2026 after leadership discovered widespread ‘cheating’ and vanity usage.
-
Hacker News discussion of Suresh’s essay — https://news.ycombinator.com/item?id=48962963
↩GitHub Copilot suggested a non-existent service was down, and the team spent hours convincing operations that the system did not exist… any time saved by AI is frequently ‘swallowed up tenfold’ by the need to debug hallucinated code.
-
HMA Academy — ‘WISeR is losing in Washington but still running in six states’ — https://hmacademy.com/insights/all-insights/artificial-intelligence/wiser-is-losing-in-washington-but-still-running-in-six-states
↩Washington State Hospital Association survey found procedure completion times for covered services stretched from a historical average of two weeks to between four and eight weeks
-
Becker’s Payer Issues — CMS corrective action on Virtix — https://www.beckerspayer.com/payer/cms-orders-corrective-action-plan-for-ai-vendor-in-medicare-prior-authorization-pilot/
↩CMS placed Virtix Health on a corrective action plan following complaints regarding its inability to meet the mandated 72-hour turnaround time for clinical determinations
-
Forvis Mazars advisory — ‘CMS WISeR Model: What providers need to know’ — https://www.forvismazars.us/forsights/2025/12/cms-wiser-model-what-providers-need-to-know-by-january-2026
↩vendors are compensated based on a percentage of the spending their reviews avert… the vendor share [is] 10% to 20% of averted expenditures
-
Healthcare Dive — EFF v. CMS FOIA suit — https://www.healthcaredive.com/news/electronic-frontier-foundation-sues-cms-medicare-ai-prior-authorization-wiser/816087/
↩The Electronic Frontier Foundation filed a FOIA lawsuit against CMS seeking transparency regarding the algorithms’ training data and safeguards against bias or ‘hallucinations’
-
Thompson Coburn LLP — ‘Class actions highlight AI-assisted payer denials’ — https://www.thompsoncoburn.com/insights/class-actions-highlight-ai-assisted-payer-denials-102jebl/
↩Plaintiffs allege the [nH Predict] algorithm has a 90% error rate, citing that nine out of ten denials are overturned when patients actually pursue an appeal
-
The Lund Report — ‘Medicare’s new AI experiment sparks alarm’ — https://www.thelundreport.org/content/medicares-new-ai-experiment-sparks-alarm-among-doctors-lawmakers
↩ ↩2On July 16, 2026, the U.S. Senate narrowly voted 46-50 to preserve the pilot after a Democratic-led effort to defund it failed… a ‘horrendous’ clinical review process in states like Oklahoma
-
Investing.com / Bloomberg — https://www.investing.com/news/stock-market-news/lila-sciences-in-talks-for-2b-funding-at-85b-valuation—bloomberg-93CH-4724927
↩Lila Sciences in talks for $2B funding at $8.5B valuation
-
Startup Fortune — https://startupfortune.com/lila-sciences-is-testing-how-much-investors-will-pay-for-automated-labs/
↩ ↩2Lila Sciences is testing how much investors will pay for automated labs — a valuation that surged from $1.3 billion to an estimated $8.5 billion in under a year reflects a market pricing an option on a hypothetical outcome rather than a proven business model.
-
Endpoints News (Jeff Coller commentary) — https://endpoints.news/lila-sciences-is-near-a-300m-fundraise-to-back-autonomous-science-source/
↩ ↩2 ↩3Coller described the ‘Move 37’ branding as a ‘reach,’ suggesting that the startup’s results likely represent ‘meaningful incremental gains on well-understood biology’ rather than an alien-like leap… beating a public baseline does not necessarily prove that Lila has outperformed the commercial-grade technology used by established pharmaceutical companies.
-
Shanaka Perera, ‘The Lab Without Scientists’ (Substack) — https://shanakaanslemperera.substack.com/p/the-lab-without-scientists
↩Berkeley’s A-Lab claimed 41 novel materials in 17 days; Palgrave and Schoop showed the AI’s automated XRD analysis failed to account for compositional disorder, mistaking known materials or mixtures for entirely new compounds — a ‘beginner-level’ understanding of solid-state chemistry.
-
PMC / NIH review on AI reproducibility — https://pmc.ncbi.nlm.nih.gov/articles/PMC12402953/
↩Identical code can yield divergent results due to variations in hardware architecture… random initialization, stochastic gradient descent, and specific software library versions can alter outputs even when training data remains constant.
-
R&D World — self-driving labs feature — https://www.rdworldonline.com/self-driving-cars-are-hitting-the-streets-is-your-lab-up-next-for-automation/
↩Emerald Cloud Lab relies on human operators — often recruited from logistics giants like Amazon — to perform non-automated tasks like moving large containers, treating them as specialized ‘pickers’ within a software-driven workflow.