Databricks hits $188B, OpenAI skips SWE-Bench Pro, Patreon blocks scrapers
Databricks fundraises pre-IPO, OpenAI skips a losing benchmark, and Patreon blocks scrapers — three defensive plays against AI-boom pressure.
Databricks hits $188B, OpenAI skips SWE-Bench Pro, Patreon blocks scrapers
TL;DR
- Databricks raises $3B at $188B, a ~40% step-up from February’s $134B round.
- OpenAI pushes per-task pricing to reframe GPT-5.6’s 15-point SWE-Bench Pro shortfall.
- Patreon retires robots.txt for Cloudflare’s network-level AI-bot blocking.
- GPU-backed lenders extend a $400M loan secured by inference silicon, not training chips.
- San Francisco orders Apple and Google to pull AI nudify apps from their stores.
Today’s three lead stories are all defensive plays against pressures the AI boom itself created. Databricks raises $3B at a $188B valuation — a ~40% step-up in five months — while quietly disclosing that gross margins have slid from 80%+ to 74% as agentic workloads burn resold GPU capacity faster than analytics ever did. OpenAI’s new Useful Intelligence per Dollar scorecard reframes agent value as cost per finished task, a metric that flatters GPT-5.6 Sol’s token efficiency and skips SWE-Bench Pro, where it trails Claude Mythos 5 by 15 points. And Patreon retires robots.txt for network-level Cloudflare blocking, betting its creators’ content is worth defending against scrapers that have stopped self-identifying.
The round-ups widen the aperture on the same defensive posture: Apple’s trade-secrets suit rippling into OpenAI’s IPO timing, Anthropic reversing course to keep Fable 5 bundled in Max plans, a $400M loan collateralized by inference silicon, and San Francisco pushing Apple and Google to pull AI nudify apps.
Databricks raises $3B at $188B as AI compute erodes margins
Source: techcrunch-ai · published 2026-07-17
TL;DR
- Databricks raised $3B at $188B, a ~40% step-up from February’s $134B, led by Coatue.
- Gross margins fell from 80%+ to 74% as agentic workloads burn resold GPU capacity faster than analytics ever did.
- Ghodsi called 2026 “a terrible year to go public,” with $200B+ of IPO demand about to clear the market.
- 52%+ of Snowflake customers now also run Databricks, mediated by Apache Iceberg — coexistence, not conquest.
The number is real — and so is what it cost
Coatue led a $3B strategic round at a $188B valuation on July 17, roughly 40% above the $134B mark set just five months earlier, with Andreessen Horowitz, Thrive, MGX, GIC and Insight Partners all rolling their positions 1. Growth justifies the price on its face: Databricks crossed a $6.9B annualized revenue run-rate in June 2026, accelerating to 80% YoY, and its AI-specific product line alone hit $1.7B ARR — up from $1.4B in February — with net revenue retention above 140% 23. That NRR figure is the tell. It means existing customers are spending 40%+ more year-over-year, which is what you’d expect when automated agents, not humans, are generating the queries.
The margin picture is less flattering. Gross margins have compressed from above 80% in 2024 to 74% in mid-2026, because Databricks resells hyperscaler GPU capacity and agentic workloads consume dramatically more compute per dollar of revenue than traditional Spark analytics 2. The company is trading gross-margin structure for top-line velocity — a defensible bet if the AI ARR keeps compounding, a much worse one if it plateaus.
Why still private
Ghodsi has been unusually candid about the IPO calendar. He publicly called 2026 “a terrible year to go public,” arguing that SpaceX, OpenAI and Anthropic are lined up to absorb more than $200B in IPO capital and turn every other listing into a “side show” 4. Read against that, the $3B round is bridge liquidity: tender offers for employees, M&A dry powder, and multi-year GPU commitments to hold the AI ARR curve intact until the mega-IPO window clears in late 2026 or 2027. This is not a company that needs growth capital. It’s a company timing the exit.
Snowflake: coexistence, not conquest
The “AI’s favorite second act” framing implies Databricks is eating Snowflake. The data says otherwise. More than 52% of Snowflake customers now also run Databricks, and Apache Iceberg is increasingly the shared table format between them 5. The pattern enterprises are settling into is Snowflake for BI, finance and SQL analytics; Databricks for ML training, feature engineering and agent workloads. That’s excellent for Databricks’ wallet share but complicates any narrative in which one platform routs the other.
The standing bear case
The dissent worth taking seriously is architectural. Practitioner analysis argues Unity Catalog’s governance and lineage don’t propagate cleanly outside the Databricks perimeter, and that a “Delta Lake-first” design bias slows Iceberg parity and treats non-Databricks systems as “second-class citizens” 6. Combined with opaque DBU-based pricing, that’s the lock-in critique institutional buyers will keep pressing on — quietly during fundraising rounds, loudly at IPO.
The $188B is what investors will pay to own the token path. The 74% gross margin is what Databricks is paying to keep it.
OpenAI’s GPT-5.6 wins on tokens, loses on SWE-Bench Pro
Source: openai-blog · published 2026-07-17
TL;DR
- Sarah Friar’s “Useful Intelligence per Dollar” scorecard reframes AI value around task completion, not per-token pricing.
- GPT-5.6 Sol used 54% fewer output tokens than Claude Fable 5 on DeepSWE v1.1, at $8.39 per run vs. $21.63.
- OpenAI omits SWE-Bench Pro, where Sol scores 64.6% against Claude Mythos 5’s 80.3% — a 15-point gap.
- 88% of enterprise agent pilots never reach production, making the “cost per successful task” metric hard to actually measure.
The scorecard, in one metric
OpenAI CFO Sarah Friar wants CFOs to stop asking about price-per-token and start asking about “Useful Intelligence per Dollar” — four pillars covering work completion, cost per successful task, dependability, and return on compute. It arrives bundled with the GPT-5.6 family (Sol, Terra, Luna), a ChatGPT Work environment, and a Broadcom “Jalapeño” inference chip. Read charitably, it’s an argument that frontier models are cheaper than they look once you count human review time. Read less charitably, it’s the pricing narrative a pre-IPO OpenAI needs when 85% of companies still can’t prove positive ROI on AI 7.
The efficiency claim actually holds up
The token-efficiency headline survives independent scrutiny. Artificial Analysis clocked Sol at roughly 54% fewer output tokens than Claude Fable 5 on DeepSWE v1.1 while staying in a statistical tie on quality, with per-run costs of $8.39 vs. $21.63 8. That’s a reproducible ~2.6× cost delta on long-horizon engineering tasks, not marketing math. LayerLens confirms the Pareto story for Sol and Terra — both hold above 89% on MRCR v2 8-needle recall — but flags a cliff Friar’s post doesn’t mention: Luna collapses to 41.3% on the same benchmark 9. The “Useful Intelligence per Dollar” pitch quietly depends on routing hard work to the expensive tier.
The benchmark OpenAI didn’t mention
Sol’s dominance isn’t universal, and the omission is telling. On SWE-Bench Pro, Sol scores 64.6% versus Claude Mythos 5’s 80.3% 10 — a 15-point deficit on the benchmark most closely tracking real codebase remediation. This directly inverts Friar’s core claim. For deep programming work, Anthropic’s denser reasoning wins the outcome even while burning more tokens, which means the cheaper-per-successful-task math flips in exactly the workload OpenAI’s coding numbers imply it owns. Choosing DeepSWE v1.1 for the blog and skipping SWE-Bench Pro isn’t a rounding decision; it’s the whole thesis.
The measurement problem
The framing is also getting read as strategy, not methodology. Hacker News commenters called the scorecard a “self-interested” move to shift evaluation from objective benchmarks toward vendor-defined “useful work” — a “beautifying lie meant to placate CFOs nervous about the lack of measurable ROI” 11. The critique lands because the pillar most central to the pitch — “cost per successful task” including the human tax of review and retries — is exactly what enterprises say they cannot yet quantify. Dataiku puts 88% of agent pilots as never reaching production, with governance and infrastructure, not model IQ, as the failure mode 7.
Meanwhile the surface where developers judge these claims has gotten worse. ChatGPT Work users report confusion tracking Agentic Credits across parallel Chat, Work, and Codex threads, and longtime Codex users say their terminal-first environment has been demoted for a business UI optimized for decks and Sites 12. Friar’s scorecard asks buyers to trust that agents complete useful work autonomously. The tooling they’d use to verify that just got harder to read.
Patreon switches from robots.txt to Cloudflare bot blocking
Source: techcrunch-ai · published 2026-07-17
TL;DR
- Patreon turned on Cloudflare’s Crawl Control to network-block AI training bots, retiring
robots.txtas its main defense. - Cloudflare will make “block AI training” the default for all new ad-supported domains by 15 September 2026.
- Reddit’s suit against Anthropic alleges 100,000+ crawls after the company publicly said it had stopped.
- Network blocks miss residential-proxy and human-mimicking scrapers that never self-identify as bots.
From polite request to network block
Patreon’s CEO Jack Conte isn’t burying the lede. His framing — “creators deserve credit, compensation, and consent. If that’s not on the table, the crawlers can stay the f*** off Patreon” 13 — is being quoted verbatim across coverage because it’s the actual policy. The technical change is that Patreon is no longer relying on robots.txt, an honor-system text file, and has switched to Cloudflare’s Crawl Control, which sits at the network edge and drops requests before they reach origin.
Under the hood, this is the same stack that gained Cloudflare’s Precursor engine in July 2026, shifting bot detection from one-shot header checks to continuous behavioral session monitoring 14. Patreon isn’t building anything bespoke — it’s adopting Cloudflare’s default posture, which flips to “block AI training” for every new ad-supported domain on the network by 15 September 2026 15.
| Layer | Mechanism | Enforceable? |
|---|---|---|
robots.txt | Text file requesting crawlers stay out | No — honor system |
| Cloudflare Crawl Control | Edge block on identified AI bots | Yes, if bot self-identifies |
| Pay-Per-Crawl (forthcoming) | Priced access via NET Dollar stablecoin | Yes, and monetized 15 |
The enforcement wave behind the move
Patreon is one node in a rapidly hardening pattern. Reddit’s parallel case against Anthropic — remanded to California state court as a User Agreement breach — cites logs allegedly showing Anthropic bots hit the platform more than 100,000 times after the company publicly claimed to have stopped 16. That’s the exhibit that killed robots.txt as an evidentiary layer: platforms need a block they can point to in discovery, not a request they hoped would be honored.
Independent audits reinforce the case. ByteDance’s Bytespider, per one crawler-behavior tracker, “frequently ignores robots.txt directives entirely” and has triggered unintentional DoS conditions with hundreds of requests per minute 17. Against that baseline, Patreon’s escalation reads as table stakes rather than vanguard.
What the block doesn’t catch
Not everyone is celebrating. TechEchelon points out the sharp edge of network-level defenses:
Network-level blocking primarily catches bots that self-identify, potentially leaving the platform vulnerable to pirate scrapers using residential proxies or bots that mimic human behavior. 18
That caveat is the whole ballgame. The bots that announce themselves as GPTBot or ClaudeBot are also the ones most likely to negotiate license deals; the ones running through residential IP pools and headless Chrome are the ones Patreon most wants to stop, and they’re the hardest to catch. Cloudflare’s Precursor is aimed at exactly that population 14, but the arms race is live.
The tollbooth question
The strategic tension worth watching: Cloudflare is simultaneously the bouncer and, via its forthcoming Pay-Per-Crawl marketplace settled in a USD-backed NET Dollar stablecoin, the toll collector 15. Patreon has not opted into monetized crawl access, and Conte’s rhetoric suggests it won’t. But the same edge that enforces today’s absolutist block is the infrastructure that would meter tomorrow’s paid feed. “Block by default” is quietly becoming “price by default” one config flag away — and the platforms holding the line on consent will have to decide, publicly, whether that line moves when the invoice arrives.
Round-ups
Apple’s trade-secrets suit lands as OpenAI eyes IPO
Source: the-verge-ai, techcrunch-ai, techcrunch-ai
Apple’s complaint alleges a pattern of misconduct reaching OpenAI’s chief hardware officer and notes that more than 400 former Apple employees now work there. The timing threatens OpenAI’s reported IPO plans, and analysts are debating whether the alleged conduct is genuinely actionable or industry-standard poaching.
Anthropic keeps Claude Fable 5 in Max and Team plans
Source: simon-willison
Starting July 20, Fable 5 stays bundled with Max and Team Premium subscriptions at 50% of standard limits, reversing a plan to move it to API-only pricing. Pro and Team Standard users get a one-time $100 credit as competition from GPT-5.6 Sol intensifies.
GPU lenders back $400M loan secured by inference chips
Source: techcrunch-ai
The financiers who pioneered GPU-backed lending are extending the model to inference silicon in a $400 million deal, signaling that dedicated inference hardware is now bankable collateral and hinting at the next wave of AI infrastructure financing beyond training clusters.
San Francisco orders Apple, Google to pull nudify apps
Source: ars-technica-ai
The city attorney is demanding Apple and Google remove AI nudify apps from their stores, with officials estimating the two platforms have taken in millions of dollars in fees from apps used to generate non-consensual intimate imagery.
TikTok tests opt-in AI likeness detection for creators
Source: the-verge-ai
The tool scans uploads for AI-generated likenesses and lets creators flag them to TikTok, initially rolling out to a subset of US creators. It mirrors a similar system YouTube has been building as platforms scramble to give talent recourse against synthetic impersonation.
Google-backed FireSat satellites launch amid US-Canada smoke crisis
Source: ars-technica-ai
The FireSat constellation is designed to detect wildfires that existing satellites miss, using AI-tuned infrared sensors to catch small ignitions early. Launch coincides with a smoke event choking swaths of the US and Canada, giving the system an immediate operational test.
Weather-data sabotage emerges as new AI-era threat vector
Source: mit-tech-review-ai
Forecasts drive decisions for airlines, grid operators, and farmers, and the sensor networks feeding them are increasingly exposed to tampering. Corrupted inputs would ripple through AI-driven weather models that industries now treat as ground truth for billion-dollar operational calls.
Footnotes
-
SiliconANGLE — funding round coverage — https://siliconangle.com/2026/07/17/databricks-raising-new-funding-188b-valuation/
↩Coatue Management led the $3 billion strategic round at a $188B valuation, a ~40% jump from $134B just five months earlier, with continued participation from Andreessen Horowitz, Thrive Capital, MGX, GIC and Insight Partners.
-
mlq.ai — Databricks revenue analysis — https://mlq.ai/news/databricks-revenue-hits-69b-annualized-as-80-growth-comes-with-shrinking-margins/
↩ ↩2Databricks crossed $6.9B annualized revenue in June 2026 with growth accelerating to 80% YoY, but gross margins fell from above 80% in 2024 to 74% in mid-2026 as AI agents generate significantly more compute demand than traditional analytics.
-
tomtunguz.com — analyst take — https://tomtunguz.com/databricks-widens-lead/
↩Databricks’ AI-specific products reached a $1.7B ARR run-rate by June 2026, up from $1.4B in February, with net revenue retention above 140% — capturing the ‘token path’ where revenue scales with automated agent queries rather than human users alone.
-
The Next Web — Ghodsi on IPO timing — https://thenextweb.com/news/databricks-ceo-calls-2026-a-terrible-year-to-go-public-as-spacex-anthropic-and-openai-prepare-to-absorb-200-billion-in-ipo-capital
↩Databricks CEO Ali Ghodsi called 2026 ‘a terrible year to go public,’ citing a crowded market where SpaceX, Anthropic and OpenAI are expected to absorb over $200 billion in IPO capital, leaving other listings as ‘side shows.’
-
definite.app — Databricks vs Snowflake 2026 — https://www.definite.app/blog/databricks-vs-snowflake-2026
↩Over 52% of Snowflake customers now also use Databricks, up sharply from prior years; enterprises increasingly deploy Snowflake as the analytical SQL platform while using Databricks as the AI/engineering engine, mediated by Apache Iceberg.
-
dsstream.com — Mosaic AI critique — https://www.dsstream.com/post/databricks-vs-snowflake
↩Unity Catalog’s governance and lineage do not easily propagate beyond the Databricks perimeter, and its ‘Delta Lake-first’ design bias causes lag in supporting open standards like Apache Iceberg, treating non-Databricks systems as ‘second-class citizens.’
-
Dataiku enterprise-AI-agents guide — https://www.dataiku.com/blog/enterprise-ai-agents-guide-for-modern-businesses
↩ ↩2Approximately 88% of AI agent pilots fail to reach full production, typically due to governance gaps and infrastructure limitations rather than the intelligence of the models themselves; 85% of companies still cannot prove positive ROI on AI investments.
-
Artificial Analysis — ‘GPT-5.6 has landed’ — https://artificialanalysis.ai/articles/gpt-5-6-has-landed
↩Sol used roughly 54% fewer output tokens than Claude Fable 5 while maintaining a statistical tie in quality; DeepSWE v1.1 runs cost $8.39 vs. $21.63 for Fable 5.
-
LayerLens benchmark review — https://layerlens.ai/blog/gpt-5-6-benchmark-review-sol-terra-luna
↩While Sol and Terra maintain 8-needle context recall above 89% on MRCR v2, the entry-level Luna tier drops sharply to 41.3% — cost-effective at 24 benchmark points per dollar but unsuitable for large-codebase reasoning.
-
Kingy.ai — Claude Fable 5 vs GPT-5.6 Sol — https://kingy.ai/ai/ai-guides/claude-fable-5-vs-gpt-5-6-sol/
↩On SWE-Bench Pro, Sol’s score of 64.6% trailed significantly behind Claude Mythos 5 at 80.3%, suggesting that for specific deep-programming tasks, larger ‘dense’ reasoning may still outperform Sol’s ‘efficient’ approach.
-
Hacker News discussion — https://news.ycombinator.com/item?id=43720374
↩Critics characterized the scorecard as a ‘self-interested’ marketing framing designed to move the goalposts away from objective technical benchmarks… a ‘beautifying lie’ meant to placate CFOs nervous about the lack of measurable ROI.
-
Novaedge — ChatGPT Agent Mode 2026 guide — https://www.novaedgedigitallabs.tech/blog/chatgpt-agent-mode-complete-guide-2026
↩Users struggle to track consumption across parallel ‘Chat,’ ‘Work,’ and ‘Codex’ threads… the once-dedicated Codex environment now feels secondary to the ‘Work’ interface, which prioritizes deck creation over technical terminal control.
-
404 Media — https://www.404media.co/patreon-cloudflare-partnership-ai-crawlers/
↩Creators deserve credit, compensation, and consent. If that’s not on the table, the crawlers can stay the f*** off Patreon.
-
Cloudflare press release (Precursor) — https://www.cloudflare.com/press/press-releases/2026/cloudflare-introduces-precursor-one-click-behavioral-defense-against-modern-bots/
↩ ↩2Precursor shifts detection from point-in-time checks to continuous session monitoring, one-click behavioral defense against modern bots.
-
CryptoBriefing — https://cryptobriefing.com/patreon-cloudflare-block-ai-scraping-bots/
↩ ↩2 ↩3By September 15, 2026, Cloudflare plans to set ‘block AI training’ as the default for all new ad-supported domains, with a Pay-Per-Crawl marketplace settling in the USD-backed NET Dollar stablecoin.
-
RedditWatch — Reddit v. Anthropic tracker — https://redditwatch.org/issues/reddit-v-anthropic-scraping-lawsuit-2025
↩Reddit’s discovery logs allegedly show Anthropic bots accessed the platform over 100,000 times after the company publicly claimed to have stopped; a federal judge remanded the dispute to California state court as a breach of the User Agreement.
-
AICrawlerCheck — Bytespider report — https://aicrawlercheck.com/blog/bytespider-aggressive-ai-scrapers
↩Bytespider frequently ignores robots.txt directives entirely, especially when they are grouped with other bots, and has triggered unintentional DoS conditions with hundreds of requests per minute.
-
↩Network-level blocking primarily catches bots that self-identify, potentially leaving the platform vulnerable to pirate scrapers using residential proxies or bots that mimic human behavior.