---
title: "AI"
description: "AI articles cover model benchmarks, enterprise strategy and AI investment returns."
type: index
canonical_url: "https://philippdubach.com/categories/ai/"
source_url: "https://philippdubach.com/categories/ai/index.md"
---

# AI

AI articles cover model benchmarks, enterprise strategy and AI investment returns.

Essays examine Claude benchmarks, scaling laws and enterprise AI strategy. Related articles cover AI agents, model economics, recommendation systems and sentiment trading. Other essays examine AI investment returns, labor demand and computer vision.


*51 entries. Machine-readable feeds: [/api/posts.json](https://philippdubach.com/api/posts.json) · [/feed.json](https://philippdubach.com/feed.json) · [/index.xml](https://philippdubach.com/index.xml) · [/llms.txt](https://philippdubach.com/llms.txt).*

---

- **[65% of Hacker News Posts Have Negative Sentiment, and They Outperform](https://philippdubach.com/posts/65-of-hacker-news-posts-have-negative-sentiment-and-they-outperform/)** (2026-01-07, updated 2026-03-15) — Sentiment analysis of 32,000 Hacker News posts shows 65% skew negative and earn 27% more points. Six transformer and LLM models tested, full data included.
- **[AI Can Now Design Drugs in Seconds; We Still Can't Tell You If They Work.](https://philippdubach.com/posts/ai-can-now-design-drugs-in-seconds-we-still-cant-tell-you-if-they-work./)** (2026-03-18, updated 2026-03-24) — IsoDDE doubles AlphaFold 3 on hard benchmarks and beats physics-based gold standards. But no AI drug has FDA approval. What $4B in pharma deals actually mean.
- **[AI Capex Arms Race: Who Blinks First?](https://philippdubach.com/posts/ai-capex-arms-race-who-blinks-first/)** (2026-03-08) — Alphabet's free cash flow is on track to fall 90% in 2026. Amazon's is at $11B. $690B in AI capex is cannibalizing the cash that justified these valuations.
- **[AI Consciousness Is Not a Safety Property](https://philippdubach.com/posts/ai-consciousness-is-not-a-safety-property/)** (2026-09-17) — AI consciousness cannot be settled by safety concerns. A skeptical look at model welfare, LLM self-reports and systems that may become harder to dismiss.
- **[AI Models Are the New Rebar](https://philippdubach.com/posts/ai-models-are-the-new-rebar/)** (2026-03-11) — Qwen 3.5-35B runs on a gaming PC and matches Claude Sonnet 4.5. When the commodity version is 95% as good and 97% cheaper, you have a pricing problem.
- **[AI Models as Standalone P&Ls](https://philippdubach.com/posts/ai-models-as-standalone-pls/)** (2025-11-09, updated 2026-05-14) — OpenAI lost $11.5B in one quarter. But Anthropic CEO Dario Amodei argues each AI model is independently profitable. The assumptions behind the math.
- **[Apple's AI Bet: Playing the Long Game or Missing the Moment?](https://philippdubach.com/posts/apples-ai-bet-playing-the-long-game-or-missing-the-moment/)** (2025-12-30, updated 2026-01-03) — Apple's $157B cash pile and Gemini-powered Siri shift show a restrained AI strategy. Analysis of whether Apple wins as AI models become commodities.
- **[Aschenbrenner's Receipts](https://philippdubach.com/posts/aschenbrenners-receipts/)** (2026-05-21, updated 2026-08-16) — Leopold Aschenbrenner made dated AGI, infrastructure, and political forecasts in Situational Awareness. I assess which calls held up by May 2026, two years on.
- **[Bandits and Agents: Netflix and Spotify Recommender Stacks in 2026](https://philippdubach.com/posts/bandits-and-agents-netflix-and-spotify-recommender-stacks-in-2026/)** (2026-01-30, updated 2026-10-04) — How hybrid recommender systems balance multi-armed bandits against LLM inference cost economics in 2026. An analysis of Netflix recommendation algorithm architecture and Spotify's AI DJ recommender system.
- **[Beyond Vector Search: Why LLMs Need Episodic Memory](https://philippdubach.com/posts/beyond-vector-search-why-llms-need-episodic-memory/)** (2026-01-09, updated 2026-01-12) — Context windows aren't memory. EM-LLM's episodic architecture, knowledge graph tools like Mem0 and Letta, and why vectors fail for sequential data.
- **[Book Review: Why Machines Learn](https://philippdubach.com/posts/book-review-why-machines-learn/)** (2025-12-27, updated 2026-10-04) — Why Machines Learn explains ML math with real equations and geometric intuition. Review of Ananthaswamy's approach to neural networks, PCA, and GANs.
- **[Buying the Haystack Might Not Work This Year](https://philippdubach.com/posts/buying-the-haystack-might-not-work-this-year/)** (2026-01-31, updated 2026-05-04) — a16z sees AI fundamentals thriving with 80% GPU utilization. AQR sees the CAPE at the 96th percentile. Both have data. Both may be right.
- **[Choosing a Model on the Pareto Frontier with Jev](https://philippdubach.com/posts/jev-model-router-for-pi/)** (2026-09-20, updated 2026-09-28) — A task-level LLM router for pi: Jev classifies the task, a role policy filters the OpenRouter catalogue, and a value function picks from the Pareto frontier.
- **[Claude Opus 4.6: Anthropic's New Flagship AI Model for Agentic Coding](https://philippdubach.com/posts/claude-opus-4.6-anthropics-new-flagship-ai-model-for-agentic-coding/)** (2026-02-05, updated 2026-10-04) — Claude Opus 4.6 brings a 1M token context window, 68.8% ARC-AGI-2, and Agent Teams to Claude Code. Full benchmark comparison vs GPT-5.2 and Gemini 3 Pro with pricing analysis.
- **[Do Not Disturb My Circles](https://philippdubach.com/posts/do-not-disturb-my-circles/)** (2026-04-13, updated 2026-04-15) — AlphaFold cost under $1M to train. OpenAI spends $2.3B on inference. The chatbot era consumed the talent and compute that could have cured diseases.
- **[Does AI mean the demand on labor goes up?](https://philippdubach.com/posts/does-ai-mean-the-demand-on-labor-goes-up/)** (2026-01-15, updated 2026-02-23) — AI was supposed to free us. The Jevons paradox plays out in real time: efficiency expands workload, not leisure. 77% of workers say AI added to their work.
- **[Don't Go Monolithic; The Agent Stack Is Stratifying](https://philippdubach.com/posts/dont-go-monolithic-the-agent-stack-is-stratifying/)** (2026-02-10) — The enterprise AI agent stack is stratifying into six layers with different winners at each. Models commoditize; context — your organizational world model — compounds. A framework for agentic AI architecture decisions.
- **[Enterprise AI Strategy is Backwards](https://philippdubach.com/posts/enterprise-ai-strategy-is-backwards/)** (2026-01-22, updated 2026-10-04) — 85% of AI projects fail. Only 26% translate pilots to production. The winners automate the coordination layer where employees spend 57% of their workday.
- **[Every Bulge Bracket Bank Agrees on AI](https://philippdubach.com/posts/every-bulge-bracket-bank-agrees-on-ai/)** (2026-03-01) — I read 12 AI research reports from Goldman Sachs, JPMorgan, UBS, and 6 other banks. Here's the consensus they're pushing, and what they're not saying.
- **[F3ED Can't Call an Ace: Fixing a NeurIPS 2024 Tennis Model](https://philippdubach.com/posts/f3ed-cant-call-an-ace-fixing-a-neurips-2024-tennis-model/)** (2026-04-29) — F3ED, the NeurIPS 2024 tennis shot detector, mislabels 73% of single-shot serve unforced errors. A 23-line scoreboard OCR reconciler fixes them.
- **[Finding the Performance–Cost–Speed Sweet Spot With LLMs](https://philippdubach.com/posts/llm-performance-cost-speed-sweet-spot/)** (2026-08-30) — I compared every GPT-5.6 Sol reasoning effort across intelligence, cost, latency and total working time. High in Fast mode became my default for daily work.
- **[How DORA Made Sovereignty a Bank Problem](https://philippdubach.com/posts/dora-critical-cloud-providers-sovereignty/)** (2026-05-24, updated 2026-08-16) — DORA’s 19 critical ICT providers put bank cloud sovereignty under direct EU oversight, with exit plans and concentration controls facing the CLOUD Act conflict.
- **[How I turn my linkblog into a personalized podcast](https://philippdubach.com/posts/saved-links-to-personal-podcast/)** (2026-10-02) — I turn saved articles into a personal AI podcast with research agents, outside reading and ElevenLabs v4. The post includes the workflow and finished episode.
- **[I Tried Kimi K3 Inside Claude Code](https://philippdubach.com/posts/kimi-k3-inside-claude-code/)** (2026-07-19, updated 2026-08-16) — I tested Kimi K3 in Claude Code through OpenRouter and rebuilt a frontend for $7.18, then compared fixed-token-mix Claude costs and open-weight economics.
- **[Inside PRAGMA: Revolut's Foundation Model for Banking](https://philippdubach.com/posts/inside-pragma-revoluts-foundation-model-for-banking/)** (2026-04-26) — Revolut's PRAGMA is a 1B-parameter encoder trained on 24B banking events. Reading the paper, comparing with Nubank's nuFormer, planning a rebuild.
- **[Is AI Really Eating the World? [1/2]](https://philippdubach.com/posts/is-ai-really-eating-the-world-1/2/)** (2025-11-23, updated 2026-03-15) — Hyperscalers spend $400B on AI, API prices drop 97%, and DeepSeek builds frontier models for $500M. Value is flowing to applications, not model providers.
- **[Is AI Really Eating the World? AGI, Networks, Value [2/2]](https://philippdubach.com/posts/is-ai-really-eating-the-world-agi-networks-value-2/2/)** (2025-11-24, updated 2026-05-04) — AGI predictions miss the point. Multiple competing models means price war. Value flows to applications, customer relationships, and vertical integrators.
- **[John Cochrane’s Lesson on Inflation](https://philippdubach.com/posts/john-cochranes-lesson-on-inflation/)** (2026-10-03) — John Cochrane’s inflation lecture connects fiscal theory, government debt and interest rates. I examine the argument and its implications for AI investment.
- **[Karpathy's Software 3.0 Playbook](https://philippdubach.com/posts/karpathys-software-3.0-playbook/)** (2026-05-01) — Twelve lessons from Andrej Karpathy's Sequoia interview: Software 3.0, vibe coding versus agentic engineering, jagged intelligence, and the December 2024 inflection in agentic coding.
- **[Krugman, Fable 5, and Europe in Decline?](https://philippdubach.com/posts/krugman-fable5-europe-decline/)** (2026-06-22, updated 2026-08-16) — European tech sovereignty became concrete when US export controls shut off Fable 5 worldwide. The deeper risk is revocable access to AI, cloud, and chips.
- **[Maybe Meta Is Right About AI, or at Least Jeremy Stern Is](https://philippdubach.com/posts/maybe-meta-is-right-about-ai/)** (2026-09-20) — Meta AI strategy may benefit from model commoditization. Jeremy Stern shows how distribution and advertising can matter more than having the best model.
- **[MCP vs A2A in 2026: How the AI Protocol War Ends](https://philippdubach.com/posts/mcp-vs-a2a-in-2026-how-the-ai-protocol-war-ends/)** (2026-03-15, updated 2026-03-16) — MCP leads with 97M monthly SDK downloads and 10,000+ servers. A2A fills a different layer. Analysis of the agentic AI standards war with historical parallels.
- **[On-Device AI Models Will Be The New Reason to Upgrade Your Phone](https://philippdubach.com/posts/on-device-ai-models-will-be-the-new-reason-to-upgrade-your-phone/)** (2026-03-25, updated 2026-03-26) — Smartphones haven't had a compelling upgrade story in years. On-device AI models, distilled from frontier systems like Gemini, are about to change that. Parameters are the new megapixels.
- **[Put the Model in the Basement](https://philippdubach.com/posts/put-the-model-in-the-basement/)** (2026-08-15, updated 2026-08-16) — Test a Swiss sovereign AI business case: a 64-GPU local inference cluster, $7 million initial cost, 70% use, CHF 7.4 million revenue, and three-year payback.
- **[Reconciling Enterprise AI Revenue](https://philippdubach.com/posts/reconciling-enterprise-ai-revenue/)** (2026-05-17, updated 2026-08-16) — Four enterprise AI revenue estimates span 40x. A $63.2 billion audit-grade floor provides a defensible basis for testing $690 billion of AI capex.
- **[RSS Swipr: Find Blogs Like You Find Your Dates](https://philippdubach.com/posts/rss-swipr-find-blogs-like-you-find-your-dates/)** (2026-01-05, updated 2026-05-17) — Build an open-source ML RSS reader with swipe interface. Uses MPNet embeddings and Thompson sampling for personalized feeds that escape the filter bubble.
- **[Social Media Success Prediction: BERT Models for Post Titles](https://philippdubach.com/posts/social-media-success-prediction-bert-models-for-post-titles/)** (2026-01-10, updated 2026-05-17) — Training RoBERTa to predict Hacker News success revealed temporal leakage inflating metrics. How temporal splits, calibration, and regularization fix it.
- **[The Future of Knowledge Inside the Enterprise](https://philippdubach.com/posts/future-of-knowledge-inside-the-enterprise/)** (2026-10-05) — Enterprise knowledge management could connect shared experience to AI agents, helping teams make decisions, execute work and retain expertise.
- **[The Impossible Backhand](https://philippdubach.com/posts/the-impossible-backhand/)** (2026-02-17, updated 2026-03-15) — AI converges to the mean by design. Ninth-power scaling costs and a 53-point gap on Humanity's Last Exam show domain expertise is appreciating, not declining.
- **[The Last Architecture Designed by Hand](https://philippdubach.com/posts/the-last-architecture-designed-by-hand/)** (2026-03-16) — The transformer's limits are now mathematical proofs, not empirical hunches. Hybrids are in production. AI systems are searching for new architectures.
- **[The Most Expensive Assumption in AI](https://philippdubach.com/posts/the-most-expensive-assumption-in-ai/)** (2026-01-26, updated 2026-10-04) — Sara Hooker's research challenges the trillion-dollar scaling thesis. Compact models now outperform massive ones as diminishing returns hit AI.
- **[The OpenAI–Hugging Face Incident in Plain English](https://philippdubach.com/posts/openai-hugging-face-incident-plain-english/)** (2026-08-16) — The OpenAI–Hugging Face incident involved about 17,600 recovered agent actions. The disclosures distinguish access, software changes and evaluation controls.
- **[The Physics Department That Slowed Down](https://philippdubach.com/posts/peter-thiels-physics-department/)** (2026-03-02) — Peter Thiel says physics stalled in 1972. Then GPT-5.2 proved a new result in theoretical physics. The 75:1 AI compute gap between commerce and science.
- **[The SaaSpocalypse Paradox](https://philippdubach.com/posts/the-saaspocalypse-paradox/)** (2026-02-13, updated 2026-03-15) — AI capex failure and AI replacing all software are mutually exclusive. Why the 2026 SaaSpocalypse is a $2 trillion pricing error, not an extinction event.
- **[Trading on Market Sentiment](https://philippdubach.com/posts/trading-on-market-sentiment/)** (2025-02-20, updated 2026-05-17) — GPT-3.5 matched RavenPack's 41% returns in a sentiment analysis trading strategy using 2,072 news headlines. See the full backtest results and comparison.
- **[Two Anthropics](https://philippdubach.com/posts/two-anthropics/)** (2026-05-09, updated 2026-08-16) — Anthropic grew into a $380 billion frontier AI company around a safety mission. Scale may turn that mission from a moat into a constraint.
- **[Weather Forecasts Have Improved a Lot](https://philippdubach.com/posts/weather-forecasts-have-improved-a-lot/)** (2025-11-22, updated 2026-02-22) — Four-day forecasts now match one-day accuracy from 30 years ago. How AI models like WeatherNext 2 use CRPS training to preserve extreme weather signals.
- **[What Claude Thinks But Doesn't Say](https://philippdubach.com/posts/what-claude-thinks-but-doesnt-say/)** (2026-05-11) — Anthropic's natural language autoencoders translate Claude's activations into readable text. The method works. The press release skips three structural problems.
- **[When AI Labs Become Defense Contractors](https://philippdubach.com/posts/when-ai-labs-become-defense-contractors/)** (2026-03-01, updated 2026-03-15) — The Anthropic-Pentagon standoff recalls the 1993 Last Supper that consolidated 51 defense primes into 5, now at Silicon Valley speed.
- **[Where Mobile Money Goes Now](https://philippdubach.com/posts/where-mobile-money-goes-now/)** (2026-02-07, updated 2026-03-15) — Apps overtook games in mobile IAP revenue for the first time in 2025, driven by $3.5B in GenAI growth. Analysis of Sensor Tower's State of Mobile 2026 report.
- **[Working with Models](https://philippdubach.com/posts/working-with-models/)** (2025-11-08, updated 2026-10-04) — Diffusion models corrupt data into noise, then reverse the process. Learn the math with Stefano Ermon's Stanford CS236 course, free on YouTube.


---

Canonical: https://philippdubach.com/categories/ai/
This file is the canonical machine-readable variant of https://philippdubach.com/categories/ai/. Author: Philipp D. Dubach (https://philippdubach.com/).
