Essays examine Claude benchmarks, scaling laws and enterprise AI strategy. Related articles cover AI agents, model economics, recommendation systems and sentiment trading. Other essays examine AI investment returns, labor demand and computer vision.
AI
AI articles cover model benchmarks, enterprise strategy and AI investment returns.
- 65% of Hacker News Posts Have Negative Sentiment, and They Outperform
Sentiment analysis of 32,000 Hacker News posts shows 65% skew negative and earn 27% more points. Six transformer and LLM models tested, full data included.
- AI Can Now Design Drugs in Seconds; We Still Can't Tell You If They Work.
IsoDDE doubles AlphaFold 3 on hard benchmarks and beats physics-based gold standards. But no AI drug has FDA approval. What $4B in pharma deals actually mean.
- AI Capex Arms Race: Who Blinks First?
Alphabet's free cash flow is on track to fall 90% in 2026. Amazon's is at $11B. $690B in AI capex is cannibalizing the cash that justified these valuations.
- AI Consciousness Is Not a Safety Property
AI consciousness cannot be settled by safety concerns. A skeptical look at model welfare, LLM self-reports and systems that may become harder to dismiss.
- AI Models Are the New Rebar
Qwen 3.5-35B runs on a gaming PC and matches Claude Sonnet 4.5. When the commodity version is 95% as good and 97% cheaper, you have a pricing problem.
- AI Models as Standalone P&Ls
OpenAI lost $11.5B in one quarter. But Anthropic CEO Dario Amodei argues each AI model is independently profitable. The assumptions behind the math.
- Apple's AI Bet: Playing the Long Game or Missing the Moment?
Apple's $157B cash pile and Gemini-powered Siri shift show a restrained AI strategy. Analysis of whether Apple wins as AI models become commodities.
- Aschenbrenner's Receipts
Leopold Aschenbrenner made dated AGI, infrastructure, and political forecasts in Situational Awareness. I assess which calls held up by May 2026, two years on.
- Bandits and Agents: Netflix and Spotify Recommender Stacks in 2026
How hybrid recommender systems balance multi-armed bandits against LLM inference cost economics in 2026. An analysis of Netflix recommendation algorithm architecture and Spotify's AI DJ recommender system.
- Beyond Vector Search: Why LLMs Need Episodic Memory
Context windows aren't memory. EM-LLM's episodic architecture, knowledge graph tools like Mem0 and Letta, and why vectors fail for sequential data.
- Book Review: Why Machines Learn
Why Machines Learn explains ML math with real equations and geometric intuition. Review of Ananthaswamy's approach to neural networks, PCA, and GANs.
- Buying the Haystack Might Not Work This Year
a16z sees AI fundamentals thriving with 80% GPU utilization. AQR sees the CAPE at the 96th percentile. Both have data. Both may be right.
- Choosing a Model on the Pareto Frontier with Jev
A task-level LLM router for pi: Jev classifies the task, a role policy filters the OpenRouter catalogue, and a value function picks from the Pareto frontier.
- Claude Opus 4.6: Anthropic's New Flagship AI Model for Agentic Coding
Claude Opus 4.6 brings a 1M token context window, 68.8% ARC-AGI-2, and Agent Teams to Claude Code. Full benchmark comparison vs GPT-5.2 and Gemini 3 Pro with pricing analysis.
- Do Not Disturb My Circles
AlphaFold cost under $1M to train. OpenAI spends $2.3B on inference. The chatbot era consumed the talent and compute that could have cured diseases.
- Does AI mean the demand on labor goes up?
AI was supposed to free us. The Jevons paradox plays out in real time: efficiency expands workload, not leisure. 77% of workers say AI added to their work.
- Don't Go Monolithic; The Agent Stack Is Stratifying
The enterprise AI agent stack is stratifying into six layers with different winners at each. Models commoditize; context — your organizational world model — compounds. A framework for agentic AI architecture decisions.
- Enterprise AI Strategy is Backwards
85% of AI projects fail. Only 26% translate pilots to production. The winners automate the coordination layer where employees spend 57% of their workday.
- Every Bulge Bracket Bank Agrees on AI
I read 12 AI research reports from Goldman Sachs, JPMorgan, UBS, and 6 other banks. Here's the consensus they're pushing, and what they're not saying.
- F3ED Can't Call an Ace: Fixing a NeurIPS 2024 Tennis Model
F3ED, the NeurIPS 2024 tennis shot detector, mislabels 73% of single-shot serve unforced errors. A 23-line scoreboard OCR reconciler fixes them.
- Finding the Performance–Cost–Speed Sweet Spot With LLMs
I compared every GPT-5.6 Sol reasoning effort across intelligence, cost, latency and total working time. High in Fast mode became my default for daily work.
- How DORA Made Sovereignty a Bank Problem
DORA’s 19 critical ICT providers put bank cloud sovereignty under direct EU oversight, with exit plans and concentration controls facing the CLOUD Act conflict.
- How I turn my linkblog into a personalized podcast
I turn saved articles into a personal AI podcast with research agents, outside reading and ElevenLabs v4. The post includes the workflow and finished episode.
- I Tried Kimi K3 Inside Claude Code
I tested Kimi K3 in Claude Code through OpenRouter and rebuilt a frontend for $7.18, then compared fixed-token-mix Claude costs and open-weight economics.
- Inside PRAGMA: Revolut's Foundation Model for Banking
Revolut's PRAGMA is a 1B-parameter encoder trained on 24B banking events. Reading the paper, comparing with Nubank's nuFormer, planning a rebuild.
- Is AI Really Eating the World? [1/2]
Hyperscalers spend $400B on AI, API prices drop 97%, and DeepSeek builds frontier models for $500M. Value is flowing to applications, not model providers.
- Is AI Really Eating the World? AGI, Networks, Value [2/2]
AGI predictions miss the point. Multiple competing models means price war. Value flows to applications, customer relationships, and vertical integrators.
- John Cochrane’s Lesson on Inflation
John Cochrane’s inflation lecture connects fiscal theory, government debt and interest rates. I examine the argument and its implications for AI investment.
- Karpathy's Software 3.0 Playbook
Twelve lessons from Andrej Karpathy's Sequoia interview: Software 3.0, vibe coding versus agentic engineering, jagged intelligence, and the December 2024 inflection in agentic coding.
- Krugman, Fable 5, and Europe in Decline?
European tech sovereignty became concrete when US export controls shut off Fable 5 worldwide. The deeper risk is revocable access to AI, cloud, and chips.
- Maybe Meta Is Right About AI, or at Least Jeremy Stern Is
Meta AI strategy may benefit from model commoditization. Jeremy Stern shows how distribution and advertising can matter more than having the best model.
- MCP vs A2A in 2026: How the AI Protocol War Ends
MCP leads with 97M monthly SDK downloads and 10,000+ servers. A2A fills a different layer. Analysis of the agentic AI standards war with historical parallels.
- On-Device AI Models Will Be The New Reason to Upgrade Your Phone
Smartphones haven't had a compelling upgrade story in years. On-device AI models, distilled from frontier systems like Gemini, are about to change that. Parameters are the new megapixels.
- Put the Model in the Basement
Test a Swiss sovereign AI business case: a 64-GPU local inference cluster, $7 million initial cost, 70% use, CHF 7.4 million revenue, and three-year payback.
- Reconciling Enterprise AI Revenue
Four enterprise AI revenue estimates span 40x. A $63.2 billion audit-grade floor provides a defensible basis for testing $690 billion of AI capex.
- RSS Swipr: Find Blogs Like You Find Your Dates
Build an open-source ML RSS reader with swipe interface. Uses MPNet embeddings and Thompson sampling for personalized feeds that escape the filter bubble.
- Social Media Success Prediction: BERT Models for Post Titles
Training RoBERTa to predict Hacker News success revealed temporal leakage inflating metrics. How temporal splits, calibration, and regularization fix it.
- The Future of Knowledge Inside the Enterprise
Enterprise knowledge management could connect shared experience to AI agents, helping teams make decisions, execute work and retain expertise.
- The Impossible Backhand
AI converges to the mean by design. Ninth-power scaling costs and a 53-point gap on Humanity's Last Exam show domain expertise is appreciating, not declining.
- The Last Architecture Designed by Hand
The transformer's limits are now mathematical proofs, not empirical hunches. Hybrids are in production. AI systems are searching for new architectures.
- The Most Expensive Assumption in AI
Sara Hooker's research challenges the trillion-dollar scaling thesis. Compact models now outperform massive ones as diminishing returns hit AI.
- The OpenAI–Hugging Face Incident in Plain English
The OpenAI–Hugging Face incident involved about 17,600 recovered agent actions. The disclosures distinguish access, software changes and evaluation controls.
- The Physics Department That Slowed Down
Peter Thiel says physics stalled in 1972. Then GPT-5.2 proved a new result in theoretical physics. The 75:1 AI compute gap between commerce and science.
- The SaaSpocalypse Paradox
AI capex failure and AI replacing all software are mutually exclusive. Why the 2026 SaaSpocalypse is a $2 trillion pricing error, not an extinction event.
- Trading on Market Sentiment
GPT-3.5 matched RavenPack's 41% returns in a sentiment analysis trading strategy using 2,072 news headlines. See the full backtest results and comparison.
- Two Anthropics
Anthropic grew into a $380 billion frontier AI company around a safety mission. Scale may turn that mission from a moat into a constraint.
- Weather Forecasts Have Improved a Lot
Four-day forecasts now match one-day accuracy from 30 years ago. How AI models like WeatherNext 2 use CRPS training to preserve extreme weather signals.
- What Claude Thinks But Doesn't Say
Anthropic's natural language autoencoders translate Claude's activations into readable text. The method works. The press release skips three structural problems.
- When AI Labs Become Defense Contractors
The Anthropic-Pentagon standoff recalls the 1993 Last Supper that consolidated 51 defense primes into 5, now at Silicon Valley speed.
- Where Mobile Money Goes Now
Apps overtook games in mobile IAP revenue for the first time in 2025, driven by $3.5B in GenAI growth. Analysis of Sensor Tower's State of Mobile 2026 report.
- Working with Models
Diffusion models corrupt data into noise, then reverse the process. Learn the math with Stefano Ermon's Stanford CS236 course, free on YouTube.