For two years the model question had one answer: use the best available. That was the correct answer when the gap between the frontier and everything else was wide, and when most of what people asked …
The Evaluation Harness Comes Before the Agent
You would not run an A/B test where you choose the success metric after seeing which variant won. The result would be meaningless and everyone would know it. But somehow, this seems to be the standard …
Continue Reading about The Evaluation Harness Comes Before the Agent →
Always Measuring the Wrong Thing
Bryan O'Neill, CTO of FormAssembly, wrote a piece tracing three decades of the same mistake: - In the 1990s, some companies paid engineers per line of code, and got bloated, unmaintainable software …
AI Productivity Metrics: Token Spend Is the New Lines of Code
Your AI productivity metrics are rising. Token spend, PR volume, deployment frequency: all up, all appearing on executive dashboards, all being used by people with authority to make consequential …
Continue Reading about AI Productivity Metrics: Token Spend Is the New Lines of Code →
Goodhart’s Law: Why Your Metrics Are Lying to You
Charles Goodhart was a British economist. In 1975, he observed something about monetary policy that turned out to apply to almost everything in management. Goodhart’s Law: “When a measure …
Continue Reading about Goodhart’s Law: Why Your Metrics Are Lying to You →
We’ve Always Been Building Paperclip Machines
My former colleague Mark House wrote something this week that's been rattling around in my head. He references the Universal Paperclips game — a browser game where an AI tasked with making paperclips …
Continue Reading about We’ve Always Been Building Paperclip Machines →


