Innovation Matters
Essays on AI, developer tools, and thoughtful software.
An eval score is only half the result
Your agent's score improved. Its inference bill increased 15 times. Before you celebrate, decide what the improvement is actually worth.
Your agent eval is more than its result
Your agent completed the task and every grader passed. What happened before the score reached 100%?
Is your eval lying to you?
Your eval passes on every commit. But what did its graders actually prove?
Not every app in an enterprise is an enterprise app
An app used in an enterprise and an enterprise app may sound like the same thing. Confuse them, and you'll badly underestimate what it takes to replace one.
A knowledge cutoff is a ceiling, not a snapshot
A model's knowledge cutoff looks precise. But how much does that date actually tell you about what it knows?