Innovation Matters
Essays on AI, developer tools, and thoughtful software.
Your eval should describe the user's success, not your product's behavior
Your eval checks whether the agent called the tool you built, and every run passes. What does that tell you about the developer who asked for help?
Stop sensationalizing
Somewhere along the way, a settings page became an experience and a fix became an unlock. We use big words for small things when there's no need to.
Why pay an agent to run a command you already know?
Asking an agent to list files takes longer than typing ls and costs tokens, only for the agent to run ls anyway. On paper, delegating a command you already know makes no sense.
An eval score is only half the result
Your agent's score improved. Its inference bill increased 15 times. Before you celebrate, decide what the improvement is actually worth.
Your agent eval is more than its result
Your agent completed the task and every grader passed. What happened before the score reached 100%?