evals
An eval score is only half the result
Your agent's score improved. Its inference bill increased 15 times. Before you celebrate, decide what the improvement is actually worth.
Your agent eval is more than its result
Your agent completed the task and every grader passed. What happened before the score reached 100%?
Is your eval lying to you?
Your eval passes on every commit. But what did its graders actually prove?
What's your minimal viable model?
A better AI model appears, so you switch. But how would you know whether its extra intelligence makes any difference to your work?
Test whether your knowledge skill actually works
Your knowledge skill gives reasonable answers to the questions you've tried. What would you find in the parts of the source you didn't think to test?