evals
What's your minimal viable model?
A better AI model appears, so you switch. But how would you know whether its extra intelligence makes any difference to your work?
Test whether your knowledge skill actually works
Your knowledge skill gives reasonable answers to the questions you've tried. What would you find in the parts of the source you didn't think to test?
Your coding agent extension needs continuous evaluation
Your extension passed its evals. Days later, its code is identical, but can you still trust the result?
Does every agent skill need an eval?
A skill can make an agent better and still be a bad trade. How much evidence should you expect before trusting someone else's skill?