agent experience
Your product didn't get worse
Customers still like your product and keep renewing. Meanwhile, their expectations may be changing in ways your usual signals won't reveal.
Your coding agent extension needs continuous evaluation
Your extension passed its evals. Days later, its code is identical, but can you still trust the result?
The red pen
AI can keep an entire working session in context. Yet changing one sentence still requires directions that sound like a confused treasure map.
Your app is losing its grip on your customers
Your customers still use your app and renew their contracts. Meanwhile, the work they once did in it might already be moving somewhere else.
Does every agent skill need an eval?
A skill can make an agent better and still be a bad trade. How much evidence should you expect before trusting someone else's skill?