agents
What if coding agents had a learning mode?
Coding agents can take us from intent to production in minutes. What happens to the understanding we used to build along the way?
Your coding agent extension needs continuous evaluation
Your extension passed its evals. Days later, its code is identical, but can you still trust the result?
Does every agent skill need an eval?
A skill can make an agent better and still be a bad trade. How much evidence should you expect before trusting someone else's skill?
What's volatile about your technology?
A model's knowledge cutoff doesn't tell you how old its knowledge of your SDK is. Technology providers know something more useful.
Your agent should start with a hypothesis, not a plan
Agents can turn a plausible first answer into a plan before current documentation gets a say. By then, correcting an outdated assumption is already expensive.