When you build an extension for an AI coding agent, it’s tempting to include its invocation in your evaluation criteria. You built a tool to help the agent. Surely, if the agent uses it, that’s a good thing? Not necessarily. Your eval should describe the user’s success, not your product’s behavior.

Say you’re evaluating a scenario where a developer asks an agent for architectural guidance. You provide a documentation search tool that you believe will help, so one of your evaluation criteria could be:

The agent uses the documentation search tool.

But does the developer actually care? Probably not. They care whether the architectural guidance is correct and complete for their situation, and perhaps whether getting there was reasonably cheap. If the agent can provide perfect guidance from what it already knows, calling your tool makes the experience worse, because it consumes additional tokens without improving the result. By scoring the invocation, you’ve made your eval prefer your product even when the user doesn’t.

Invocation isn’t success

Suppose the agent calls your tool. Great. What happened next? Did it use the result? Did the result change its answer, and did the answer become better? Was the tool even necessary? Invocation alone answers none of these questions.

Worse, treating invocation as success can hide exactly the information your eval is supposed to uncover. Imagine the agent needs current information to complete a task. If your criterion says:

The agent calls our tool for current information.

a failure tells you only that your tool wasn’t called. A criterion such as:

The answer uses the current version of the API.

tells you whether the agent has the capability the user needs.

Perhaps the agent already knows the answer, or it finds the information another way. Agents can be surprisingly resourceful. If either path produces the right result, the user got what they needed and your extension wasn’t necessary for that run. An invocation criterion would report that run as a failure, even though it just showed you a part of the task where your extension adds nothing.

Start with the job, not the tool

When defining an eval, start with the user’s need. What job are they trying to do, and how would they determine whether they succeeded? Those answers should become your evaluation criteria. Notice, that none of them mention your extension.

Run that eval against the agent without your extension first. You might discover that the agent already handles parts of the scenario perfectly well. That tells you where an extension isn’t needed and lets you focus on the gaps where it can actually improve the experience.

Only then introduce your extension and run the same eval again. For every scenario, look at two dimensions: quality and cost. Your extension might improve quality, reduce task cost, or both. It can also make either one worse. Without the baseline, you don’t know whether your extension solved a problem, moved it around, or introduced one.

Put invocation where it belongs

None of this means you shouldn’t measure invocation. You absolutely should. If you’ve established that your extension improves the outcome at an acceptable cost, you want the agent to use it. After all, the user went through the trouble of discovering and acquiring your extension, and an extension that consumes context but rarely gets invoked isn’t particularly useful.

But treat invocation as a diagnostic that explains the score rather than a part of it. Capture it in the trajectory. If your evaluation system supports unscored checks, surface invocation there so you don’t have to dig through every trajectory manually. Then you can separate two very different questions: did the user get a better result? and did the agent use the thing we built to help it get there? The first belongs in your score. The second helps you explain and improve it.

If your extension creates measurable lift when invoked but isn’t invoked reliably, you have a discovery or invocation problem, and that’s worth improving. If it’s invoked reliably but improves neither quality nor cost, you have a much more fundamental problem. Don’t let a high invocation rate disguise it.

Beware the metric you want to grow

Tool owners have skin in the game. Once an extension exists, it’s easy for its discovery and invocation to become product metrics. More installations and more invocations look like growth, and growth demonstrates that the product is valuable and justifies further investment.

Then that thinking creeps into the eval. Invocation becomes an evaluation criterion, so higher invocation produces a higher score, and the higher score becomes evidence that the extension works. You’ve built a measurement system that validates the assumption it was supposed to test.

So make the extension earn its place. Does it improve Agent Experience for the scenario? In other words, can the agent produce a better result, or the same result for less money? Those are tangible improvements. And ultimately, everyone understands dollars.

Measure invocation to understand how your extension contributes to those improvements. Optimize invocation once you’ve demonstrated that using the extension actually helps. But don’t confuse your product being used with the user’s problem being solved. Your eval should describe the user’s success, not your product’s behavior.