You can't judge an AI agent until you know what the task is worth


We keep asking whether AI coding agents are good or bad, worth it or not. But good compared to what? Until you know what a task is actually worth to you, you have no way to tell whether the output was a bargain or a rip-off.

Say your agent scaffolds a new project and it costs you $1. Is that good? Most people would shrug and say sure, a dollar is nothing. But what about $3? Or $5? At some point your gut tells you it’s too much, and yet you’d struggle to explain where that line is or why it sits there. That’s the problem. We judge the price without ever pricing the task.

The missing number

Consider a task at hand: what would you pay to have this done, right now, without doing it yourself? Scaffolding is a great example because you can often do it for $0. Run the CLI, answer a few prompts, and you have a working skeleton in seconds. If the agent charges you a dollar to produce roughly the same thing, you didn’t save money, you spent it. The interesting question isn’t “was a dollar cheap?” It’s “was a dollar cheaper than the alternative I already had?” And for scaffolding, you usually did have one.

Now flip it. You let the CLI scaffold the project for free, then hand the boring part to the agent: rename things to match your conventions, wire up the config you always forget, swap in the logging setup you copy from repo to repo. That’s work you’d otherwise do by hand for twenty minutes. Paying a dollar to skip it starts to look like a good trade. Same dollar, completely different verdict, because the task underneath it was worth something different.

Why our pricing instincts fail here

We’re good at pricing coffee. We’ve bought thousands of cups, we know what a fair one costs, and we notice instantly when a café charges double. We have years of repetition to calibrate against. Transport, groceries, a haircut: we’ve priced them so many times the judgment is automatic.

LLMs don’t get that treatment. They’re new, the unit of work is fuzzy, and the same prompt can cost wildly different amounts depending on context length and how many times the agent loops. And it’s hard to build intuition on something that moves that much and that you’ve only been paying for a couple of years.

The pricing itself makes it worse. When you buy coffee, you feel the money leave your hand and you decide, right there, whether it was worth it. With an agent, the payment is delayed. You see a token counter, maybe, but the bill lands weeks later. Very often it’s your employer’s card, not yours, so the signal that normally trains your judgment never reaches you at all. You get the output, but you don’t feel the cost, and you move on.

That’s fine for one task, and in the grand scheme of things a $1 or even $5 doesn’t matter that much. The thing is though, that these small amounts compound. A dollar here, three there, a handful of retries you didn’t think about, repeated across every developer, every day. None of them felt like a decision. Then the invoice arrives at the end of the month and it’s a real number, and nobody can point to the moment they agreed to it.

Two things we should get better at

I don’t propose we should stop using agents or to obsess over every token. It’s to bring back the judgment that delayed, invisible payment quietly took away from us.

First, get better at estimating what a task is worth to you before you hand it off. You don’t need a spreadsheet. Just ask: if I did this myself, how long would it take, and could I do it at all? A task you could knock out in two minutes has a low ceiling. A task that would eat your whole afternoon, or that you genuinely don’t know how to do, is worth a lot more. Once you have that number in your head, even roughly, the agent’s cost stops being an abstract dollar and becomes a comparison you can actually make.

Second, evaluate the agent against the task cost, not the token count. Different models have different token prices, and token use patterns. Some models use more tokens that are cheaper. Some are more concise, other explore more. So in the end, you want to know what a task will cost you, not in tokens, but in money.

So next time your agent hands you a result and you catch yourself asking whether it was worth it, answer a different question first. What was this task worth to me? Get that number right, and “good” and “bad” stop being feelings. They become math.

Others found also helpful: