It worked because the semantic layer did the resolution, not the agent. More correctly, a human resolved the metric definition and encoded it into the semantic layer, so the agent didn’t have to guess.

Someone has to decide, and it shouldn’t be the agent

When the data supports multiple answers, but only one of them is valid, the agent is forced to choose. That decision cannot be trusted: the agent does not have the authority to make it, it cannot be held accountable for it, and it cannot be audited afterward because its reasoning leaves no durable trace. Even a perfectly deterministic LLM would produce a consistent answer that no one approved.

The easy answer is to supervise the agent with a human. This adds accountability, but also cost. Humans are slow and expensive, errors still pass through and the resolution logic remains in someone’s head rather than in the system. Worse, accountability often lands with the analyst, sales rep or revenue operator who reviewed the answer, rather than with the person who has the authority to define it.