Get the Best of Data Leadership
Stay Informed
Get Data Insights Delivered
AI agent costs are shaped by behavior. Every decision an agent makes, to reason, retrieve data, call a tool, or retry an action, can change the cost of completing a task, often without any change to the underlying model price. An agent that starts consuming three times as many tokens per conversation may also be taking more steps or carrying more context into each interaction, and by the time that shows up as an unusual number on the invoice, it may have been happening for days.
That gap between behavior and invoice is what EY's fifth US AI Pulse Survey found this summer: 82% of senior leaders at organizations investing in AI are concerned about token usage and related costs, but only 64% say their organization actively monitors usage with clear budgets and guardrails. For AI agents specifically, that gap widens as agents get more autonomous.
In May, we wrote about how to understand the cost of running AI agents. The foundation there is attribution: knowing spend by agent, user, workflow, and conversation, which answers questions a billing dashboard can't, like which agent or workflow generated the spend and which conversation consumed the tokens. But attribution only tells you where spend came from. As agents become more autonomous, teams also need to know when something changes. That's where cost anomaly detection comes in.
From attribution to detection
Behavior changes first, token usage follows, and cost follows after that. Historically, teams caught these shifts after the fact: someone noticed a strange number and dug through logs to reconstruct what happened, by which point an agent may have generated days of unnecessary spend. The reverse matters too. A sudden drop to zero usage can look like a savings story for a week, until someone notices the report it was supposed to generate never ran, usually because of a broken integration or failed deployment, not a genuine win. Cost anomaly detection continuously monitors usage patterns and flags meaningful changes in either direction, automatically, with nothing to configure first: every agent, user, and workspace gets its own expected range, derived from its own history, so detection starts as soon as there's enough data. Worth noting since these numbers get discussed in dollars: the figures behind detection are estimates for spotting trend changes, not numbers that reconcile to a provider invoice.
What should teams monitor?
A useful detection system looks beyond total spend, at the behaviors underneath it:
- Cost per conversation: more tokens or model calls for the same type of interaction
- Token usage: a jump in average input or output consumption
- Agent steps: more steps than usual to complete the same task
- Model calls: a workflow invoking a model more frequently
- Data quality: a change in inputs that increases retries or expands prompts
- Observability signals: missing traces or metadata gaps that make behavior harder to explain
- Conversation volume: usage that spikes, drops, or stops unexpectedly
- Conversation history: agents carrying substantially more context from one interaction to the next
Together, these signals answer a better question than "why did our AI bill increase?" They answer: what changed in the way our agents are behaving? The goal is to flag meaningful deviations, tie them to the agent or workflow behind them, and give the right team enough context to investigate, not to flag every time costs move.
From detection to explanation
A signal only matters if it's easy to act on. Every anomaly links straight to the conversations behind it, and because the same system already understands agent activity, teams can ask a plain-language question, like which agent cost the most this week, and get a ranked answer instead of pulling a spreadsheet. None of this lives in a separate cost silo, either: anomalies inherit the same workspace and agent-level permissions already configured, and sit alongside the same agent view teams use for everything else, including the data an agent accessed and its quality and freshness signals. For a governance lead, that means one more signal in a view they already know how to read, not a new tool to learn.
AI cost management is becoming an operational-control problem
Gartner's research backs up the urgency: 56% of AI-engaged organizations have no AI FinOps practice or tooling in place, and the firm expects 60% to encounter unforeseen cost overruns through 2029 as opaque vendor pricing collides with weak consumption tracking. The pattern holds across this research: visibility alone rarely changes behavior, but pairing automatic detection with a direct path to the team that can act on it does.
From visibility to control
Knowing what an agent spent was never really the goal. Knowing when its behavior changes, understanding why, and getting that information to the person who can act on it: that's the goal, and it's why we built Cost Anomaly Detection as the next layer on top of attribution rather than as a separate report: observe behavior, detect change, understand the cause, take action.
Cost Anomaly Detection is available now for Bigeye customers through the Agent Trust Hub. If you're trying to catch these shifts before they land on next month's invoice, reach out and we'll show you how detection reads against your own agent traffic.
Sources:
How to track AI agent costs and token usage, Bigeye, May 19, 2026
Data Intelligence Monthly: Executive Insights on AI FinOps and Tokenomics (ID G00860930), Gartner, August 2026
AI FinOps: Why Cloud Cost Optimization Recommendations Don't Get Implemented, and How AI Agents Can Fix It (ID G00852511), Gartner, August 2026
Monitoring
Schema change detection
Lineage monitoring




.png)
