Overview
- Friday’s reporting confirmed Microsoft Research found frontier models cannot reliably predict their own token consumption and identical tasks can vary widely in token spend.
- Academic work shows agentic workflows use many more tokens than simple chats, with coding agents consuming roughly 1,000 times the tokens of ordinary code assistance.
- Enterprises are already seeing material overruns with named examples of organizations exhausting annual AI allocations and imposing emergency caps on spending.
- Industry responses include adding runtime visibility, per-workflow attribution, automatic spending guardrails, mixed-model routing, and the formation of AI FinOps teams to forecast dynamic costs.
- Vendors such as Dell are pitching on-prem and deskside agent deployments with commissioned analyses claiming large token-cost cuts, but those savings are vendor-backed and should be weighed against migration, compliance, and maintenance trade-offs.