Why Your AI Bill Will Double Before It Gets Better
Listen to episode
About this episode
In this episode, we're joined by Josh Collier, FinOps Lead at Superhuman (formerly Grammarly), to explore what it really costs to run AI at scale and why the rules of the game changed faster than anyone expected.
We discuss how AI token costs dropped 80% in two years, why that trend has sharply reversed with frontier models doubling in price, and how Josh rebuilt a single LLM workflow that cost $400k a month down to $80k by rethinking the architecture. He also shares how a cost calculator built in 15 minutes transformed the way his team estimates spend before running experiments, and why research-led optimization is the only kind that works without degrading the product.
Along the way, we cover hidden costs most teams miss, the trade-off between Azure reserved capacity and OpenAI Priority Processing, why fixed subscription pricing is broken in an AI-native world, vendor lock-in risk, and what OpenAI's Guaranteed Capacity announcement really signals about where vendor relationships are heading next.
Superhuman: https://superhuman.com
Josh Collier: https://www.linkedin.com/in/josh-collier-945b7029/
Demetrios: https://www.linkedin.com/in/dpbrinkm
Timestamps:
[00:00] OpenAI Guaranteed Capacity: what's really going on
[01:04] Josh's path into AI FinOps
[02:48] Token costs: the 80% price drop
[04:16] Why costs will only go up
[05:06] External LLMs as financial risk
[07:16] Why subscription pricing is dead
[08:22] The data residency fee nobody notices
[09:33] The cost calculator built in 15 minutes
[10:24] How it changed dev team speed
[13:00] Tracking costs by service and team
[15:33] $400k workflow rebuilt for $80k
[17:13] Why only research can optimize tokens
[20:00] Speculative decoding win
[23:11] One bad query, $40k gone
[26:00] Why Azure PTU was exhausting
[28:59] Shadow traffic load testing
[29:07] Priority processing: no brainer
[31:10] Guaranteed capacity: lock-in signal?
[32:18] The danger of multi-year AI deals
[33:28] Vendor-agnostic proxy as exit strategy
More AI podcast episodes
Browse all →Want to find AI jobs?
Join thousands of AI professionals finding their next opportunity