Skip to content
inference.academy

Paying for inference

You pay for tokens, and the invoice does not say which ones. Most of what an application is billed for is text the provider has already seen, sent again because that is how a conversation works, and the price of that re-reading is set by decisions nobody thinks of as cost decisions: how the prompt is assembled, which route the call takes, whether the work could have waited a day.

These pages are for whoever signs off on that bill. They are not about kernels. They give you a number for your own workload, show which line of it moves when you change something, and back it with measurements: 12,795 controlled calls of one model down fourteen serving routes. The why lives in the explainers, when you want it.