Running an AI model on a rented GPU is a trade between speed, how many people you serve at once, and money. Describe your situation, and the page works out the trade.
I serve Pricing a model is not serving it: the second one is arithmetic over its config, never run here. on one that costs me an hour. A typical prompt is Log scale, 128 to 32 768. Everything the model reads before it answers: system prompt, history, the question. tokens and the answer A seat reserves room for prompt plus answer. ; about The share of each prompt the card has already seen: a shared system prompt, the same examples. A property of your traffic, not of the card. of each prompt repeats something the card has already seen.
Change anything dotted. Every answer below follows. A token is about three quarters of an English word.
More assumptions, and the numbers behind the answers
How fast should each word appear?
People reading a chat notice anything slower than about 50 ms per token. A batch job only needs its deadline, so it can trade speed for seats.
What does a million tokens cost at that speed?
Only tokens that arrived on time count. A faster promise means fewer people share the hourly bill, so each token costs more.
It got slow. Where do you look first?
Answer what your dashboard shows; don't know is fine. The page names the one cause that fits your operating point, and the knob for it.
I have the numbers, let me type them
Leave a field empty for No data. On kind that is most of them: the stub exports the two queue gauges and no histogram.