Every token an AI writes pushes data through physical chips, so a 500-word answer burns computation a 5-word one never touches. A bigger model moves more data for each token, too.
Providers batch many requests to keep the hardware busy, but a batch takes time to fill — cheaper compute, slower answer.
The bill never follows the model's name — it follows the work: how many tokens, how big the model, how crowded the servers.