LLM token

A token is the basic unit a large language model, or LLM, reads and writes. A token is typically a word fragment, not a whole word. Both input and output tokens count toward usage. The prompt, any retrieved knowledge-base content, and the model’s reply all add up to the total. Usage-based AI pricing is commonly tied to token volume, so conversation length, system prompts, and retrieved context all affect cost. Example: passing a long call transcript to a model for quality scoring consumes far more input tokens than a short chat exchange does. A ten-minute call transcript might use thousands of tokens before the model even produces a response. Every model also has a fixed context window, which is the maximum number of tokens it can process at once. That limit controls how much transcript or history the model can consider, and it factors directly into latency.

A token is the basic unit a large language model, or LLM, reads and writes. A token is typically a word fragment, not a whole word. Both input and output tokens count toward usage. The prompt, any retrieved knowledge-base content, and the model’s reply all add up to the total. Usage-based AI pricing is commonly tied to token volume, so conversation length, system prompts, and retrieved context all affect cost. Example: passing a long call transcript to a model for quality scoring consumes far more input tokens than a short chat exchange does. A ten-minute call transcript might use thousands of tokens before the model even produces a response. Every model also has a fixed context window, which is the maximum number of tokens it can process at once. That limit controls how much transcript or history the model can consider, and it factors directly into latency.

Support teams are no longer a cost centre.

See how AI Support Agent, Agent Assist, and VoC Analyst work on your own ticket volume.