LLM token
Related terms
Large language model
A large language model is the model class behind modern conversational AI, summarisation, and reply drafting. Trained on broad text corpora, it predicts language fluently but holds no inherent knowledge of a specific business, which is why grounding and retrieval are required before answering customer questions.
Latency
Latency is the delay between a customer’s message and the first useful response. In live chat and voice it directly shapes perceived quality, and it is one of the few AI quality measures a customer notices immediately, well before judging whether the answer was correct.
Prompt
A prompt is the instruction set that defines how an AI agent behaves in a given task, covering role, tone, permitted actions, and output format. Prompts express intent but do not enforce it, which is why hard limits belong in guardrails rather than in prompt wording.
Retrieval-augmented generation
Retrieval-augmented generation grounds answers by retrieving an organisation’s own documents at the moment of answering and passing them to the model as context. It keeps responses current without retraining, and it makes each answer traceable to the specific source the system actually consulted.

Support teams are no longer a cost centre.
See how AI Support Agent, Agent Assist, and VoC Analyst work on your own ticket volume.