Building customer agents that fix, not just answer

Eight principles we build customer agents by, each one earned from something that broke in production.

·

13 min read

Summarize this page with your favourite AI assistant
ChatGPTClaudeGoogle GeminiGrokPerplexity

We build AI agents that talk to customers on chat, email, voice and inside apps. Getting one to work in a demo takes an afternoon. A demo has three tools, a tidy knowledge base and a customer who asks the question you rehearsed.

Production is different. Two help articles disagree, and the agent picks one with full confidence. A customer explains their problem, gets handed to a person, and explains it again. Someone edits the agent on Tuesday and breaks what was fixed on Monday. The line goes quiet while a lookup runs, and the caller asks whether anyone is still there.

None of this shows up in a demo, and most of it isn’t a model problem. Over the past year we have watched our agents fail in each of these ways, changed what we build because of it, and settled on a short list of principles.

This post is that list. Each principle starts with a failure we saw, then what we built, how we check it and what it costs us. At the end there is a ten-question audit you can run on your own agent in an afternoon. We hope it is useful if you are building your own.

What do we mean by an agent, and by fixed?

Anthropic draws a useful line between workflows, where a model follows predefined code paths, and agents, where the model decides which steps to take and which tools to use [1]. Customer agents sit on both sides of that line. Checking an order status is close to a workflow. Working out why a refund failed, and doing something about it, needs an agent.

We also need a definition of done. Most support tooling counts a conversation as resolved when a reply goes out and the customer stops writing. We count it as fixed when the problem stops: the customer doesn’t come back about it, and whatever caused it has changed. Our shorthand is closed means it worked. Everything below follows from taking that definition seriously.

Figure 1. The agent in context. Conversations arrive from chat, email, voice and apps. The agent answers from the knowledge base, the customer’s history and the tools it can call, and hands over to a person, with everything so far, when it should stop. After a conversation ends, a grader with more information than the agent scores it. When something is wrong, a fix is proposed, a person approves it, and it is written back to its source: the knowledge, the agent’s instructions or a tool. Every change runs through a test suite before it reaches a customer.

1. Context is the job, and it has a budget

What broke. An agent was connected to two internal APIs by pasting each API’s full specification into the tool setup, a few thousand lines each. Every endpoint became a tool, and every tool definition went to the model on every turn of every conversation. Nothing errored. Replies got slower, and the agent had dozens of ways to call the wrong thing before it had read the customer’s first word.

The opposite failure is more familiar. An agent that can’t see the order talks about orders in general, confidently, and the customer gives up.

What we built. Agents reach the systems a problem lives in through tools the team adds in the agent builder: authenticated API calls, connections to other systems, and code tools (more on those in principle 4). Conversations carry the customer’s details, the attributes the business routes on, and the customer’s previous tickets, so the agent starts from what is already known.

How we check it. We look at what goes into a turn, not only what comes out: which tools were offered, which were called, and how much of the context the answer needed. A tool that is never called gets removed. Tool definitions stay short and stable, because every word in them is read by the model on every turn and steers what it does [2].

The trade-off. Research on long inputs shows models use information at the start and end of a long context far better than information in the middle [3]. More context isn’t free accuracy. Each system you connect is also a system you have to secure and keep in sync. We would rather give an agent five tools it uses well than fifty it might use.

2. Keep the knowledge honest

What broke. Two articles covered the same refund policy. One had been updated and one hadn’t. The agent picked the old one and answered with full confidence. It wasn’t a hallucination. The source was wrong, and nothing told anyone.

The second failure was quieter. A support agent got a bad AI suggestion, pressed thumbs down, and the suggestion disappeared. The article behind it stayed exactly as it was, so the next customer who asked got the same wrong answer.

What we built. Every time an article is added or changed, it is checked against the rest of the knowledge base for contradictions and outdated content before it reaches a conversation. When someone flags an answer, Brain Guard traces it back to the article behind it, shows the gap or the contradiction, and proposes a fix. A person approves it, and the source changes, not just the reply. Knowledge base analytics show which articles the agent answers from, and where it has nothing to answer from.

How we check it. Two things we watch: how many conflicts are open in the knowledge base, and how long a flagged answer waits before its source changes. The target for the second is the same day.

The trade-off. Some contradictions are deliberate: two regions, two policies. Detection can’t know which one is right, so the person who owns the content decides. We chose approval over automatic edits. It is slower, and it means the agent improves from fixes a person approved, not from guesses the system made.

3. Remember the customer, and only that customer

What broke. A caller reported a problem, hung up, and called back an hour later. The agent greeted them as a stranger and asked for everything again, including what they had said on the first call. To the customer it was the same company. To the agent it was a new conversation.

What we built. Agents recall a customer’s recent conversations across channels, within a window the team sets: how many past tickets, and how far back. Voice agents greet a returning caller by name. When a customer opens live chat, it opens straight into their most recent conversation instead of a blank start.

How we check it. Test scenarios include returning customers: someone who called yesterday, someone who wrote by email and now calls. We check that the agent uses what it knows, and that nothing from one customer’s history appears in another customer’s conversation.

The trade-off. Memory costs storage and governance, and a longer window isn’t automatically better: an old ticket can mislead as easily as it helps. The bigger point is that recognising a returning customer is not the same as verifying them. A phone number or a signed-in session tells you who is probably there. We treat memory as context for the conversation, and keep it apart from anything that authorises a sensitive action.

4. Give the agent hands it can’t improvise with

What broke. A customer asked for a refund. The agent explained the refund policy, accurately and politely. It had no way to issue one, so it described the job instead of doing it. From the customer’s side that is an answer, not a fix, and no amount of better knowledge would have changed it. The agent was missing a tool.

What we built. Three kinds of hands. Code tools: our team writes the code, the agent decides when to call it, and the code runs the same way every time. The model chooses the action; it doesn’t write the action at runtime. Operators: AI that reads conversations and acts on them, from filling in the fields people never fill in, to reassigning, escalating, alerting a team or starting a workflow in another system. And interactive widgets inside the conversation, from order tracking and a cart to a full-screen app, so the customer finishes the task where they asked for help.

How we check it. Code tools are tested like any other code. Every action an agent can take has scenarios in its test suite, including the ones where it shouldn’t act.

The trade-off. Every action is also a risk. An agent that can issue a refund can issue the wrong refund. Deciding which actions run on their own and which wait for a person is a decision per action, not per agent, and today we make it in how each tool is built. We come back to this under what is still hard.

5. Never leave the line silent

What broke. Partway through a call, a caller asked: “where did you go?” The speech pipeline was fast. The model was still generating. Its reasoning step was on for every turn, including simple ones, and the slowest turn of the call was a long read-back confirming an appointment. In chat, the same root cause once showed up differently: an agent sent a customer its reasoning instead of its answer.

What we built. Non-blocking tools, so a voice agent keeps the conversation going while a slow lookup runs instead of leaving dead air. Hold phrases the team configures, for the moments the agent is waiting on a system. Streamed answers in chat. Fallback models for speech-to-text, text-to-speech and the language model, so one provider having a bad minute doesn’t become the customer’s bad minute. Reasoning stays inside the system; only the answer reaches the customer.

How we check it. We measure the gap before every agent turn on real calls, and we listen to the long ones. A turn that is slow because the agent is doing something useful is a different problem from a turn that is slow because it is thinking about nothing.

The trade-off. Turning reasoning down makes simple turns faster and can make hard turns worse. Hold phrases get irritating when overused. Fallbacks are a second path to test. Latency is a product decision, not an infrastructure statistic.

6. Know when to stop, and hand over everything

What broke. An outbound agent offered to transfer a customer to a person, on a campaign where no team was there to take the call. The more common version comes up on almost every customer call we have: the agent collects the details, decides it can’t help, hands over, and the person who picks up asks for the details again.

What we built. A confidence threshold decides who finishes the conversation. On voice, a warm transfer briefs the person before the customer joins, or a cold transfer passes the call straight through. Supervisors can listen to a live call, step in, or steer the AI mid-conversation and let it carry on from there. Live call warnings flag a conversation that needs attention while it is still happening, and can start a workflow. Escalations follow the team’s SLAs.

How we check it. Test suites include handovers, and the question we ask is what the person sees in their first ten seconds: what was said, what was tried, and why the agent stopped.

The trade-off. Set the threshold low and people handle conversations the agent could have finished. Set it high and the agent finishes conversations it shouldn’t have. There is no right number, only the one that matches what a wrong answer costs in that conversation. Narrow agents need narrow rules: an outbound agent shouldn’t offer a transfer nobody can take, and an agent built for one topic should say what it is for instead of guessing.

7. Test every change before a customer hears it

What broke. A voice agent ended a call on its own. The tool it uses to hang up had been described for a different agent, an outbound one, whose rules included ending the call if the person who answered wasn’t the intended customer. When a caller asked our support agent to confirm a name, the model applied that rule and hung up. One sentence, copied from one agent to another, changed what the second agent did.

Tool descriptions are part of the prompt. Copying an agent copies its intent.

What we built. Edits land as a draft, and nothing changes for customers until someone publishes. Every change is saved as a version that can be rolled back, and voice agents keep a record of who changed what. Supervisors build test suites once (an angry customer, a price objection, the wrong language, whatever breaks their agent) and the suites run on every edit. In evaluations, one AI plays the customer and another plays the agent, running off the real instructions, and the report shows where the conversation went wrong. The agent also reviews its recent real calls and points out what to fix.

How we check it. Pass rates before publishing, and consistency across runs, because one success proves little. The τ-bench study introduced pass^k, the chance an agent succeeds on all of k attempts at the same task. It found that leading function-calling agents succeeded on fewer than half of its tasks, and that pass^8 in its retail domain was under 25% [4]. Consistency is the property customers notice.

The trade-off. Suites take time to build and to keep current, and simulated customers are more polite and more predictable than real ones. That is why real failures, like the call above, become new test cases.

8. Count the fix, not the reply

What broke. In a review with a customer’s leadership team, one slide put a deflection rate next to a resolution rate. Half the room read them as the same number. They aren’t. Deflection says the conversation didn’t reach a person. Resolution should say the problem stopped. A conversation where the customer goes quiet and comes back two days later on another channel counts as a success on the first measure and a failure on the second.

What we built. Reopen rate sits next to resolution in support analytics, so a resolution that doesn’t hold shows up. What we are building toward is an outcome for every AI-handled conversation: resolved and confirmed by the customer; resolved and assumed, because they went quiet after a relevant answer; not resolved; or no real request, with greetings, tests and spam left out of the rate.

How we check it. Repeat contacts on the same issue within a set window, counted per issue, not per conversation.

The trade-off. Counting honestly makes the numbers look worse at first. Assumed resolutions are the uncomfortable category, and they are the one worth watching.

How do we evaluate customer agents?

Three layers, because none of them is enough on its own.

Before a change ships: test suites and agent-to-agent evaluations, as in principle 7.

After a conversation ends: a grader scores it. The grader gets more than the agent had: the full transcript, the whole knowledge base, every tool the agent could have called, and no latency budget. So it can work out the answer the agent should have given and point to the moment the conversation went wrong.

People review the graders. Strong language models used as judges agree with human preferences more than 80% of the time, about as often as people agree with each other, and they also show biases towards position, length and their own style of answer [5]. So people review a sample of grades and every grade that is disputed, and their corrections calibrate the grader. Human review still catches what automation misses: a policy that changed last week, or a reply that is technically correct and still wrong for that customer.

What is still hard?

Consistency. An agent that handles a task well once and badly the third time is the failure customers remember. pass^k is a better target than a single success, and it is harder to move.

Trust per action. Agents that act need a way to decide, action by action, what runs alone, what waits for a person, and how each action is undone. Today we draw that line in how each tool is built. It needs to become something a CX team can set and see for itself.

Handover in chat. On voice, a warm transfer briefs the person first. In chat, the person gets the full conversation. A summary they can take in within ten seconds is what we are building toward.

Mixed languages. Multilingual customers switch language mid-sentence. A borrowed word isn’t a request to change language. Our rule: switch immediately when the customer asks, or after a few turns in the other language. It is a rule, and rules have edges.

Knowing a fix worked. Proof that a problem stopped arrives days later, not when the conversation ends. Evaluations are fast. Outcomes are slow.

Conclusion

None of these principles is about the model. A bigger model doesn’t make two articles agree, doesn’t stop a copied sentence from changing an agent’s behaviour, and doesn’t tell the person who picks up why the agent stopped.

What does is the work around the model: context with a budget, knowledge that is checked, memory with limits, hands that can’t improvise, a line that doesn’t go silent, handovers that carry everything, tests on every change, and a count of what actually got fixed.

We will keep writing about each of these as we build them. If you are building your own agent, the audit below is a good place to start.

Appendix: audit your own agent in an afternoon

Ten yes-or-no questions. Each “no” points back to a principle above.

Download the one-page checklist (PDF)

  1. Can your agent look up the specific order, account or ticket the customer is talking about, without asking for something they already gave? (1)

  2. Do you know how much goes into your agent’s context on each turn, and how much of it is tool definitions? (1)

  3. When an article is added or changed, is it checked against the rest of the knowledge base before customers see it? (2)

  4. When someone flags a wrong answer, does the source change the same day? (2)

  5. Does a returning customer get recognised, and is that memory limited to that customer and kept apart from anything that authorises an action? (3)

  6. For every action your agent can take, can you say who allowed it, how it is tested and how it is undone? (4)

  7. Does your voice agent keep the conversation going while a slow lookup runs? (5)

  8. When the agent hands over, does the person see what was said, what was tried and why the agent stopped, within ten seconds? (6)

  9. Does every edit run against a fixed set of scenarios before it reaches a customer, including every scenario that has broken you before? (7)

  10. Do you count a conversation as resolved only when the customer doesn’t come back about the same issue? (8)

References

  1. Anthropic, “Building effective agents”, 19 December 2024. https://www.anthropic.com/engineering/building-effective-agents

  2. Anthropic, “Writing effective tools for AI agents, using AI agents”, 11 September 2025. https://www.anthropic.com/engineering/writing-tools-for-agents

  3. Liu et al., “Lost in the Middle: How Language Models Use Long Contexts”, arXiv 2307.03172, 2023. https://arxiv.org/abs/2307.03172

  4. Yao et al., “τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains”, arXiv 2406.12045, 2024. https://arxiv.org/abs/2406.12045

  5. Zheng et al., “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena”, arXiv 2306.05685, 2023. https://arxiv.org/abs/2306.05685

FAQ

What makes a customer agent different from a chatbot?

A chatbot replies. A customer agent can see the systems the problem lives in, take actions through tools, remember the customer within limits, and hand over to a person with everything so far. It is judged on whether the problem stops, not on whether a reply went out.

How do you test an AI customer agent before it goes live?

Build a suite of scenarios that have broken agents before, run it on every edit, use simulated customers to talk to the agent, and measure consistency across repeated runs, not a single pass. Then grade real conversations after they end with a grader that has more information than the agent had.

Should an AI agent remember customers?

Yes, within a window you set and scoped to that customer. Memory should make the next conversation shorter. It should never be the thing that authorises a sensitive action.