The first time I trained the agent on a real question, I fed it our help center and walked away. That was the mistake. The bot learned our articles, but it also learned how a stranger reads them. Shoppers asked the way they ask, not the way we wrote. I rewrote the system prompt around how people actually talked - clipped abbreviations like "plz" and "pls ref," brand names they spelled wrong because they were typing on a phone, and the fishing-phrase "just checking if I can get a refund" that almost always meant a real complaint hiding underneath.

- Tone. Short sentences, no marketing voice, no "we'd love to help." Match the agent on the floor.
- Refusals and escalation. Pricing disputes, account changes, and complaints a human agent go straight to a person, and any message with "refund," "cancel," or "lawyer" hands off in one click. I retested every Friday against the same five shopper questions and patched the prompt when an answer drifted. One Friday I caught it repeating a shipping cutoff that we'd already moved up by two hours, and another time it kept applying a "first order 10% off" code twice because the test cart and the live cart were wired differently. Those catches became a checklist I run before each weekly patch.
Launch Day, Where the Handoff Matters Most
The first launch I did was on the busiest page of the site. That was the mistake. The second time around, the agent went onto a low-traffic returns page first and watched from there for a week. I read every transcript, not the summary. When a draft looked safe, a human sent it. When it didn't, the bot stayed quiet.

The first month taught me more than the build did. The widget was live, the tickets were answered, and I still felt like I was missing something. Then a shopper asked the bot a question it had nailed in testing, and it made up a return window that did not exist. I caught it because I was reading every transcript, not the summary dashboard. A useful agent is not a launch event; it is a habit.
The metrics I watch are the boring ones.
- Handoff rate. If it climbs on a Tuesday, a policy page probably changed and nobody told the bot.
- Unanswered rate. The questions the widget shrugs at are the next training set, not noise to ignore.
- Resolution without a follow-up. A shopper who gets an answer and walks away is the only score that matters. The transcripts are the gold. I read them on Friday mornings with coffee, looking for the moment a shopper rephrased the same question twice because the bot kept misreading them. Those rephrases become new training examples, and the prompt gets a small patch. The system stays alive because I keep feeding it the questions real people actually ask. The thing I wish I'd known on day one is that watching is the job, and the month after launch is probation.