Putting an AI Agent on a Website: What I'd Build First, Train Second, and Watch Forever

An AI agent from real tickets, launching it on a quiet returns page, and treating the first month as probation.

0:00

The first time I trained the agent on a real question, I fed it our help center and walked away. That was the mistake. The bot learned our articles, but it also learned how a stranger reads them. Shoppers asked the way they ask, not the way we wrote. I rewrote the system prompt around how people actually talked - clipped abbreviations like "plz" and "pls ref," brand names they spelled wrong because they were typing on a phone, and the fishing-phrase "just checking if I can get a refund" that almost always meant a real complaint hiding underneath.

e-commerce customer support
The system prompt became the rulebook, and the rules came from the tickets.

  • Tone. Short sentences, no marketing voice, no "we'd love to help." Match the agent on the floor.
  • Refusals and escalation. Pricing disputes, account changes, and complaints a human agent go straight to a person, and any message with "refund," "cancel," or "lawyer" hands off in one click. I retested every Friday against the same five shopper questions and patched the prompt when an answer drifted. One Friday I caught it repeating a shipping cutoff that we'd already moved up by two hours, and another time it kept applying a "first order 10% off" code twice because the test cart and the live cart were wired differently. Those catches became a checklist I run before each weekly patch.

Launch Day, Where the Handoff Matters Most

The first launch I did was on the busiest page of the site. That was the mistake. The second time around, the agent went onto a low-traffic returns page first and watched from there for a week. I read every transcript, not the summary. When a draft looked safe, a human sent it. When it didn't, the bot stayed quiet.

AI Replies or Human Takeover outbound campaign features
By the end of that week I trusted the handoff more than the answers, which is the part that actually matters on launch day. A confident wrong reply burns trust faster than a slow honest one, and the bot staying quiet when a human takes the chat is the signal that the wiring under the widget is sound.

The first month taught me more than the build did. The widget was live, the tickets were answered, and I still felt like I was missing something. Then a shopper asked the bot a question it had nailed in testing, and it made up a return window that did not exist. I caught it because I was reading every transcript, not the summary dashboard. A useful agent is not a launch event; it is a habit.

The metrics I watch are the boring ones.

  • Handoff rate. If it climbs on a Tuesday, a policy page probably changed and nobody told the bot.
  • Unanswered rate. The questions the widget shrugs at are the next training set, not noise to ignore.
  • Resolution without a follow-up. A shopper who gets an answer and walks away is the only score that matters. The transcripts are the gold. I read them on Friday mornings with coffee, looking for the moment a shopper rephrased the same question twice because the bot kept misreading them. Those rephrases become new training examples, and the prompt gets a small patch. The system stays alive because I keep feeding it the questions real people actually ask. The thing I wish I'd known on day one is that watching is the job, and the month after launch is probation.