Going Live, Watching the First Hundred Chats, and Fixing What Broke
The first hundred chats taught me more than the sandbox ever did. By chat forty I had a list of pricing hallucinations, two missed escalations, and one very polite refund that should have been a handoff. The agent was useful. It was also, in places, confidently wrong. The operational checklist I wish I'd printed and taped to the wall:
- Sandbox first, with real questions. Pull last week's hardest tickets and run them through before flipping the widget on.
- Always-on human fallback. A button or keyword that hands the thread to a person, with the full transcript, not a clipped screenshot.
- Watch the first hundred conversations live. Track resolution rate, escalations, and any reply where the agent names a price or policy.
- Fix failure modes in order. Wrong answers first, then missed escalations, then tone, then latency.
- Iterate on prompts and knowledge weekly. Each batch of real chats is new training data. The pricing hallucination showed up on chat twenty-three. A customer asked about a bundle discount the help center never mentioned, and the agent quoted a number. I caught it, deleted the answer from the knowledge base, and added an explicit guardrail: any price the agent isn't sure about must be confirmed against the live source before it ships. The missed escalation was worse. A refund dispute sat in the agent's hands for four minutes before the customer typed "human please," and even handoff dropped the thread context. I rebuilt the escalation rule to fire the moment sentiment tipped negative, regardless of who said it first. That fix mattered more than any prompt rewrite I did that week. Looking back, the first hundred chats taught me that the agent only gets useful once you start treating its wrong answers like bugs in production, and the AI Agent for Website: A Guide to Picking, Training, and Going Live walks through that afternoon the same way I learned it - by watching what broke and writing the fix before the next shift.
What I'd Do Differently Next Time
The honest list is short, and none of it is about the tool.
- Start with the messy sources, not the polished ones. I'd hand the vendor the 400-page help center with stale articles on day one and watch what happens before any demo. That's the real test.
- Write the guardrails first, the brand voice second. Tone is a polish step. Guardrails belong in the prompt before any style guide.
- Build the handoff before the outbound campaign. I'd wire the always-on human fallback - the button, the keyword, the full transcript - and test it with a fake angry customer before I ever sent the first proactive message.
- Cap the agent's authority on price and policy. Anything the agent isn't sure about must be checked against the live source before it ships.
- Treat the knowledge base as a product. Stale articles taught the agent stale answers. I'd put someone on rotation to delete, rewrite, and add to it every week. Looking back, training day was the stretch I would not rush again, and the AI Agent for Website: A Guide to Picking, Training, and Going Live says it the same way - that whole walk-through is worth reading before you sign anything.