I'd ship a smaller first version. The temptation is to turn on every action and every channel on day one. I did, and half of them broke under real load. The Calendly action worked; the WhatsApp handoff didn't, because I hadn't set up the routing queue. Cut over the widget first, watch the transcripts, then add a channel only when the previous one is boring. I'd also write the handoff reasons down before launch, not after. I had a vague "escalate when confused" rule and it escalated constantly. The day I wrote the four reasons on a sticky note and pasted them into the guardrails, my team's inbox went quiet for the first time in months. And I'd keep a "things the agent got wrong this week" doc open from day one. The agent improves only as fast as I read its transcripts and rewrite the instructions that caused the failures. That doc is the training loop. Everything else is just configuration. If you're standing where I was - widget ready, training done, demo behind you - that's left is the boring kind. Read the transcripts. Fix the same three failures on repeat. Don't ship a new channel until the last one is quiet. The agent gets sharper not because the model got smarter, but because you finally started listening to the people on the other side of it.
What an AI agent actually does once it's on the site
I left the widget in shadow mode for four days before I cut traffic over, and I sat with the transcripts open in one window and a notebook in the other. The first hundred chats taught me the difference I had missed in every demo. A shopper on the order-tracking page typed, "where is my package?" The agent read the page she was reading, looked up the order, saw the carrier had scanned it at the sort facility that night, and replied with the right tracking link. That wasn't new - my old widget could have same. Then a different shopper typed, "I need to return this, it's the wrong size." The agent read the order, saw it was inside the return window, generated a prepaid label, and emailed it to her. No human touched it. By chat forty, a third shopper asked for a refund she wasn't entitled to, and the agent stopped, drafted the reason, and paged my support lead with a clean transcript attached. The lead answered in eight minutes and closed the ticket. That handoff was the part no demo had shown me.
Picking the platform without falling for the demo
The demos all look the same after a while. Friendly avatar, slick one-liner about "next-gen conversations," a planted question landing perfectly. I sat through four in a week before I made myself a checklist. It wasn't about which model was smartest in a vacuum. It was whether the platform could plug into the systems I already run, and whether the price made sense if half my tickets actually got deflected. I started with the boring questions. Where do my files live? How does the agent authenticate a real shopper before it pulls an order? Can I write the handoff rule myself, or am I paging support? Does the pricing page actually call Zendesk, Stripe, and my booking tool, or is "integrations" a roadmap slide? The answers sorted the field faster than any benchmark video. The pieces I weighed, in roughly this order:
- Data sources on day one: help docs, Notion, URLs, catalog, past tickets.
- Real actions in my stack: Zendesk, Stripe refunds, Calendly, Slack pings, custom APIs.
- Identity check before the agent sees a shopper's data.
- Channels beyond the widget: WhatsApp, email, Slack, voice.
- Guardrails I write: tone, escalation triggers, no-guess topics, reason code on handoff.
- Analytics: resolution rate, deflection, sentiment, repeat topics. One demo stood out, and it wasn't the one with the best small talk. The rep opened the Actions panel and pointed at Stripe, Calendly, Slack, and a Custom Action for any API endpoint. They let me write a handoff rule out loud - legal, damaged-on-arrival, name mismatch - and the rule took. A couple of competitors flunked the same test. One had no real actions beyond "send to email." Another gated half its integrations behind enterprise. A third's identity check was a checkbox that said "trust me," which is the opposite of identity.
Going live, watching the first hundred chats, and fixing what broke
I didn't flip the switch. I left the agent in shadow mode for four days, watching every transcript while my old widget kept answering. Then I cut over to 10% of traffic for a weekend, then half, then full. The launch I'd imagined - a single toggle, a Slack emoji, done - never happened. The real launch was a slow bleed, and the first hundred chats taught me more than the whole training phase did. The early transcripts looked almost boring, which is how I knew something was off. One shopper asked about a return window and the agent pulled the right order, generated the label, and emailed it before I'd finished my coffee. Another asked the same question three different ways and got three different answers because I'd written three guardrail notes that contradicted each other. I had to pick one, delete the other two, and watch the next ten chats to see which phrasing survived.