I almost booked three demos in a single afternoon, then caught myself. Every vendor on the shortlist would have shown me a polished flow, a clean dashboard, a happy-path conversation - none of them would have shown me what their agent did on my messiest support ticket. So I built a different filter, starting with a one-page scorecard and refusing to read it until I had tried each platform against the same three jobs: grounding answers in our own help center, taking a real action through an API, and handing a confused customer to a person with the transcript attached. Demos could not tell me that. Trial accounts could.

- Sources it accepts out of the box, and whether Notion, raw URLs, and custom Q&A were first-class or bolted on.
- Whether Actions went beyond calendars and Slack into real API calls I could wire to our order system.
- How clean the handoff to Zendesk or Slack looked, and whether the transcript actually followed the customer.
- How it handled identity, because a logged-in customer asking about their invoice is a different conversation from a stranger on the pricing page.
- What the analytics gave me on day one, before I had any traffic to optimize. The pricing page and the feature grid told me almost nothing. The trial told me everything, because I was feeding the same five awkward questions into each one and watching what came back. Two platforms passed cleanly. One passed with a workaround. The fourth could not ground an answer in our help center without me rewriting every article first, so it was out. The vendor I ended up on was not the cheapest and not the one with the loudest launch announcement. It was the one where the first hour of building matched the first hour of the lifecycle I had already planned. That mattered more than the demo I almost sat through.
Training it
The first hour of Build taught me more about our support knowledge than a year of ticket tagging. I pointed the agent at our help center, our Notion workspace, and a folder of internal how-tos, and started asking it the questions our customers actually ask, in the words they actually use. Half of them came back wrong. Not because the model was weak - because our writing was. Some articles contradicted each other on return windows. One Notion page from 2023 still said we ship to Canada. Another described a subscription plan we had retired. Training, in practice, is editing your own knowledge until the agent can read it without flinching. I went article by article, cut the prose that had drifted out of date, rewrote the answers in the voice we wanted the agent to use, and added a short list of things it should never claim. That last bit mattered more than I expected. Telling the agent what not to say - no order ETAs outside the carrier's tracking link, no refund promises without a ticket, no medical or legal advice even if the help center mentions it - kept it from filling silence with confidence. I treated custom Q&A as a scratch pad for the questions that did not deserve a full article: edge cases, seasonal exceptions, the two or three things we tell every customer on chat. If a question came up three times in a week, it earned a Q&A before it earned a help center article. After two days of editing, I re-ran the same five awkward questions from the scorecard against the agent. Three of them were clean. The other two surfaced gaps I would never have caught by reading the help center top to bottom, because I had written those articles and could not see their seams anymore. That is the real return on a tight training loop: the agent reads your content the way a stranger does, and it tells you which pages have rotted. Launch day was less dramatic than I expected. I scheduled the embed for a Tuesday morning, watched the first five chats come through, and realized the real work had just started. I kept a tab open for the rest of the week and let the agent answer real customers while I sat behind it like a supervisor at a new hire's first shift. I let it run on every channel at once - website chat, WhatsApp, email - and watched where it stumbled. Three patterns showed up in the first hundred chats, and each one needed a different fix.
- Confident wrong answers. About one in ten replies sounded right and was not. Most were date-sensitive promises - shipping windows, sale end times, subscription pricing - that had drifted in the help center. I added a quarterly source review to the calendar so those did not rot again.
- Quiet over-escalation. A handful of questions the agent could have answered got punted to a human because my guardrails were too cautious. I trimmed the never-claim list down to the cases that actually cost us, and the false handoffs dropped within a day.
- Channel drift. Tone that worked on chat sounded stiff over email, and the WhatsApp formatting broke two delivery notifications. I wrote a short per-channel style note and the seam stopped showing. I also wired two outbound flows to match how customers actually ask. Routine follow-ups - order shipped, subscription renewed, return received - went through AI Replies or Human Takeover outbound campaign features, where the same widget that answered questions could send a clean follow-up without a human in the loop. Anything that needed a real person's tone - a complaint, a goodwill credit, a sensitive timing question - got handed to a teammate with the transcript already attached, so the customer never repeated themselves. Treating those as one feature would have been the easy mistake; splitting them kept the tone right and the inbox sane. The first week taught me three things I would carry into the next launch. Soft launch on a small slice of traffic before turning the widget on for everyone. Keep a human in the loop for the first hundred conversations, even if most of them look routine. Read the unresolved chats on Friday the way I read server logs - they tell you where the knowledge base is rotting faster than any dashboard will. That loop is not glamorous. It is that turns a demo-ready agent into one customers stop noticing, which is the only metric I have come to trust.