I learned the hard way that demos are theater. The first vendor I evaluated put a beautiful widget on my screen, typed a polished question into it, watched the reply land like a movie line, and asked me what I thought. I thought: this tells me nothing about how the agent will do on Tuesday morning when a customer in a hurry asks whether a return label covers duties. So before I sit through another sales call, I run my own trial on the vendor's own playground, with three short scenes that mirror the work I actually need the agent to do. Scene one is the policy corner. I paste in a sentence from my own return policy, then I ask the agent, in plain language, the exact question a visitor would type at 11 p.m. If the answer is fuzzy, I move on. Scene two is the cross-document question. A real visitor rarely asks about one page; they ask about shipping and the help center and the FAQ at the same time. I ask a question that requires pulling from two places and I watch for a clean stitch instead of two stitched-together half answers. Scene three is the no-answer. I ask something the corpus cannot help with and I read what the agent says when it has nothing. A clean handoff to a human beats a confident guess, every time. Two filters sit on top of those scenes. First, corpus fit. If the agent's onboarding assumes a single help center and my site has a shop, a knowledge base, and three product docs, the demo will flatter the agent and the rollout will punish me. I look for an agent that ingests the messy stack I already have, not one that asks me to clean the house first. Second, the human handoff. I want a visible takeover path, because the first week of transcripts will have questions the agent should not answer alone, and I need to see them routed to a person without losing the thread. An outbound campaign queue that surfaces those AI replies queued for human takeover is the simplest signal that the vendor has thought about the first week, not just the demo. That is the only way I trust a shortlist: I run my own three scenes on the vendor's playground, I check whether the stack fits the corpus I already keep, and I confirm the human handoff is a button I would actually press. The lesson is unglamorous and it works. A demo can hide a weak corpus fit and a missing handoff, but a ten-minute trial on a real corner of your own site will surface both before you sign.
Training on the Corpus You Already Have
Training day, for me, looks less like a launch and more like unpacking. The first thing I do is resist the urge to write a fresh knowledge base. The agent is only as good as what I already keep, and what I already keep is a shop, a help center, three product docs, and a FAQ that has not been touched in eighteen months. So I treat the corpus the way a new hire would treat it: I lay it out, I admit which drawer is messy, and I clean the drawers that matter. Scene one is the page the agent should know cold. I take the return policy, I ask the agent the question a customer types in at 11 p.m., and if the answer is fuzzy I do not blame the agent. I blame the page, because the page is the corpus. A clean return window written in plain English will out-train any prompt I could write. Scene two is the cross-document stitch. Real visitors do not ask about one page. They ask about shipping and the help center and the FAQ in the same breath, and a stitched answer is the difference between a tool and a toy. I ask a question that pulls from two corners of the corpus and I read for a clean join instead of two half answers glued together. When the stitch is clean, the visitor feels answered. When it is not, the visitor feels routed. Scene three is the no-answer rehearsal. I ask something the corpus cannot help with, on purpose, and I read what the agent says when it has nothing. A clean handoff to a person, with the thread kept, is the only answer I will accept. This is also the moment I confirm the human handoff is a button I would actually press, because the first week of transcripts will surface questions the agent should not answer alone. Two filters sit on top of those scenes. First, corpus fit. If the onboarding assumes a single help center and my site has a shop, a knowledge base, and three product docs, the demo will flatter the agent and the rollout will punish me. I look for an agent that ingests the messy stack I already have, not one that asks me to clean the house first. Second, a visible handoff. The lesson is unglamorous and it works. The agent does not get smarter than the pages I feed it. Clean the drawers that matter, rehearse the stitches, rehearse the no-answer, and the rollout has a chance.
Going Live
The first week live is not the training scene. The first week live is the Tuesday morning scene, with a real visitor in a hurry and a real question about a return label and duties. So I treat going live the way I treated the trial: I run a few short scenes on my own site before I open the door, and I read the transcripts the way I would read a new hire's first week. Scene one is the live policy corner. I take the return policy page, the same one I cleaned in training, and I watch the agent answer a real visitor's question in real time. If the reply lands clean, I leave the page alone. If it stalls, I know the page is the corpus and the page is what I fix.
