Putting an AI Agent on a Website: What I Wish I'd Known on Day One
Live with an AI agent on a small e-commerce site: the widget moment, picking the right tool, training it on real sources, setting guardrails, and reading the first hundred chats.
The moment the widget went live and the first confused question hit
I had told myself I would wait until Friday to flip the switch. I did not.
The embed code went into the header on a Tuesday morning, the widget rendered in the bottom-right corner, and the first notification arrived before my coffee was cool. It was a billing question in broken English, asking whether a refund would land before a Friday deadline.
The agent read the visitor's account, confirmed the refund was already queued, and gave a precise date. The second chat was harder.
Someone asked for a discount the agent was not authorized to offer. It explained the policy, offered a related bundle, and asked whether they would like to talk to a person about a custom order.
The visitor never replied.
I went in thinking I was shopping for a chatbot. Within an hour I realized I was shopping for a teammate who would sit in the corner of my site and either earn their keep or quietly drive visitors away.
The first thing I did was write down what I wanted the agent to handle: live chat support for billing and shipping, lead capture for visitors who looked interested but did not buy, product guidance for the handful of items that always confused people, and outbound follow-up for campaigns that needed a nudge back to the site. Every vendor could do one of those beautifully and the others badly.
An AI agent, when you stop reading the marketing, is a language model, a retrieval layer over your own content, tool calls that check an order or send a Slack alert, and a handoff path to a human. If a pitch skips one of those four, they are selling a smarter FAQ page.
I ranked candidates by , not the price. A pure support agent was wrong for product guidance and worse for outbound. A generalist that claimed everything was good at nothing in my trial. The one I kept had a clean support mode, a separate product guidance mode, and a clear story about handing a hot lead to my inbox without losing the thread.
That seam is where most of the early embarrassment happens.
I stopped comparing agents on benchmarks and started comparing them on the ugly cases: an unauthorized request, a language it was not trained on, a 2 a.m. Sunday with nobody watching. The vendor who had rehearsed those was the one I trusted.
"AI agent for website" is also a category, not one product - a widget, an outbound engine, and an inbox tool. Most platforms sell the widget and call the rest an upsell. If you need all three, you want one platform, not three invoices. By the end of that first week I had a shortlist of two.
The reason I picked one was not better small talk or a prettier avatar. It was the vendor who could say, in plain words, what the agent would refuse, where it would hand off, and which of my real questions it had already broken in rehearsal. This is the walk-through I wished I had read before I started comparing AI agents for a website.
I started with the obvious sources: help docs, the FAQ page, the return policy. Then I dumped in the product catalog and the last six months of support transcripts, which I almost immediately regretted. The agent started repeating the apologetic filler my team had leaned on during a rough patch - "totally understand your frustration" three times in a row, even on a shipping question. My voice had become its voice.
What fixed it was curating, not just uploading. I trimmed transcripts to the resolved ones, dropped the ones where a customer walked away angry, and rewrote the brand-voice instructions in plain words: short answers, no throat-clearing.
Source
What it taught the agent
What it cost me
FAQ PDF
Clean, on-brand answers to the questions I already had copy for
Nothing - the safest starting point
Helpdesk export (resolved tickets)
Real phrasing for real customer questions, with answers that had actually worked
A few hours of trimming to drop the angry threads
Notion (internal notes and policies)
Edge cases the public docs never mentioned, plus the off-menu rules my team used daily
Rewriting the voice once so the agent stopped inheriting the filler
I trusted them in roughly that order. The FAQ PDF was the floor. The helpdesk export was the biggest lift, once I stopped handing the agent every angry thread and kept only the resolved ones. Notion was where the edge cases lived.
The thing I wish I had done earlier was write the brand-voice rules before I uploaded a single transcript. Without rules in plain words - short answers, no apologetic filler, no off-brand jokes - the agent just imitated whatever it had read most recently, which on a bad week meant imitating me at my worst.
Guardrails went down on paper before I touched the dashboard. No discounts above the published ceiling, no medical advice, no off-brand jokes, no promises about shipping windows my warehouse could not meet. The rule that mattered more than the rest: when in doubt, hand off. Live chat escalated if they asked for a human, swore, or repeated themselves. Email escalated after two reads, because a slow inbox punishes hesitation. WhatsApp and Instagram needed the lightest touch - mostly order-status pings I wanted closed in one reply.
The first hundred chats are the rehearsal you should have run but did not, except now the audience is real and typing at 1 a.m. I sat with the inbox open for a full day, flagging every reply that felt off. Three patterns dominated:
Visitors still repeating themselves, usually because the agent answered a related question instead of the one asked.
The agent apologizing when the visitor had not complained.
Silent abandonment: the agent answered, the visitor stopped replying, and I never found out whether the answer was right.
That silence is the part you chase. e-commerce customer support
Resolution rate was the number I watched first, then stopped trusting. A reply that ends a chat is not a reply that solved it, and three cheerful "anything else?"s can hit a beautiful number while the visitor leaves. The metric that changed my behavior was escalations - what the agent said right before it handed off. Reading those in order showed me where training was thin and where my guardrails were too loose.