How to train an AI chatbot on your website content

October 5, 2026

You don't retrain a model to teach it your site. You index your pages so it reads them before every reply. Here's how, and why it costs cents.

N0VA Mockup

Everyone asks how to train a chatbot on their website. Almost nobody should.

Training changes the model itself, and it's the slow, expensive route. What you want is a chatbot that reads your pages before it answers. That process is called indexing, and it turns your site into something the model can search in milliseconds, every time a visitor asks a question.

The difference matters more than the vocabulary. A trained model remembers a snapshot of your site and forgets nothing it got wrong. An indexed one reads the current page, quotes it, and stops when the page runs out. This guide covers the setup step by step, with the real costs and a way to test the result. It's the method behind N0VA, the assistant on our own site.

Key takeaways

  • You don't retrain a model to teach it your website. You index your pages so the model reads them before every reply.
  • How you split pages into chunks drives accuracy. In Anthropic's 2024 tests, better chunk preparation cut failed retrievals from 5.7% to 1.9%.
  • Indexing a typical 50-page site costs a fraction of a cent. Keeping it current is a publish webhook.
  • Test with 20 or more known questions, half of them outside your content, before a visitor sees the bot.

Do you actually train a chatbot on your website?

No. For website content, you connect the model to your pages instead of retraining it. OpenAI's own guide to Optimizing LLM Accuracy splits the problem in two. When a model lacks proprietary or current knowledge, you fix the context it receives. When its tone or format is off, you fix the model.

Your pricing, your services, your shipping policy and your returns window are knowledge problems. The model has never seen them, and they change. That makes them a job for retrieval-augmented generation (RAG), the method formalized by Lewis et al. in 2020: retrieve the relevant passages first, then generate an answer from them.

Fine-tuning has a real place. It shapes how a model behaves, which is useful once you need a very specific voice or output format at scale. It's a poor way to store facts, because every page edit means another training run. OpenAI notes that many of its largest customer deployments ran on prompt engineering and RAG alone, with no fine-tuning at all.

So when a vendor says a chatbot is "trained on your website", read it as "indexed". Then ask how often the index refreshes.

How does a chatbot learn your website content?

It learns in two separate moments. Once, and again on every publish, your pages are cleaned, split into short passages, turned into numbers and stored in an index. Then, on every visitor question, the system searches that index and hands the closest passages to the model, which answers from them or declines.

ON EVERY PUBLISH Your pages become a searchable index Site pages pages, CMS, PDFs Clean strip nav, footers Chunk split by heading Embed text to vectors Index ON EVERY QUESTION The model reads before it replies Visitor question "Do you ship to Canada?" Retrieve passages closest chunks from index Answer + link to the source page No match: decline, hand off to a person searched at question time
How a website chatbot "learns" your content. Pages are indexed on every publish; each question searches that index before the model replies.

The first lane is the "training" people mean. It runs in the background and nobody sees it. The second lane is what visitors experience, and it's only as good as the first.

Each stage in the top lane has one job:

  1. Clean. Strip navigation, footers, cookie banners and repeated boilerplate, so only real content is left.
  2. Chunk. Split each page into passages short enough to match one question.
  3. Embed. Convert each passage into a vector, a list of numbers that captures its meaning. Similar meanings land close together.
  4. Index. Store the vectors so the system can find the nearest ones to a question in milliseconds.

A website chatbot trained on site content works by retrieval, a method described by Lewis et al. in 2020. Pages are cleaned, chunked, embedded and indexed ahead of time. For each question, the system retrieves the closest passages and the model answers from them, or declines when nothing matches.

Which pages should the chatbot read?

Only the pages you'd stand behind if a customer quoted them back to you. Start with service and product pages, pricing, FAQs and policies. Leave out anything outdated or contradictory, because the model will repeat whatever it retrieves. The Air Canada chatbot case showed the stakes: a British Columbia tribunal held the airline liable for a refund policy its bot described in 2024.

Then fix the pages themselves. This is the step most teams skip and the one that pays back most.

  • One fact, one place. If your delivery times appear on four pages with three different numbers, the bot will eventually quote the wrong one.
  • Headings that say something. "Pricing for retainers" gives the chunk its own context. "More info" gives it none.
  • Write the answer, then the detail. Passages that open with the answer get retrieved and quoted cleanly.
  • Name things consistently. If your product has two names across the site, pick one.
Content prepared for a chatbot is content prepared for AI search. The same clean, answer-first pages that help your own assistant are the ones ChatGPT and Perplexity find easiest to cite.
‍What is generative engine optimization

How should you split your pages into chunks?

Split by meaning, then check the size. OpenAI's file search tool defaults to chunks of 800 tokens with a 400-token overlap, roughly 600 words with half repeated into the next chunk. That works for long documents. For a website, splitting at each heading usually works better, because a heading already marks where one answer ends and the next begins.

Chunking sounds like plumbing. It decides what the model can find.

RETRIEVAL FAILURE RATE (TOP-20 CHUNKS) How you prepare chunks changes what the bot can find 0%1.5%3%4.5%6% Standard embeddings Contextual embeddings + Contextual BM25 + Reranking 5.7% 3.7% 2.9% 1.9% Lower is better. Source: Anthropic, Introducing Contextual Retrieval, September 2024.
Share of questions where the right passage was missing from the top 20 results. Source: Anthropic, Introducing Contextual Retrieval, September 2024.

In September 2024, Anthropic published Introducing Contextual Retrieval, measuring how often the right passage failed to appear in the top 20 results. Standard embeddings missed it 5.7% of the time. Adding a short line of context to every chunk before embedding it, such as which page and section it came from, cut that to 3.7%. Pairing that with keyword search brought it to 2.9%, and a reranking step brought it to 1.9%.

You don't need every technique on day one. Two habits capture most of the gain:

  1. Prefix each chunk with its page title and heading. A passage that reads "Starts at $2,500" means nothing alone. "Retainers > Pricing: Starts at $2,500" means something.
  2. Keep tables and lists whole. A pricing table split across two chunks is a pricing table the bot can't read.

How much does it cost to index a website?

Almost nothing. OpenAI prices its text-embedding-3-small model at $0.02 per million tokens in 2026. A 50-page site at around 1,000 words a page is roughly 65,000 tokens. Even with heavy chunk overlap doubling that, a full re-index costs about a quarter of a cent.

That arithmetic is why the running cost of a site assistant sits almost entirely in the answers. N0VA, the assistant on our own site, runs on this architecture for under $1 a month in model costs. The two levers are which model writes the replies and how many passages it reads per question.

Where the real cost sits is in the work around the index: deciding what the bot may answer, cleaning the content, writing the refusal rule and testing it. That's a few days of focused work, done once. How much does an AI chatbot for your website cost ?

How do you keep the chatbot up to date?

Re-index whenever the site changes, automatically. Webflow's Data API supports webhooks for events including site_publish and collection_item_changed. Point one at a small function that re-crawls the changed pages and updates their chunks, and the bot knows about a new page within minutes of publishing.

This is where indexing beats training outright. A fine-tuned model is a photograph of your site on the day it was trained. An index is a mirror. Change a price at 10 a.m., publish, and the 10:05 answer is correct.

Two details keep the mirror clean:

  • Delete as well as add. When a page is unpublished, remove its chunks. Old pages are the most common source of confidently wrong answers.
  • Store the source URL with every chunk. The bot can then link each answer to the page it came from, and you can trace any bad answer to the page that caused it.

Most sites built on other platforms offer the same pattern through a CMS webhook or a scheduled nightly crawl. Webflow design and development

How do you know the training worked?

Test it against questions with known answers before visitors do. OpenAI's accuracy guide recommends building an evaluation set of 20 or more questions with ground-truth answers, then reviewing every failure before changing anything. For a website bot, half the set should be questions your site answers and half should be questions it doesn't.

The second half is the one that matters. A bot that answers everything is a bot that invents. Here's what to log for each test question:

  1. Did retrieval find the right page? If not, the fix is in chunking or content.
  2. Did the answer match the page? If not, the fix is in the instructions to the model.
  3. Did it link the source? Every answer should point somewhere a visitor can check.
  4. Did it decline when it should have? An out-of-scope question should end in a handoff to a person.

Rerun the same set after every significant content change. It takes ten minutes and catches regressions before a customer does.

The takeaway

Nobody needs to train a model to teach it their website. The work is closer to editing than to engineering: choose the pages you'd stand behind, write them so each answer lives in one clear place, split them where the meaning splits, and re-index every time you publish. Get those right and the model has very little room to be wrong. The chatbot is only ever as accurate as the content it reads, so the best investment in your assistant is usually an afternoon spent on your own site.

‍

Can I train ChatGPT on my website?

You can connect a model like ChatGPT to your website content, which is what most people mean. The model reads your indexed pages at question time instead of being retrained. OpenAI's own guidance treats missing proprietary knowledge as a context problem, solved with retrieval rather than fine-tuning.

Does the chatbot learn from visitor conversations?

Not by default, and it shouldn't. A retrieval-based bot answers only from your indexed pages, so a visitor can't teach it something false. You can review conversation logs to find questions your site doesn't answer yet, then add that content to the site and re-index.

‍

How long does it take to train a chatbot on a website?

The indexing itself takes minutes. A 50-page site can be cleaned, chunked, embedded and indexed in one run for under a cent at current embedding prices. The time goes into preparing content and testing answers, usually a few days for a small business site.

What happens when I update a page?

With a publish webhook in place, the changed page is re-indexed within minutes and the next answer reflects it. Webflow supports this through site_publish and CMS item webhooks. Without automation, the bot keeps quoting the old version until someone re-indexes manually.

DOPAMINE STUDIO

Your site already knows the answers

We'll map which questions your content can answer today and what an assistant built on it would cost to run.

UK Flag
English