Nobody is training a model here, despite the word everyone uses. You are building a searchable index of your content and letting the assistant read from it. That distinction matters, because it tells you what to fix when answers are wrong: the content, not the AI.
Start with less than you think
The instinct is to index everything. Resist it. A site with 200 blog posts and 8 service pages usually gives better answers when only the service pages, the FAQ and the contact page are indexed.
The reason is competition. When someone asks about your prices, every blog post that mentions cost competes with your actual pricing page. Fewer, better sources means the right passage wins.
A sensible starting set
- Your service or product pages. What you do, for whom, at what price.
- Your FAQ, if you have one. It is already written in customer language.
- Contact and about pages.
- Products, if you sell.
- Twenty to fifty Q&A pairs you write yourself, in the words customers actually use.
Add blog posts later, and mark them as background reading so the assistant explains them without turning them into promises about your services.
Write for the question, not the page
Content written for search engines is often terrible for a chatbot: long preambles, the answer buried in paragraph six. A page that states the answer plainly near a clear heading gets retrieved and quoted accurately.
The single highest-value hour you can spend is writing Q&A pairs for the twenty questions you already answer by email every week.
Then check it
- Ask it the ten questions you get most. Any wrong answer is a content gap, and usually an obvious one.
- Look at what it retrieves, not just what it says. If the right passage is not being found, no amount of prompt tuning helps.
- Read the list of questions it could not answer after the first week. That list is your content roadmap.
Questions
How often do I need to re-index?
With a plugin that watches for changes, never manually. Publishing an edit re-indexes that page automatically, and unchanged content costs nothing to re-check.
Can it read PDFs?
Yes, if the PDF contains real text. A scan of a printed page is an image, and needs optical character recognition before it is any use.