Turn your back catalogue into a knowledge base fans can query
You already made the content. The problem is nobody can search it. A practical guide to preparing videos, PDFs and podcasts so an AI can answer from them accurately.
Most creators are sitting on more material than they realise. Three years of weekly videos is roughly a hundred hours of speech — call it eight hundred thousand words. A newsletter twice a month for two years adds another fifty thousand. Nobody, including you, can find anything in there.
Turning that archive into something answerable is less about technology than about preparation. Here is what actually matters, in the order it matters.
Start with the 20% that carries the questions
The instinct is to upload everything. Resist it. Retrieval quality is inversely proportional to noise: the more mediocre material sits in the index, the more often a mediocre passage wins the search and becomes the answer.
Instead, list the ten questions you are asked most. Then find the content where you answered each one best. That is your first upload. It will cover the majority of real traffic and it will be accurate, which buys you trust early.
Everything else can come later, once you can see from real conversations what is actually missing.
Format by format: what to expect
Text (articles, newsletters, notes). The easiest and the most reliable. No transcription loss, clean structure. If you only have time for one format, use this.
PDFs. Good, with one caveat: a PDF that is really a scan of a page is an image, not text. If you cannot select the words in your PDF reader, an extractor cannot read them either — it needs OCR first.
Video and podcasts. Transcription is very good now, but it is not free of errors, especially with proper nouns, brand names and technical jargon. Budget ten minutes to skim the transcript of your most important episodes and fix the recurring mistakes. A model that has learned your product name wrong will repeat it wrong forever.
Comments and DMs. Tempting, because it is where the real questions live, but be careful: it also contains other people's words and personal information. If you use it, extract your answers, not the whole thread.
Structure beats volume
The single biggest quality lever is not how much you upload, but whether each piece stands on its own.
A retrieval system works by finding a passage and handing it to the model. If that passage says "as I explained above, do the same thing but with the other variable", it is useless out of context — and out of context is exactly how it will be retrieved.
Two habits fix most of it:
- Use real headings. They tell the chunker where ideas begin and end.
- Repeat the subject occasionally. "For sourdough, the hydration should be…" survives extraction. "For it, it should be…" does not.
You do not have to rewrite your archive. Applying this to new content is enough; the old material simply performs slightly worse, and you will see which pieces in the conversation logs.
Say what it does not know
Every knowledge base has holes, and the holes are where trust dies. Decide in advance what happens when someone asks about a topic you never covered.
The answer should be a clean "I have not covered that" — optionally with a pointer to what you have covered nearby. What it must never be is an improvised answer that sounds like you.
Same for topics you covered but changed your mind about. If your 2023 advice contradicts your 2026 advice, the system will happily retrieve the old one. Either remove the outdated piece from the index or add a newer, clearer version that will win the search.
Maintenance is a fifteen-minute habit
Once it is live, the loop is short:
- Read the conversations weekly. Look for two things: answers that are wrong, and questions with no good answer.
- Fix the wrong ones at the source — correct the transcript, remove the outdated article.
- Turn the unanswered ones into content. This is the part that pays for itself. Your audience is handing you a ranked list of what to make next, in their own words.
That third point is worth more than the subscription revenue for a lot of creators. Most content research is guesswork about what people want. This is a transcript of them asking.
The realistic outcome
Done properly, a well-prepared knowledge base answers the large majority of routine questions correctly, cites where it got each answer, and admits the rest.
That is not magic, and it does not need to be. It just means the four hundredth person asking which camera you use gets a good answer immediately — and you get your evening back.
Keep reading
How to monetize your audience with an AI character (without burning out)
Most creators hit a ceiling: more followers means more messages, and there is no more time. Here is how an AI character trained on your content breaks that trade-off.
AI character vs chatbot: why the difference decides whether fans trust it
A chatbot generates plausible text. An AI character answers from your actual content and shows its sources. Here is what separates them, and how to tell which one you are being sold.
La Méditerranée Tamuda Bay by Robuchon: the table built out over the sea
In M'diq, the legacy of the most Michelin-starred chef in history meets the Moroccan catch of the day. Style, signature dishes, real guest reviews — plus a Kirikat you can question directly.