Back to blog
4 min readcontentworkflowai

Turn your back catalogue into a knowledge base fans can query

You already made the content. The problem is nobody can search it. A practical guide to preparing videos, PDFs and podcasts so an AI can answer from them accurately.

Most creators are sitting on more material than they realise. Three years of weekly videos is roughly a hundred hours of speech — call it eight hundred thousand words. A newsletter twice a month for two years adds another fifty thousand. Nobody, including you, can find anything in there.

Turning that archive into something answerable is less about technology than about preparation. Here is what actually matters, in the order it matters.

Start with the 20% that carries the questions

The instinct is to upload everything. Resist it. Retrieval quality is inversely proportional to noise: the more mediocre material sits in the index, the more often a mediocre passage wins the search and becomes the answer.

Instead, list the ten questions you are asked most. Then find the content where you answered each one best. That is your first upload. It will cover the majority of real traffic and it will be accurate, which buys you trust early.

Everything else can come later, once you can see from real conversations what is actually missing.

Format by format: what to expect

Text (articles, newsletters, notes). The easiest and the most reliable. No transcription loss, clean structure. If you only have time for one format, use this.

PDFs. Good, with one caveat: a PDF that is really a scan of a page is an image, not text. If you cannot select the words in your PDF reader, an extractor cannot read them either — it needs OCR first.

Video and podcasts. Transcription is very good now, but it is not free of errors, especially with proper nouns, brand names and technical jargon. Budget ten minutes to skim the transcript of your most important episodes and fix the recurring mistakes. A model that has learned your product name wrong will repeat it wrong forever.

Comments and DMs. Tempting, because it is where the real questions live, but be careful: it also contains other people's words and personal information. If you use it, extract your answers, not the whole thread.

Structure beats volume

The single biggest quality lever is not how much you upload, but whether each piece stands on its own.

A retrieval system works by finding a passage and handing it to the model. If that passage says "as I explained above, do the same thing but with the other variable", it is useless out of context — and out of context is exactly how it will be retrieved.

Two habits fix most of it:

  • Use real headings. They tell the chunker where ideas begin and end.
  • Repeat the subject occasionally. "For sourdough, the hydration should be…" survives extraction. "For it, it should be…" does not.

You do not have to rewrite your archive. Applying this to new content is enough; the old material simply performs slightly worse, and you will see which pieces in the conversation logs.

Say what it does not know

Every knowledge base has holes, and the holes are where trust dies. Decide in advance what happens when someone asks about a topic you never covered.

The answer should be a clean "I have not covered that" — optionally with a pointer to what you have covered nearby. What it must never be is an improvised answer that sounds like you.

Same for topics you covered but changed your mind about. If your 2023 advice contradicts your 2026 advice, the system will happily retrieve the old one. Either remove the outdated piece from the index or add a newer, clearer version that will win the search.

Maintenance is a fifteen-minute habit

Once it is live, the loop is short:

  1. Read the conversations weekly. Look for two things: answers that are wrong, and questions with no good answer.
  2. Fix the wrong ones at the source — correct the transcript, remove the outdated article.
  3. Turn the unanswered ones into content. This is the part that pays for itself. Your audience is handing you a ranked list of what to make next, in their own words.

That third point is worth more than the subscription revenue for a lot of creators. Most content research is guesswork about what people want. This is a transcript of them asking.

The realistic outcome

Done properly, a well-prepared knowledge base answers the large majority of routine questions correctly, cites where it got each answer, and admits the rest.

That is not magic, and it does not need to be. It just means the four hundredth person asking which camera you use gets a good answer immediately — and you get your evening back.

Keep reading