Back to blog
5 min readproduct

Turn 800,000 words into a knowledge base fans can query

Three years of videos ≈ 800,000 words nobody can search. The 5-step method — and the mistake almost every creator makes on day one.

Written with AI assistance by Kirikat — facts verified against dated public sources.

Most creators are sitting on more material than they realise. Three years of weekly videos is roughly a hundred hours of speech — eight hundred thousand words. A newsletter twice a month for two years adds another fifty thousand. Nobody, including you, can find anything in there.

Turning that archive into something answerable is less about technology than about preparation. Here is what actually matters, in the order it matters — with the questions we get asked most at the end.

Where should you start? The 20% that carries the questions

Start with the ten questions you are asked most — not with your whole catalogue. Find the content where you answered each one best: that is your first upload. It will cover the majority of real traffic and it will be accurate, which buys you trust early.

The opposite instinct — uploading everything — is the most common mistake. Retrieval quality is inversely proportional to noise: the more mediocre material sits in the index, the more often a mediocre passage wins the search and becomes the answer.

Everything else can come later, once real conversations show you what is actually missing.

Which formats work best?

Text first, then PDFs, then audio and video — in that order of reliability. Here is why, format by format:

Text (articles, newsletters, notes). The simplest and most reliable. No transcription loss, clean structure. If you only have time for one format, pick this one.

PDF. Fine, with one caveat: a PDF that is just a scanned page is an image, not text. If you cannot select the words in your reader, OCR is needed first — Kirikat runs it automatically on scanned PDFs, but a born-digital document will always be more precise.

Video and podcasts. Automatic transcription is excellent today, but not error-free — especially on proper nouns, brand names and technical jargon. Budget ten minutes to proofread the transcripts of your most important episodes and fix recurring mistakes. A model that learned your product's name wrong will repeat it wrong forever.

Comments and DMs. Tempting, because that is where the real questions live, but be careful: they also contain other people's words and personal data. If you use them, extract your answers, not the whole thread.

Why does structure beat volume?

Because a retrieval system never reads your archive end to end: it extracts passages of roughly 300-400 words and answers from those. In Kirikat, every imported document is split into fragments of that size, indexed individually. If a fragment says "as explained above, do the same but with the other variable", it is useless out of context — and out of context is exactly where it will be retrieved.

Two habits solve most of it:

  • Use real headings. They mark where ideas start and end — the splitting follows them.
  • Restate the subject now and then. "For sourdough, hydration should be…" survives extraction. "For that, you need…" does not.

No need to rewrite your archive. Applying this to new content is enough; the old material will simply perform a little worse, and the conversation logs will tell you which pieces matter.

What should it say when it doesn't know?

"I haven't covered that" — never an improvisation that sounds like you. Every knowledge base has gaps, and gaps are where trust goes to die. This is a design decision: a proper AI character refuses to invent and cites the source of every answer.

The same applies to topics where you changed your mind. If your 2023 advice contradicts your 2026 advice, the system will happily retrieve the old one. Remove the outdated piece from the index, or add a newer, sharper version that wins the search.

How much maintenance does it take?

Fifteen minutes a week, in three moves:

  1. Read the conversations. Look for two things: wrong answers, and questions with no good answer.
  2. Fix wrong answers at the source — correct the transcript, remove the outdated article.
  3. Turn unanswered questions into content. This is the part that pays for itself: your audience hands you a prioritised list of what to produce next, in their own words. For many creators that list is worth more than the subscription revenue.

What results should you realistically expect?

Well prepared, a knowledge base answers the large majority of common questions correctly, cites where every answer comes from, and owns the rest. It is not magic, and it does not need to be. It simply means the four-hundredth person asking which camera you use gets a good answer instantly — and you get your evening back.

Frequently asked questions

How much content do I need to start? Ten good answers are enough. A base of ten accurate documents beats a thousand noisy ones: start small and let real conversations drive the next uploads.

Which formats can I import into Kirikat? Plain text, web articles (by URL), PDFs — including scanned ones, OCR is automatic —, Office documents, images (described automatically), and audio or video via transcription. Native text remains the most reliable format.

Can the AI make up answers? It is designed for the opposite: every answer is grounded in a passage of your content and the source is cited. When no passage matches, it says so — that is the fundamental difference from a generic chatbot.

Is my content used to train other models? No. Your knowledge base only powers your own AI character; it is neither shared between creators nor used to train third-party models.

How long does the initial import take? A few minutes per document: splitting and indexing are automatic. The only step worth your time is proofreading the transcripts of your most important content.

Your archive is years of answers you already wrote. Make them searchable →

Keep reading