A RAG knowledge base answers questions from your own documents. Whether one is worth building depends on four things: whether RAG fits the questions people ask, what the documents are like, the risks of sending them through the pipeline, and how the answers will be checked. This checklist covers each in turn, following our series on RAG (starting with what problem does RAG actually solve?).
How does the checklist work?
Twelve questions in four groups, about eight minutes. Each answer shows guidance and a link to further reading, and the result names the area to start with.
The questions
Fit
1. What should users get back when they ask a question?
A written answer drawn from several documents
RAG fits: the system retrieves the relevant passages and the model writes the answer from them.
Ask a few likely users what they would do with an answer. If they would check the source anyway, search may be enough.
2. How often does the knowledge change?
Weekly or more often
This suits RAG: updating an index costs far less than retraining a model.
Rarely; it is stable, foundational material
Stable knowledge (house style, a domain's reasoning, how a task should behave) can suit fine-tuning, since the retraining cost is paid rarely. RAG still fits if answers must cite their sources.
Decide between extracting the text and embedding the pages as images before indexing anything; the choice is hard to reverse once the corpus is indexed.
5. Do typical questions need facts from more than one document?
Rarely; the answer is usually in one place
Flat chunks and similarity search handle this well.
Often
Connecting facts that never appear in the same passage is where flat chunks struggle. A knowledge graph, or a reference compiled from the sources, handles it.
Embedding sends every indexed document to the provider at ingestion, whether or not anyone queries it. Keep the sensitive part on infrastructure you control.
A naive pipeline retrieves the most relevant passage whoever it belongs to, and puts it in the answer. Rebuild the permission model in the index before launch.
Estimate it from the corpus: documents, times chunks per document, times the cost of an embedding call, for every re-index. The retrieved context also sets the cost of each query.
Generated question-and-answer pairs make evaluation practical at scale. Have a domain expert review a sample, since a judge model misses specialist errors.
A system that answers questions from your own documents. It retrieves the passages most likely to hold the answer and gives them to a language model, which writes the answer from them and can cite where it came from.
Should we use RAG or fine-tuning?
RAG suits knowledge that changes often and answers that must cite a source. Fine-tuning suits stable knowledge, such as house style or how a task should behave. Many questions are better served by plain search, which returns the source itself.
Can a RAG system keep our documents private?
Only if embedding and generation both run on infrastructure you control. A standard pipeline sends every indexed document to the embedding provider at ingestion, and the retrieved passages to the generation provider on every query.
How do you measure whether a RAG system works?
Stage by stage: the embedding model, the chunking, retrieval, then generation. A single score cannot say which stage failed, and each stage fails in its own way.
left to answer.
All questions answered. The reading list below collects the articles linked from your answers.