RAG Knowledge Base Checklist

A RAG knowledge base answers questions from your own documents. Whether one is worth building depends on four things: whether RAG fits the questions people ask, what the documents are like, the risks of sending them through the pipeline, and how the answers will be checked. This checklist covers each in turn, following our series on RAG (starting with what problem does RAG actually solve?).

How does the checklist work?

Twelve questions in four groups, about eight minutes. Each answer shows guidance and a link to further reading, and the result names the area to start with.

The questions

Fit

1. What should users get back when they ask a question?

A written answer drawn from several documents

RAG fits: the system retrieves the relevant passages and the model writes the answer from them.

The right document or passage, which they read themselves

A well-built search index may serve this better and more reliably. Returning the source text avoids the risks that summarising it introduces.

Not sure yet

Ask a few likely users what they would do with an answer. If they would check the source anyway, search may be enough.

2. How often does the knowledge change?

Weekly or more often

This suits RAG: updating an index costs far less than retraining a model.

Rarely; it is stable, foundational material

Stable knowledge (house style, a domain's reasoning, how a task should behave) can suit fine-tuning, since the retraining cost is paid rarely. RAG still fits if answers must cite their sources.

We do not know

Find out; the answer decides between RAG and fine-tuning.

3. Do answers need to cite their source?

Yes

RAG provides this as a by-product, because the system knows which passage it retrieved.

No

Then fine-tuning is also an option. Weigh it on how often the knowledge changes.

Data

4. What are the source documents mostly like?

Prose: contracts, policies, correspondence, reports

The mature text pipeline suits it. Start with fixed-size or semantic chunking.

Tables, forms, scans or diagrams

Decide between extracting the text and embedding the pages as images before indexing anything; the choice is hard to reverse once the corpus is indexed.

A database more than documents

Structured data works better through a documented schema and generated queries than through a pipeline built to chunk prose.

A mixture we have not catalogued

Catalogue it first. The right preparation method depends on what the documents look like.

5. Do typical questions need facts from more than one document?

Rarely; the answer is usually in one place

Flat chunks and similarity search handle this well.

Often

Connecting facts that never appear in the same passage is where flat chunks struggle. A knowledge graph, or a reference compiled from the sources, handles it.

Not sure

Collect twenty real questions and check where each answer lives. The pattern decides the method.

6. Is it clear which version of each document is current?

Yes

Then the index can be kept in step with the source.

Mostly

Outdated copies are retrieved as readily as current ones. Decide which source is authoritative and index only that.

No

Fix this first. A model given an outdated passage writes an outdated answer with the same confidence.

Risks

7. Can the documents be sent to third-party embedding and generation services?

Yes; none of it is sensitive

Then hosted services are an option. Budget for embedding the whole corpus, as well as the queries.

No, or only some of them

Embedding sends every indexed document to the provider at ingestion, whether or not anyone queries it. Keep the sensitive part on infrastructure you control.

We have not checked

Check before indexing anything. Ingestion is the point where documents leave.

8. Do different users have permission to see different documents?

No; everyone can see everything

Then permission-blind retrieval is not a risk. Check again if the audience grows.

Yes, and retrieval will filter by permission

Store access metadata with each chunk and filter at retrieval time; the index will not inherit the source system's permissions.

Yes, and we have not planned for it

A naive pipeline retrieves the most relevant passage whoever it belongs to, and puts it in the answer. Rebuild the permission model in the index before launch.

9. Have you budgeted for re-indexing as the corpus grows and changes?

Yes

Ingestion is a recurring cost tied to corpus size, and the budget treats it as one.

We budgeted per query only

Queries are the easier half of the bill. Every re-index (changed documents, a better embedding model) repeats the ingestion cost in full.

No budget yet

Estimate it from the corpus: documents, times chunks per document, times the cost of an embedding call, for every re-index. The retrieved context also sets the cost of each query.

Evaluation

10. Do you have a set of real questions with known correct answers?

Yes

That set is what every stage of the pipeline gets measured against.

We could generate one from our documents

Generated question-and-answer pairs make evaluation practical at scale. Have a domain expert review a sample, since a judge model misses specialist errors.

No

Without one, a fluent wrong answer and a correct one look the same. Build the set before choosing tools.

11. Will you measure retrieval separately from the answers?

Yes, stage by stage

Then a falling score points at the stage that broke: chunking, retrieval or generation.

We will check the answers

Answer quality alone cannot locate a fault. A model given the wrong passage still writes a confident answer.

Not planned

Retrieval can fail silently, on top of anything the model gets wrong. Plan to measure it.

12. Who will check the knowledge base once it is live?

A named owner, re-running the evaluation on a schedule

Scores decay as documents are added, edited and retired, with no change to the model. A schedule catches it.

Users will report problems

Users are poorly placed to spot an unsupported claim, because it reads like a correct one. Add sampled review by someone who knows the material.

Nobody yet

Name an owner before launch. Drift in the corpus is separate from drift in the model, and needs its own check.

Your results

Scoring runs in your browser and needs JavaScript. The printable checklist below lists every question and answer.

Can I print the RAG knowledge base checklist?

Yes. The version below lists every question with the points for each answer.

  1. What should users get back when they ask a question?
    • [ ] A written answer drawn from several documents (2)
    • [ ] The right document or passage, which they read themselves (0)
    • [ ] Not sure yet (1)
  2. How often does the knowledge change?
    • [ ] Weekly or more often (2)
    • [ ] Rarely; it is stable, foundational material (1)
    • [ ] We do not know (0)
  3. Do answers need to cite their source?
    • [ ] Yes (2)
    • [ ] No (1)
  4. What are the source documents mostly like?
    • [ ] Prose: contracts, policies, correspondence, reports (2)
    • [ ] Tables, forms, scans or diagrams (1)
    • [ ] A database more than documents (1)
    • [ ] A mixture we have not catalogued (0)
  5. Do typical questions need facts from more than one document?
    • [ ] Rarely; the answer is usually in one place (2)
    • [ ] Often (1)
    • [ ] Not sure (0)
  6. Is it clear which version of each document is current?
    • [ ] Yes (2)
    • [ ] Mostly (1)
    • [ ] No (0)
  7. Can the documents be sent to third-party embedding and generation services?
    • [ ] Yes; none of it is sensitive (2)
    • [ ] No, or only some of them (1)
    • [ ] We have not checked (0)
  8. Do different users have permission to see different documents?
    • [ ] No; everyone can see everything (2)
    • [ ] Yes, and retrieval will filter by permission (2)
    • [ ] Yes, and we have not planned for it (0)
  9. Have you budgeted for re-indexing as the corpus grows and changes?
    • [ ] Yes (2)
    • [ ] We budgeted per query only (1)
    • [ ] No budget yet (0)
  10. Do you have a set of real questions with known correct answers?
    • [ ] Yes (2)
    • [ ] We could generate one from our documents (1)
    • [ ] No (0)
  11. Will you measure retrieval separately from the answers?
    • [ ] Yes, stage by stage (2)
    • [ ] We will check the answers (1)
    • [ ] Not planned (0)
  12. Who will check the knowledge base once it is live?
    • [ ] A named owner, re-running the evaluation on a schedule (2)
    • [ ] Users will report problems (1)
    • [ ] Nobody yet (0)

Score out of 24: 0-11, groundwork first; 12-19, ready to prototype; 20-24, ready to build.

www.bayis.co.uk/checklists/rag-knowledge-base.html

Frequently asked questions

What is a RAG knowledge base?

A system that answers questions from your own documents. It retrieves the passages most likely to hold the answer and gives them to a language model, which writes the answer from them and can cite where it came from.

Should we use RAG or fine-tuning?

RAG suits knowledge that changes often and answers that must cite a source. Fine-tuning suits stable knowledge, such as house style or how a task should behave. Many questions are better served by plain search, which returns the source itself.

Can a RAG system keep our documents private?

Only if embedding and generation both run on infrastructure you control. A standard pipeline sends every indexed document to the embedding provider at ingestion, and the retrieved passages to the generation provider on every query.

How do you measure whether a RAG system works?

Stage by stage: the embedding model, the chunking, retrieval, then generation. A single score cannot say which stage failed, and each stage fails in its own way.