RAG Explained: How Retrieval-Augmented Generation Works and How to Build One
Ask ChatGPT for the capital of India and you get the right answer instantly. Ask it how much your company sold last Saturday and it either admits it has no idea or, worse, invents a confident-sounding number. The model is not broken. It simply never saw your data.
This is the gap that Retrieval-Augmented Generation (RAG) closes. RAG lets a language model look up information from your own documents at the moment a question is asked, and answer based on what it finds. It is the technique behind most “chat with your documents” tools, company knowledge assistants, and documentation bots.
In this guide you will learn what RAG is, how each part of the pipeline works, how to build a working RAG system in Python with ChromaDB and OpenAI, and how to connect it to a mobile app.
What is RAG?
Retrieval-Augmented Generation is a technique that combines information retrieval with a large language model (LLM) to produce answers that are more accurate, current, and grounded in real sources.
A normal LLM answers only from what it learned during training. With RAG, the system first searches an external knowledge source (PDFs, text files, databases, or APIs), pulls out the passages most relevant to the question, and hands them to the model along with the question. The model then writes its answer using that material.
In short: instead of asking the model to remember, you let it look things up.
Why large language models need RAG
Public chatbots are trained on huge amounts of public text such as web pages, books, and encyclopedias. That gives them broad knowledge, but three problems remain:
- No access to private data. A model does not know your sales figures, internal policies, meeting notes, or product specifications.
- Outdated knowledge. Training stops at a cutoff date. Anything that changed afterwards is invisible to the model.
- Hallucination. When a model does not know something, it may still produce a fluent answer that is wrong.
A practical example: paste a Docker error message into a general chatbot and you may get a vague or incorrect fix. Paste the same error into an assistant that searches Docker’s official documentation and it can return the correct answer along with links to the pages it used. The difference is not a smarter model. It is retrieval.
RAG solves all three problems at once. Your data stays in your own store, you can update it whenever you like without retraining anything, and you can instruct the model to answer only from the retrieved material.
How RAG works, step by step
Every RAG system has two stages: a one-time (or occasional) indexing stage that prepares your documents, and a query stage that runs each time a user asks something. The query stage has three phases, which give the technique its name: retrieval, augmentation, and generation.
Stage 1: Indexing your documents
- Load your documents (PDFs, Markdown, text files, and so on).
- Split them into smaller pieces called chunks.
- Convert each chunk into a vector, a list of numbers that captures its meaning, using an embedding model.
- Store the vectors, along with the original text and metadata, in a vector database.
Stage 2: Answering a question
- Retrieval. The user’s question is converted into a vector using the same embedding model. The vector database finds the stored chunks whose vectors are closest to it. This is called semantic search.
- Augmentation. The retrieved chunks are added to the prompt as context, together with the user’s question and instructions on how to behave.
- Generation. The LLM reads the prompt and writes the final answer, based on the supplied context.
Where the “no hallucination” behaviour comes from: the system prompt. If you tell the model to answer only from the provided context and to say “I don’t know” when the answer is missing, it will decline questions your documents cannot answer, instead of guessing. You will see this in the demo below.
Key concepts: embeddings, vector databases and chunking
Tokens and embeddings
Language models do not read raw text. Text is first split into tokens (a process called tokenization), and tokens are turned into vectors, also called embeddings. An embedding is simply a long array of numbers. Texts with similar meanings end up with vectors that point in similar directions, which is what makes meaning-based search possible.
The length of the array, called the vector dimension, depends on the embedding model you use:
| Embedding model | Dimensions |
|---|---|
| all-MiniLM-L6-v2 (Hugging Face / Sentence Transformers) | 384 |
| nomic-embed-text (runs locally with Ollama) | 768 |
| OpenAI text-embedding-3-small | 1,536 |
| OpenAI text-embedding-3-large | 3,072 |
Whichever model you choose, you must use the same one for indexing documents and for embedding questions. Vectors from different models cannot be compared, and your database must be configured for the matching dimension.
Vector databases
A vector database stores embeddings and can quickly find the ones closest to a query vector. Popular options include Pinecone, Weaviate, Milvus, and Chroma. If you already run PostgreSQL, the pgvector extension lets you use it as a vector database too. For learning and small projects, Chroma is a good choice because it runs locally with almost no setup.
Semantic search vs keyword search
Traditional SQL or keyword search matches exact words. Semantic search matches meaning. A question like “Can I bring my dog to work?” will find a paragraph titled “Pet policy” even though the word “dog” may not appear in it, because the vectors are close in meaning.
Chunking
Documents are split into chunks because whole documents are too large to embed as a single meaningful unit and too large to fit in a prompt. Two settings matter most:
- Chunk size: how much text goes in each chunk.
- Chunk overlap: how much text is repeated between neighbouring chunks, so a sentence cut at a boundary is not lost.
A recursive character splitter, which tries to break at paragraph and sentence boundaries first, is a sensible default.
Build a RAG app with Python, ChromaDB and OpenAI
Now let’s build a small but complete RAG system. It will answer questions about the contents of PDF files you provide. The example uses a guide about growing vegetables, but any PDFs will work.
The project has two scripts: one that fills the database, and one that answers questions.
Step 1: Set up the project
Create a folder with this structure:
rag-project/
data/ # put your PDF files here
fill_db.py # indexes the PDFs
ask.py # asks questions
requirements.txt
.env # holds your API key
chromadb
openai
python-dotenv
langchain-community
langchain-text-splitters
pypdf
Install the dependencies:
pip install -r requirements.txt
OPENAI_API_KEY=your-api-key-here
Keep your key private. Never commit the .env file to Git. Add it to .gitignore.
Step 2: Fill the vector database
This script loads every PDF in the data folder, splits the text into chunks, and stores the chunks in a persistent Chroma collection. Chroma embeds the text automatically using its built-in default embedding model (all-MiniLM-L6-v2), so you do not need to call an embedding API yourself.
import os
import chromadb
from langchain_community.document_loaders import PyPDFDirectoryLoader
from langchain_text_splitters import RecursiveCharacterTextSplitter
DATA_PATH = "data"
CHROMA_PATH = "chroma_db"
# 1. Create (or open) the vector database
chroma_client = chromadb.PersistentClient(path=CHROMA_PATH)
collection = chroma_client.get_or_create_collection(name="knowledge_base")
# 2. Load all PDFs from the data folder
loader = PyPDFDirectoryLoader(DATA_PATH)
raw_documents = loader.load()
# 3. Split the documents into chunks
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=300,
chunk_overlap=100,
length_function=len,
is_separator_regex=False,
)
chunks = text_splitter.split_documents(raw_documents)
# 4. Prepare documents, metadata and IDs for Chroma
documents = []
metadatas = []
ids = []
for i, chunk in enumerate(chunks):
source = os.path.basename(chunk.metadata.get("source", "unknown"))
documents.append(chunk.page_content)
metadatas.append({"source": source, "page": chunk.metadata.get("page", 0)})
ids.append(f"{source}-chunk-{i}")
# 5. Store everything (upsert avoids duplicates if you run the script again)
collection.upsert(documents=documents, metadatas=metadatas, ids=ids)
print(f"Stored {len(documents)} chunks in the database.")
What each part does:
PersistentClientsaves the database to disk in thechroma_dbfolder, so you only need to index once.get_or_create_collectioncreates the collection the first time and reuses it afterwards.- The IDs let you update or delete specific chunks later.
- The metadata records where each chunk came from, so answers can cite their source.
Run it:
python fill_db.py
The first run downloads the small default embedding model, so it needs an internet connection. When it finishes, a chroma_db folder appears in your project.
Step 3: Ask questions
This script takes a question, retrieves the most relevant chunks, builds a prompt around them, and sends it to OpenAI.
import chromadb
from dotenv import load_dotenv
from openai import OpenAI
load_dotenv() # loads OPENAI_API_KEY from the .env file
CHROMA_PATH = "chroma_db"
openai_client = OpenAI()
chroma_client = chromadb.PersistentClient(path=CHROMA_PATH)
collection = chroma_client.get_or_create_collection(name="knowledge_base")
SYSTEM_PROMPT = """You are a helpful assistant. Answer the user's question
using ONLY the context provided below. Do not use your own knowledge and do
not make things up. If the answer is not in the context, reply exactly:
"I don't know."
Context:
{context}
"""
def ask(question, n_results=3):
# Retrieval: find the chunks closest in meaning to the question
results = collection.query(query_texts=[question], n_results=n_results)
chunks = results["documents"][0]
sources = results["metadatas"][0]
# Augmentation: put the retrieved text into the prompt
context = "\n\n".join(chunks)
system_prompt = SYSTEM_PROMPT.format(context=context)
# Generation: let the LLM write the answer
response = openai_client.chat.completions.create(
model="gpt-4o", # any chat model works
messages=[
{"role": "system", "content": system_prompt},
{"role": "user", "content": question},
],
)
answer = response.choices[0].message.content
return answer, sources
if __name__ == "__main__":
user_question = input("What do you want to know? ")
answer, sources = ask(user_question)
print("\nAnswer:", answer)
print("\nSources:")
for s in sources:
print(f"- {s['source']} (page {s['page']})")
Notice how the three RAG phases map directly to the code: collection.query is retrieval, filling the system prompt with context is augmentation, and the chat completion call is generation.
Testing your RAG
Run the script and try two kinds of questions.
1. A question your documents can answer. For the vegetable guide, asking “What are the benefits of composting?” returns an answer based on the composting section of the PDF: that compost acts as an organic fertilizer and soil conditioner for infertile native soils.
2. A question your documents cannot answer. Ask something unrelated, such as “What is the capital of India?” A well-configured RAG replies “I don’t know,” because that information is not in the retrieved context and the system prompt forbids guessing.
If you want to understand what is happening internally, print the context and system_prompt variables before the API call. Seeing exactly which chunks the database returned is the fastest way to debug a RAG system.
Settings worth experimenting with
n_resultscontrols how many chunks are retrieved. One is the simplest and most focused. Two to five gives the model more context but may add noise.- Chunk size and overlap change what each retrieved piece contains. Smaller chunks are more precise; larger chunks carry more surrounding context.
- The system prompt shapes tone, format, and how strictly the model sticks to the context.
- The model. You can swap
gpt-4ofor another OpenAI model or for a locally hosted one.
Run RAG fully locally
You do not have to send your data to a cloud API. The same architecture works with models running on your own machine. With a tool like Ollama you can run both the embedding model (for example nomic-embed-text) and the answering model locally, so your documents never leave your computer. This is useful for confidential data such as internal company documents or client records.
To switch, keep the vector database and pipeline the same and change only the embedding and generation calls to point at your local model server. Remember that if you change the embedding model, you must re-index your documents, since the old vectors will not match the new model.
Using RAG in a mobile app
Mobile apps built with Flutter, React Native, Swift, or Kotlin can offer “ask our documentation” or “ask your files” features using RAG, but the RAG pipeline itself should not run inside the app. The correct design is a small backend that the app talks to.
Recommended architecture
Mobile app --HTTPS--> Your backend API --> Vector database
(Flutter, (FastAPI, Flask, --> LLM (OpenAI, or a
React Native, ...) Node, ...) self-hosted model)
Why a backend and not direct calls from the app:
- Security. An API key embedded in an app can be extracted by anyone who inspects it. Keep keys on the server.
- Cost control. On the server you can add rate limits, authentication, and usage caps.
- Updates without app releases. You can change prompts, chunking, or models on the server without waiting for an app store review.
- Performance. Vector search and document processing belong on a server, not on a phone.
A minimal backend endpoint
Here is a small FastAPI wrapper around the ask function from earlier. Install it with pip install fastapi uvicorn.
from fastapi import FastAPI
from pydantic import BaseModel
from ask import ask
app = FastAPI()
class Question(BaseModel):
question: str
@app.post("/ask")
def ask_endpoint(body: Question):
answer, sources = ask(body.question)
return {"answer": answer, "sources": sources}
Start it with:
uvicorn server:app --host 0.0.0.0 --port 8000
Your mobile app then sends a POST request to /ask with a JSON body like {"question": "How do I reset my password?"} and displays the answer field, optionally listing the sources beneath it. Any HTTP client in your mobile framework can do this. In Flutter that is the http or dio package; in React Native it is the built-in fetch.
Mobile-specific tips
- Show a loading state. RAG answers take a second or more. Use a spinner or, better, stream the response so text appears as it is generated.
- Handle “I don’t know” gracefully. Show a friendly message and offer to contact support or rephrase the question.
- Show sources. Displaying where an answer came from builds user trust.
- Plan for poor connectivity. Handle timeouts and offline states clearly.
- Protect the endpoint. Require authentication and rate-limit requests before going to production.
RAG vs fine-tuning
A common question is whether to use RAG or to fine-tune a model on your data. They solve different problems.
| RAG | Fine-tuning | |
|---|---|---|
| Best for | Answering from changing facts and documents | Changing the model’s style, format, or specialised behaviour |
| Updating knowledge | Add or edit documents; no retraining | Requires another training run |
| Source citations | Easy, since you know which chunks were used | Difficult |
| Upfront effort | Moderate (build a pipeline) | Higher (prepare data, train, evaluate) |
| Data privacy | Data stays in your own store | Data is baked into model weights |
For most “answer questions about our content” use cases, RAG is the simpler and more maintainable starting point. The two can also be combined.
Best practices and common mistakes
- Use one embedding model consistently. Mixing models between indexing and querying gives meaningless search results.
- Write a strict system prompt. Tell the model to use only the supplied context and to say when it does not know.
- Tune chunking on your own content. There is no universal best chunk size. Test with real questions.
- Store metadata. Source file and page number make citations and debugging possible.
- Inspect what was retrieved. When answers are poor, the problem is usually retrieval, not the model.
- Clean your documents. Scanned PDFs without text, broken tables, and repeated headers all reduce quality.
- Keep the index fresh. Re-run indexing when documents change, using stable IDs so old chunks are replaced.
- Consider hybrid search for larger projects. Combining keyword search (such as BM25) with semantic search can improve results when exact terms matter.
Frequently asked questions
What does RAG stand for?
Retrieval-Augmented Generation. The system retrieves relevant information, augments the prompt with it, and then generates an answer.
Does RAG stop hallucinations completely?
No, but it greatly reduces them. Grounding answers in retrieved documents and instructing the model to say “I don’t know” when the context is missing makes wrong answers much less likely. Quality still depends on good retrieval and good source documents.
Do I need to train a model to use RAG?
No. RAG works with an existing pre-trained model. You only prepare and index your documents.
What is a vector database?
A database built to store embeddings and find the ones most similar to a given query vector. Examples include Chroma, Pinecone, Weaviate, Milvus, and PostgreSQL with pgvector.
Can I use RAG with PDFs?
Yes. PDFs are one of the most common sources. Text-based PDFs work directly; scanned PDFs need OCR first to extract the text.
Can RAG run without the cloud?
Yes. With locally hosted embedding and language models, the entire pipeline can run on your own hardware, which keeps sensitive data private.
How many chunks should I retrieve?
Start with one to three and adjust. Fewer chunks are more focused; more chunks give broader context but can introduce irrelevant text.
Can I add RAG to an existing mobile app?
Yes. Build the RAG pipeline as a backend API and call it from the app over HTTPS. Do not place API keys or the vector database inside the app itself.
Conclusion
RAG gives language models what they lack on their own: access to your knowledge, kept up to date, with answers you can trace back to a source. The core idea is simple. Index your documents as vectors, retrieve the closest matches to each question, and let the model answer from them.
The Python example above is small enough to understand in one sitting, yet it contains every real component: loading, chunking, embedding, vector storage, retrieval, prompt construction, and generation. From here you can swap in different models, add more documents, expose it through an API, and bring it into your mobile app. Start with a handful of your own PDFs, try questions your documents can and cannot answer, and refine from there.