Key takeaways
- RAG retrieves relevant information at question time and gives it to the LLM as context.
- It lets a model use private or recent data without retraining it.
- Good chunking, retrieval and evaluation matter as much as the choice of LLM.
Large Language Models (LLMs) are impressive, but on their own they have two big limits: they only know what was in their training data, and they sometimes produce confident answers that are simply wrong. Retrieval-Augmented Generation, or RAG, is one of the most widely used techniques for solving both problems.
The problem RAG solves
- Knowledge cut-off: a model does not know about events or documents created after its training data was collected.
- No access to private data: it has never seen your company’s policies, product manuals or internal wiki.
- Hallucinations: when it does not know something, it may still generate a plausible but incorrect answer.
What is RAG?
RAG is an approach where the system first retrieves relevant pieces of information from an external knowledge source, then adds them to the prompt, and finally lets the LLM generate an answer based on that context. The term comes from a 2020 research paper by Patrick Lewis and colleagues, “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”.
A simple way to think about it: instead of asking the model to answer from memory, you give it an open book and ask it to answer from the right pages.
How RAG works, step by step
Part 1: Indexing (done ahead of time)
- Load your documents: PDFs, web pages, database records, tickets and so on.
- Chunk them into smaller passages, because models and search work better with focused pieces of text.
- Embed each chunk: an embedding model turns text into a vector (a list of numbers) that captures its meaning.
- Store the vectors, with the original text and metadata, in a vector index or vector database.
Part 2: Retrieval and generation (at question time)
- Embed the question using the same embedding model.
- Search for the chunks whose vectors are most similar to the question.
- Build a prompt that contains the question plus the retrieved chunks, with instructions to answer only from that context.
- Generate the answer with the LLM, ideally citing which sources it used.
Here is a simplified sketch of that flow. The function names are placeholders, not a specific library:
question = "What is our refund policy?"
query_vector = embed(question)
chunks = vector_store.search(query_vector, top_k=4)
context = "nn".join(chunk.text for chunk in chunks)
prompt = f"""Answer using only the context below.
If the answer is not in the context, say you don't know.
Context:
{context}
Question: {question}"""
answer = llm.generate(prompt)
Key components of a RAG system
- Embedding model: converts text into vectors so that similar meanings end up close together.
- Vector store: stores and searches vectors. Options range from libraries such as FAISS to databases such as Chroma, Pinecone, Weaviate, Milvus and PostgreSQL with pgvector.
- Retriever: the logic that finds relevant chunks, often combining vector search with keyword search.
- Reranker (optional): a second model that reorders retrieved chunks by relevance.
- LLM: generates the final answer from the question and context.
- Orchestration: frameworks such as LangChain and LlamaIndex help wire these pieces together, though you can also build RAG with plain code.
RAG vs fine-tuning
RAG
- What changes
- The information given to the model at question time.
- Best for
- Answering from specific, changing or private documents.
- Updating knowledge
- Add or edit documents and re-index them.
- Traceability
- Answers can cite the source passages.
Fine-tuning
- What changes
- The model’s weights, through extra training.
- Best for
- Teaching a consistent style, format or specialised behaviour.
- Updating knowledge
- Requires another round of training.
- Traceability
- Hard to trace where an answer came from.
The two are not rivals. Many production systems use RAG for knowledge and fine-tuning or careful prompting for behaviour.
Common problems and how to fix them
- Chunks that are too big or too small: experiment with chunk size and overlap, and split on natural boundaries such as headings.
- Relevant documents not retrieved: try hybrid search (keyword plus vector), better embeddings or query rewriting.
- Too much noise in the context: add a reranker, retrieve fewer chunks, or filter by metadata such as date or department.
- Model ignores the context: tighten the instructions and ask it to say “I don’t know” when the answer is not present.
How to evaluate a RAG application
Evaluate retrieval and generation separately. For retrieval, check whether the right chunks appear in the top results. For generation, check whether answers are faithful to the retrieved context and actually relevant to the question. Build a small test set of real questions with expected answers, and use evaluation tools such as Ragas or your own scripts to track quality as you change the system.
Skills you need to build RAG apps
- Python and working with APIs
- Embeddings and vector search
- Prompt design
- Evaluation and testing
- Deploying an app, for example with FastAPI and Docker
Frequently asked questions
Does RAG completely stop hallucinations?
No. It reduces them by grounding answers in real sources, but a model can still misread or go beyond the context. Clear instructions, citations and evaluation help keep answers accurate.
Do I always need a vector database?
Not always. Small projects can use an in-memory index or a library such as FAISS. A dedicated vector database becomes useful as your data, traffic and filtering needs grow.
Is RAG a good first Generative AI project?
Yes. A question-answering app over a set of documents touches almost every core skill: data preparation, embeddings, retrieval, prompting, evaluation and deployment.
Our 3-month generative AI course in Pune takes you from Python to working AI applications, including RAG and AI agents. It runs online and comes with 1-to-1 interview preparation and up to 1 year of placement support.


