You have a folder of PDFs and you want to ask it questions in plain English. The short version: extract the text, cut it into overlapping chunks, embed each chunk, and at question time embed the question, grab the closest chunks, and hand them to an LLM with an instruction to answer only from them. That is a RAG chatbot. This post builds one in 55 lines and then breaks it on purpose, because the breakage is where you learn something.
If you want the concepts first, how RAG works covers them. If you want to compare vector stores, FAISS vs Chroma vs Pinecone does that. This post is the middle bit: a small working thing, plus what goes wrong.
What you need, and what I used
I ran everything below on 2026-10-08 in a fresh Python 3.12 virtual environment on a Mac. I had no API keys available, so I picked options that need none:
- pypdf to extract text from PDFs.
- fastembed with
BAAI/bge-small-en-v1.5for embeddings. It runs locally on CPU, and the model download was about 64 MB on disk. - NumPy for the search. No vector database.
python -m venv .venv && source .venv/bin/activate
pip install pypdf fastembed numpy
One honest limit up front: the retrieval half of this post is fully run, with real output. The generation half is not. With no key, my script prints the exact prompt it would send to the LLM, and the hosted-API version further down is untested. I'll mark it again where it appears.
My test PDFs are four short, fictional documents for an invented company called Northwind: a returns policy, a travel policy with a small table, a security handbook, and a "scanned" parental leave policy that is only an image. I made them myself so I knew the right answers. Use your own PDFs; the code doesn't care.
The whole chatbot in 55 lines
import sys
from pathlib import Path
import numpy as np
from fastembed import TextEmbedding
from pypdf import PdfReader
CHUNK_CHARS, OVERLAP = 500, 100
embedder = TextEmbedding("BAAI/bge-small-en-v1.5") # local, no API key
def load_chunks(folder):
chunks = []
for pdf in sorted(Path(folder).glob("*.pdf")):
for page_no, page in enumerate(PdfReader(pdf).pages, start=1):
text = " ".join((page.extract_text() or "").split())
if not text:
print(f"WARNING: no text on {pdf.name} p{page_no} (scanned?)")
continue
step = CHUNK_CHARS - OVERLAP
for i in range(0, len(text), step):
chunks.append({"text": text[i:i + CHUNK_CHARS],
"source": f"{pdf.name} p{page_no}"})
return chunks
def embed(texts):
vecs = np.array(list(embedder.embed(texts)))
return vecs / np.linalg.norm(vecs, axis=1, keepdims=True)
def retrieve(question, chunks, vecs, k=3):
q = next(iter(embedder.query_embed(question)))
q = q / np.linalg.norm(q)
scores = vecs @ q
top = np.argsort(scores)[::-1][:k]
return [(float(scores[i]), chunks[i]) for i in top]
def build_prompt(question, hits):
context = "\n\n".join(f"[{h['source']}]\n{h['text']}" for _, h in hits)
return ("Answer using ONLY the context below. Cite the [source] you used. "
"If the context does not contain the answer, say \"Not in the documents.\"\n\n"
f"CONTEXT:\n{context}\n\nQUESTION: {question}")
if __name__ == "__main__":
folder, question = sys.argv[1], sys.argv[2]
chunks = load_chunks(folder)
vecs = embed([c["text"] for c in chunks])
print(f"{len(chunks)} chunks indexed\n")
hits = retrieve(question, chunks, vecs)
for score, h in hits:
print(f"{score:.3f} {h['source']}\n {h['text'][:110]}...")
print("\n--- prompt sent to the LLM ---\n" + build_prompt(question, hits))
Run it with python rag.py ./pdfs "How many days do I have to return a tent?". My output, trimmed:
WARNING: no text on parental-leave-scan.pdf p1 (scanned?)
8 chunks indexed
0.818 returns-policy.pdf p1
Northwind Customer Returns Policy Return window Unused items can be returned within 45 days of delivery for a ...
0.706 returns-policy.pdf p1
f the item arrived damaged or the wrong item was sent. Warranty Northwind backpacks carry a 5 year warranty ag...
0.638 travel-policy.pdf p1
nses Submit your expense report within 30 days of returning. Reports submitted after 30 days are paid in the next ...
The top chunk contains the right answer: 45 days, and the sentence that a tent that has been set up counts as used. The third result is noise from the travel policy ("30 days of returning"). That's normal. You always get k results whether or not k things are relevant, and that fact is behind two of the three failures below.
What each piece is doing
Chunking is the blunt text[i:i+500] slice with a 100-character overlap. It's deliberately dumb so you can see its effects. Normalising the vectors to length 1 means the dot product vecs @ q is cosine similarity, which is why the scores sit between 0 and 1 for this model. query_embed versus embed: bge-style models expect queries to be embedded slightly differently from documents, and fastembed handles that for you. If you roll your own, check the model card.
Failure 1: the scanned PDF that silently vanished
The script printed WARNING: no text on parental-leave-scan.pdf. That warning is the only reason I know. Without it, the file would simply contribute nothing, and nothing in the system would complain.
Here's what happens when someone asks about it anyway. The question "How many weeks of parental leave do primary caregivers get?" returned:
0.592 travel-policy.pdf p1 :: nses Submit your expense report within 30 days of returning...
0.548 travel-policy.pdf p1 :: ily meal allowance. The allowance depends on the destination tier...
0.527 returns-policy.pdf p1 :: f the item arrived damaged or the wrong item was sent...
Three chunks, all irrelevant, and an LLM handed that context plus a lazy prompt may well improvise a number. The answer exists in the folder (16 weeks, in the image) but your pipeline can't see it.
Fix: treat zero-text pages as an error you surface, not a skip. Then either OCR them or tell the user which files are unreadable. I haven't run OCR here, so I won't hand you code for it; tools like Tesseract or the OCR modes in document parsers are the usual routes. Also watch for the partial case: a PDF with text on most pages and scanned appendices. Log characters per page, not per file.
Failure 2: the chatbot that never says "I don't know"
Look at the scores for questions where I knew the answer was absent or unreadable, against questions where it was present. These are all real runs against the same 8 chunks:
| Question | Answer in the PDFs? | Top score |
|---|---|---|
| Warranty on zippers? | Yes (2 years) | 0.779 |
| Daily meal allowance in Madrid? | Yes (55 EUR) | 0.708 |
| Rotate password every 90 days? | Yes (no, only after compromise) | 0.700 |
| Business class to Zurich? | Yes (economy under 6 hours) | 0.691 |
| Parental leave for primary caregivers? | Only in the scan | 0.592 |
| Dress code in the office? | No | 0.563 |
Retrieval gave the dress-code question three chunks like everything else. On this tiny set, a cutoff around 0.65 would have separated answerable from unanswerable. Don't copy that number. It came from 6 questions and 8 chunks, and scores shift with the embedding model, the chunk size and the corpus. The method is what transfers: write 20 questions you know are answerable and 10 you know aren't, print the top score for each, and look for the gap.
Add the guard in two places:
MIN_SCORE = 0.65 # tune on your own questions; this value is from my 8-chunk toy set
hits = [(s, h) for s, h in retrieve(question, chunks, vecs) if s >= MIN_SCORE]
if not hits:
print("Not in the documents.")
And keep the prompt line that allows "Not in the documents." A threshold alone is brittle; a prompt alone gets ignored when the context looks plausible. Using both is cheap.
Failure 3: the answer lives on a chunk boundary
My chunker cut the returns policy right through a sentence. The first chunk ends with Warranty Northwind backpacks carry a 5 year w and the next begins with f the item arrived damaged. The overlap rescued it here: the second chunk contains the whole warranty paragraph, and the zipper question scored 0.779 against it. With less overlap or a longer sentence, neither chunk would hold the complete fact, and the model would get half a sentence.
The same naive slicing produced a second, sillier bug I didn't plan. The returns policy's last chunk is a 20-character runt, , 9:00 to 17:00 CET., and it ranked second (0.541) for the Madrid meal allowance question. A fragment with no meaning took a slot in the top 3 that a useful chunk could have had.
Fixes, in order of effort:
- Drop chunks under about 100 characters, or merge them into the previous chunk.
- Split on paragraph or sentence boundaries first, and only fall back to a hard character cut for oversized paragraphs.
- Prepend the section heading to each chunk so a fragment still says what it's about. Anthropic's contextual retrieval write-up takes this idea further by having an LLM write a short context line per chunk; they report the top-20 retrieval failure rate dropping from 5.7% to 3.7% with contextual embeddings alone. That's their number on their data, so treat it as a direction, not a promise.
One thing that worked better than I expected: the travel policy's small table survived extraction as plain text (Tier 2 Berlin, Madrid, Toronto 55 EUR) and the Madrid question found it. Don't generalise. Tables with merged cells, multi-page tables and scanned tables are a different story and I haven't tested those.
Swapping in a hosted LLM and embeddings
The prompt builder returns a string, so the generation step is a single call. This block is not run, because I had no key; check the parameter names against your provider's current docs before trusting it.
# UNTESTED: needs `pip install anthropic` and ANTHROPIC_API_KEY set.
import anthropic
client = anthropic.Anthropic()
def answer(question, hits):
msg = client.messages.create(
model="<a current model id from Anthropic's docs>",
max_tokens=500,
messages=[{"role": "user", "content": build_prompt(question, hits)}],
)
return msg.content[0].text
For hosted embeddings, replace embed() and the query embedding in retrieve() with calls to your provider's embedding endpoint. Two rules that bite people:
- Re-embed everything. Vectors from different models aren't comparable, and dimensions usually differ, so the matrix product will either error or, worse, quietly return nonsense. Store the model name next to the index.
- Same model for documents and queries. Always.
My honest view: the local model is fine for a first build, and a fair test is whether retrieval puts the right chunk in the top 3 for your 20 real questions. Switch only when it doesn't. The embedding models comparison helps pick candidates.
Where this stops being enough
Index rebuilds here are instant because the corpus is eight chunks. At a few thousand chunks the NumPy approach is still fine; the embedding step is the slow part, so cache vectors to disk instead of recomputing on every run (I didn't, in the interest of line count). Past that, or once you need filtering by document, date or team, move to a real store.
Once the pipeline runs, the next improvements are query rewriting and reranking, covered in prompt engineering for RAG pipelines, and letting the model decide when to retrieve again, covered in agentic RAG. For measuring whether any of that helped, you need a set of questions with known answers; building golden test sets walks through it. The RAG lesson has the underlying theory.
Checklist before you trust it with real documents
- Every PDF contributes chunks. Print chunk counts per file.
- Zero-text pages are logged loudly.
- No chunk is a runt.
- You have a minimum score, tuned on real answerable and unanswerable questions.
- The prompt allows "Not in the documents" and requires a source citation.
- You can print the retrieved chunks for any answer. When the bot is wrong, read those first. All three failures above were retrieval problems, visible before any model was involved.



