An interviewer asks: "Your RAG bot gave a wrong answer yesterday. Walk me through what you do." Candidates who start with "I'd tune the prompt" lose a round. Candidates who start with "I'd pull the retrieved chunks for that query" usually pass it. Most of this list is that difference, repeated.
These are 34 questions in five groups. Each has an answer outline (what a strong answer contains, not a script) and a follow-up that pushes past the textbook version. I kept it to what I'm confident is accurate rather than trying to be exhaustive. Where I cite a figure, it comes from the linked primary source. For prompting-specific questions, there's a separate list in prompt engineering interview questions with answers; this one is about building systems.
How to use it: cover the outline, answer aloud, then read what I flagged. The "great" signal is almost always a measurement or a failure mode you've personally seen.
RAG questions
1. Walk me through a RAG pipeline and where it fails. Outline: ingest, chunk, embed, index; at query time embed, retrieve, optionally rerank, build prompt, generate. Name a failure per stage: bad extraction, bad chunk boundaries, embedding mismatch, missing the right chunk, right chunk ignored by the model. Follow-up: Which failure do you check first? Retrieval. Look at the retrieved chunks before touching the prompt. If the right chunk isn't there, no prompt fixes it.
2. A user says the answer is wrong. How do you debug it? Outline: log the query, retrieved chunks with scores, final prompt and answer. Classify: was the right chunk retrieved? If no, retrieval problem (chunking, query, embedding). If yes but ignored, generation or context-order problem. If the source itself is wrong or stale, data problem. Follow-up: How do you stop it recurring? Add the case to the eval set.
3. How do you choose chunk size? Outline: there's no universal number. Trade-off: small chunks are precise but lose context; large chunks carry context but dilute the embedding and cost more tokens. Prefer structure-aware splitting (headings, paragraphs) over fixed character cuts, add overlap, and decide by measuring retrieval recall on your own questions. Follow-up: What do you do with tables and code? Keep them whole, or convert to a form the embedder handles. Be honest that this is where simple extraction breaks.
4. What is the difference between dense, sparse and hybrid retrieval? Outline: dense uses embeddings and catches paraphrase; sparse (BM25) matches exact terms and is strong on IDs, error codes, names. Hybrid combines both, often by fusing the two ranked lists. Follow-up: How do you fuse them? Reciprocal rank fusion is a common, tuning-light choice (the original paper used a constant of 60); weighted score mixing needs score normalisation because the scales differ.
5. Why add a reranker, and what does it cost? Outline: first-stage retrieval is fast and approximate. A cross-encoder reads query and chunk together, scoring more accurately but slower, so you rerank only the top 20 to 100 candidates and pass the best few on. Follow-up: How do you know it's worth it? Compare recall or answer quality with and without on your eval set, and compare added latency.
6. How do you evaluate retrieval separately from generation? Outline: build questions with labelled relevant chunks or documents. Measure recall@k (is the right chunk in the top k?) and ranking quality such as MRR. Then evaluate answers separately for correctness and faithfulness to the retrieved context. Follow-up: Why separate them? Otherwise you can't tell whether a score change came from the retriever or the model.
7. What is contextual retrieval and when would you use it? Outline: chunks lose meaning out of context ("the company's revenue grew 3%" doesn't say which company). Contextual retrieval has an LLM prepend a short description of where the chunk sits in its document before embedding. Anthropic's write-up reports the top-20 retrieval failure rate falling from 5.7% to 3.7% with contextual embeddings, 2.9% adding contextual BM25, and 1.9% with reranking on top. Those are their results on their datasets. Follow-up: What's the catch? An LLM call per chunk at indexing time, so cost scales with corpus size and re-indexing frequency. Prompt caching can reduce it.
8. The model ignores the retrieved context. Why? Outline: candidates: context is long and the relevant part sits in the middle (the "lost in the middle" result found performance is often best when relevant information is at the start or end of the input, Liu et al. 2023); too many near-duplicate chunks; prompt doesn't make grounding the priority; the model's prior conflicts with the context. Follow-up: What do you try? Fewer, better chunks; put the strongest first or last; explicit grounding and citation instructions; check by ablation, not guesswork.
9. How do you handle "the answer isn't in the documents"? Outline: similarity search always returns k results. Combine a minimum score threshold (tuned on answerable and unanswerable questions), a prompt that permits abstaining, and evaluation cases where the right answer is "not found". Follow-up: Why is a threshold alone fragile? Score distributions shift with model, chunking and corpus, so a threshold has to be re-tuned when any of them change.
10. How do you keep an index fresh and handle deletions? Outline: track document versions and content hashes, re-embed only what changed, delete old chunks by document ID, and make sure access-control metadata updates too. Version the embedding model name with the index. Follow-up: What happens if you switch embedding models? Re-embed everything; vectors from different models aren't comparable.
11. When would you not use RAG? Outline: if the corpus fits in the context window and changes rarely, long context or caching may be simpler; if you need a style or format, fine-tuning or prompting fits better; if the question is analytical over structured data, query the database. See RAG vs long context vs compaction. Follow-up: What does long context not fix? Cost and latency per call, and the position effects above.
If RAG is your weak spot, build a small pipeline yourself and break it; you'll have real stories to tell.
Agents and tool use
12. What's the difference between a workflow and an agent? Outline: Anthropic's Building effective agents defines workflows as LLMs and tools orchestrated through predefined code paths, and agents as systems where the LLM dynamically directs its own process and tool use. Workflow patterns it names: prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer. Follow-up: Which do you start with? The simplest thing that works, often a single well-prompted call, adding structure or autonomy only when measured results justify the extra latency and cost.
13. Describe the basic agent loop. Outline: model receives goal and tool definitions, decides to call a tool or answer, your code executes the tool, the result goes back into the context, repeat until done or a limit is hit. Follow-up: What are your stop conditions? Max steps, max tokens or cost budget, a repeated-call detector, and a clean failure path. "Until the model says it's done" is not enough.
14. How do you design good tools for an agent? Outline: clear names and descriptions, tight parameter schemas, examples and edge cases in the description, boundaries versus similar tools, helpful error messages the model can act on. Anthropic's article suggests making mistakes harder ("poka-yoke"); their example is requiring absolute file paths. More in designing tools for AI agents. Follow-up: How do you test tool design? Run many varied inputs and read where the model picks the wrong tool or arguments.
15. How many tools is too many? Outline: there's no magic number. Selection accuracy and prompt size both degrade as tool lists grow and overlap. Remedies: fewer, broader-purpose tools; routing to a sub-agent or tool subset; retrieving tool definitions on demand. Follow-up: How do you know you've hit the limit? Tool selection accuracy on your eval set drops, or similar tools get confused.
16. A tool call fails or returns garbage. What should happen? Outline: return a structured error to the model, retry only idempotent calls with backoff, cap retries, validate outputs before feeding them back, and escalate to a human or a safe fallback. Follow-up: What about non-idempotent tools? Charging a card or sending an email twice is the classic bug. Use idempotency keys and confirmations.
17. How do you handle prompt injection in an agent that reads external content? Outline: treat retrieved and fetched content as untrusted data. Limit tool permissions to least privilege, require human approval for consequential actions, separate read-only from write tools, and don't let the content that was read dictate privileged actions. Read prompt injection in tool-using agents. Follow-up: Can a better system prompt solve it? No. Instructions help but aren't a security boundary; permissions and architecture are.
18. What is MCP and how does it differ from function calling? Outline: function calling is how a model asks your code to run a function you defined in the request. MCP is a protocol for exposing tools and resources from separate servers so any compatible client can use them. They operate at different layers. See MCP vs function calling vs tool use. Follow-up: What's the risk of connecting many MCP servers? More tool definitions in context and a larger attack surface from servers you don't control.
19. How does an agent remember things across a long task? Outline: the context window is the working memory. Options: summarise or compact old turns, write notes to external storage and read them back, retrieve relevant memories. Each has information-loss risk. Follow-up: What's lost when you summarise? Exact values and decisions' reasoning, which is why you pin critical facts rather than rely on a summary. More in agent memory architectures.
20. When is multi-agent worth it? Outline: for work that parallelises or needs separate contexts or permissions. Costs: more tokens, coordination bugs, harder debugging. Justify with a measured gain over a single agent. Follow-up: What's the simplest multi-agent design? An orchestrator that delegates bounded subtasks and receives compact results. See when to use multiple agents.
21. How do you make an agent safe to run unattended? Outline: sandboxed execution, scoped credentials, budget and step limits, approval gates on irreversible actions, full tracing, and an audit log. Follow-up: What would you never automate? Be concrete: irreversible financial or destructive actions without a human gate.
Evaluation
22. How do you start evaluating an LLM feature with no data? Outline: collect real or realistic inputs (20 to 50 is a start), write the expected answer or a pass/fail rubric, run them on every change, and grow the set from production failures. Details in building golden test sets. Follow-up: Where do the first examples come from? Real user queries, support tickets, and cases you deliberately make hard.
23. When is LLM-as-judge acceptable? Outline: for subjective qualities you can't check by string match. Use a clear rubric with narrow questions, compare against human labels on a sample, and watch for biases such as preferring longer answers or the first option in pairwise comparisons. Follow-up: How do you validate the judge? Measure agreement with humans on a labelled subset; investigate disagreements.
24. Which metrics for a RAG system? Outline: retrieval: recall@k, MRR. Generation: correctness against reference, faithfulness (is every claim supported by the context?), answer relevance, plus rate of correct abstention. Follow-up: Which one do you trust least? Any automated faithfulness score that hasn't been checked against human review.
25. How do you evaluate an agent, as opposed to a single call? Outline: judge the end state (did the task succeed?), not only the final text; also steps taken, tool-call correctness, cost and latency. Run multiple trials because outputs vary. See AI agent evaluation and the evaluating agents lesson. Follow-up: What does pass^k add? The τ-bench paper proposes pass^k to measure reliability across repeated trials; it reports that even strong function-calling agents were highly inconsistent, with pass^8 under 25% in its retail domain. One success is not reliability.
26. How do you catch regressions when you change a model or prompt? Outline: pin the eval set, run before and after, compare per-category rather than only the average, and gate deploys on thresholds. Keep prompts and model IDs versioned. Follow-up: Average score is flat but users complain. Why? A regression in one category hidden by gains elsewhere, or an eval set that doesn't reflect real traffic.
27. How do you evaluate in production? Outline: sample and review traces, collect explicit feedback and implicit signals (retries, abandons), monitor drift in inputs, alert on cost and latency, and feed failures back into the eval set. Follow-up: What is risky about user thumbs-up data? It's sparse and biased toward extremes.
Cost, latency and reliability
28. Your LLM bill doubled. What do you check? Outline: break spend down by feature, model, input vs output tokens. Typical causes: longer prompts or retrieved context, extra agent steps, retries, a model upgrade, a new traffic source. Then act: shorter context, caching, a cheaper model for easy cases, batching. Follow-up: Where do you look first? Per-request token counts in your traces, not the invoice.
29. How does prompt caching help, and when doesn't it? Outline: it reuses the processed prefix of a prompt across requests, so a long stable prefix (system prompt, documents) is cheaper and faster to reuse. It requires the prefix to be identical, so volatile content must go after the stable part. Check the provider's docs for pricing and minimum lengths; they differ and change. See prompt caching with the Claude API. Follow-up: What breaks a cache? Changing anything early in the prefix, such as a timestamp in the system prompt.
30. How do you reduce latency? Outline: stream output, shorten prompts, retrieve and call tools in parallel where independent, use a smaller model for routing and simple steps, cache, and cut agent steps. Measure time to first token and total time separately. Follow-up: What does streaming not fix? Total time and tool-call round trips.
31. How do you pick a model for a task? Outline: define the quality bar with your eval set, then test candidates from cheap to expensive and choose the cheapest that clears it. Re-test when models change. Don't rely on public leaderboards alone. Follow-up: When do you route across models? When a cheap model handles most traffic and a classifier or confidence check sends hard cases up.
32. How do you make structured outputs reliable? Outline: use the provider's structured output or schema-constrained mode where available, validate with a schema library anyway, retry with the validation error on failure, and keep schemas simple. See structured outputs across providers. Follow-up: Does schema-valid mean correct? No. Valid JSON can hold wrong values; evaluate content separately.
Judgment and system design
33. Design a support bot over 5,000 internal documents. Outline: clarify users, accuracy needs and permissions first. Then ingestion with metadata and access control, hybrid retrieval plus reranking, grounded prompt with citations and abstention, an eval set built from real tickets before launch, tracing, and a human handoff path. Follow-up: What would you cut for a first version? Agent loops and fancy reranking; ship the measured baseline first.
34. Tell me about an AI system that failed and what you changed. Outline: pick a real one. State the symptom, how you found the cause (traces, retrieved chunks, eval case), the fix, and how you verified it. Name what you'd do differently. Follow-up: What did you add to prevent recurrence? A test case, a guard, a monitor.
What the strongest candidates do differently
They ask clarifying questions before designing. They name what they'd measure before what they'd build. They say "I don't know, here's how I'd find out" without flinching. And they have at least one failure they diagnosed themselves, with numbers. If you can't produce that yet, build something small this week, break it and write down what you saw. For the career side, see the AI engineering roadmap and the evaluation framework walkthrough.



