The questions on most "prompt engineering interview" lists are the ones you can answer by reciting a definition: what's zero-shot, what's few-shot, what's chain-of-thought. Those come up, usually in the first five minutes. The part that decides the interview is the follow-up: "OK, and your classifier is now wrong 15 percent of the time on a new customer. What do you do?"
So this list gives you both. Each question has a short outline of a strong answer, then the follow-up an interviewer tends to ask. They're grouped from fundamentals to evaluation. The answer outlines are what a solid answer should cover, not scripts. Say them in your own words and with your own examples, or you'll sound rehearsed.
A note on sourcing: these come from the technique literature (Wei et al. on chain-of-thought, Wang et al. on self-consistency, Yao et al. on ReAct, Greshake et al. on indirect prompt injection) and from standard practice. I haven't scraped real interview transcripts, so don't treat this as a leaked question bank. Where a point depends on a specific model or vendor feature, check the current docs before you repeat it in a room.
Fundamentals (1-10)
1. What is prompt engineering, in one sentence you'd say to a non-technical manager? Designing the instructions, context and examples that get a model to do a task reliably, then testing that it does. The word "reliably" is the point. Follow-up: How is it different from just chatting with the model? Chatting is one-off and judged by feel. Engineering means the prompt runs on many inputs, is versioned, and has a measurable pass rate.
2. What are the parts of a good prompt? Task, context, constraints, input data, output format, and sometimes examples and a role. Separate instructions from data clearly (headings or XML-style tags) so the model doesn't confuse them. See prompt elements. Follow-up: Which part do people most often leave out? The output format and the "what to do when the input doesn't fit" case.
3. Zero-shot vs few-shot: when do you use each? Start zero-shot with a clear instruction. Add examples when the model misreads the format, tone or edge cases that are hard to describe. Few-shot prompting teaches by demonstration. Follow-up: What can go wrong with few-shot examples? The model copies surface features (length, wording, label order), so examples must vary, cover edge cases, and not leak the test set.
4. How many examples should you include? No magic number. Add one at a time and measure on a held-out set; stop when gains flatten. Diverse examples beat many similar ones, and long examples cost tokens on every call. Follow-up: How do you choose which examples? Pull from real failures, or retrieve the most similar examples per input at runtime.
5. Explain temperature. It rescales the model's next-token probabilities before sampling. Lower values make high-probability tokens even more likely; higher values flatten the distribution so unlikely tokens get picked more often. See LLM settings. Follow-up: Is temperature 0 deterministic? Not guaranteed. It reduces variation, but infrastructure and model updates can still change outputs, so test stability rather than assuming it.
6. Temperature vs top-p? Temperature reshapes the whole distribution. Top-p (nucleus sampling) cuts the candidate set to the smallest group of tokens whose cumulative probability reaches p, then samples inside it. Common advice is to tune one, not both. Some providers restrict or ignore these parameters on certain models, so read the API reference. Follow-up: When would you want a high temperature? Brainstorming, varied phrasing, or deliberately sampling several reasoning paths for self-consistency.
7. What is a context window, and what happens when you exceed it? The maximum number of tokens (input plus output) the model can handle per request. Over the limit, the API returns an error or you must truncate, summarize or retrieve. Follow-up: Does a bigger window mean you should put everything in? No. More tokens cost more and add latency, and models can use information in long contexts unevenly (the "lost in the middle" finding from Liu et al.). Put what matters in a prominent position and cut the rest.
8. What is a token? A chunk of text (word piece, punctuation, whitespace) the model reads and writes. Cost, limits and latency are counted in tokens, not characters. Follow-up: Why does a prompt in another language often cost more? Tokenizers are usually trained mostly on English-heavy text, so other scripts often split into more tokens. Measure with the provider's tokenizer.
9. What does a system prompt do compared with a user message? It sets persistent behavior, role and rules for the whole conversation, and models are typically trained to give it higher priority. See system prompts. Follow-up: Is a system prompt a security boundary? No. Users and injected text can still override or extract it, so never put secrets in it.
10. What causes hallucinations, and what do you do about them in a prompt? The model produces plausible text, not verified facts. Mitigate by grounding answers in supplied sources, allowing "I don't know", requiring quotes or citations, and checking claims afterward. See avoiding hallucinations. Follow-up: Can a prompt eliminate hallucinations? No. It reduces them. You still measure the rate on your own data.
Techniques and reasoning (11-20)
11. What is chain-of-thought prompting and when does it help? Asking the model to produce intermediate reasoning before the answer (Wei et al., 2022). It helps on multi-step arithmetic, logic and planning. It adds tokens and latency and can make easy tasks worse. See the chain-of-thought lesson. Follow-up: Do you still need it with reasoning models? Often less. Many reasoning models think internally and vendors advise against heavy step-by-step scaffolding, but check the model's guidance and test on your task.
12. Is the visible reasoning trace a faithful explanation of why the model answered? Not necessarily. Research has shown written reasoning can diverge from what actually drove the answer, so don't treat it as a guarantee, and don't ship it to users as an audit trail. Follow-up: So how do you verify an answer? Check it against ground truth or tools, not against the model's own justification.
13. What's self-consistency? Sample several reasoning paths at non-zero temperature and take the majority answer (Wang et al., 2022). It improves accuracy on tasks with a checkable final answer at a multiple of the cost. See self-consistency. Follow-up: Where does it not work? Open-ended generation, where there's no single answer to vote on.
14. What is ReAct? A pattern where the model alternates reasoning with actions (tool calls) and observations (Yao et al., 2022). It is the basis of many agent loops. See ReAct prompting. Follow-up: What breaks in a ReAct loop in production? Infinite loops, repeated identical calls, bad tool arguments, and context bloat. You need step limits, error handling and logs.
15. What is prompt chaining and why use it? Split a big task into smaller prompts where each output feeds the next. You get easier debugging, per-step validation and the option to use cheaper models for easy steps. See prompt chaining. Follow-up: What's the cost? More calls, more latency, and errors can compound across steps.
16. Tree-of-thought vs chain-of-thought? ToT explores several branches and evaluates them, backtracking if needed; CoT follows one line. ToT is more expensive and rarely worth it outside search-like puzzles. See tree of thought. Follow-up: Have you used it in production? Be honest. Most people haven't; say what you'd try first (self-consistency or a verifier step).
17. How do you make a model follow a strict output format? Use the provider's structured output or JSON-schema mode where available, otherwise give a schema and an example, validate with a parser, and retry with the error message on failure. See structured prompting. Follow-up: Schema valid but content wrong. Now what? Validity of format doesn't mean validity of content. Add semantic checks and evals.
18. Role prompting: does "you are an expert" help? Sometimes it shifts tone and vocabulary. It isn't a reliable accuracy booster, and a specific task description usually does more. Test it; don't assume. See the role prompting playbook. Follow-up: Show me a case where it hurt. Personas can make the model overconfident or push style over substance.
19. What is meta-prompting or using an LLM to write prompts? Asking a model to draft, critique or improve prompts. Useful for generating variants quickly. The output must still be evaluated on real data. See meta-prompting. Follow-up: What's the risk? Optimizing to a handful of examples and overfitting.
20. When would you use fine-tuning instead of prompting? When you have many labeled examples, need a consistent style or format, or want to cut prompt length or latency, and prompting plus retrieval has plateaued. Start with prompting, since it's cheaper to iterate. See fine-tuning vs prompting. Follow-up: Does fine-tuning add knowledge? It is better at behavior and format than facts. For fresh or private facts use retrieval.
Context, RAG and tools (21-28)
21. How does retrieval-augmented generation work? Retrieve relevant chunks from your data, put them in the prompt, and instruct the model to answer from them. See RAG and how RAG works. Follow-up: The bot gives a wrong answer. How do you tell if retrieval or generation failed? Inspect the retrieved chunks. If the answer isn't in them, retrieval failed; if it is and the model ignored it, generation failed.
22. How do you write the prompt for a RAG answer step? Tell it to use only the provided context, to say when the context doesn't contain the answer, to cite chunk IDs, and keep context clearly delimited from instructions. Follow-up: What if two chunks conflict? Define a rule (newest wins, or flag the conflict) and test it.
23. Chunking: how do you choose chunk size? Trade-off between precision (small chunks) and context (large). Respect document structure such as headings, add overlap where needed, and tune against retrieval metrics on real questions. Follow-up: Table or code in the document? Naive splitting breaks them; chunk by structure.
24. What is context engineering and how does it relate to prompt engineering? Deciding everything that enters the context: instructions, retrieved docs, memory, tool results, history. Prompt wording is one piece. See context engineering. Follow-up: What do you cut first when the context is too full? Stale history and low-relevance retrieval, usually before shortening the instructions.
25. What is function (tool) calling? The model outputs a structured request to call a function you defined; your code executes it and returns the result. The model never runs code itself. See function calling. Follow-up: How do you write good tool descriptions? Clear names, precise descriptions of when to use and not use the tool, typed parameters, and examples.
26. A model keeps picking the wrong tool. What do you check? Overlapping tool descriptions, vague names, too many tools, missing guidance on when not to call. Then reproduce with logged inputs and fix the description before touching the model. Follow-up: How do you prevent a destructive tool call? Permissions, confirmation steps and least privilege in code, not just a prompt instruction.
27. How do you handle long documents? Chunk and summarize hierarchically, retrieve relevant parts, or use a long context if cost allows; ask for quotes to verify. See working with long documents. Follow-up: Long context vs RAG? Long context is simpler, RAG is cheaper and scales to more data. Test both on your questions.
28. What is prompt caching? Many providers let you reuse a processed static prefix to cut cost and latency on repeated calls. Details and pricing differ per provider, so check the docs. Follow-up: How do you structure a prompt to benefit? Put stable content (instructions, examples, reference) first and variable user input last.
Safety and robustness (29-35)
29. What is prompt injection? Untrusted text that contains instructions the model follows. Direct: the user types it. Indirect: it hides in a webpage, email or document the model reads (Greshake et al., 2023). See prompt injection. Follow-up: Can you fully prevent it? No known prompt-only fix is complete. Layer defenses.
30. What defenses do you apply? Separate and label untrusted data, limit tool permissions, require human approval for risky actions, filter inputs and outputs, and monitor. See prompt injection explained. Follow-up: Why isn't "ignore instructions in the document" enough? Because the model can't reliably distinguish instruction from content.
31. What is jailbreaking vs prompt injection? Jailbreaking targets the model's own safety policy, typically by the user. Injection targets an application by smuggling instructions through data. See jailbreaking. Follow-up: Which matters more for a customer support bot? Usually injection and data leakage, since it reads untrusted input and may have tools.
32. How do you reduce bias in outputs? Test across demographic variations, avoid leading wording, use structured criteria, and audit outcomes. See bias mitigation prompts. Follow-up: How do you measure it? Counterfactual tests: swap names or attributes and compare outputs.
33. What do you do about sensitive data in prompts? Minimize it, redact before sending, check the vendor's data retention and training terms, and log carefully. Follow-up: Who needs to approve this? Security and legal in most companies. Know your policy.
34. How do you handle a request the model should refuse? Give explicit scope and refusal guidance in the system prompt, test with adversarial cases, and make the refusal useful (say what it can do). Follow-up: Over-refusal? Measure both false refusals and false compliances.
35. How do you test a prompt against adversarial inputs? Build a red-team set: injections, odd formats, empty input, huge input, other languages. See red-teaming your prompts. Follow-up: How often? On every prompt or model change.
Evaluation and production (36-46)
36. How do you know if a prompt change made things better? Run both versions on the same labeled test set and compare metrics. Anecdotes and spot checks mislead. See prompt testing and evaluation. Follow-up: How big should the test set be? Big enough that the difference you care about exceeds the noise; start with dozens of real cases, grow with every failure.
37. How do you build an evaluation set? Collect real inputs, label expected outputs, include edge cases and past failures, and keep it separate from the examples in the prompt. See golden test sets. Follow-up: Where do labels come from? Domain experts, existing records, or careful review of model output with human correction.
38. LLM-as-judge: when is it OK? For qualities that are hard to check by rule (helpfulness, tone), with a clear rubric. Validate the judge against human labels first; known biases include favoring longer answers and its own outputs. Follow-up: How do you reduce judge bias? Pairwise comparison with position swapping, rubrics with anchors, and spot human audits.
39. Which metrics would you use for a classification prompt vs a summarization prompt? Classification: accuracy, precision, recall, per-class confusion. Summarization: no single metric; use rubric scoring, faithfulness checks against the source, and human review. Follow-up: Accuracy is 95 percent. Is it good? Depends on class balance and cost of errors.
40. A prompt works on your 5 examples and fails in production. Why? Test set too small or unrepresentative, input distribution differs, edge cases absent, or the prompt is overfit. Pull production failures into the eval set. Follow-up: What's the first thing you do? Look at 20 failing real inputs and categorize them before changing anything.
41. The model suddenly changed behavior and you didn't touch the prompt. What happened? Possibly the provider updated the model behind an alias. Pin a dated model version where offered, keep regression tests, and monitor. Follow-up: How do you detect it quickly? Scheduled eval runs and alerting on pass-rate drops.
42. How do you version and manage prompts? Store them in source control as code or config, tag versions, link each to eval results, and make rollback trivial. See prompt versioning. Follow-up: Who can change a production prompt? Same review process as code.
43. How do you reduce cost and latency? Shorter prompts, caching, smaller models for easy steps, routing, limiting output length, batching where offered. Measure quality after each change. Follow-up: What would you try first? Usually model routing and trimming context, because those give the biggest changes for the least risk.
44. How would you debug a prompt that returns invalid JSON 3 percent of the time? Log failing outputs, categorize (trailing text, missing field, truncated output), use native structured outputs if available, raise max tokens if truncated, validate and retry with the error message, and track parse-failure rate as a metric. Follow-up: Retry loops? Cap retries and alert on the rate.
45. Describe a prompt you improved and how you proved it. Use a real story: the baseline, the failure categories, the change, the before/after on a fixed test set. If you lack professional experience, build a small project and measure it. Follow-up: What would you do differently? Have one ready. It signals judgment.
46. What would you do in your first 30 days in this role? Read existing prompts and failures, set up or review the eval set, find the highest-volume failure category, and ship one measured improvement. See the learning path if you're building skills first. Follow-up: What if there's no eval set? Make one. That's usually the highest-value first task.
How to use this list
Don't memorize 46 answers. Pick 10 questions, answer each aloud in 60 seconds, then answer the follow-up. If you stall on the follow-up, that's the gap to study. To rehearse out loud with an AI interviewer, use the setup in AI interview prep prompts. To see how these roles are advertised and paid in India, read prompt engineering salary in India.



