A hiring manager opens your GitHub link between two meetings. They give it maybe a minute. They see a repo called ai-chatbot, a README that says "A chatbot built with LangChain and OpenAI," and no way to tell whether it works. They close the tab.
The projects that get interviews aren't the most impressive ones. They're the ones where a stranger can tell in sixty seconds what the thing does, whether it works, and how you know. This post gives you eight AI portfolio project ideas, each with a scope you can finish, one metric that proves it, and the README detail reviewers look for. First, what the minute is spent on.
One caveat on method: what reviewers look for here is my judgment from how technical hiring generally works, not a survey, and I have no data on hit rates. Treat it as a checklist to apply, not a guarantee.
What does a reviewer check in the first minute?
Most reviewers scan in roughly this order. Design for it.
- The repo name and the first screen of the README. Is it clear what problem this solves, in a sentence?
- Is there proof it runs? A GIF, a short video, screenshots, or a live link.
- Is there evidence it works well? A number from a test set beats an adjective.
- Can they run it? A single command, an
.env.examplefile, no hidden steps. - Is the code readable? Skim of the folder structure, a couple of files, commit history that looks like a person building something over time.
- Do you know its limits? A "Known failures" section is a strong signal. Almost nobody writes one.
Item 3 is where AI projects differ from ordinary web apps. Anyone can wire an LLM to a UI. Few people can show how often it's wrong. The AI engineer interview questions post lists what interviewers ask next, and a good share of it is about evaluation and failure.
Which AI portfolio projects are worth building?
| # | Project | Shows | Metric to report | Finish in |
|---|---|---|---|---|
| 1 | RAG over a real document set, with evals | Retrieval, evaluation | Retrieval hit rate and answer correctness on your question set | 1 to 2 weekends |
| 2 | Structured extraction from messy documents | Schemas, validation | Field-level accuracy against hand-labelled samples | 1 to 2 weekends |
| 3 | Text-to-SQL analyst with guardrails | Tool use, safety | Execution accuracy on a question set; number of blocked unsafe queries | 2 weekends |
| 4 | Support or FAQ bot that refuses well | Grounding, abstention | Correct refusal rate on out-of-scope questions | 1 to 2 weekends |
| 5 | MCP server for a tool you use | Protocol, tool design | Task success on a small scripted set | 1 weekend |
| 6 | PR review or issue triage bot | Real integration | Agreement with a human on a labelled set | 2 weekends |
| 7 | Agent eval harness | Testing mindset | Pass rate across repeated trials | 1 to 2 weekends |
| 8 | A tool that solves one real problem for one real person | Product sense | Hours saved, or a quote from the user (real, with permission) | Varies |
Timeframes are rough guesses for someone who already codes, not measurements. Now each in turn.
1. RAG over a real document set, with evals
The basic version is everywhere, so the evaluation is the project. Use documents with real stakes: a government scheme's circulars, a library's docs, your university's regulations. Write 30 to 50 questions with the page where the answer lives. Measure two things separately: did retrieval find the right chunk, and did the final answer match? The beginner RAG chatbot tutorial gets you the working base, and building evaluation datasets covers how to write the question set.
README must show: the question set (committed), the hit-rate table, and three failures with a diagnosis of each.
2. Structured extraction from messy documents
Pick a document type with real variety: invoices, rental agreements, résumés, lab reports. Define a schema, extract with a model, validate the output, and handle the failures. Hand-label 30 documents for ground truth. See AI document processing with PDF invoice extraction for a worked pipeline.
Report accuracy per field (total amount, date, vendor), not one overall figure. Reviewers like seeing that dates fail more than totals, or the reverse. Scanned and low-quality documents are your honest failure case. Don't use real people's documents, and strip personal data from any samples you commit.
3. Text-to-SQL analyst with guardrails
A natural-language interface to a database you can legally share (a public dataset loaded into SQLite or Postgres). The interesting part isn't the SQL generation, it's the guardrails: read-only credentials, a statement allowlist, row limits, and what happens when the question is ambiguous. The data analyst agent tutorial is a good base.
Report execution accuracy on 30 questions you wrote, split by difficulty. Include a list of attacks you tried ("drop the table", "show me all emails") and what stopped them.
4. A support bot that refuses well
Take a company's public help centre (check its terms before scraping, or write your own fictional product docs). The measure is not how many answers it gives, but how it behaves when it shouldn't answer. Include out-of-scope questions, questions with a wrong premise, and prompts that try to override its instructions. The no-hallucinations support agent post shows the pattern.
Report: of N out-of-scope questions, how many did it decline correctly, and what did the wrong ones look like?
5. An MCP server for a tool you use
Wrap something you actually use (your notes folder, a spreadsheet, a calendar) as an MCP server with three or four well-described tools. The MCP server tutorial in Python covers the build. Reviewers care about tool design: descriptions the model can act on, sensible error messages, and permission limits. Add a short note on what the server can't do and what you restricted, drawing on prompt injection in tool-using agents.
6. PR review or issue triage bot
A GitHub Action that comments on pull requests or labels issues. It's a real integration with real inputs, which sets it apart from notebook demos. See the GitHub automation post. Run it on 20 to 30 old PRs or issues from a public repo, label them yourself, and report agreement. Show a false positive and explain why it happened.
7. An agent eval harness
Instead of building another agent, build the thing that tests agents: tasks, graders that check the end state, tool-call assertions, repeated trials. It demonstrates the habit teams want most and candidates have least. The agent evaluation frameworks post is a starting map. Run your harness against two prompt versions and report which wins and by how much, with the number of trials stated. If two runs of the same agent give different results, say so; that's a finding.
8. A tool for one real person
Find someone with a repetitive task: a shop owner reconciling payments, a teacher making worksheets, a relative filing forms. Build the smallest thing that helps, watch them use it, and record what broke. This is the project with no tutorial to copy, which is why it stands out. Ask permission before you name anyone or quote them, and never claim usage you don't have. "Used by my sister's tuition centre for four weeks, saved her roughly an hour a day by her estimate" is fine only if it's true.
What goes in the README?
Use this skeleton. The order matters because of the minute.
# Project name: one-line description of the problem it solves

## What it does
Two or three sentences. Who it's for, what goes in, what comes out.
## Results
| Metric | Value | How measured |
|---|---|---|
| Retrieval hit rate@4 | 0.00 (fill in) | 40 questions in /eval/questions.jsonl |
## Run it
1. `cp .env.example .env` and add your key
2. `pip install -r requirements.txt`
3. `python app.py`
## How it works
One diagram or five bullet points. Name the model, the chunk size, the
prompt file. Link to the code that matters.
## Known failures
- Scanned PDFs return empty text (see issue #3)
- Fails on questions that need two documents
## Decisions and trade-offs
Why this approach over the obvious alternative. What you'd do next.
## Which parts were AI-assisted
Honest note, short.
Placeholders such as "0.00 (fill in)" are exactly that: placeholders. Replace them with numbers you measured, or delete the row.
Mistakes that sink an otherwise good project
- No evaluation. "It works well" with no test set. This is the most common gap.
- Secrets in the repo. Anyone can search a public commit history for keys, and bots do. Rotate any key that was ever committed, since deleting the file doesn't remove it from history.
- One giant notebook. Fine for exploration, weak as a deliverable. Move the logic into modules with a thin entry point.
- Unpinned dependencies and no run instructions. If it takes twenty minutes to start, it won't be started.
- Tutorial clone. If the repo is identical to a popular tutorial, say that and show what you changed. Better, change something that matters: the data, the evaluation, the failure handling.
- Inflated claims. "Production-ready" for a weekend project invites hard questions. "Prototype with 40-question eval" doesn't.
- A deployed demo with an open API key. Put a spend cap on the key and rate-limit the endpoint, or don't deploy.
Choosing: a quick decision rule
If you're moving from backend or QA, pick 7 and 1: they reuse testing instincts. The AI engineering roadmap for SDETs and backend developers goes into that path. If you're a fresher with no work history, pick 1 and 8, since 8 gives you something to talk about that isn't a tutorial. If you want agent roles, pick 3 or 6 plus 7.
Whatever you pick, plan the interview. For each project, prepare a two-minute walkthrough, the three decisions you'd defend, and one thing you'd change. Practising that with a model is useful; the mock interviewer prompt can be pointed at your own repo. Paste in the README and ask it to probe the weakest claim.



