You asked a chatbot for "five peer-reviewed papers on remote work and productivity," and got a tidy list with authors, years, journals and DOIs. You paste the first DOI into a browser. Not found. You try the second. It opens a paper on something else entirely.
That's the central problem with using AI for academic papers, and it has a boring fix: let the model read what you give it, never trust anything it recalls, and check every reference in a database before it goes in your work. This post covers prompts for reading, summarising and critiquing papers, a script that checks DOIs against Crossref (with output I ran), and a list of what to keep for yourself. For general research habits, see how to use AI for research and AI research workflows. This post is narrower: the paper-level work.
What is safe to ask and what isn't?
The dividing line is where the information comes from.
| Task | Source of facts | Risk | Verdict |
|---|---|---|---|
| Summarise a PDF you uploaded | The PDF | Low to medium: it can miss tables, supplements, caveats | Safe with page-number quotes |
| Explain a method or term in a paper you've pasted | The PDF plus general knowledge | Low | Safe |
| Compare two papers you uploaded | Both PDFs | Medium: it may blur which paper said what | Safe, make it attribute each point |
| "Give me papers about X" from memory | The model's training data | High: invented or garbled references | Don't use as a source list |
| "What did Smith (2019) find?" with no PDF | Memory | High | Don't |
| Suggest search terms and databases | General knowledge | Low | Safe, then search yourself |
| Write the literature review text from a topic | Memory | High, and it's your argument | Don't |
If a search-enabled tool such as a deep-research mode finds real URLs, that's better than memory, but you still open each source. The deep research guide covers what those tools do well and where they slip.
Prompt 1: Read a paper without losing the caveats
Upload the PDF (or paste the text), then use this untested template. I haven't run it against a model for this post, so check the first output yourself against the paper.
You are helping me read a paper. Use only the attached document. Do not use
outside knowledge about this paper or its authors.
1. In two sentences: what question does the paper ask and what did it find?
2. Method: design, sample (N, who, where, when), measures, analysis method.
3. Main results: each as one sentence with the number, and the page or table
it comes from.
4. Limitations: what the authors say, with a quote and page number. Then
limitations they did NOT state that you can see from the design.
5. What the paper does not show. List claims a reader might wrongly take away.
If something isn't in the document, write "not stated". Never fill a gap.
Two parts earn their place. "Not stated" gives the model a legitimate way out instead of a guess. And item 5 forces the part most summaries skip: the gap between the abstract's confidence and what the design supports. Summaries of papers tend to inherit the abstract's tone, so ask separately for the weaknesses.
Quote requirements make checking cheap. A claim with "p. 7, Table 2" takes ten seconds to confirm. A claim with no location takes ten minutes.
Prompt 2: Critique the paper like a reviewer
A critique prompt works best when it names the standard. Different fields care about different flaws, so tell it yours.
Act as a sceptical peer reviewer in [FIELD]. Use only the attached paper.
Assess, with a quote or table reference for each point:
- Does the design support the causal language used? (association vs cause)
- Sample: size, selection, who is missing
- Measurement: could the key variable be measured badly?
- Statistics: what was tested, were multiple comparisons handled, are effect
sizes reported or only p-values?
- Alternative explanations the authors didn't rule out
- Which conclusions are stronger than the evidence
Then list the three questions you'd ask the authors. Do not praise.
If the paper doesn't give enough information to judge a point, say so.
Don't treat the output as a verdict. It's a set of leads. A model can raise a real concern about sampling, and it can also raise a generic one that doesn't apply because the authors handled it in an appendix you didn't upload. Open the paper and decide.
Prompt 3: Compare several papers without blending them
When models merge sources, they attribute one paper's finding to another. Make it build a table with per-cell evidence.
I've attached [N] papers, labelled A to [N]. Build a table with one row per
paper: question, method, sample size, main finding, main limitation.
Rules: every cell must come from that paper only. Add a page reference in
each cell. If a cell can't be filled from the paper, write "not stated".
After the table, list where the papers disagree and which method differences
might explain it. Do not reconcile them if you can't tell why they differ.
The last line stops the model from smoothing over real disagreement, which is often the interesting part of your review.
How do I verify citations?
This is the step people skip, so make it mechanical.
- Get the reference list as data. Whether it came from a model or your own notes, put each one on a line with DOI, claimed title and year.
- Resolve every DOI. Crossref has a free public REST API for this; the Crossref documentation says no sign-up is required. Not every scholarly item is in Crossref. Preprints on arXiv, for example, are looked up through arXiv's own API, as I show below.
- Compare title and year. A DOI that resolves to a different paper is a particularly sneaky error, because it looks real.
- Open the paper and find the claim. A real paper doesn't prove it says what you cited it for.
Here's a script that does step 2 and 3. It compares claimed titles and years against Crossref. I ran it; the output is below. It uses only the standard library, and falls back to curl because the Python on my Mac failed certificate verification (a setup issue on that machine, not in the API).
import json
import subprocess
import urllib.error
import urllib.parse
import urllib.request
# (doi, title the AI claimed, year the AI claimed)
claims = [
("10.1038/nature14539", "Deep learning", 2015),
("10.1038/nature14540", "Deep learning", 2015),
("10.9999/this.doi.does.not.exist", "Prompt tuning for everyone", 2023),
]
def lookup(doi):
url = "https://api.crossref.org/works/" + urllib.parse.quote(doi)
try:
with urllib.request.urlopen(url, timeout=20) as r:
return json.load(r)["message"]
except urllib.error.HTTPError as e:
if e.code == 404:
return None
raise
except urllib.error.URLError:
# Some Python installs on macOS lack CA certs; fall back to curl.
out = subprocess.run(
["curl", "-s", "-w", "\n%{http_code}", url],
capture_output=True, text=True, check=True,
).stdout
body, code = out.rsplit("\n", 1)
return json.loads(body)["message"] if code == "200" else None
for doi, title, year in claims:
rec = lookup(doi)
if rec is None:
print(f"{doi}: NOT FOUND in Crossref")
continue
real_title = rec["title"][0]
real_year = rec["issued"]["date-parts"][0][0]
ok = title.lower() in real_title.lower() and year == real_year
print(f"{doi}: {'MATCH' if ok else 'MISMATCH'} -> '{real_title}' ({real_year})")
Output from my run:
10.1038/nature14539: MATCH -> 'Deep learning' (2015)
10.1038/nature14540: MISMATCH -> 'Reinforcement learning improves behaviour from evaluative feedback' (2015)
10.9999/this.doi.does.not.exist: NOT FOUND in Crossref
The middle line is the one to learn from. I changed one digit of a real DOI, and it resolves to a real, unrelated paper from the same year. If you only checked "does the link work," that reference would pass. The three outcomes mean different things: MATCH is a lead worth reading, MISMATCH means the AI's metadata is wrong, and NOT FOUND means either a fabricated or mistyped DOI, or an item Crossref doesn't hold.
That last case matters. I also tried searching Crossref by title for Hinton's well-known knowledge-distillation paper, and the top two results were two other, unrelated papers. A title that isn't in Crossref isn't proof of fabrication. For preprints, ask arXiv directly. I queried its API for ID 1706.03762 and it returned the expected paper title ("Attention Is All You Need") with the authors and the 2017 submission date. Google Scholar, PubMed and your library's discovery tool are the other places to check.
A script only checks metadata. Step 4, reading the passage, is still yours.
Can AI help me find papers at all?
Yes, for the search, not for the list. Ask it for search terms, synonyms, adjacent fields and the names of the databases people use in your area. Then run the searches yourself.
I'm researching [TOPIC] for a [paper/thesis/review]. Suggest:
1. Ten search strings (with Boolean operators) for Google Scholar and
[PubMed / arXiv / SSRN / IEEE Xplore, as relevant].
2. Synonyms and older terms for the same idea that earlier papers might use.
3. Three neighbouring fields that study the same problem under other names.
4. Types of study I should look for to cover opposing views.
Do not list specific papers or authors.
"Do not list specific papers" is deliberate. Once you've got real papers from a database, snowball: read the references and citing papers of the best two. Some services, such as Semantic Scholar, offer an API for programmatic lookups, but I couldn't confirm its current limits from official pages in this session, so check its documentation before building on it.
What should I never delegate?
- The argument. What your review claims, and why paper A matters more than paper B, is the thing you're being assessed on. A model can't have read the field the way you have.
- Reading the key papers. The five or six you build on, you read in full. Summaries miss the footnote that changes the meaning.
- The reference list. Every entry verified by you, in a database, by title and DOI.
- Judgments of quality. Whether a journal is credible, or a sample is adequate.
- Quotes. Copy from the PDF yourself. Models paraphrase while sounding like they quote.
- Data analysis you can't explain. If you can't defend the statistics in a viva, don't use them.
- Disclosure decisions. The ICMJE recommendations, which many biomedical journals follow, say AI can't be an author, authors should disclose AI use, and humans are accountable for the material, including citations and attribution. Your university or target journal may have stricter rules. Read them before you start, not after.
A workflow that fits in an afternoon
- Search yourself with AI-suggested strings. Download 8 to 15 real PDFs.
- Run Prompt 1 on each. Save the outputs with page references.
- Read the abstract, methods and limitations of the best four in full.
- Run Prompt 3 across your shortlist for a comparison table.
- Run Prompt 2 on the two papers your argument depends on. Check each concern against the text.
- Build your own reference list from the database records, not from the model's output.
- Run the Crossref script on the final list, then open each paper to find the sentence you're citing.
If you study with a source-grounded notebook tool, how to use NotebookLM for study and research covers one way to keep answers tied to your uploads. For the failure mode underneath all of this, see the hallucinations deep dive.



