You added an MCP server because a blog post said it would let your assistant read Notion pages. Thirty seconds later it's running on your laptop with your user account's permissions, reading config from your shell environment, and returning text that your model will treat as context. Nobody reviewed it. That's the normal way MCP servers get installed.
This is the checklist I'd run before connecting one, and before shipping one. It's built from the MCP specification's own security pages and the Claude Code MCP docs, which I read on 2026-10-08, plus a small config linter I wrote and ran so you can see what the checks catch. The spec revision I read was dated 2026-07-28 (what latest pointed to that day). Older revisions differ in places, which I flag where it matters.
Not covered here: how the protocol works (MCP complete guide), which servers to pick (best MCP servers), or how to write one (build an MCP server in Python). For the model-side half of the problem, see prompt injection for tool-using agents.
What actually gets exposed when you install an MCP server?
Three things, and people usually only think about the first.
- Your machine. A local server is a program that runs with your privileges. The MCP security best-practices page lists the consequences plainly: arbitrary code execution, no visibility into what's running, data exfiltration, data loss. Its example of a malicious startup command is a
curlthat posts~/.ssh/id_rsato an attacker, tacked on after a legitimate-lookingnpxcall. - Your credentials. Whatever you put in the server's
envor headers, it can read, log or send somewhere. For stdio servers this is the whole credential story, because the spec says stdio implementations should not use the OAuth flow and should take credentials from the environment instead. - Your model's context. Everything the server returns goes into the prompt. The Claude Code docs say it directly: servers that fetch external content can expose you to prompt injection. A server doesn't need to be malicious for this to hurt. A Notion page, an issue, or a search result is enough.
So "is this server safe" really means three questions: what can it run, what can it read, and what can it make my model do.
Before you install: five checks
1. Read the launch command, all of it. The spec requires clients that offer one-click local server setup to show the exact command, untruncated, and get explicit approval. Do the same manually. Be suspicious of sudo, rm -rf, curl, &&, ;, and anything wrapped in bash -c.
2. Pin the version. npx some-package and uvx some-package fetch whatever is current at launch. A server you reviewed on Monday can be a different program on Friday. Use package@1.4.2 or a lockfile, and update on purpose.
3. Check who maintains it. Official vendor servers and servers from projects you already depend on are different from a repo with one commit and a polished README. I won't pretend there's a number for this. The question is whether you'd run an unsigned script from that author, because you are.
4. Decide the scope before you add it. Claude Code has three: local (default, just you in this project), project (shared through .mcp.json in version control) and user (all your projects). Project scope is convenient and also the way a cloned repo can propose servers to everyone who opens it. The docs say Claude Code prompts for approval before using project servers from .mcp.json in interactive sessions. In non-interactive modes (claude -p, Agent SDK, cloud sessions) it can't prompt and loads them without asking, and the docs list disabledMcpjsonServers and --strict-mcp-config for blocking. If you run agents headless against repos you didn't write, that's the setting to look at.
5. Ask whether you need it at all. A shell script or a plain API call from your own code has no long-lived server, no tool descriptions for the model to be confused by, and nothing new to patch.
Permissions: give the server the least you can
Filesystem servers are the classic case. Passing / or your home directory as the allowed path turns "read my project notes" into "read everything". Pass the one folder.
For anything stateful or destructive, split capabilities. Use a read-only database role for a query server. Use a GitHub token scoped to one repo and read access unless you need writes. If a vendor offers fine-grained tokens, use them even though the setup takes ten more minutes.
If you're on the building side, the spec's scope-minimization guidance is worth copying even for non-OAuth servers:
- Start with a minimal set of low-risk scopes (read and discovery).
- Ask for more via a
WWW-Authenticatechallenge with ascopevalue when a privileged tool is first used, instead of publishing every scope up front. - Don't use wildcard scopes like
*,allorfull-access, and don't bundle unrelated privileges to avoid future prompts. - Log scope elevation events with correlation IDs.
And a rule that's easy to skip: the spec lists "treating claimed scopes in token as sufficient without server-side authorization logic" as a common mistake. A scope on the token says what was granted. Your handler still has to check that this user may touch this record.
Credentials: where secrets go and where they must not
For servers you install:
- Put secrets in environment variables, not literal values in a config that you'll commit. Claude Code expands
${VAR}and${VAR:-default}incommand,args,env,urlandheaders, so the file can hold a reference while the value stays in your shell or secret manager. - Never put a token in a URL query string. The authorization spec says access tokens must not be in the URI query string, and they end up in proxy and server logs.
- One credential per server, not your personal god-token reused everywhere. When a server is compromised you want to revoke one thing.
One detail from the Claude Code docs I liked: in a remote server's url and headers, it reads certain credential variables (the page names ANTHROPIC_API_KEY, NPM_TOKEN and a few others) as empty, so a project config can't make your client send them to a server it names. If you really need to forward one, copy it into a variable with your own name.
For servers you build over HTTP:
- Validate that every incoming token was issued for your server (the audience). The spec says MCP servers must reject tokens that don't include them as an intended recipient.
- Never pass the client's token through to an upstream API. The spec says an MCP server must not pass through the token it received. If you call GitHub or Slack, use a separate token you obtained from them.
- Use short-lived access tokens; the spec says authorization servers should issue them, and must rotate refresh tokens for public clients.
- If you're an OAuth proxy with a static client ID and you allow dynamic client registration, get consent per client before forwarding to the third party. The spec describes the confused deputy attack that skips consent through a cookie from an earlier approval, and says proxies must keep a per-user registry of approved clients and use exact
redirect_urimatching, no wildcards. - If you hand out state handles (cart IDs, workflow IDs), bind them to the authenticated user on your side. The newer spec revision I read says possession of a handle must not be treated as authentication. If you're on an older revision with server-assigned session IDs, the same principle applies and the older version of that page covers session hijacking.
Honest caveat: I haven't implemented the full OAuth flow for this post and you probably shouldn't hand-roll it either. Use a library or your platform's auth layer, and use this section as the list of things to verify it does.
Untrusted output: treat tool results as hostile text
The spec puts obligations on both sides. Servers must validate all tool inputs, apply access controls, rate limit invocations and sanitize outputs. Clients should show tool inputs to the user before calling, validate results before passing them to the model, set timeouts, and log tool usage.
What that means in practice:
- Server authors: validate arguments against a real schema and reject extras. If a tool takes a path, resolve it and check it's inside the allowed root after resolving
..and symlinks. If it takes a URL, apply the SSRF rules from the next section. Return only the fields the model needs, not whole rows or raw HTML. - Server authors: keep tool descriptions short, static and honest. The description is text the model reads as part of its context, so it's an injection surface too.
- Everyone: don't trust annotations. The tools spec says clients must consider annotations untrusted unless the server is trusted. A
readOnlyHint: truefrom a stranger's server is a claim, not a fact. - Agent builders: keep a human approval step for anything irreversible or that sends data out, outside the model's control. The spec says there should always be a human in the loop with the ability to deny invocations. The longer discussion of gates is in the prompt injection post.
If you only do one thing here, make the server's blast radius small enough that a fully hijacked model can't do much. You can't reliably stop every injected instruction. You can control what the tools are allowed to do when one lands.
If your client fetches URLs, block the internal ones
This one is for people building clients or hosting server-side MCP clients, not for desktop users. During OAuth discovery, a client fetches URLs supplied by the server: the resource_metadata URL, authorization server URLs, token and authorization endpoints. A malicious server can point those at http://169.254.169.254/ (cloud metadata), localhost services, or private ranges.
The spec's mitigations: require HTTPS for OAuth URLs in production, block private and link-local ranges, validate every redirect hop instead of following blindly, and consider an egress proxy. It also warns against writing your own IP validation, because octal, hex and IPv4-mapped IPv6 encodings trip up hand-rolled parsers. Do not skip that note. Use a vetted library or proxy.
A related must from the same page: clients must only open http and https authorization URLs, reject javascript:, data:, file: and similar, and must not open URLs via a shell.
Audit: a config linter you can run
None of the above helps if nobody looks at the config. Below is a 63-line Python script that reads an mcpServers JSON file and flags the patterns above. It's a heuristic, not a security product: it catches the lazy mistakes, not a determined attacker. It uses only the standard library.
"""Lint an MCP client config (the {"mcpServers": {...}} shape) for common red flags.
Usage: python mcp_audit.py mcp.json
"""
import json
import re
import sys
from pathlib import Path
SECRET_KEY = re.compile(r"(key|token|secret|password|authorization)", re.I)
SECRET_VALUE = re.compile(r"^(sk-|ghp_|github_pat_|xox[bp]-|AKIA|eyJ)")
SHELL_WORDS = {"sh", "bash", "zsh", "cmd", "cmd.exe", "powershell", "pwsh"}
BROAD_PATHS = {"/", "~", "$HOME", "${HOME}", "C:\\", "/Users", "/home"}
def is_env_ref(v):
return isinstance(v, str) and re.fullmatch(r".*\$\{?[A-Z_][A-Z0-9_]*(:-[^}]*)?\}?.*", v) is not None
def audit(name, cfg):
out = []
cmd = cfg.get("command", "")
args = [str(a) for a in cfg.get("args", [])]
full = " ".join([cmd, *args])
if Path(cmd).name in SHELL_WORDS:
out.append("HIGH launches through a shell; the spec flags shell-chained startup commands")
if re.search(r"\b(sudo|curl|wget|rm -rf)\b|&&|\|\||;", full):
out.append("HIGH command contains sudo/curl/wget/rm -rf or shell chaining")
if cmd in ("npx", "uvx", "pipx") and not any(re.search(r"@\d|==\d", a) for a in args):
out.append("MED unpinned package; a later release runs with your privileges")
if cmd == "npx" and "-y" in args:
out.append("MED npx -y skips the install prompt")
for a in args:
if a in BROAD_PATHS:
out.append(f"HIGH filesystem-style arg grants a very broad path: {a}")
for k, v in {**cfg.get("env", {}), **cfg.get("headers", {})}.items():
if isinstance(v, str) and (SECRET_KEY.search(k) or SECRET_VALUE.match(v)):
if not is_env_ref(v):
out.append(f"HIGH literal secret in config: {k}")
url = cfg.get("url", "")
if url.startswith("http://") and not re.match(r"http://(localhost|127\.0\.0\.1|\[::1\])", url):
out.append("HIGH remote server over plain http")
if "?" in url and re.search(r"(token|key|secret)=", url, re.I):
out.append("HIGH credential in URL query string")
return out
def main(path):
servers = json.loads(Path(path).read_text()).get("mcpServers", {})
bad = 0
for name, cfg in servers.items():
findings = audit(name, cfg)
print(f"{name}: {'ok' if not findings else ''}")
for f in findings:
print(" ", f)
bad += any(f.startswith("HIGH") for f in findings)
print(f"\n{len(servers)} servers, {bad} with HIGH findings")
sys.exit(1 if bad else 0)
if __name__ == "__main__":
main(sys.argv[1])
I ran it on 2026-10-08 against a deliberately bad sample config with five invented servers (package names are fake). Output:
docs:
MED unpinned package; a later release runs with your privileges
MED npx -y skips the install prompt
HIGH literal secret in config: DOCS_API_KEY
notes: ok
helper:
HIGH launches through a shell; the spec flags shell-chained startup commands
HIGH command contains sudo/curl/wget/rm -rf or shell chaining
remote:
HIGH remote server over plain http
HIGH credential in URL query string
files:
MED unpinned package; a later release runs with your privileges
HIGH filesystem-style arg grants a very broad path: /
5 servers, 4 with HIGH findings
The notes entry passed because it pins @1.4.2, scopes the path to one folder, and references its token as ${NOTES_TOKEN}. That's the shape you want. The script exits non-zero when it finds a HIGH, so you can run it in CI against a committed .mcp.json.
Things it will not catch: a pinned package that is itself malicious, a legitimate server with an over-broad token, or a tool description that carries an injection. Those need human review and the permission limits above.
Monitoring: what to log and what to review
You want to be able to answer "what did the model ask this server to do last Tuesday" without guessing.
- Log each tool call with its full arguments and the caller identity (derived from the verified token, not a client-supplied field), plus a correlation ID. The spec's client guidance says to log tool usage for audit purposes, and its scope guidance says to log elevation events.
- Alert on a tool that's normally idle suddenly firing, or arguments that fall outside their usual shape (a new recipient domain, a path outside the project).
- Rate-limit per tool, not just per server. A runaway
send_messageshouldn't takelist_channelsdown with it. - Re-review the server list on a schedule. Servers accumulate the way browser extensions do. Run the linter, remove what you haven't used in a month, and rotate the credentials of anything you keep.
The checklist, condensed
Copy this into your repo's docs.
| Stage | Check |
|---|---|
| Install | Read the full launch command; no shell wrappers, chaining, sudo or curl |
| Install | Version pinned; source and maintainer reviewed |
| Install | Scope chosen deliberately (local, project, user); headless runs use a strict config |
| Permissions | Filesystem limited to one folder; database role read-only unless writes are needed |
| Permissions | Tokens fine-grained, one per server, shortest lifetime available |
| Credentials | Secrets referenced as environment variables; none in committed files or URLs |
| Credentials (building) | Audience validated; no token passthrough; per-user binding on state handles |
| Output | Inputs schema-validated, paths and URLs checked after resolving; outputs trimmed |
| Output | Annotations not used as a trust signal; human approval on irreversible actions |
| Client (building) | HTTPS for OAuth URLs; private ranges blocked; no shell to open URLs |
| Operations | Full-argument tool-call logs; per-tool rate limits; scheduled review |
What I did not verify
I read the MCP specification and the Claude Code MCP documentation but did not test any real third-party server for vulnerabilities, and I did not exercise the OAuth flows. The Claude Code page was truncated by my fetch tool at about 100,000 of 117,000 characters, so I haven't read its last section. Spec behavior changes between revisions (the version I read is stateless, with no protocol-level sessions, and older ones are not), so check the revision your SDK implements. If you work in a different client, its settings will be named differently, but the questions are the same ones.



