All posts
What Is the X-Robots-Tag? How to Noindex Pages You Can't Add a Meta Tag To

What Is the X-Robots-Tag? How to Noindex Pages You Can't Add a Meta Tag To

The X-Robots-Tag lets you noindex PDFs, images, and other files where a meta robots tag won't work. How it works, how to set it up, and how to confirm it's live.

September 28, 2026 · 5 min read

If Google Search Console shows one of your PDFs, images, or API-served pages sitting in the index when it shouldn't be, a meta robots tag can't fix it — you can't inject <meta name="robots"> into a PDF's <head>, because a PDF doesn't have one. That's what the X-Robots-Tag is for.

What the X-Robots-Tag Actually Does

The X-Robots-Tag is an HTTP response header that carries the same instructions as a meta robots tag — noindex, nofollow, noarchive, and so on — but it's attached to the server response itself instead of embedded in HTML. A response might look like this:

HTTP/1.1 200 OK
Content-Type: application/pdf
X-Robots-Tag: noindex

Because it lives at the HTTP layer, it works on any file type a crawler can request: PDFs, images, JSON responses, downloadable spreadsheets, even HTML pages if you'd rather manage indexing rules at the server level than edit templates. Google, Bing, and other major crawlers all honor it the same way they honor the meta tag equivalent.

X-Robots-Tag vs. the Meta Robots Tag: When to Use Which

The two directives do the same job but at different layers, and that difference decides which one you reach for:

| Situation | Use | |---|---| | Noindexing an HTML page you can edit | Meta robots tag | | Noindexing a PDF, image, or non-HTML file | X-Robots-Tag | | Applying one rule to an entire folder (e.g., /downloads/) at once | X-Robots-Tag via server config | | A CMS where you can't touch server headers but can edit page <head> | Meta robots tag |

A common mistake is trying to set both and assuming they stack additively — they don't conflict, but maintaining the same rule in two places just means twice the chance one gets left stale after a redesign. If your server setup allows it, picking one source of truth per content type is simpler to audit later.

How to Add an X-Robots-Tag (Apache, Nginx, and CDN Examples)

The exact method depends on where the file is served from.

Apache (.htaccess):

<FilesMatch "\.pdf$">
  Header set X-Robots-Tag "noindex, nofollow"
</FilesMatch>

Nginx:

location ~* \.pdf$ {
  add_header X-Robots-Tag "noindex, nofollow";
}

Vercel / Next.js, via next.config.js headers:

async headers() {
  return [
    {
      source: "/downloads/:path*",
      headers: [{ key: "X-Robots-Tag", value: "noindex" }],
    },
  ];
}

Cloudflare or another CDN: most let you set custom response headers per path rule in the dashboard, without touching origin server config at all — useful if you don't control the backend directly.

Whichever route you use, the rule applies to the response, not the file itself, so it only takes effect once it's actually deployed and served — a header added in a config file that hasn't shipped yet won't do anything, no matter how correct the syntax is.

How to Check Whether Your X-Robots-Tag Is Actually Working

Don't assume a deployed header is a working header. Confirm it three ways:

  1. Check the raw response. Run curl -I https://yoursite.com/file.pdf and look for x-robots-tag in the output. If it's missing, the header isn't reaching that path — check your matching rule (file extension, folder path) before assuming Google is ignoring it.
  2. Use Google's URL Inspection tool in Search Console. It shows the indexing status and will flag noindex detected via HTTP header the same way it flags a meta tag.
  3. Watch the Page Indexing report. Pages correctly excluded this way show up under "Excluded by 'noindex' tag" — the same report covers both the HTTP header and meta tag methods, so don't be surprised to see them grouped together. If you're not sure what that report is telling you more broadly, the excluded by noindex tag breakdown covers the other ways a page can land there.

Deindexing isn't instant either way — expect the removal to show up in Search Console over days to a couple of weeks as Google recrawls the URL, not immediately after the header goes live.

Common X-Robots-Tag Mistakes That Quietly Deindex Pages

The failure mode that actually costs traffic is the header ending up somewhere you didn't intend:

  • A wildcard rule that's too broad. A path match meant for /downloads/*.pdf that accidentally also catches /downloads/case-study.html will quietly noindex a page you wanted ranked, with no warning beyond a slow ranking drop.
  • Staging config that leaked to production. Staging environments are often blanket-noindexed at the header level for good reason — the risk is that config surviving a deploy and noindexing the live site. This is a more aggressive version of the soft 404 problem: instead of Google guessing a page is low-value, you've told it directly not to index it.
  • CDN and origin server disagreeing. If your CDN sets one X-Robots-Tag value and the origin server sets another, behavior depends on which one the crawler actually receives — check with curl -I against the live URL, not just each config file in isolation.
  • Forgetting it also blocks non-search consumers. Some tools that fetch your pages via HTTP (link preview generators, certain monitoring services) respect the same header, so a rule that's correct for search engines can have side effects elsewhere. Worth a quick check if an integration suddenly stops rendering previews for noindexed paths.

Regularly auditing what's set to noindex at the header level — not just in your HTML — is part of the same indexing hygiene as checking your robots.txt file for accidental blocks. Both are easy to set once, forget about, and only notice when traffic to a page you actually wanted indexed quietly disappears. If you're tracking indexing status across a growing set of URLs, Traffic Monitor pulls your Search Console data automatically so a stray noindex header shows up as a change worth investigating, not a mystery three months later.

Track your blog rankings automatically

Connect Google Search Console and get daily alerts, keyword research, and AI diagnostics — free to start.

Start free →