The noindex tag is a robots directive that tells search engines not to include a specific page in their search index. It lets you keep a URL crawlable and its links followable while preventing it from appearing in results. Marketers and engineers use it to control precisely which pages enter the index without blocking crawl access entirely.
What Is the Noindex Tag?
The noindex tag is a page-level instruction placed in the HTML head that asks search engines to drop a page from their index. The most common form is a meta robots tag. When a crawler sees it, it may still fetch the page and follow the links on it, but it should not show that URL in search results.
The standard markup looks like this, and any tag example in this guide is shown as escaped text so it renders as code rather than live markup:
<meta name="robots" content="noindex">You can also target specific crawlers, for example <meta name="googlebot" content="noindex">. The key nuance is that noindex controls indexing, not crawling. The page remains reachable, and its outbound links still pass discovery signals to the crawler even though the page itself stays out of the index.
How Is Noindex Different from Robots.Txt Disallow?
These two controls are frequently confused, yet they operate at very different layers. The table below maps the most common instruments so you can pick the right one.
| Directive | What It Controls | Page Can Rank? | Best Use Case |
|---|---|---|---|
| meta noindex | Removes a fetched page from the index | No | Thin, low-value, or duplicate-ish pages you still want crawled |
| X-Robots-Tag | Sends noindex in the HTTP response header | No | Non-HTML files such as PDFs and images |
| robots.txt disallow | Blocks crawling of a URL path | No (and cannot be indexed) | Large or private areas you never want fetched |
| canonical tag | Points to the preferred version of duplicate content | Yes (the canonical) | Near-duplicate pages that consolidate to one URL |
| nofollow | Stops link equity from flowing through links | Maybe | Untrusted or paid links you do not want to endorse |
| password protection | Requires auth before content is served | No | Genuinely private account and member areas |
The most dangerous mistake is combining robots.txt disallow with a noindex tag. A disallowed page is never fetched, so the crawler can never see the noindex directive on it. The result is silent: the page is merely uncrawled, not deindexed, and any already-indexed URL may linger in results indefinitely. If you need deindexing, keep the URL crawlable.
When Should You Use Noindex?
Noindex is the right tool for pages that serve a real user but add no search value. Typical candidates include:
- Internal search results pages, which create near-infinite low-quality entry points.
- Thin faceted or filtered URLs produced by ecommerce and listing filters.
- Thank-you and confirmation pages shown after a form or purchase.
- Paid landing pages that reuse duplicate copy from your main site.
- Staging and preview environments that must stay out of public search.
- Low-value tag and category archives with little unique content.
- Gated or account pages behind a login that should never be indexed.
- Duplicate print views and printer-friendly versions of existing articles.
When Should You NOT Use Noindex?
Noindex is not a universal "hide this" switch. In several situations another instrument is clearly better, and using noindex will cost you rankings or links.
For pages that already have backlinks or existing rankings, noindexing destroys the equity those URLs carry. A 301 redirect to the correct page preserves it instead. For near-duplicate pages that should consolidate, a canonical tag tells search engines which version to index while keeping the others accessible. For content you only want hidden temporarily, noindex is risky because it is easy to forget; a temporary auth gate or a robots.txt rule reviewed on a schedule is safer. For paginated series, noindexing every page often wastes crawl budget and hides valuable deep content; letting the series index or using view-all is usually stronger. When a page is truly gone, a 410 signals permanent removal far more cleanly than a lingering noindex.
How Do You Implement Noindex with the X-Robots-Tag HTTP Header?
Some resources are not HTML, so a meta tag cannot be placed in them. The X-Robots-Tag HTTP header solves this by carrying the directive at the response level, which works for PDFs, images, video, and any file served over HTTP.
The header is straightforward:
X-Robots-Tag: noindexYou typically add it in server configuration such as an Apache or Nginx block, or in your application framework's response pipeline. Because it lives at the server layer, you can apply per-file-type rules, for example sending noindex only to *.pdf or to a specific media directory, without touching page markup. You can also combine directives in one header, such as X-Robots-Tag: noindex, nofollow, when you want to stop both indexing and link discovery on the resource.
How Do You Verify a Noindex Is Actually Working?
Trusting the rendered page is not enough, because client-side code may hide or inject tags after crawl time. Verification should always inspect the raw signals a crawler receives.
Start with the URL Inspection live test in Google Search Console, which reports how Googlebot actually saw the page. Separately, curl the raw response headers and the HTML source to confirm the directive is present in the served document, not only in the rendered DOM. Watch the "Excluded by noindex tag" report in the Coverage section to see that Google has registered the exclusion. Allow time for recrawl, since deindexing is not instant, and confirm the URL is not also blocked in robots.txt, which would prevent the noindex from ever being seen. A page that lingers in the index despite a correct tag is usually a disallow conflict or a stale cache.
What Goes Wrong with Noindex in Practice?
Most noindex failures are process failures rather than syntax errors. The classic incident is a staging noindex rule shipped to production, quietly removing an entire site from search. A related trap is noindex injected only by JavaScript after render; many crawlers still parse raw HTML first, so the directive never registers. Teams also leave noindex on after a page becomes valuable, kneecapping its potential traffic.
The noindex plus disallow conflict described earlier is common and silent. Mass noindex from a plugin update or template change can sweep thousands of pages out of the index in minutes. Finally, long-term noindexed pages tend to be treated as nofollow for link discovery, so links on them stop contributing to site discovery and the pages around them lose crawl paths. Treat noindex as a deliberate, monitored setting, not a default.
Key Takeaways
- Noindex controls indexing, not crawling; the page can still be fetched and its links followed.
- Never pair noindex with robots.txt disallow, or the directive can never be seen and deindexing silently fails.
- Use noindex for low-value, thin, or duplicate-ish pages, and prefer canonical, 301, or 410 for other cases.
- For non-HTML files, use the X-Robots-Tag HTTP header with per-file-type server rules.
- Verify with raw headers, curl, and Search Console rather than the rendered DOM, and allow recrawl time.
Frequently Asked Questions
Can a Noindexed Page Still Be Crawled?
Yes. The noindex directive tells search engines not to index the page, but it does not block crawling unless you also add a robots.txt disallow. The crawler can still fetch the URL and follow links on it, which is why noindex is useful for keeping link discovery active while keeping a page out of results. If you disallow the URL, the noindex is never seen.
How Long Does It Take for Noindex to Remove a Page from Search?
After the directive is live and crawlable, removal is not immediate. Search engines must recrawl the URL, detect the noindex, and then process the exclusion, which typically takes days to a few weeks depending on crawl frequency. Monitor the "Excluded by noindex tag" report in Search Console and confirm the raw headers show the directive so you know the signal reached the crawler.
Should I Use Noindex or Canonical for Duplicate Content?
Use a canonical tag when the pages are near-duplicates that should consolidate to one preferred URL, because canonical keeps the main page ranking while signaling the others are alternates. Use noindex when a page is genuinely low value and should not appear at all, such as internal search results. Picking the wrong one either wastes crawl signals or throws away ranking equity you wanted to keep.
Does X-Robots-Tag Noindex Work on Pdfs and Images?
Yes. Because PDFs and images are not HTML, they cannot carry a meta robots tag, so the X-Robots-Tag HTTP header is the correct instrument. Set X-Robots-Tag: noindex at the server layer, scoped by file type or directory, and you can combine it with nofollow. Verify by curling the response headers of the file rather than opening it in a browser, which will not show the header.