SPA Crawlability and Indexation Guide: Getting Search Engines to See Your Content

Google Search Console shows hundreds of pages as "Discovered - currently not indexed." Your users see every page load perfectly. Your search traffic flatlines while your content library grows. This is the signature symptom of a single page application with crawlability failures -- and this spa crawlability indexation guide walks you through diagnosing and fixing every common cause.

Crawlability and indexation are prerequisites, not optimizations. A page that cannot be crawled cannot be indexed. A page that cannot be indexed cannot rank. Everything else you do for SEO -- keyword targeting, content quality, link building -- is wasted if search engines never see the page in the first place. For the full rendering and architecture context, see the complete single page application SEO guide.

How to Diagnose SPA Crawlability Problems

Start with data, not assumptions. These diagnostic steps reveal exactly what search engines can and cannot see on your SPA.

Check the Raw HTML Response

Run curl -s https://yourdomain.com/target-page | head -100 and examine the output. If you see a root div, script tags, and no meaningful content, search engines are receiving the same empty shell. This is the most basic and most important test.

Compare this against the rendered output. Open the URL Inspection tool in Google Search Console, request a live test, and compare the "crawled HTML" tab against the "rendered HTML" tab. The crawled HTML is what Googlebot receives. The rendered HTML is what the Web Rendering Service produces after executing JavaScript.

Any content that exists in the rendered HTML but not in the crawled HTML depends on JavaScript rendering -- and is therefore subject to rendering delays, failures, and incompatibility with non-Google search engines.

Audit the Coverage Report

Google Search Console's Pages report (formerly Index Coverage) categorizes your URLs into four states: Indexed, Not Indexed, Excluded, and Error. For SPAs, focus on these specific statuses:

  • Discovered - currently not indexed: Google found the URL but has not crawled it. This often indicates crawl budget exhaustion or low perceived value.
  • Crawled - currently not indexed: Google crawled the URL but chose not to index it. For SPAs, this often means the crawled HTML was empty or thin, and the rendering pass either has not happened or produced insufficient content.
  • Duplicate without user-selected canonical: Multiple URLs resolve to the same content. SPAs with improper routing often serve identical content across different URL patterns.

Validate Internal Link Structure

Use Screaming Frog or a similar crawler configured to execute JavaScript. Run a crawl from your homepage and compare the URLs discovered against your sitemap. Any URL in your sitemap that the crawler cannot reach through internal links has a link discovery problem.

Common SPA link discovery failures include: navigation menus rendered via JavaScript event handlers instead of anchor tags, pagination controls that use click events instead of linked URLs, and filter/facet interfaces that modify content without updating the URL.

Test Non-Google Crawlers

Google is the most capable JavaScript renderer among search engines, but it is not the only search engine. Test your pages with Bing's URL Inspection tool and check social media preview debuggers (Facebook Sharing Debugger, Twitter Card Validator, LinkedIn Post Inspector). If these tools show missing content, your SPA is invisible to their platforms.

How to Fix SPA Crawlability Issues

Each fix targets a specific failure mode. Implement them in order of impact.

Implement Server-Side Rendering for Indexable Routes

The highest-impact fix is ensuring that every page you want indexed returns complete HTML in the server response. Server-side rendering delivers direct SEO benefits by removing the dependency on JavaScript execution for content visibility.

If a full SSR implementation is not feasible immediately, use dynamic rendering as a bridge. Serve pre-rendered HTML to known crawler user agents while serving the standard SPA to regular users. This is not cloaking -- Google explicitly supports this approach as an interim solution.

Fix Internal Link Markup

Replace every navigation element that uses JavaScript event handlers with standard anchor tags containing href attributes. This is non-negotiable.

Wrong: <button @click="router.push('/products')">Products</button>

Right: <a href="/products">Products</a> (with client-side navigation handled by the router intercepting the click)

Most SPA frameworks (Next.js Link, Nuxt.js NuxtLink, Angular routerLink) provide components that render standard anchor tags while still enabling client-side navigation. Use these components for every internal link.

Generate and Submit a Comprehensive Sitemap

Your XML sitemap should include every URL you want indexed. For SPAs with dynamic routes (product pages, user profiles, articles), generate the sitemap from your database or API at build time or on a schedule.

Include <lastmod> dates to help search engines prioritize re-crawling updated pages. Submit the sitemap through Google Search Console and reference it in your robots.txt file.

Do not include URLs that return empty HTML shells. If a URL in your sitemap leads to a page that requires JavaScript rendering to show content, and you have not implemented SSR for that URL, the sitemap is directing crawlers to empty pages -- which wastes crawl budget and signals low quality.

Implement Proper Canonical URLs

SPAs frequently create duplicate content through multiple URL patterns:

  • /products/widget and /products/widget/ (trailing slash)
  • /products/widget and /products/widget?ref=homepage (query parameters)
  • /products/widget and /#/products/widget (hash routing)

Set a canonical URL on every page that points to the single authoritative version. Server-side rendering makes this straightforward -- the canonical tag is in the initial HTML response and does not depend on JavaScript execution.

Handle Status Codes Correctly

SPAs often return HTTP 200 for every request, including URLs that should return 404 (page not found) or 301/302 (redirects). When your SPA serves the same JavaScript shell for every URL and handles routing client-side, the server always responds with 200.

This confuses search engines. A URL that shows "Page Not Found" in the rendered DOM but returns HTTP 200 is a soft 404. Google may eventually detect it, but the ambiguity wastes crawl budget and can lead to low-quality pages in the index.

Configure your server to return appropriate HTTP status codes. With SSR frameworks, this is handled in the routing layer. Without SSR, configure your server or CDN to return 404 for known non-existent routes.

Common Mistakes That Block Indexation

These mistakes are specific to SPAs and compound over time.

Lazy-loading above-the-fold content. If your hero section, primary heading, or main content area depends on lazy loading triggered by viewport intersection, crawlers may never see it. The WRS renders a snapshot of the initial viewport state and does not scroll. Load all primary content eagerly.

Infinite scroll without paginated URLs. If your content list (products, articles, search results) uses infinite scroll as the only navigation method, Googlebot can only see the first page of results. Provide paginated URL alternatives (/products?page=2, /products?page=3) that are crawlable and linked from within the DOM.

Client-side redirects. JavaScript redirects (window.location.href = '/new-url') may or may not be followed by crawlers. Use server-side 301 redirects for permanent URL changes. This ensures all search engines process the redirect immediately and transfer ranking signals.

Render-blocking third-party scripts. Analytics, chat widgets, A/B testing scripts, and tag managers that load synchronously can delay or break the rendering of your primary content. Load third-party scripts asynchronously and after your main content has rendered.

Missing robots meta tags on non-indexable pages. SPA routes for login pages, account settings, and other non-public content should include <meta name="robots" content="noindex"> to prevent indexation. Without this, search engines may index pages that provide no value to searchers and dilute your site's quality signals.

Crawlability and Indexation Checklist

Run through this checklist for every indexable route in your SPA.

  • [ ] curl returns complete HTML content, not just a JS shell
  • [ ] Title tag and meta description are in the initial HTML response
  • [ ] Canonical URL is set and points to the correct version
  • [ ] Open Graph and Twitter Card tags are present in initial HTML
  • [ ] All internal links use <a href="..."> markup
  • [ ] JSON-LD structured data is in the initial HTML response
  • [ ] Page returns correct HTTP status code (200, 404, 301 as appropriate)
  • [ ] URL is included in the XML sitemap
  • [ ] Page is reachable through internal links from the homepage within 3 clicks
  • [ ] No render-blocking resources prevent content from appearing
  • [ ] JavaScript rendering does not delay critical content
  • [ ] Non-indexable pages have noindex robots meta tag
  • [ ] Pagination uses crawlable URL-based navigation, not infinite scroll alone
  • [ ] Images have alt attributes and are not loaded exclusively via JavaScript
  • [ ] robots.txt does not block CSS or JS files needed for rendering

Frequently Asked Questions

How long does it take Google to index a new SPA page with SSR?

For pages with server-rendered HTML, proper internal linking, and sitemap inclusion, indexing typically happens within 1-7 days. High-authority domains with frequent crawling may see indexing within hours. Without SSR, the timeline extends to days or weeks because the page must pass through both the crawl queue and the rendering queue.

Why does Google show my SPA page as indexed but with the wrong title?

This usually means Google indexed the page based on the initial HTML response before JavaScript rendering completed. If your title tag is set client-side, Google may have used the default title from your HTML shell (often the app name or a generic title). Implementing server-side rendering fixes this by placing the correct title in the initial response.

Can I use a CDN in front of my SSR application to improve crawlability?

Yes. CDN caching reduces TTFB and ensures consistent response times for crawlers. Configure your CDN to cache server-rendered HTML for an appropriate duration (based on how frequently your content changes). Use cache invalidation when content updates to prevent stale content in the index.

Do single page applications need a separate mobile site for SEO?

No. Responsive design with a single URL structure is Google's recommended approach. Your SSR implementation should serve the same HTML to all devices, with CSS handling the layout adaptation. Separate mobile URLs (m.domain.com) create duplicate content management overhead and are no longer necessary with responsive design and mobile-first indexing.

Key Takeaways

  • Crawlability problems in SPAs are diagnosed by comparing raw HTML responses (curl) against rendered output in Google Search Console's URL Inspection tool.
  • Server-side rendering is the most reliable fix for indexation failures because it eliminates the dependency on JavaScript execution for content visibility.
  • Every internal link must use standard anchor tags with href attributes; JavaScript event handlers are invisible to crawlers and block URL discovery.
  • Generate XML sitemaps from your data source and ensure every listed URL returns complete, server-rendered content -- do not direct crawlers to empty JS shells.
  • Use the crawlability checklist on every indexable route before launch to catch issues that compound silently as your site grows.