Many website owners use crawling and indexing as interchangeable terms. They are not β€” they are two distinct steps in how Google processes web content, and the difference matters practically. A page can be crawled without being indexed, and the reasons and fixes for each failure are completely different.

Quick Answer: Crawling means Googlebot visited your page and downloaded its content. Indexing means Google processed that content and added the page to its search index β€” making it eligible to appear in search results. Every indexed page was crawled first, but not every crawled page gets indexed. Google can crawl a page and then decide not to index it based on quality, duplication, or blocking signals.

What happens during crawling?

Crawling is the discovery and downloading phase. Googlebot β€” Google's web crawler β€” follows links from page to page across the internet, downloading each page's HTML content to Google's servers. When Googlebot visits your page:

  • It sends an HTTP request to your server
  • Your server responds with the HTML content
  • Googlebot reads the HTML, follows any links it finds, and adds those URLs to its crawl queue
  • The downloaded content is sent to Google's processing systems

Crawling does not guarantee the page will appear in search results. It simply means Google now has a copy of the page to evaluate.

Audit Focus AreaCrawl ImpactOrganic Ranking WeightVerification Tool
Crawl Budget GovernanceCritical for sites >100 URLsHigh (Ensures fresh indexation)SEOLinkScan Broken Link Checker
Canonicalization ConsistencyPrevents duplicate dilutionHigh (Consolidates PageRank)Inspect self-referential canonical tags
Robots.txt & Sitemap SyncGuards crawl accessibilityFoundational bot directiveGoogle Search Console Crawl Stats
Mobile-First Viewport ScalingSeamless responsive layoutOfficial Google mobile ranking factorLighthouse Mobile Performance

What happens during indexing?

Indexing is the evaluation and storage phase. After crawling, Google's systems process the downloaded content to determine:

  • What the page is about (topic, language, entities)
  • Whether it meets quality thresholds for inclusion in the index
  • Whether it is a duplicate of another page already indexed
  • Which version is canonical if multiple versions exist

If the page passes these evaluations, it is added to Google's search index and becomes eligible to appear in search results. If it fails β€” due to thin content, explicit blocking, or duplication β€” it is excluded and the reason appears in Search Console's Coverage report.

How do you know if a page has been crawled but not indexed?

Google Search Console tells you directly. In the Coverage report, pages with status "Crawled β€” currently not indexed" have been visited by Googlebot but excluded from the index. The report shows the specific reason for exclusion. Common reasons include: "Discovered β€” currently not indexed" (in the crawl queue but not yet visited), "Crawled β€” currently not indexed" (visited but quality standards not met), and "Page with redirect" (crawled but follows a redirect).

What causes pages to be crawled but not indexed?

  • Thin or low-quality content β€” Google crawled the page but judged it insufficient quality for inclusion in the index
  • Noindex meta tag β€” a <meta name="robots" content="noindex"> tag explicitly instructs Google not to index the page after crawling
  • Duplicate content β€” Google crawled the page but found essentially the same content already indexed at another URL
  • Soft 404 β€” the page returned a 200 status code but contains content that looks like an error or empty page
  • Canonicalisation β€” the page's canonical tag points to a different URL, so Google indexes that other URL instead

What prevents a page from being crawled at all?

  • Robots.txt Disallow β€” blocking Googlebot from accessing the URL in robots.txt prevents crawling entirely
  • Server errors (5xx) β€” if your server returns errors when Googlebot visits, it cannot crawl the page
  • No links pointing to the page β€” orphan pages with no internal or external links may never be discovered by crawling
  • Crawl budget exhaustion β€” on large sites, lower-priority pages may not be crawled if your crawl budget is consumed by higher-priority pages

Use the free SEOLinkScan Site Scanner to identify pages returning server errors or caught in redirect chains that prevent successful crawling.

Technical Diagnostic Checklist: Technical SEO Architecture
  • ✓
    Conduct Complete Monthly Site Audits: Scan your domain with SEOLinkScan to detect broken links, redirect loops, and server error codes.
  • ✓
    Validate Self-Referential Canonical Tags: Prevent duplicate content indexation by confirming every URL features a verified canonical tag.
  • ✓
    Review XML Sitemap Hygiene: Ensure your sitemap exclusively contains clean 200 OK indexable URLs with zero 404s or 301 redirects.
  • ✓
    Inspect Bot Log Files: Analyze server logs to confirm Googlebot and Bingbot are spending crawl budget on cornerstone commercial content.

Strategic Action Plan: Mastering Technical SEO Architecture

  • Crawling = Googlebot visits and downloads your page content
  • Indexing = Google evaluates the content and adds it to its searchable index
  • A page can be crawled without being indexed β€” quality, duplication, or blocking signals prevent indexing
  • Check Search Console Coverage report to see which pages are in which state and why
  • Fix crawl failures with our scanner; fix indexing failures by addressing the specific reason shown in Search Console

Continue reading: How do you improve your website's E-E-A-T score?