Many website owners use crawling and indexing as interchangeable terms. They are not β they are two distinct steps in how Google processes web content, and the difference matters practically. A page can be crawled without being indexed, and the reasons and fixes for each failure are completely different.
What happens during crawling?
Crawling is the discovery and downloading phase. Googlebot β Google's web crawler β follows links from page to page across the internet, downloading each page's HTML content to Google's servers. When Googlebot visits your page:
- It sends an HTTP request to your server
- Your server responds with the HTML content
- Googlebot reads the HTML, follows any links it finds, and adds those URLs to its crawl queue
- The downloaded content is sent to Google's processing systems
Crawling does not guarantee the page will appear in search results. It simply means Google now has a copy of the page to evaluate.
| Audit Focus Area | Crawl Impact | Organic Ranking Weight | Verification Tool |
|---|---|---|---|
| Crawl Budget Governance | Critical for sites >100 URLs | High (Ensures fresh indexation) | SEOLinkScan Broken Link Checker |
| Canonicalization Consistency | Prevents duplicate dilution | High (Consolidates PageRank) | Inspect self-referential canonical tags |
| Robots.txt & Sitemap Sync | Guards crawl accessibility | Foundational bot directive | Google Search Console Crawl Stats |
| Mobile-First Viewport Scaling | Seamless responsive layout | Official Google mobile ranking factor | Lighthouse Mobile Performance |
What happens during indexing?
Indexing is the evaluation and storage phase. After crawling, Google's systems process the downloaded content to determine:
- What the page is about (topic, language, entities)
- Whether it meets quality thresholds for inclusion in the index
- Whether it is a duplicate of another page already indexed
- Which version is canonical if multiple versions exist
If the page passes these evaluations, it is added to Google's search index and becomes eligible to appear in search results. If it fails β due to thin content, explicit blocking, or duplication β it is excluded and the reason appears in Search Console's Coverage report.
How do you know if a page has been crawled but not indexed?
Google Search Console tells you directly. In the Coverage report, pages with status "Crawled β currently not indexed" have been visited by Googlebot but excluded from the index. The report shows the specific reason for exclusion. Common reasons include: "Discovered β currently not indexed" (in the crawl queue but not yet visited), "Crawled β currently not indexed" (visited but quality standards not met), and "Page with redirect" (crawled but follows a redirect).
What causes pages to be crawled but not indexed?
- Thin or low-quality content β Google crawled the page but judged it insufficient quality for inclusion in the index
- Noindex meta tag β a
<meta name="robots" content="noindex">tag explicitly instructs Google not to index the page after crawling - Duplicate content β Google crawled the page but found essentially the same content already indexed at another URL
- Soft 404 β the page returned a 200 status code but contains content that looks like an error or empty page
- Canonicalisation β the page's canonical tag points to a different URL, so Google indexes that other URL instead
What prevents a page from being crawled at all?
- Robots.txt Disallow β blocking Googlebot from accessing the URL in robots.txt prevents crawling entirely
- Server errors (5xx) β if your server returns errors when Googlebot visits, it cannot crawl the page
- No links pointing to the page β orphan pages with no internal or external links may never be discovered by crawling
- Crawl budget exhaustion β on large sites, lower-priority pages may not be crawled if your crawl budget is consumed by higher-priority pages
Use the free SEOLinkScan Site Scanner to identify pages returning server errors or caught in redirect chains that prevent successful crawling.
- Conduct Complete Monthly Site Audits: Scan your domain with SEOLinkScan to detect broken links, redirect loops, and server error codes.
- Validate Self-Referential Canonical Tags: Prevent duplicate content indexation by confirming every URL features a verified canonical tag.
- Review XML Sitemap Hygiene: Ensure your sitemap exclusively contains clean 200 OK indexable URLs with zero 404s or 301 redirects.
- Inspect Bot Log Files: Analyze server logs to confirm Googlebot and Bingbot are spending crawl budget on cornerstone commercial content.
Strategic Action Plan: Mastering Technical SEO Architecture
- Crawling = Googlebot visits and downloads your page content
- Indexing = Google evaluates the content and adds it to its searchable index
- A page can be crawled without being indexed β quality, duplication, or blocking signals prevent indexing
- Check Search Console Coverage report to see which pages are in which state and why
- Fix crawl failures with our scanner; fix indexing failures by addressing the specific reason shown in Search Console
Continue reading: How do you improve your website's E-E-A-T score?