Skip to content
SEO 3 min read Sajid Aslam

How Google Crawls and Indexes a Website

Four stages, four different failure modes. Most SEO confusion comes from treating them as one process.

The four stages between discovery and ranking in Google Search

Short answer

Discovery, crawling, rendering, indexing, then ranking. A page can pass one stage and fail the next, and Search Console tells you which. Most "not ranking" problems are actually indexing problems, and most indexing problems are quality or duplication signals rather than technical blocks.

Understanding the stages is worth more than memorising tactics, because it tells you which tactic is even relevant.

1. Discovery

Google has to learn the URL exists. It does that from:

  • links on pages it already crawls — by far the most important route
  • your XML sitemap
  • a manual submission in Search Console

A page with no internal links and no sitemap entry may sit undiscovered indefinitely. This is what "orphan page" means, and it is why a proper internal linking system matters more than it appears to.

2. Crawling

Googlebot requests the URL. Things that go wrong here:

  • robots.txt disallows it
  • the server returns an error or is too slow
  • the page redirects, possibly in a chain
  • crawl budget is being spent elsewhere

Crawl budget only becomes a real constraint on large sites. If you have two hundred pages, you do not have a crawl budget problem, whatever a tool told you. What you might have is a crawl *efficiency* problem — Googlebot spending its visits on parameter URLs, paginated archives and redirect chains instead of your actual content.

3. Rendering

Google runs the page's JavaScript to see the final content. This is a separate, queued stage — it can happen well after the initial crawl.

The practical implication: content that only exists after JavaScript runs is indexed more slowly and less reliably than content in the initial HTML.

This is one reason server-rendered sites have an advantage. It is also why blocking CSS and JavaScript in robots.txt is harmful — Google renders the page badly and may conclude it is broken or unfriendly on mobile.

A related, subtler version: content that exists in the HTML but is hidden by CSS until JavaScript runs. Crawlers read it, but if a rendering failure leaves it invisible to users, you have a page that reads well to a bot and blank to a person. This site had that exact pattern in its scroll animations, fixed with a no-JS fallback.

4. Indexing

Google decides whether to store the page. It can decline. Reasons it declines:

  • noindex
  • it considers the page a duplicate and picked a different canonical
  • it judged the page not worth indexing — usually thin, or near-identical to something else

That last one shows in Search Console as "Crawled, currently not indexed", and it is a quality judgement rather than a technical fault. More pages will not fix it. A better page might.

5. Ranking

Only now does the page compete for queries. Everything people usually call "SEO" applies at this stage, and none of it matters if you failed an earlier one.

Using Search Console properly

The Pages report tells you which stage each URL is stuck at. That is the single most useful diagnostic view available, and it is free.

URL Inspection on a specific page shows: when it was last crawled, whether it is indexed, what the rendered HTML looked like to Google, and the canonical Google chose versus the one you declared. When those two disagree, that is your answer.

What this means practically

  • Get the sitemap right, and make sure it is generated from reality rather than baked at build time from a source that might be unavailable
  • Link internally with intent, so importance is inferrable from structure
  • Put important content in the initial HTML
  • Do not block assets
  • Fix duplicates rather than hoping canonicals sort it out
  • Publish fewer, better pages rather than more thin ones

The SEO website audit checklist turns this into something you can work through, and the SEO service starts exactly here.

Related services

Related reading