What indexation is and why it matters

A search engine does not serve pages directly — it works from an index built in advance. The index
is a vast database in which every page is stored alongside a set of keywords, authority metrics and
technical characteristics.

A page’s route into the results looks like this:

Discovery  →  Crawling  →  Rendering  →  Indexation  →  Ranking
    ^             ^            ^              ^
 sitemap /     Googlebot    JavaScript      search
 internal     downloads     executed        index
  links         HTML

A page that fails any one of those stages will not appear in the results, however good its content.

What blocks indexation

Technical blocks

robots.txt — the file of instructions for bots. Disallow: /page forbids crawling but does not
guarantee absence from the index: a page can still be indexed through external links without the
bot ever visiting it.

The noindex meta tag — the reliable way to keep a page out of the index:

<meta name="robots" content="noindex, nofollow">

Rendering problems — if the content is generated by JavaScript and the bot did not wait for it
to execute, the page enters the index empty.

Structural problems

Problem Description Fix
Orphan page No inbound internal links Add links from relevant sections
Duplication Several URLs with identical content Canonical tag on the master version
Pagination /page=2, /page=3 indexed separately rel=”canonical” or rel=”next/prev”
Parameter URLs ?sort=price&filter=new — millions of variants Disallow the parameters in robots.txt

Indexation in e-commerce

Online stores face problems of their own:

Faceted search pages — /category?color=red&size=M&brand=Nike can generate thousands of pages
with overlapping content. Indexing them all burns crawl budget and creates duplication.

Out-of-stock product pages — remove them from the index or keep them? If the item is
temporarily unavailable, keep it. If it is discontinued for good, 301-redirect to the closest
alternative or to the category.

Product pages with no description — the bot indexes a page holding one image and an SKU. The
SEO value is zero and the crawl budget is spent all the same.

Important: engines prioritise indexation differently. Some are noticeably more conservative
than Google, and new pages can wait longer for their first visit. A sitemap with <lastmod>
helps speed up reindexation when content changes.

Monitoring index status

Google Search Console — the Coverage report shows how many pages are indexed, how many are
excluded and what the errors are. URL Inspection checks the status of one specific page.

Other webmaster consoles — Bing Webmaster Tools and Yandex Webmaster (Yandex being the dominant
search engine in Russia and several CIS markets) show index history and offer recrawl requests for
changed pages.

Regular monitoring is what tells you in good time that new pages have stopped being indexed, or
that an important section has dropped out of the index.