What indexation is and why it matters
A search engine does not serve pages directly — it works from an index built in advance. The index
is a vast database in which every page is stored alongside a set of keywords, authority metrics and
technical characteristics.
A page’s route into the results looks like this:
Discovery → Crawling → Rendering → Indexation → Ranking
^ ^ ^ ^
sitemap / Googlebot JavaScript search
internal downloads executed index
links HTML
A page that fails any one of those stages will not appear in the results, however good its content.
What blocks indexation
Technical blocks
robots.txt — the file of instructions for bots. Disallow: /page forbids crawling but does not
guarantee absence from the index: a page can still be indexed through external links without the
bot ever visiting it.
The noindex meta tag — the reliable way to keep a page out of the index:
<meta name="robots" content="noindex, nofollow">
Rendering problems — if the content is generated by JavaScript and the bot did not wait for it
to execute, the page enters the index empty.
Structural problems
| Problem | Description | Fix |
|---|---|---|
| Orphan page | No inbound internal links | Add links from relevant sections |
| Duplication | Several URLs with identical content | Canonical tag on the master version |
| Pagination | /page=2, /page=3 indexed separately | rel=”canonical” or rel=”next/prev” |
| Parameter URLs | ?sort=price&filter=new — millions of variants | Disallow the parameters in robots.txt |
Indexation in e-commerce
Online stores face problems of their own:
Faceted search pages — /category?color=red&size=M&brand=Nike can generate thousands of pages
with overlapping content. Indexing them all burns crawl budget and creates duplication.
Out-of-stock product pages — remove them from the index or keep them? If the item is
temporarily unavailable, keep it. If it is discontinued for good, 301-redirect to the closest
alternative or to the category.
Product pages with no description — the bot indexes a page holding one image and an SKU. The
SEO value is zero and the crawl budget is spent all the same.
Important: engines prioritise indexation differently. Some are noticeably more conservative
than Google, and new pages can wait longer for their first visit. A sitemap with<lastmod>
helps speed up reindexation when content changes.
Monitoring index status
Google Search Console — the Coverage report shows how many pages are indexed, how many are
excluded and what the errors are. URL Inspection checks the status of one specific page.
Other webmaster consoles — Bing Webmaster Tools and Yandex Webmaster (Yandex being the dominant
search engine in Russia and several CIS markets) show index history and offer recrawl requests for
changed pages.
Regular monitoring is what tells you in good time that new pages have stopped being indexed, or
that an important section has dropped out of the index.