Site Indexing: Why Your Pages Never Reach Search
A page cannot rank if it is not in the index. That is obvious, yet it is the last thing anyone checks: first the texts get rewritten, links get bought and the design gets changed. Meanwhile, on an average site somewhere between 15 and 40% of pages are outside the index, and some of them are important.
How to check in five minutes
In Yandex.Webmaster open "Indexing → Pages in search" and compare the number with the real page count. The same section lists the excluded pages with the reason — the fastest source of truth there is. In Search Console the indexing report gives you the same picture. Additionally, check a specific URL through the inspection tool — it shows what the crawler sees right now. The logic behind these reports is described in the Yandex help.
Blocks: the site shut itself out
- A Disallow directive in robots.txt that closes a needed section along with the junk
- A noindex meta tag forgotten after a move from the staging server
- X-Robots-Tag in the server headers — invisible in the page source, look in the server response
- A canonical pointing to another page: the crawler treats yours as a duplicate and does not index it
- A block via .htaccess or basic auth left over from the development stage
Refusals: the crawler did not like the page
The "low-value page" category in Webmaster means the crawler arrived but decided not to add the page. Usually these are pages with almost no content, duplicate product cards, filtered selections with no products, or pagination pages with no substance. The cure is editorial, not technical: either fill the page, or merge it, or close it deliberately. How to separate useful filters from junk ones is covered in the article on ecommerce SEO.
Crawl budget: the robot never got there
If the site serves tens of thousands of URLs and a thousand of them are useful, the crawler spends its visits on junk. The symptom is new pages entering the index in 4–8 weeks instead of 2–7 days. It is solved by cleaning up duplicates, dropping sorting and parameter pages from indexing, speeding up the server response and building sensible internal linking: a page that takes five clicks from the home page gets crawled less often. The order of checks is in the audit checklist.
The orphan page
It has not a single internal link; it exists only in the sitemap. Pages like that index badly and rank even worse, because they receive no internal weight. The rule: every important page must be reachable from the menu, from the catalog or from related content — and preferably no deeper than three clicks from the home page.
What to do once you find the cause
- Remove the block and submit the page for recrawling through Webmaster — usually 1–7 days
- Add internal links to the page from sections that are already indexed
- Update the sitemap and make sure it contains only canonical URLs returning 200
- If the page was judged low-value, improve the content rather than resubmitting it for recrawling in circles
How long to wait
A new page on a healthy site enters the index in 2–7 days; on a neglected one it takes a month or more. After indexing, positions take another 2–6 weeks to settle. If a month after removing the block the page is still missing, the problem is not crawl speed but the quality assessment. General timing benchmarks are in the piece on how long SEO takes. The file settings that most often break indexing are covered in the article on robots.txt and the sitemap.
Checking indexing is 20 minutes once a month that save months of confusion. If doing it yourself is awkward, we run this diagnostic separately from a full audit — write to us through contacts.
Need help with a project?
Let's discuss your task and propose a solution — from a website to SaaS and security.
Get in touch