XML sitemaps: what to include, what to leave out, and why
A sitemap is a list of the URLs you are confident about. Padding it with redirects and noindex pages teaches Google to ignore it.
Crawl budget is how many URLs a search engine will fetch from your site in a given period, determined by crawl capacity (what your server can handle) and crawl demand (how much the content is worth re-fetching). It matters mainly for sites with tens of thousands of URLs; for smaller sites, non-indexing is almost always a quality or duplication issue rather than a budget one.
A page can fail at any stage, and Search Console reports different exclusion reasons for each. Diagnosing the stage is the whole job — the fix for a crawl problem and an indexing problem have nothing in common.
| Status | Meaning | Typical fix |
|---|---|---|
| Discovered – currently not indexed | Known but not crawled yet | Improve internal linking and page value |
| Crawled – currently not indexed | Fetched but judged not worth indexing | Improve content depth and uniqueness |
| Duplicate, Google chose different canonical | Seen as a duplicate | Consolidate or differentiate the page |
| Soft 404 | Returns 200 but looks empty | Return a real 404, or add content |
| Blocked by robots.txt | Crawling disallowed | Fix the robots rule |
| Excluded by noindex | Explicitly excluded | Remove the tag if unintentional |
| Alternate page with proper canonical | Working as intended | None |
For sites of a few hundred to a few thousand pages, the honest answer to 'how do I get indexed' is usually to make fewer, better pages and link to them properly. Technical fixes only help when there is a technical obstruction.
Google has indicated it is generally a concern above roughly 10,000 URLs, or for sites with rapidly changing content. Below that, indexing problems are almost always about value and duplication.
Not directly — the page must still be crawled to see the tag. To prevent crawling entirely, disallow in robots.txt, but note that disallowed pages can still be indexed without content if they are linked externally.
New pages on an established, well-linked site are often indexed within days. A new site with few links can take weeks. Persistent non-indexing after a month indicates a value or duplication problem.
Google's Indexing API is officially limited to job postings and livestream content. For everything else, use sitemaps and URL Inspection requests.
ROVQIX Growth
SEO & growth team, ROVQIX
The ROVQIX growth team handles technical SEO, Core Web Vitals and AI-search visibility for the sites we build. Recommendations here are the ones we apply to client projects and to rovqix.in itself.
ROVQIXdesigns and builds production web platforms — Next.js front ends, Node.js APIs and the infrastructure behind them. Tell us what you're building and we'll scope it with you.
A sitemap is a list of the URLs you are confident about. Padding it with redirects and noindex pages teaches Google to ignore it.
Internal linking is the one ranking factor you fully control. Most sites still link almost entirely from the navigation.
Google renders JavaScript. It renders it late, inconsistently, and other crawlers often do not render it at all.
No spam. Just the occasional case study and craft breakdown.