Crawl budget and indexing: why your pages are not in the index
'Discovered – currently not indexed' is Google saying your page is not worth the crawl. That is a content problem wearing a technical costume.
An XML sitemap lists the canonical URLs you want indexed, helping search engines discover pages that internal linking might not surface quickly. Include only 200-status, indexable, self-canonical URLs, keep lastmod accurate, split into a sitemap index beyond 50,000 URLs or 50MB, and reference it from robots.txt.
| URL type | Include? |
|---|---|
| Canonical, indexable pages | Yes |
| Pages with noindex | No |
| Non-canonical duplicates | No |
| Redirecting URLs | No |
| 404 or 410 pages | No |
| Paginated pages beyond the first | Optional; generally yes if indexable |
| Pages blocked in robots.txt | No |
| Login and account pages | No |
// app/sitemap.ts — generated from the same source the pages use
export default async function sitemap(): Promise<MetadataRoute.Sitemap> {
const articles = await getAllArticles();
return [
{ url: "https://rovqix.in", lastModified: new Date(), changeFrequency: "weekly", priority: 1 },
{ url: "https://rovqix.in/services", lastModified: new Date(), priority: 0.9 },
...articles.map((a) => ({
url: `https://rovqix.in/blog/${a.slug}`,
lastModified: new Date(a.updatedAt ?? a.publishedAt),
changeFrequency: "monthly" as const,
priority: 0.7,
})),
];
}Because it is generated from the content source, publishing an article adds it to the sitemap automatically. Manual sitemap maintenance always drifts, and drift is what makes sitemaps unreliable.
# robots.txt
User-agent: *
Allow: /
Sitemap: https://rovqix.in/sitemap.xmlUnder a few hundred well-linked pages, discovery usually happens through internal links anyway. It costs nothing to have one, and it makes Search Console reporting clearer.
Automatically, whenever content changes. Generating it at build or request time from the content source means it is always current.
For media-heavy sites where image or video search is a meaningful traffic source, yes. For a typical business site, image markup on the page is sufficient.
It speeds up discovery, particularly for new or poorly-linked pages. It does not guarantee or accelerate the decision to index.
ROVQIX Growth
SEO & growth team, ROVQIX
The ROVQIX growth team handles technical SEO, Core Web Vitals and AI-search visibility for the sites we build. Recommendations here are the ones we apply to client projects and to rovqix.in itself.
ROVQIXdesigns and builds production web platforms — Next.js front ends, Node.js APIs and the infrastructure behind them. Tell us what you're building and we'll scope it with you.
'Discovered – currently not indexed' is Google saying your page is not worth the crawl. That is a content problem wearing a technical costume.
Technical SEO is mostly engineering work: correct status codes, canonical URLs, crawlable HTML and fast pages. None of it requires a keyword tool.
Hand-writing meta tags per page guarantees drift. Here is how we wire the Metadata API once so no page can ship without a canonical, an OG image and a description.
No spam. Just the occasional case study and craft breakdown.