Generative Engine Optimisation: making your brand legible to AI systems
GEO is less about ranking and more about whether an AI system can describe your company accurately when someone asks about it.
AI companies operate distinct crawlers for training, live retrieval and user-initiated fetching, each with its own user agent that can be allowed or blocked in robots.txt. llms.txt is a proposed convention for offering a markdown summary of your site to language models; it is not an official standard and is not widely honoured yet.
| User agent | Operator | Purpose |
|---|---|---|
| GPTBot | OpenAI | Training data collection |
| OAI-SearchBot | OpenAI | Search indexing for its products |
| ChatGPT-User | OpenAI | Fetching a page on a user's request |
| Google-Extended | Controls use in Gemini and AI training | |
| ClaudeBot | Anthropic | Training data collection |
| PerplexityBot | Perplexity | Search index for answers |
| CCBot | Common Crawl | Open dataset used by many models |
| Applebot-Extended | Apple | Controls use in Apple AI training |
Note that Google-Extended does not affect Googlebot or your search rankings — it only controls AI training and Gemini use. Blocking Googlebot itself would remove you from search entirely, which is a different decision.
# robots.txt — allow retrieval and citation, block training
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: CCBot
Disallow: /
# Allow the crawlers that power citations in AI answers
User-agent: OAI-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: *
Allow: /
Sitemap: https://rovqix.in/sitemap.xml| Your situation | Suggested stance |
|---|---|
| Marketing site seeking visibility | Allow everything |
| Publisher monetising via traffic | Block training, allow retrieval |
| Paid content or courses | Block training; gate content behind auth |
| Documentation you want cited | Allow everything |
| Proprietary research | Block training, consider blocking retrieval |
| Ecommerce | Allow — product discovery via AI is growing |
The genuine trade-off: training use may reduce the need for anyone to visit you, while retrieval use is how you get cited today. Blocking everything is a defensible commercial position for a publisher and a poor one for a services business trying to be found.
A proposed convention: a markdown file at /llms.txt providing a curated, machine-friendly overview of your site — what it is, and links to the most important pages with brief descriptions. The intent is to give language models a clean summary instead of requiring them to parse your navigation.
# ROVQIX
> A website development studio building business websites, SaaS platforms
> and web applications with Next.js, React and Node.js.
## Services
- [Services overview](https://rovqix.in/services): what we build and for whom
- [Pricing](https://rovqix.in/pricing): package structure and typical ranges
## Guides
- [Next.js caching explained](https://rovqix.in/blog/nextjs-caching-explained)
- [Technical SEO checklist](https://rovqix.in/blog/technical-seo-checklist-developers)
## Contact
- [Contact](https://rovqix.in/contact): rovqix@gmail.comNo. It is separate from Googlebot and only controls use of your content for AI training and Gemini. Regular search crawling and ranking are unaffected.
Technically yes — it is a convention, not enforcement. Major operators publish their user agents and state that they honour it, but nothing prevents non-compliant scraping.
It takes an hour and cannot hurt. Do not expect measurable results, and do not let it displace work on page structure and content quality.
Filter your server access logs by user agent. This is more reliable than analytics, which typically excludes bots entirely.
ROVQIX Growth
SEO & growth team, ROVQIX
The ROVQIX growth team handles technical SEO, Core Web Vitals and AI-search visibility for the sites we build. Recommendations here are the ones we apply to client projects and to rovqix.in itself.
ROVQIXdesigns and builds production web platforms — Next.js front ends, Node.js APIs and the infrastructure behind them. Tell us what you're building and we'll scope it with you.
GEO is less about ranking and more about whether an AI system can describe your company accurately when someone asks about it.
AI search measurement is genuinely immature. The honest approach is a few reliable signals and a manual log, not a dashboard implying precision.
Technical SEO is mostly engineering work: correct status codes, canonical URLs, crawlable HTML and fast pages. None of it requires a keyword tool.
No spam. Just the occasional case study and craft breakdown.