What GEO is

GEO stands for Generative Engine Optimization. It is not a content marketing strategy. It is the set of technical signals that determine whether AI search engines - ChatGPT, Perplexity, Google AI Overview, Claude - can crawl, parse, and cite your content.

Most of the gaps are in code. A robots.txt rule that blocks GPTBot. A React component that sets the page title inside a useEffect. A JSON-LD entity with no @id. These are not content problems. They are code problems.

Traditional SEO focused on keyword placement and link authority. AI search engines add a layer on top: the crawler needs access, the parser needs clean structure, and the knowledge graph needs entities it can anchor. If any of those three conditions fail, the content is invisible - not just ranked low, but absent.

GEO is the practice of making sure those conditions hold. And because the signals live in code, the right place to check them is before the code ships.

The four signal categories

Every GEO gap falls into one of four categories. Understanding each one makes it clear where to look and what to fix.

01

Crawler access

If GPTBot, PerplexityBot, ClaudeBot, or Google-Extended is blocked in robots.txt, that site gets zero citations regardless of content quality. The block is total. The AI engine never sees the page.

02

Entity resolution

AI engines build knowledge graphs to understand who and what a page is about. Organization and Person entities in JSON-LD need an @id (a canonical URL that identifies the entity) and sameAs links (to Wikidata, LinkedIn, Crunchbase) so the engine can anchor the entity across sources. Without these, the entity is opaque to GraphRAG pipelines.

03

Crawler-visible metadata

AI crawlers that do not execute JavaScript never see metadata set inside useEffect, useState, or any client-side lifecycle. The title, description, og:image, and canonical tag must be present in the server-rendered HTML. If they are not, the crawled version of the page has no metadata.

04

Passage structure

RAG systems chunk content for citation. A chunk typically follows a semantic boundary - a heading, a paragraph break, a list. Content under an h2 or h3 not wrapped in article, section, or aside may split across chunk boundaries in unpredictable ways. The first sentence after a heading should be a direct answer, not a preamble.

Crawler access in detail

The major AI search crawlers each have a named user agent. If any of these appears in a Disallow rule in robots.txt, the corresponding engine cannot index the site.

robots.txt - blocks GPTBot and ClaudeBot
# blocks ChatGPT from crawling
User-agent: GPTBot
Disallow: /

# blocks Claude from crawling
User-agent: ClaudeBot
Disallow: /

# blocks Perplexity from crawling
User-agent: PerplexityBot
Disallow: /

# blocks Google AI Overview
User-agent: Google-Extended
Disallow: /

This pattern appears more often than it should. It is usually introduced by copying a robots.txt from a template, or by a blanket "block all bots" rule added years before these crawlers existed.

Entity resolution in detail

An Organization entity with no @id cannot be linked across sources. The engine sees a name but cannot confirm it is the same entity as the one on Wikidata or LinkedIn. Add an @id using the canonical URL of the entity, and add sameAs links to known authority sources.

JSON-LD - entity without @id (invisible to GraphRAG)
{
  "@context": "https://schema.org",
  "@type": "Organization",
  "name": "Acme Corp",
  "url": "https://acme.example"
  /* no @id, no sameAs - engine cannot resolve this entity */
}
JSON-LD - entity with @id and sameAs (resolvable)
{
  "@context": "https://schema.org",
  "@type": "Organization",
  "@id": "https://acme.example/#organization",
  "name": "Acme Corp",
  "url": "https://acme.example",
  "sameAs": [
    "https://www.wikidata.org/wiki/Q12345",
    "https://www.linkedin.com/company/acme-corp",
    "https://www.crunchbase.com/organization/acme-corp"
  ]
}

Crawler-visible metadata in detail

A React app that sets the document title like this:

React - title set client-side only
useEffect(() => {
  document.title = 'How to configure Acme Corp | Acme Docs';
}, []);

The crawled HTML will have whatever title the server template sets - often a generic fallback or nothing at all. The same applies to meta name="description", og:image, and link rel="canonical". Set them server-side, in a framework's head mechanism, or via a static site generator.

Passage structure in detail

RAG citation quality depends on how cleanly a passage answers a question. Two habits improve it:

  • BLUF (Bottom Line Up Front). The first sentence after a heading should state the answer directly. "Server components run on the server and return HTML." Not "In this section we will explain how server components work."
  • Semantic wrapping. Wrap each major content section in an article, section, or aside element. This gives the chunker a reliable boundary to split on.

A chunk that begins mid-sentence because a heading was the last thing in the previous chunk is a failed citation. The model either skips it or paraphrases it inaccurately.

Why these are code problems

Each of these issues is introduced by a code change. Not a content edit - a code change.

  • Someone adds a User-agent: * / Disallow: / block to robots.txt while setting up a staging environment and forgets to remove it before deploying to production.
  • A developer migrates the site from a server-rendered framework to a client-side SPA and moves all head tags into a useEffect.
  • A backend engineer refactors the JSON-LD generation script and removes the @id field because it was not used in any template.
  • A design system update removes article and section wrappers in favor of generic divs.

None of these show up in a content audit. A content audit checks what is written, not what the crawler sees. These issues are invisible to anyone reviewing prose.

They show up in pull requests. The change is visible in the diff. The robots.txt edit, the removed @id, the useEffect that replaced a server-side title - all of these are one or two lines in a diff.

That is the right place to catch them: before the change lands in production, when the fix is trivial.

What SEOCode catches

We built SEOCode to catch these issues at the PR boundary. Every pull request that touches HTML, robots.txt, JSON-LD, or a sitemap is reviewed against a set of GEO rules. If something is wrong, SEOCode posts a comment with the exact fix.

The GEO rules we enforce:

01

geo/robots-blocks-ai-crawlers

Detects Disallow rules that block GPTBot, PerplexityBot, ClaudeBot, or Google-Extended.

02

geo/jsonld-missing-id

Flags Organization and Person entities in JSON-LD that are missing the @id field.

03

geo/jsonld-missing-sameas

Flags Organization entities that have no sameAs links to authority sources.

04

geo/client-side-title

Detects document.title assignments inside React lifecycle hooks and event handlers.

05

geo/client-side-meta

Detects meta tag DOM manipulation (description, og:image, canonical) done client-side.

06

geo/missing-canonical

Flags pages with no canonical link in the server-rendered head.

07

geo/no-semantic-section-wrapper

Flags h2/h3 headings that are not inside an article, section, or aside element.

08

geo/no-bluf-paragraph

Detects headings followed immediately by a preamble sentence rather than a direct answer.

These rules run on every pull request, automatically. The full rule reference is in the docs.