Every check we run

We publish the full rule set — 25 checks across 5 areas. If a report tells you something is wrong, you can read here exactly what we looked for and what would fix it.

The engine's code is not open source. The rules are: same words the report uses, taken straight from the engine itself, so this page can't drift from what actually runs.

Crawlability · 25% of the score

Homepage could not be fetched CRWL-001 · critical

If a plain HTTP request fails, search engines and AI systems can't get the content either — the page effectively doesn't exist for them.

How to fix: Check the server and CDN configuration to make sure ordinary clients (no cookies, no JavaScript) are not blocked.

robots.txt blocks AI crawlers CRWL-002 · high

Blocked crawlers can't bring your content into their index, so they won't cite it in answers. This is the most common reason a well-written page goes unseen.

How to fix: Allow these crawlers if you want to be cited. If you only want to opt out of training, unblock the retrieval bots and keep the training-only ones restricted.

No robots.txt CRWL-003 · medium

Without robots.txt, crawlers have no way to learn your intent (what may be crawled, where the sitemap is). Most default to allowing everything, but you lose the chance to declare it.

How to fix: Add a robots.txt declaring your sitemap and stating your policy for AI crawlers explicitly.

Sitemap is unreachable CRWL-004 · medium

The sitemap is one of the main paths crawlers use to discover deeper pages; without it, new content gets indexed more slowly.

How to fix: Generate sitemap.xml and declare its location in robots.txt.

No llms.txt CRWL-005 · low

llms.txt is still an emerging convention for telling models which pages matter. Missing it doesn't break crawling — it's an optional addition.

How to fix: Optionally publish an llms.txt listing your key pages. Don't expect it to guarantee citations.

Training-only crawlers are restricted CRWL-006 · info

This restriction only affects model training, not retrieval or citation. Listed separately so it isn't mistaken for being blocked.

How to fix: If your goal is to be cited rather than trained on, this setting is fine as it is.

Some site resources could not be verified CRWL-007 · info

These requests failed even after a retry, so we can't tell whether the files exist. Listed so a fetch failure isn't misread as a missing file.

How to fix: Re-run the check later. If it keeps failing, look for CDN or WAF rules that block non-browser clients on those paths.

Homepage has no static links to other pages CRWL-008 · high

Crawlers that don't run JavaScript can only follow links that exist in the HTML. With none on the homepage, how much of your site gets discovered depends entirely on whether a sitemap is reachable.

How to fix: Render the main navigation as real <a href> links in the server-side HTML rather than JavaScript click handlers, and publish a sitemap so deeper pages have a second path in.

Basics · 25% of the score

Page has no title META-001 · high

The title is the primary heading used in search results and AI summaries; without it, systems have to guess what the page is about.

How to fix: Give every page a unique title with the distinctive words in the first 60 characters.

Title is too long META-001b · low

Overlong titles get truncated in results and summaries, so the tail of the message is never read.

How to fix: Trim to about 60 characters, with the most distinguishing words first.

No meta description META-002 · medium

AI systems often reuse the description as the page summary. Without one they excerpt the body, and the result is unpredictable.

How to fix: Write a 120–160 character description per important page saying what the page answers.

No canonical URL META-003 · medium

Without a canonical, parameterised copies (e.g. ?utm_source=…) can be treated as separate pages, splitting the weight of one piece of content.

How to fix: Declare a self-referencing canonical on every page.

Page is marked noindex META-004 · critical

noindex removes the page from search indexes, and AI systems generally won't cite it either.

How to fix: Confirm this is intentional. If the page should be findable, remove the directive.

Wrong number of H1 headings META-005 · high

The H1 is the clearest statement of what a page is about. Zero or many makes it harder for systems to tell.

How to fix: Keep a single H1 as the main heading and demote the rest to H2.

Open Graph data is incomplete META-006 · low

Many AI and social systems prefer Open Graph fields for display. Missing them makes presentation unpredictable.

How to fix: Add og:title, og:description and og:image.

Content readability · 25% of the score

Almost no readable text without JavaScript EXTRACT-001 · critical

Most AI crawlers don't execute JavaScript. If the static HTML has no body text, the page looks empty to them — however good the content is.

How to fix: Render the core content server-side or at build time. If client rendering is required, at least put the key facts of the first screen into the static HTML.

Too much code relative to content EXTRACT-002 · high

Scripts and styles dominate the payload, leaving little room for content. Crawl budgets are finite; pages like this are more likely to be judged not worth the time.

How to fix: Reduce inline scripts and styles, and output the first screen of content server-side.

The root URL is a redirect stub EXTRACT-003 · critical

This page contains nothing but a meta-refresh directive. Crawlers that don't run JavaScript and don't follow meta refresh read nothing here, and usually never reach the target page.

How to fix: Serve real content at the root, or use an HTTP 301. Don't use meta refresh for the first hop.

Images missing alt text EXTRACT-004 · medium

AI systems can't see images; the alt attribute is the only textual entry point. Missing alt text means that information simply doesn't exist to them.

How to fix: Add alt text describing what each meaningful image conveys (not just "image"). Decorative images can keep an empty alt.

Structured data · 15% of the score

No JSON-LD structured data STRUCT-001 · high

Structured data is the most direct way to hand facts to machines: who you are, what the page is, what it costs. Without it, AI has to infer from prose.

How to fix: Declare at least Organization and WebSite; add Article, Product or FAQPage per page type.

JSON-LD cannot be parsed STRUCT-002 · high

Syntactically invalid JSON-LD is discarded by crawlers — it adds weight without adding meaning.

How to fix: Fix the JSON syntax (commonly a trailing comma or an unescaped quote) and re-validate.

No Organization / WebSite declaration STRUCT-003 · medium

These two types tell systems which entity a site belongs to, which underpins being recognised and cited correctly in AI answers.

How to fix: Add Organization (name, logo, profiles) and WebSite markup to the homepage.

<html> has no lang attribute STRUCT-004 · low

The language declaration affects how content is interpreted and matched; without it, language may be misdetected.

How to fix: Declare the language on <html>, for example lang="en" or lang="zh-CN".

Other technical checks · 10% of the score

Slow response OTHER-001 · low

Crawlers reduce how often they visit slow sites, which affects indexing and freshness over time.

How to fix: Check time-to-first-byte and origin performance; add caching where possible.

Too many redirects OTHER-002 · low

Every hop spends crawl budget, and long chains can be abandoned part-way.

How to fix: Serve the homepage directly with a 200 and keep redirects to at most one.

What we deliberately don't do

Every finding in a report carries its own evidence (status codes, HTML fragments, byte counts), so you can verify it with curl without trusting us.


GeoAISEO — see what AI crawlers actually read from a page. Home · Check a site · Checks we run