Site report: 1 measured page. Overall grade: Fair.
Bring the findings to the people who can fix them. Copy a ready-to-send summary for Slack, Teams or your next planning meeting.
This is a public report. Your team can open it without an account.
Areas group the checks shown in this report. Their scores are weighted summaries, not individual pages.
Is there substance an agent can read?
Homepage does not serve Markdown via content negotiation
Serve a Markdown representation of your pages when agents request "Accept: text/markdown". Agents like Claude Code and Cursor ask for it, and Markdown cuts token usage by roughly 80% against HTML. Cloudflare ("Markdown for Agents") and Vercel can enable this without code changes.
No <link rel="alternate" type="text/markdown"> fallback found on the homepage
If you cannot enable content negotiation, advertise a Markdown version with <link rel="alternate" type="text/markdown" href="/index.md"> so agents can discover it.
No JSON-LD structured data found
Add a <script type="application/ld+json"> block in your HTML <head> with schema.org structured data describing your site, organization, or person.
Sparse server-rendered content (24 words, 172 chars)
Render at least the main page content server-side. Many AI crawlers (GPTBot, ClaudeBot, CCBot) do not execute JavaScript and will see only the static HTML.
Text-to-markup ratio is healthy (6.3% of structural markup)
Only 2 semantic landmark(s) found
Use semantic HTML tags so AI agents can understand page structure: <header>, <nav>, <main>, <article>, <section>, <footer>.
Single <h1> heading: "Canada.ca"
8/8 <img> tags have alt attributes
<title> is too short (9 chars): "Canada.ca"
Lengthen the title to 20-70 characters with a clear topic indicator.
Meta description length 158 chars
No <link rel="canonical"> found
Add <link rel="canonical" href="https://your-site.com/page"> so agents have an unambiguous URL to cite even when crawled via a redirect or query-string variant.
<html lang="en">
UTF-8 charset declared
Viewport meta present: "width=device-width,initial-scale=1"
4/4 interactive elements have an accessible name
8/8 images, frames or videos have no declared dimensions
Undeclared dimensions shift the layout as media loads. An agent working from a screenshot clicks where the button was a moment ago. Set width and height, or aspect-ratio.
Method note: this reads markup, not a rendered accessibility tree
Can an agent retrieve it at all?
Homepage is marked noindex
A noindex page is invisible to every search-grounded assistant, because they all cite from a search index. If this is deliberate, nothing else in this check matters. If it is not, remove the directive.
GPTBot is refused by Akamai while robots.txt permits it
Blocking it keeps your content out of OpenAI model training. It does not affect ChatGPT search citations. This edge verifies crawlers by IP range or signature, so the refusal may be correct anti-spoofing rather than a policy block. ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.
ClaudeBot is refused by Akamai while robots.txt permits it
Blocking it keeps your content out of Claude model training. This edge verifies crawlers by IP range or signature, so the refusal may be correct anti-spoofing rather than a policy block. ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.
Meta-ExternalAgent is refused by Akamai while robots.txt permits it
Blocking it keeps your content out of Meta AI training and indexing. This edge verifies crawlers by IP range or signature, so the refusal may be correct anti-spoofing rather than a policy block. ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.
Bytespider is refused by Akamai while robots.txt permits it
Collects training data for ByteDance models. Compliance with robots.txt is disputed. This edge verifies crawlers by IP range or signature, so the refusal may be correct anti-spoofing rather than a policy block. ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.
CCBot is refused by Akamai while robots.txt permits it
Builds the open Common Crawl corpus that many models train on. Blocking it is the single broadest training opt-out. This edge verifies crawlers by IP range or signature, so the refusal may be correct anti-spoofing rather than a policy block. ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.
OAI-SearchBot is refused by Akamai while robots.txt permits it
Blocking it removes your site from ChatGPT search answers and citations. This edge verifies crawlers by IP range or signature, so the refusal may be correct anti-spoofing rather than a policy block. ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.
Claude-SearchBot is refused by Akamai while robots.txt permits it
Blocking it removes your site from the index Claude cites when it searches the web. This edge verifies crawlers by IP range or signature, so the refusal may be correct anti-spoofing rather than a policy block. ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.
PerplexityBot is refused by Akamai while robots.txt permits it
Blocking it removes your site from Perplexity answers and citations. Not used for model training. This edge verifies crawlers by IP range or signature, so the refusal may be correct anti-spoofing rather than a policy block. ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.
ChatGPT-User is refused by Akamai while robots.txt permits it
Fetches a page because a ChatGPT user asked for that URL. Blocking it breaks link-following in conversations. This edge verifies crawlers by IP range or signature, so the refusal may be correct anti-spoofing rather than a policy block. ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.
9 crawler probe(s) could not be settled from outside
ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.
A missing page redirects (302) instead of returning 404
A redirect on a nonexistent path hides the error. Return 404 or 410 so a client can tell the difference.
Homepage answers after 1 redirect
HEAD requests are supported
Content-Type: text/html with charset
Site is served over HTTPS
HTTP requests redirect to HTTPS
HSTS max-age=31536000
HSTS does not include subdomains
Add includeSubDomains to apply HSTS across api., docs., etc. Required for preload list submission.
HSTS lacks the preload directive
Add preload (and ensure max-age >= 31536000 + includeSubDomains) and submit the domain at https://hstspreload.org for browser-built-in HTTPS enforcement.
Response compressed with gzip
Cache validator present (Last-Modified)
Conditional request returns 304 Not Modified
Homepage size is reasonable (8.8 KB decompressed)
Roughly 43 tokens of content in 2,258 tokens of response (98% markup)
An agent pays to receive the markup and then discards it. Serving Markdown on Accept negotiation is the direct fix.
Homepage took 3148ms to respond
Agents crawl on tighter timeouts than browsers and rarely retry. A slow page is not a slow page to them, it is a missing one.
Can an agent find what you publish?
/robots.txt exists
No core AI crawlers explicitly configured
Add User-agent entries for core AI crawlers in your robots.txt. For each crawler, add: User-agent: <name> followed by Allow: / on the next line.
No Sitemap directive found
Add a Sitemap directive to your robots.txt: Sitemap: https://your-site.com/sitemap.xml
No Content-Signal directive found (optional)
Declare how crawlers may use your content after access with the Content Signals Policy, e.g.: Content-Signal: search=yes, ai-train=no. Known signals: search, ai-input, ai-train, plus the optional use=immediate|reference|full. Generate yours at contentsignals.org.
0/57 known AI crawlers have explicit rules
Add explicit User-agent entries for more AI crawlers to maximize discoverability.
/llms.txt exists
/llms.txt Content-Type is "text/html"
Configure your server to serve /llms.txt as text/plain so AI agents parse it correctly.
Missing H1 heading (first line should start with "# ")
Add an H1 heading as the first line of your llms.txt file, e.g.: # Your Site Name
No blockquote description found ("> ...")
Add a blockquote description after the H1 heading, e.g.: > A brief summary of your site for AI agents.
No section headings found (## ...)
Organize your llms.txt content with ## section headings (e.g., ## About, ## API, ## Documentation).
No Markdown links found
Add Markdown links to relevant pages: [Page Title](https://example.com/page). This helps AI agents navigate your site.
/llms-full.txt also available (bonus)
No rel="describedby" link to llms.txt
llms.txt v2 uses this relation so a page can name the file that covers it. Add <link rel="describedby" href="/llms.txt"> or the equivalent Link header, so an agent that landed on a deep page does not have to guess that an index exists.
No per-page Markdown mirror found
llms.txt v2 documents appending .md to a URL for its Markdown version. The index tells an agent which pages exist; the mirrors are what make reading them cheap.
Consumer note: llms.txt is read by coding agents, not by search
Only 3/7 security headers present
Add security headers like Strict-Transport-Security, X-Content-Type-Options, X-Frame-Options, and Referrer-Policy to your server response.
No Link header for AI discovery (llms.txt, Agent Card)
Add a Link response header pointing to your AI discovery files: Link: </llms.txt>; rel="alternate"; type="text/plain", </.well-known/agent-card.json>; rel="alternate"; type="application/json"
CORS enabled on .well-known resources
No machine-readable discovery relations beyond llms.txt and the Agent Card
Advertise what you publish with Link relations so agents stop guessing paths. Add the ones that apply, for example: Link: </llms.txt>; rel="describedby", </.well-known/api-catalog>; rel="api-catalog". Informational in 3.x: this does not affect your score.
Sitemap located: https://canada.ca/sitemap.xml
Content-Type is XML (application/xml)
Sitemap-index references 294 child sitemap(s)
3/3 sample child sitemap(s) reachable
Sample yielded 138 URL(s) across 3 child(ren)
Newest <lastmod> is recent (2 day(s) ago)
What can an agent call?
No agent resource catalog found
A catalog is one document listing everything an agent can call here — agent cards, MCP servers, APIs, skills — so a client stops probing four conventions to find out. Worth publishing once you have more than one of those. Informational: both ai-catalog.json and ard.json are still drafts, so this never affects your score.
No agent-facing surface — an Agent Card does not apply to this site
/.well-known/api-catalog served as "text/html"
Serve the catalog as application/linkset+json.
No API surface — API discovery does not apply to this site
No MCP server — MCP discovery does not apply to this site
No developer-facing surface — skills do not apply to this site
No commerce surface — agentic-commerce discovery does not apply to this site
Nothing on this site requires authorization — auth discovery does not apply
No forms and no WebMCP code — nothing here for an agent to invoke as a tool
What usage rights are declared?
No RSL license discovery found
Declare machine-readable licensing terms for your content with Really Simple Licensing. Add to robots.txt: License: https://your-site.com/license.xml — then publish the RSL document. See https://rslstandard.org.
No machine-readable usage policy declared
State your terms where they can be read without a lawyer. The lowest-effort option is a Content-Signal line in robots.txt: Content-Signal: search=yes, ai-input=yes, ai-train=no. Absence is neutral, not permission — but it also gives you nothing to point at.
/.well-known/security.txt exists
Required field "Contact" missing (RFC 9116)
Add "Contact:" to your security.txt. Use a mailto: or https: URI, e.g., Contact: mailto:security@example.com
Required field "Expires" missing (RFC 9116)
Add "Expires:" to your security.txt. Use an ISO 8601 date, e.g., Expires: 2026-12-31T23:59:59.000Z
No optional fields (Canonical, Preferred-Languages, Policy, etc.)
Consider adding Canonical: (canonical URL), Preferred-Languages: (e.g., en), and Policy: (link to your security policy).
Give your team full reports, prioritized fixes and monitoring, so improvements last beyond the next release.