Measured on the home page. Overall grade: Fair.
Is there substance an agent can read?
Homepage does not serve Markdown via content negotiation
Serve a Markdown representation of your pages when agents request "Accept: text/markdown". Agents like Claude Code and Cursor ask for it, and Markdown cuts token usage by roughly 80% against HTML. Cloudflare ("Markdown for Agents") and Vercel can enable this without code changes.
No <link rel="alternate" type="text/markdown"> fallback found on the homepage
If you cannot enable content negotiation, advertise a Markdown version with <link rel="alternate" type="text/markdown" href="/index.md"> so agents can discover it.
No JSON-LD structured data found
Add a <script type="application/ld+json"> block in your HTML <head> with schema.org structured data describing your site, organization, or person.
318/318 interactive elements have an accessible name
6/6 form controls are labelled
1 link(s) lead nowhere without JavaScript
A link with no destination cannot be followed by a fetch-only agent and cannot be opened in a new tab by a browsing one. Give it a real href, or make it a <button> if it is an action rather than a destination.
1/1 iframes have no title
An untitled frame is an opaque region. A title tells an agent whether it is worth entering.
15/16 images, frames or videos have no declared dimensions
Undeclared dimensions shift the layout as media loads. An agent working from a screenshot clicks where the button was a moment ago. Set width and height, or aspect-ratio.
Method note: this reads markup, not a rendered accessibility tree
Server-rendered content detected (1862 words, 12120 chars of visible text)
Text-to-markup ratio is healthy (8.9%)
Semantic landmarks present (main, article, section, header, footer, nav)
No <h1> heading found
Add a single <h1> describing the page. Agents and search engines treat the H1 as the primary topic indicator.
15/15 <img> tags have alt attributes
Generator: Drupal 10 (https://www.drupal.org)
<title> is too long (78 chars)
Shorten the title to 20-70 characters. Many agents and search engines truncate beyond ~70.
Meta description length 151 chars
Canonical URL: https://www.irs.gov/
<html lang="en">
UTF-8 charset declared
Viewport meta present: "width=device-width, initial-scale=1, shrink-to-fit=no"
8 hreflang alternate(s) but no x-default
Add <link rel="alternate" hreflang="x-default" href="..."> as a fallback for unmatched locales.
Can an agent retrieve it at all?
A missing page redirects (301) instead of returning 404
A redirect on a nonexistent path hides the error. Return 404 or 410 so a client can tell the difference.
Homepage answers after 1 redirect
HEAD requests are supported
Content-Type: text/html with charset
Meta-ExternalAgent is allowed in robots.txt but its User-Agent is refused
Blocking it keeps your content out of Meta AI training and indexing. Your firewall rejects this crawler token even though robots.txt permits it — the block is invisible to you but fatal for the agent. Check your firewall rules and AI-bot toggles (for example Cloudflare "Block AI Crawlers").
Site is served over HTTPS
HTTP requests redirect to HTTPS
HSTS max-age=31536000
HSTS does not include subdomains
Add includeSubDomains to apply HSTS across api., docs., etc. Required for preload list submission.
HSTS lacks the preload directive
Add preload (and ensure max-age >= 31536000 + includeSubDomains) and submit the domain at https://hstspreload.org for browser-built-in HTTPS enforcement.
Homepage is indexable
No directive restricts how AI assistants may use this page
Response compressed with Brotli (br)
Cache validator present (ETag)
Conditional request returns 304 Not Modified
Homepage size is reasonable (132.9 KB decompressed)
Roughly 3,030 tokens of content in 34,001 tokens of response (91% markup)
An agent pays to receive the markup and then discards it. Serving Markdown on Accept negotiation is the direct fix.
Homepage responded in 71ms
Can an agent find what you publish?
/llms.txt not found
Create a /llms.txt file at your site root following the llmstxt.org specification. It should be a Markdown file starting with "# Your Site Name" and include a description, sections, and links.
No AI meta tags (ai:*) found
Add AI meta tags to your HTML <head>: <meta name="ai:summary" content="Brief description">, <meta name="ai:content_type" content="website">, <meta name="ai:author" content="Your Name">.
No rel="alternate" link to llms.txt in HTML
Add to your <head>: <link rel="alternate" type="text/plain" href="/llms.txt" title="LLM-optimized content">
No rel="alternate" link to the Agent Card in HTML
Add to your <head>: <link rel="alternate" type="application/json" href="/.well-known/agent-card.json" title="Agent Card">
No rel="me" identity links found
Add rel="me" links to verify your identity across platforms: <link rel="me" href="https://github.com/yourname">, <link rel="me" href="https://twitter.com/yourname">.
OpenGraph required tags missing: og:title, og:description, og:url, og:type
Add these meta tags: <meta property="og:title" content="...">, <meta property="og:description" content="...">, <meta property="og:url" content="...">, <meta property="og:type" content="...">.
Twitter Card required tags present (twitter:card, twitter:title, twitter:description)
/robots.txt exists
No core AI crawlers explicitly configured
Add User-agent entries for core AI crawlers in your robots.txt. For each crawler, add: User-agent: <name> followed by Allow: / on the next line.
Sitemap directive present
No Content-Signal directive found (optional)
Declare how crawlers may use your content after access with the Content Signals Policy, e.g.: Content-Signal: search=yes, ai-train=no. Known signals: search, ai-input, ai-train, plus the optional use=immediate|reference|full. Generate yours at contentsignals.org.
0/57 known AI crawlers have explicit rules
Add explicit User-agent entries for more AI crawlers to maximize discoverability.
Only 3/7 security headers present
Add security headers like Strict-Transport-Security, X-Content-Type-Options, X-Frame-Options, and Referrer-Policy to your server response.
No Link header for AI discovery (llms.txt, Agent Card)
Add a Link response header pointing to your AI discovery files: Link: </llms.txt>; rel="alternate"; type="text/plain", </.well-known/agent-card.json>; rel="alternate"; type="application/json"
No machine-readable discovery relations beyond llms.txt and the Agent Card
Advertise what you publish with Link relations so agents stop guessing paths. Add the ones that apply, for example: Link: </llms.txt>; rel="describedby", </.well-known/api-catalog>; rel="api-catalog". Informational in 3.x: this does not affect your score.
Sitemap located: https://www.irs.gov/sitemap.xml
Content-Type is XML (application/xml)
Sitemap-index references 12 child sitemap(s)
3/3 sample child sitemap(s) reachable
Sample yielded 15000 URL(s) across 3 child(ren)
Newest <lastmod> is recent (1 day(s) ago)
What can an agent call?
No agent resource catalog found
A catalog is one document listing everything an agent can call here — agent cards, MCP servers, APIs, skills — so a client stops probing four conventions to find out. Worth publishing once you have more than one of those. Informational: both ai-catalog.json and ard.json are still drafts, so this never affects your score.
2 form(s) on the page, none declared as agent tools
A declared form is called by an agent rather than driven pixel by pixel. Add toolname and tooldescription to the forms worth automating — search, filter, subscribe — and toolparamdescription to each field.
No agent-facing surface — an Agent Card does not apply to this site
No API surface — API discovery does not apply to this site
No MCP server — MCP discovery does not apply to this site
No developer-facing surface — skills do not apply to this site
No commerce surface — agentic-commerce discovery does not apply to this site
Nothing on this site requires authorization — auth discovery does not apply
What usage rights are declared?
/.well-known/security.txt not found
Create a /.well-known/security.txt file per RFC 9116. At minimum, include Contact: and Expires: fields. See https://securitytxt.org/ for a generator.
No RSL license discovery found
Declare machine-readable licensing terms for your content with Really Simple Licensing. Add to robots.txt: License: https://your-site.com/license.xml — then publish the RSL document. See https://rslstandard.org.
No machine-readable usage policy declared
State your terms where they can be read without a lawyer. The lowest-effort option is a Content-Signal line in robots.txt: Content-Signal: search=yes, ai-input=yes, ai-train=no. Absence is neutral, not permission — but it also gives you nothing to point at.