AX Rush

AI readiness of irs.gov

Measured on the home page. Overall grade: Fair.

irs.gov

Fair · 41 passing checks, 25 warnings, 6 failures · 20.6s

ax-audit@4.1.0

Content

61/100

Is there substance an agent can read?

Content Negotiation

  • Homepage does not serve Markdown via content negotiation

    Serve a Markdown representation of your pages when agents request "Accept: text/markdown". Agents like Claude Code and Cursor ask for it, and Markdown cuts token usage by roughly 80% against HTML. Cloudflare ("Markdown for Agents") and Vercel can enable this without code changes.

  • No <link rel="alternate" type="text/markdown"> fallback found on the homepage

    If you cannot enable content negotiation, advertise a Markdown version with <link rel="alternate" type="text/markdown" href="/index.md"> so agents can discover it.

Structured Data

  • No JSON-LD structured data found

    Add a <script type="application/ld+json"> block in your HTML <head> with schema.org structured data describing your site, organization, or person.

Agent Operability

  • 318/318 interactive elements have an accessible name

  • 6/6 form controls are labelled

  • 1 link(s) lead nowhere without JavaScript

    A link with no destination cannot be followed by a fetch-only agent and cannot be opened in a new tab by a browsing one. Give it a real href, or make it a <button> if it is an action rather than a destination.

  • 1/1 iframes have no title

    An untitled frame is an opaque region. A title tells an agent whether it is worth entering.

  • 15/16 images, frames or videos have no declared dimensions

    Undeclared dimensions shift the layout as media loads. An agent working from a screenshot clicks where the button was a moment ago. Set width and height, or aspect-ratio.

  • Method note: this reads markup, not a rendered accessibility tree

HTML Rendering

  • Server-rendered content detected (1862 words, 12120 chars of visible text)

  • Text-to-markup ratio is healthy (8.9%)

  • Semantic landmarks present (main, article, section, header, footer, nav)

  • No <h1> heading found

    Add a single <h1> describing the page. Agents and search engines treat the H1 as the primary topic indicator.

  • 15/15 <img> tags have alt attributes

  • Generator: Drupal 10 (https://www.drupal.org)

SEO Basics

  • <title> is too long (78 chars)

    Shorten the title to 20-70 characters. Many agents and search engines truncate beyond ~70.

  • Meta description length 151 chars

  • Canonical URL: https://www.irs.gov/

  • <html lang="en">

  • UTF-8 charset declared

  • Viewport meta present: "width=device-width, initial-scale=1, shrink-to-fit=no"

  • 8 hreflang alternate(s) but no x-default

    Add <link rel="alternate" hreflang="x-default" href="..."> as a fallback for unmatched locales.

Access

92/100

Can an agent retrieve it at all?

HTTP Hygiene

  • A missing page redirects (301) instead of returning 404

    A redirect on a nonexistent path hides the error. Return 404 or 410 so a client can tell the difference.

  • Homepage answers after 1 redirect

  • HEAD requests are supported

  • Content-Type: text/html with charset

Agent Access

  • Meta-ExternalAgent is allowed in robots.txt but its User-Agent is refused

    Blocking it keeps your content out of Meta AI training and indexing. Your firewall rejects this crawler token even though robots.txt permits it — the block is invisible to you but fatal for the agent. Check your firewall rules and AI-bot toggles (for example Cloudflare "Block AI Crawlers").

TLS / HTTPS

  • Site is served over HTTPS

  • HTTP requests redirect to HTTPS

  • HSTS max-age=31536000

  • HSTS does not include subdomains

    Add includeSubDomains to apply HSTS across api., docs., etc. Required for preload list submission.

  • HSTS lacks the preload directive

    Add preload (and ensure max-age >= 31536000 + includeSubDomains) and submit the domain at https://hstspreload.org for browser-built-in HTTPS enforcement.

AI Directives

  • Homepage is indexable

  • No directive restricts how AI assistants may use this page

Crawl Efficiency

  • Response compressed with Brotli (br)

  • Cache validator present (ETag)

  • Conditional request returns 304 Not Modified

  • Homepage size is reasonable (132.9 KB decompressed)

  • Roughly 3,030 tokens of content in 34,001 tokens of response (91% markup)

    An agent pays to receive the markup and then discards it. Serving Markdown on Accept negotiation is the direct fix.

  • Homepage responded in 71ms

Discovery

53/100

Can an agent find what you publish?

LLMs.txt

  • /llms.txt not found

    Create a /llms.txt file at your site root following the llmstxt.org specification. It should be a Markdown file starting with "# Your Site Name" and include a description, sections, and links.

Meta Tags

  • No AI meta tags (ai:*) found

    Add AI meta tags to your HTML <head>: <meta name="ai:summary" content="Brief description">, <meta name="ai:content_type" content="website">, <meta name="ai:author" content="Your Name">.

  • No rel="alternate" link to llms.txt in HTML

    Add to your <head>: <link rel="alternate" type="text/plain" href="/llms.txt" title="LLM-optimized content">

  • No rel="alternate" link to the Agent Card in HTML

    Add to your <head>: <link rel="alternate" type="application/json" href="/.well-known/agent-card.json" title="Agent Card">

  • No rel="me" identity links found

    Add rel="me" links to verify your identity across platforms: <link rel="me" href="https://github.com/yourname">, <link rel="me" href="https://twitter.com/yourname">.

  • OpenGraph required tags missing: og:title, og:description, og:url, og:type

    Add these meta tags: <meta property="og:title" content="...">, <meta property="og:description" content="...">, <meta property="og:url" content="...">, <meta property="og:type" content="...">.

  • Twitter Card required tags present (twitter:card, twitter:title, twitter:description)

Robots.txt

  • /robots.txt exists

  • No core AI crawlers explicitly configured

    Add User-agent entries for core AI crawlers in your robots.txt. For each crawler, add: User-agent: <name> followed by Allow: / on the next line.

  • Sitemap directive present

  • No Content-Signal directive found (optional)

    Declare how crawlers may use your content after access with the Content Signals Policy, e.g.: Content-Signal: search=yes, ai-train=no. Known signals: search, ai-input, ai-train, plus the optional use=immediate|reference|full. Generate yours at contentsignals.org.

  • 0/57 known AI crawlers have explicit rules

    Add explicit User-agent entries for more AI crawlers to maximize discoverability.

HTTP Headers

  • Only 3/7 security headers present

    Add security headers like Strict-Transport-Security, X-Content-Type-Options, X-Frame-Options, and Referrer-Policy to your server response.

  • No Link header for AI discovery (llms.txt, Agent Card)

    Add a Link response header pointing to your AI discovery files: Link: </llms.txt>; rel="alternate"; type="text/plain", </.well-known/agent-card.json>; rel="alternate"; type="application/json"

  • No machine-readable discovery relations beyond llms.txt and the Agent Card

    Advertise what you publish with Link relations so agents stop guessing paths. Add the ones that apply, for example: Link: </llms.txt>; rel="describedby", </.well-known/api-catalog>; rel="api-catalog". Informational in 3.x: this does not affect your score.

Sitemap

  • Sitemap located: https://www.irs.gov/sitemap.xml

  • Content-Type is XML (application/xml)

  • Sitemap-index references 12 child sitemap(s)

  • 3/3 sample child sitemap(s) reachable

  • Sample yielded 15000 URL(s) across 3 child(ren)

  • Newest <lastmod> is recent (1 day(s) ago)

Protocols

50/100

What can an agent call?

AI CatalogInformational

  • No agent resource catalog found

    A catalog is one document listing everything an agent can call here — agent cards, MCP servers, APIs, skills — so a client stops probing four conventions to find out. Worth publishing once you have more than one of those. Informational: both ai-catalog.json and ard.json are still drafts, so this never affects your score.

WebMCPInformational

  • 2 form(s) on the page, none declared as agent tools

    A declared form is called by an agent rather than driven pixel by pixel. Add toolname and tooldescription to the forms worth automating — search, filter, subscribe — and toolparamdescription to each field.

Agent Card (A2A)N/A

  • No agent-facing surface — an Agent Card does not apply to this site

OpenAPI SpecN/A

  • No API surface — API discovery does not apply to this site

MCP (Model Context Protocol)N/A

  • No MCP server — MCP discovery does not apply to this site

Agent SkillsN/A

  • No developer-facing surface — skills do not apply to this site

Commerce DiscoveryN/A

  • No commerce surface — agentic-commerce discovery does not apply to this site

Auth DiscoveryN/A

  • Nothing on this site requires authorization — auth discovery does not apply

Policy

18/100

What usage rights are declared?

Security.txt

  • /.well-known/security.txt not found

    Create a /.well-known/security.txt file per RFC 9116. At minimum, include Contact: and Expires: fields. See https://securitytxt.org/ for a generator.

RSL License

  • No RSL license discovery found

    Declare machine-readable licensing terms for your content with Really Simple Licensing. Add to robots.txt: License: https://your-site.com/license.xml — then publish the RSL document. See https://rslstandard.org.

Usage Policy

  • No machine-readable usage policy declared

    State your terms where they can be read without a lawyer. The lowest-effort option is a Content-Signal line in robots.txt: Content-Signal: search=yes, ai-input=yes, ai-train=no. Absence is neutral, not permission — but it also gives you nothing to point at.