# Prontezza IA di theguardian.com

> Misurato sulla home page. Valutazione complessiva: Discreto.

Punteggio 63/100, voto Discreto. Misurato 2026-09-04T22:00:40.373+00:00, motore ax-audit@4.1.0.

## Controlli

### Contenuto — 62/100

#### Content Negotiation — 0/100

- **FAIL** Homepage does not serve Markdown via content negotiation
  Got text/html (HTTP 200) for "Accept: text/markdown"
  Serve a Markdown representation of your pages when agents request "Accept: text/markdown". Agents like Claude Code and Cursor ask for it, and Markdown cuts token usage by roughly 80% against HTML. Cloudflare ("Markdown for Agents") and Vercel can enable this without code changes.
  https://axrush.com/guides/content-negotiation#not-supported
- **WARN** No <link rel="alternate" type="text/markdown"> fallback found on the homepage
  If you cannot enable content negotiation, advertise a Markdown version with <link rel="alternate" type="text/markdown" href="/index.md"> so agents can discover it.
  https://axrush.com/guides/content-negotiation#no-alternate

#### Operabilità per agenti — 95/100

- **PASS** 398/398 interactive elements have an accessible name
- **PASS** 9/9 form controls are labelled
- **WARN** 121/122 images, frames or videos have no declared dimensions
  Undeclared dimensions shift the layout as media loads. An agent working from a screenshot clicks where the button was a moment ago. Set width and height, or aspect-ratio.
  https://axrush.com/guides/agent-operability#unsized-media
- **PASS** Method note: this reads markup, not a rendered accessibility tree
  Labels attached by script and roles computed at runtime are invisible here, so treat low proportions as a prompt to check the real tree rather than as a count. Every finding is also a plain accessibility defect.

#### Structured Data — 0/100

- **FAIL** No JSON-LD structured data found
  Add a <script type="application/ld+json"> block in your HTML <head> with schema.org structured data describing your site, organization, or person.
  https://axrush.com/guides/structured-data#not-found

#### HTML Rendering — 80/100

- **PASS** Server-rendered content detected (3360 words, 20412 chars of visible text)
- **WARN** Low text-to-markup ratio (1.4%)
  Recommended minimum: 5%
  A very low text-to-markup ratio is a typical SPA-shell symptom. Inline more content directly into the HTML response.
  https://axrush.com/guides/html-rendering#low-ratio
- **PASS** Semantic landmarks present (main, section, header, footer, nav)
- **WARN** No <h1> heading found
  Add a single <h1> describing the page. Agents and search engines treat the H1 as the primary topic indicator.
  https://axrush.com/guides/html-rendering#no-h1
- **PASS** 121/122 <img> tags have alt attributes

#### SEO Basics — 100/100

- **PASS** <title> length 48 chars: "Latest news, sport and opinion from the Guardian"
- **PASS** Meta description length 128 chars
- **PASS** Canonical URL: https://www.theguardian.com
- **PASS** <html lang="en">
- **PASS** UTF-8 charset declared
- **PASS** Viewport meta present: "width=device-width,minimum-scale=1,initial-scale=1"

### Accesso — 87/100

#### TLS / HTTPS — 100/100

- **PASS** Site is served over HTTPS
- **PASS** HTTP requests redirect to HTTPS
- **PASS** HSTS max-age=63072000
- **PASS** HSTS includes subdomains
- **PASS** HSTS preload-eligible

#### Agent Access — 100/100

- **PASS** ClaudeBot blocked at the server — consistent with its robots.txt Disallow
  status 403
- **PASS** Meta-ExternalAgent blocked at the server — consistent with its robots.txt Disallow
  status 403
- **PASS** Amazonbot blocked at the server — consistent with its robots.txt Disallow
  status 403
- **PASS** Bytespider blocked at the server — consistent with its robots.txt Disallow
  status 403
- **PASS** CCBot blocked at the server — consistent with its robots.txt Disallow
  status 403
- **PASS** Claude-SearchBot blocked at the server — consistent with its robots.txt Disallow
  status 403
- **PASS** PerplexityBot blocked at the server — consistent with its robots.txt Disallow
  status 403

#### Direttive IA — 70/100

- **PASS** Homepage is indexable
- **WARN** noarchive excludes this page from Microsoft Copilot grounding
  <meta name="robots">, X-Robots-Tag header
  Microsoft documents that noarchive means a page is not included in Copilot answers and not linked from them. nocache is the lighter option: Copilot may use the URL, title and snippet but not the body.
  https://axrush.com/guides/ai-directives#noarchive

#### Igiene HTTP — 70/100

- **WARN** A missing page redirects (301) instead of returning 404
  Location: https://www.theguardian.com/ax-audit-probe-ldp82v1t
  A redirect on a nonexistent path hides the error. Return 404 or 410 so a client can tell the difference.
  https://axrush.com/guides/http-hygiene#soft-404-redirect
- **WARN** Homepage takes 2 redirects to answer
  301 → https://www.theguardian.com/
302 → https://www.theguardian.com/us
  Every hop is a round trip the agent pays for, and some clients cap redirects well below a browser. Collapse the chain: point the first URL straight at the final one.
  https://axrush.com/guides/http-hygiene#redirect-chain
- **PASS** HEAD requests are supported
- **PASS** Content-Type: text/html with charset

#### Crawl Efficiency — 95/100

- **PASS** Response compressed with gzip
  Brotli (br) typically compresses text 10–20% smaller — consider enabling it.
- **PASS** Cache validator present (ETag)
- **PASS** Conditional request returns 304 Not Modified
- **WARN** Homepage is on the large side (1.4 MB decompressed)
  Consider trimming inlined payloads to reduce crawl cost.
  https://axrush.com/guides/crawl-efficiency#large-page
- **WARN** Roughly 5,103 tokens of content in 355,187 tokens of response (99% markup)
  Estimated at four characters per token.
  An agent pays to receive the markup and then discards it. Serving Markdown on Accept negotiation is the direct fix.
  https://axrush.com/guides/crawl-efficiency#markup-overhead
- **PASS** Homepage responded in 63ms

### Individuabilità — 44/100

#### LLMs.txt — 0/100

- **FAIL** /llms.txt not found
  HTTP 404
  Create a /llms.txt file at your site root following the llmstxt.org specification. It should be a Markdown file starting with "# Your Site Name" and include a description, sections, and links.
  https://axrush.com/guides/llms-txt#not-found

#### Robots.txt — 38/100

- **PASS** /robots.txt exists
- **WARN** 8/12 core AI crawlers configured
  GPTBot — Blocking it keeps your content out of OpenAI model training. It does not affect ChatGPT search citations.
Google-Extended — Controls Gemini training and grounding in Gemini Apps and Vertex AI. It does NOT remove your site from AI Overviews or AI Mode, which follow Googlebot and the snippet directives.
OAI-SearchBot — Blocking it removes your site from ChatGPT search answers and citations.
ChatGPT-User — Fetches a page because a ChatGPT user asked for that URL. Blocking it breaks link-following in conversations.
  Add explicit User-agent entries for the missing crawlers with Allow: / for each one.
  https://axrush.com/guides/robots-txt#missing-crawlers
- **WARN** 17 AI crawler(s) explicitly blocked
  CCBot, PetalBot, FacebookBot, Bytespider, YouBot, PerplexityBot, ClaudeBot, Claude-SearchBot, Claude-User, Applebot-Extended, YandexAdditional, YandexAdditionalBot, meta-externalagent, Amazonbot, DuckAssistBot, Google-CloudVertexBot, Amzn-SearchBot
  These crawlers have "Disallow: /" rules. If you want AI agents to access your site, change to "Allow: /" for each blocked crawler.
  https://axrush.com/guides/robots-txt#explicitly-blocked
- **WARN** 6 assistant search crawler(s) blocked — your site cannot be cited by those assistants
  PetalBot — search index crawler
YouBot — Indexes pages for You.com search and its LLM answers.
PerplexityBot — Blocking it removes your site from Perplexity answers and citations. Not used for model training.
Claude-SearchBot — Blocking it removes your site from the index Claude cites when it searches the web.
DuckAssistBot — Fetches pages in real time for DuckDuckGo AI answers. No training use.
Amzn-SearchBot — Blocking it removes your site from Alexa and Rufus answers. No training use.
  Search crawlers build the index an assistant cites from; they are separate from the training crawlers. If the intent was to opt out of training only, allow these and block the training tokens instead.
  https://axrush.com/guides/robots-txt#blocked-search-crawlers
- **PASS** 10 training crawler(s) blocked — recorded as a deliberate policy choice
  CCBot, FacebookBot, Bytespider, ClaudeBot, Applebot-Extended, YandexAdditional, YandexAdditionalBot, meta-externalagent, Amazonbot, Google-CloudVertexBot
- **WARN** 3 robots.txt rule(s) target a retired or non-existent crawler token
  anthropic-ai — Never documented by Anthropic. Use ClaudeBot.
AwarioRssBot — Awario is a social-listening tool, not an AI crawler.
AwarioSmartBot — Awario is a social-listening tool, not an AI crawler.
  These rules have no effect. Remove them so the file reflects your actual policy.
  https://axrush.com/guides/robots-txt#legacy-tokens
- **PASS** Sitemap directive present
- **WARN** No Content-Signal directive found (optional)
  Declare how crawlers may use your content after access with the Content Signals Policy, e.g.: Content-Signal: search=yes, ai-train=no. Known signals: search, ai-input, ai-train, plus the optional use=immediate|reference|full. Generate yours at contentsignals.org.
  https://axrush.com/guides/robots-txt#missing-content-signals
- **PASS** 17/57 known AI crawlers have explicit rules

#### Meta Tags — 37/100

- **WARN** No AI meta tags (ai:*) found
  Add AI meta tags to your HTML <head>: <meta name="ai:summary" content="Brief description">, <meta name="ai:content_type" content="website">, <meta name="ai:author" content="Your Name">.
  https://axrush.com/guides/meta-tags#no-ai-meta
- **WARN** No rel="alternate" link to llms.txt in HTML
  Add to your <head>: <link rel="alternate" type="text/plain" href="/llms.txt" title="LLM-optimized content">
  https://axrush.com/guides/meta-tags#no-llms-alternate
- **WARN** No rel="alternate" link to the Agent Card in HTML
  Add to your <head>: <link rel="alternate" type="application/json" href="/.well-known/agent-card.json" title="Agent Card">
  https://axrush.com/guides/meta-tags#no-agent-alternate
- **WARN** No rel="me" identity links found
  Add rel="me" links to verify your identity across platforms: <link rel="me" href="https://github.com/yourname">, <link rel="me" href="https://twitter.com/yourname">.
  https://axrush.com/guides/meta-tags#no-rel-me
- **WARN** No OpenGraph meta tags found
  Add at minimum og:title, og:description, og:url, og:type, and og:image. Agents and link previews depend on these.
  https://axrush.com/guides/meta-tags#no-opengraph
- **WARN** Twitter Card required tags missing: twitter:card, twitter:title, twitter:description
  Add these meta tags: <meta name="twitter:card" content="...">, <meta name="twitter:title" content="...">, <meta name="twitter:description" content="...">.
  https://axrush.com/guides/meta-tags#twitter-required-missing

#### Sitemap — 100/100

- **PASS** Sitemap located: http://www.theguardian.com/sitemaps/news.xml
- **PASS** Content-Type is XML (text/xml)
- **PASS** 508 URL(s) declared
- **PASS** 100% of URLs have <lastmod>
- **PASS** Newest <lastmod> is recent (0 day(s) ago)

#### HTTP Headers — 85/100

- **PASS** All 7 security headers present
- **WARN** Link header present but does not reference AI discovery files
  Add AI discovery entries to your Link header: Link: </llms.txt>; rel="alternate"; type="text/plain", </.well-known/agent-card.json>; rel="alternate"; type="application/json"
  https://axrush.com/guides/http-headers#no-ai-discovery
- **WARN** No machine-readable discovery relations beyond llms.txt and the Agent Card
  describedby — llms.txt v2 uses this relation to point a page at the llms.txt that covers it.
api-catalog — RFC 9727: the catalog of APIs this publisher offers.
service-desc — RFC 8631: a machine-readable API description.
service-doc — RFC 8631: human documentation for the API.
ai-catalog — Draft: the AI catalog listing agent cards and MCP server cards.
c2pa-manifest — C2PA 2.4: content provenance for media on the page.
license — RSL and other machine-readable licensing terms.
  Advertise what you publish with Link relations so agents stop guessing paths. Add the ones that apply, for example: Link: </llms.txt>; rel="describedby", </.well-known/api-catalog>; rel="api-catalog". Informational in 3.x: this does not affect your score.
  https://axrush.com/guides/http-headers#discovery-relations

### Protocolli — 0/100

#### Agent Card (A2A) — 0/100

- **FAIL** /.well-known/agent-card.json not found
  Site is agent-facing (an API surface (navigation into an API or developer area)). Also tried the pre-0.3 path /.well-known/agent.json. IANA-registered well-known URI.
  This site offers something an agent could call, but nothing tells an agent what. Publish an A2A Agent Card at /.well-known/agent-card.json. A minimal 1.0 card needs name, description, version, capabilities, supportedInterfaces, defaultInputModes, defaultOutputModes and skills. Spec: https://a2a-protocol.org/latest/specification/
  https://axrush.com/guides/agent-card#not-found

#### OpenAPI Spec — 0/100

- **FAIL** API surface present but no machine-readable description found
  Evidence of an API: navigation into an API or developer area. Checked /.well-known/api-catalog, Link and <link> rel="service-desc", and /.well-known/openapi.json, /openapi.json, /openapi.yaml, /.well-known/openapi.yaml, /api/openapi.json, /v1/openapi.json, /swagger.json, /api-docs, /asyncapi.json, /arazzo.json
  The API exists; nothing describes it in a form an agent can read, so using it requires a human to read your documentation first. Serve an OpenAPI description at /openapi.json and advertise it with Link: </openapi.json>; rel="service-desc". For several APIs, publish an RFC 9727 catalog.
  https://axrush.com/guides/api-discovery#not-found

#### MCP (Model Context Protocol) — N/D

- **PASS** No MCP server — MCP discovery does not apply to this site
  A server card describes an MCP server so agents can find it. A site that runs none has nothing to advertise. Run with --profile mcp to audit as though it did.

#### Catalogo IA — 0/100 (informa, non assegna un punteggio)

- **WARN** No agent resource catalog found
  Checked robots.txt Agentmap: directive, Link header rel="ai-catalog", <link rel="ai-catalog">, well-known path and /.well-known/ai-catalog.json, /.well-known/ard.json. Both specifications are drafts.
  A catalog is one document listing everything an agent can call here — agent cards, MCP servers, APIs, skills — so a client stops probing four conventions to find out. Worth publishing once you have more than one of those. Informational: both ai-catalog.json and ard.json are still drafts, so this never affects your score.
  https://axrush.com/guides/ai-catalog#not-found

#### Competenze agente — N/D

- **PASS** No developer-facing surface — skills do not apply to this site
  No documentation links, llms.txt, or API description found. Skills describe procedures an agent follows; a site with no procedures to teach has nothing to publish.

#### WebMCP — 100/100 (informa, non assegna un punteggio)

- **WARN** 2 form(s) on the page, none declared as agent tools
  WebMCP is a W3C Community Group draft in a Chrome origin trial. It is not a standard and adoption is minimal, so this is a forward-looking note, not a defect.
  A declared form is called by an agent rather than driven pixel by pixel. Add toolname and tooldescription to the forms worth automating — search, filter, subscribe — and toolparamdescription to each field.
  https://axrush.com/guides/webmcp#no-annotations

#### Scoperta commercio — N/D

- **PASS** No commerce surface — agentic-commerce discovery does not apply to this site
  No Product or Offer structured data, cart links, or product price tags found.

#### Scoperta autenticazione — N/D

- **PASS** Nothing on this site requires authorization — auth discovery does not apply
  No API description, API catalog, MCP server card or commerce profile found.

### Criteri — 99/100

#### Security.txt — 100/100

- **PASS** /.well-known/security.txt exists
- **PASS** Required field "Contact" present
- **PASS** Required field "Expires" present
- **PASS** Expires date is in the future (2027-05-31)
- **PASS** 2/5 optional fields present

#### RSL License — 95/100

- **PASS** RSL license discovered via robots.txt License directive (1 reference(s))
- **WARN** RSL license document Content-Type is "application/xml"
  Expected one of: application/rsl+xml
  Configure your server to serve RSL license document as application/rsl+xml so AI agents parse it correctly.
  https://axrush.com/guides/rsl#wrong-content-type
- **PASS** Root <rsl> element with correct namespace
- **PASS** 1 <content> element(s), all with url attribute
- **PASS** License terms declared — permits[usage]: ai-train ai-input; prohibits[usage]: all

#### Politica d'uso — 100/100

- **PASS** 1 usage declaration(s) found
  RSL licence (https://theguardian.com/license.xml)
- **PASS** Consistent across 1 declaration(s): training on your content is denied
- **PASS** Consistent across 1 declaration(s): grounding an answer in your content is denied
- **PASS** Consistent across 1 declaration(s): indexing your content for search is denied
- **PASS** Enforcement note: only robots.txt access rules are documented as honored by major AI operators
  Content Signals, AIPREF, RSL and TDMRep are declarations. Their weight is legal — the EU AI Act opt-out provisions and the GPAI Code of Practice reference several of them — rather than technical. Pair any of them with robots.txt rules if the intent is to be enforced.

---

Rappresentazione Markdown di https://axrush.com/report/theguardian.com
