# Prontezza IA di apnews.com

> Misurato sulla home page. Valutazione complessiva: Discreto.

Punteggio 58/100, voto Discreto. Misurato 2026-09-04T22:10:36.606+00:00, motore ax-audit@4.1.0.

## Controlli

### Contenuto — 66/100

#### Content Negotiation — 0/100

- **FAIL** Homepage does not serve Markdown via content negotiation
  Got text/html (HTTP 200) for "Accept: text/markdown"
  Serve a Markdown representation of your pages when agents request "Accept: text/markdown". Agents like Claude Code and Cursor ask for it, and Markdown cuts token usage by roughly 80% against HTML. Cloudflare ("Markdown for Agents") and Vercel can enable this without code changes.
  https://axrush.com/guides/content-negotiation#not-supported
- **WARN** No <link rel="alternate" type="text/markdown"> fallback found on the homepage
  If you cannot enable content negotiation, advertise a Markdown version with <link rel="alternate" type="text/markdown" href="/index.md"> so agents can discover it.
  https://axrush.com/guides/content-negotiation#no-alternate

#### Operabilità per agenti — 50/100

- **PASS** 1792/1803 interactive elements have an accessible name
- **WARN** 2/4 form controls have no label
  <input type="text">, <input>
  An unlabelled box is one the agent has to guess the meaning of, which is how an email address ends up in a search field. Use <label for>, wrap the control in a label, or add aria-label.
  https://axrush.com/guides/agent-operability#unlabelled-controls
- **WARN** 59 link(s) lead nowhere without JavaScript
  59 with no href, 0 with a javascript: href
  A link with no destination cannot be followed by a fetch-only agent and cannot be opened in a new tab by a browsing one. Give it a real href, or make it a <button> if it is an action rather than a destination.
  https://axrush.com/guides/agent-operability#dead-links
- **WARN** 1/2 iframes have no title
  An untitled frame is an opaque region. A title tells an agent whether it is worth entering.
  https://axrush.com/guides/agent-operability#iframe-no-title
- **WARN** 1 obstacle(s) on the entry page
  a CAPTCHA widget on the entry page
  A CAPTCHA or a meta refresh on the landing page stops an agent before it reaches any content. If bot protection is needed, apply it to the actions that need it, not to reading a page.
  https://axrush.com/guides/agent-operability#entry-blockers
- **PASS** Method note: this reads markup, not a rendered accessibility tree
  Labels attached by script and roles computed at runtime are invisible here, so treat low proportions as a prompt to check the real tree rather than as a count. Every finding is also a plain accessibility defect.

#### Structured Data — 90/100

- **PASS** 2 JSON-LD block(s) found
- **PASS** @context references schema.org
- **WARN** No @graph array (single-entity only)
  Use an @graph array to define multiple entities in one JSON-LD block: { "@context": "https://schema.org", "@graph": [...] }
  https://axrush.com/guides/structured-data#no-graph
- **PASS** Key types found: Organization, WebSite
- **WARN** No BreadcrumbList found
  Add a BreadcrumbList entity to help AI agents understand your site navigation hierarchy.
  https://axrush.com/guides/structured-data#no-breadcrumb
- **WARN** No author declared in structured data
  An assistant deciding whether to cite a page weighs where it came from. Add author as a Person or Organization, not a bare string.
  https://axrush.com/guides/structured-data#no-author
- **PASS** 6 sameAs link(s) tie your entity to external identifiers
  https://www.facebook.com/APNews/, https://twitter.com/AP, https://www.instagram.com/apnews/, https://www.youtube.com/@AssociatedPress
- **WARN** No dateModified or datePublished in structured data
  Assistants weigh recency when they answer time-sensitive questions, and a page with no date cannot be weighed at all. Add dateModified to anything that changes.
  https://axrush.com/guides/structured-data#no-dates
- **PASS** Structured-data headings appear in the visible text

#### HTML Rendering — 75/100

- **PASS** Server-rendered content detected (9729 words, 62109 chars of visible text)
- **WARN** Low text-to-markup ratio (2.5%)
  Recommended minimum: 5%
  A very low text-to-markup ratio is a typical SPA-shell symptom. Inline more content directly into the HTML response.
  https://axrush.com/guides/html-rendering#low-ratio
- **PASS** Semantic landmarks present (main, footer, nav)
- **WARN** No <h1> heading found
  Add a single <h1> describing the page. Agents and search engines treat the H1 as the primary topic indicator.
  https://axrush.com/guides/html-rendering#no-h1
- **WARN** Only 49/92 <img> tags have alt attributes
  Add descriptive alt="" to every <img>. Agents use alt text to understand images they cannot process visually.
  https://axrush.com/guides/html-rendering#missing-alt

#### SEO Basics — 95/100

- **WARN** <title> is too long (75 chars)
  "Associated Press News: Breaking News, Latest Headlines and Videos | AP News…"
  Shorten the title to 20-70 characters. Many agents and search engines truncate beyond ~70.
  https://axrush.com/guides/seo-basics#long-title
- **PASS** Meta description length 148 chars
- **PASS** Canonical URL: https://apnews.com/
- **PASS** <html lang="en">
- **PASS** UTF-8 charset declared
- **PASS** Viewport meta present: "width=device-width, initial-scale=1.0, minimum-scale=1.0 user-scalable=yes"
- **PASS** 2 hreflang alternate(s) including x-default

### Accesso — 83/100

#### TLS / HTTPS — 97/100

- **PASS** Site is served over HTTPS
- **PASS** HTTP requests redirect to HTTPS
- **PASS** HSTS max-age=31536000
- **PASS** HSTS includes subdomains
- **WARN** HSTS lacks the preload directive
  Add preload (and ensure max-age >= 31536000 + includeSubDomains) and submit the domain at https://hstspreload.org for browser-built-in HTTPS enforcement.
  https://axrush.com/guides/tls-https#hsts-no-preload

#### Agent Access — 78/100

- **WARN** GPTBot receives a Cloudflare challenge instead of the page
  body matched Cloudflare challenge markup; status 200
  Challenge pages require running JavaScript. Crawlers that only fetch HTML — which is most of them — never get past one, so the page is effectively unavailable to them even though nothing is "blocked". Add a bot-management exception for verified AI crawlers. ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.
  https://axrush.com/guides/agent-access#challenge-page
- **PASS** ClaudeBot blocked at the server — consistent with its robots.txt Disallow
  status 402 with no crawler-price, x402 fields, or licence challenge
- **WARN** Meta-ExternalAgent receives a Cloudflare challenge instead of the page
  body matched Cloudflare challenge markup; status 200
  Challenge pages require running JavaScript. Crawlers that only fetch HTML — which is most of them — never get past one, so the page is effectively unavailable to them even though nothing is "blocked". Add a bot-management exception for verified AI crawlers. ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.
  https://axrush.com/guides/agent-access#challenge-page
- **WARN** Amazonbot receives a Cloudflare challenge instead of the page
  body matched Cloudflare challenge markup; status 200
  Challenge pages require running JavaScript. Crawlers that only fetch HTML — which is most of them — never get past one, so the page is effectively unavailable to them even though nothing is "blocked". Add a bot-management exception for verified AI crawlers. ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.
  https://axrush.com/guides/agent-access#challenge-page
- **WARN** Bytespider is allowed in robots.txt but its User-Agent is refused
  status 402 with no crawler-price, x402 fields, or licence challenge
  Collects training data for ByteDance models. Compliance with robots.txt is disputed. Your firewall rejects this crawler token even though robots.txt permits it — the block is invisible to you but fatal for the agent. Check your firewall rules and AI-bot toggles (for example Cloudflare "Block AI Crawlers").
  https://axrush.com/guides/agent-access#blocked-crawler
- **PASS** CCBot blocked at the server — consistent with its robots.txt Disallow
  status 402 with no crawler-price, x402 fields, or licence challenge
- **WARN** OAI-SearchBot receives a Cloudflare challenge instead of the page
  body matched Cloudflare challenge markup; status 200
  Challenge pages require running JavaScript. Crawlers that only fetch HTML — which is most of them — never get past one, so the page is effectively unavailable to them even though nothing is "blocked". Add a bot-management exception for verified AI crawlers. ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.
  https://axrush.com/guides/agent-access#challenge-page
- **PASS** Claude-SearchBot blocked at the server — consistent with its robots.txt Disallow
  status 402 with no crawler-price, x402 fields, or licence challenge
- **PASS** PerplexityBot blocked at the server — consistent with its robots.txt Disallow
  status 402 with no crawler-price, x402 fields, or licence challenge
- **WARN** ChatGPT-User receives a Cloudflare challenge instead of the page
  body matched Cloudflare challenge markup; status 200
  Challenge pages require running JavaScript. Crawlers that only fetch HTML — which is most of them — never get past one, so the page is effectively unavailable to them even though nothing is "blocked". Add a bot-management exception for verified AI crawlers. ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.
  https://axrush.com/guides/agent-access#challenge-page
- **WARN** 5 crawler probe(s) could not be settled from outside
  GPTBot, Meta-ExternalAgent, Amazonbot, OAI-SearchBot, ChatGPT-User
  ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.
  https://axrush.com/guides/agent-access#inconclusive-probe

#### Direttive IA — 70/100

- **PASS** Homepage is indexable
- **WARN** noarchive excludes this page from Microsoft Copilot grounding
  <meta name="robots">
  Microsoft documents that noarchive means a page is not included in Copilot answers and not linked from them. nocache is the lighter option: Copilot may use the URL, title and snippet but not the body.
  https://axrush.com/guides/ai-directives#noarchive

#### Igiene HTTP — 100/100

- **WARN** Could not test 404 handling — the probe was challenged
  body matched Cloudflare challenge markup; status 404
  Bot management answered a plain GET with an interstitial, so status-code honesty could not be verified from outside.
  https://axrush.com/guides/http-hygiene#probe-challenged
- **PASS** Homepage answers without a redirect
- **PASS** HEAD requests are supported
- **PASS** Content-Type: text/html with charset

#### Crawl Efficiency — 90/100

- **PASS** Response compressed with Brotli (br)
- **PASS** Cache validator present (Last-Modified)
- **PASS** Conditional request returns 304 Not Modified
- **WARN** Homepage is very large (2.4 MB decompressed)
  Large documents inflate crawl cost and token usage. Trim inlined data, split content, or serve a Markdown representation to agents (see the content-negotiation check).
  https://axrush.com/guides/crawl-efficiency#large-page
- **WARN** Roughly 15,527 tokens of content in 624,737 tokens of response (98% markup)
  Estimated at four characters per token.
  An agent pays to receive the markup and then discards it. Serving Markdown on Accept negotiation is the direct fix.
  https://axrush.com/guides/crawl-efficiency#markup-overhead
- **PASS** Homepage responded in 190ms

### Individuabilità — 52/100

#### LLMs.txt — 0/100

- **FAIL** /llms.txt not found
  HTTP 404
  Create a /llms.txt file at your site root following the llmstxt.org specification. It should be a Markdown file starting with "# Your Site Name" and include a description, sections, and links.
  https://axrush.com/guides/llms-txt#not-found

#### Robots.txt — 63/100

- **PASS** /robots.txt exists
- **WARN** 7/12 core AI crawlers configured
  Meta-ExternalAgent — Blocking it keeps your content out of Meta AI training and indexing.
Google-Extended — Controls Gemini training and grounding in Gemini Apps and Vertex AI. It does NOT remove your site from AI Overviews or AI Mode, which follow Googlebot and the snippet directives.
Bytespider — Collects training data for ByteDance models. Compliance with robots.txt is disputed.
OAI-SearchBot — Blocking it removes your site from ChatGPT search answers and citations.
ChatGPT-User — Fetches a page because a ChatGPT user asked for that URL. Blocking it breaks link-following in conversations.
  Add explicit User-agent entries for the missing crawlers with Allow: / for each one.
  https://axrush.com/guides/robots-txt#missing-crawlers
- **WARN** 9 AI crawler(s) explicitly blocked
  CCBot, GPTBot, ClaudeBot, PerplexityBot, Amazonbot, Applebot-Extended, Timpibot, Claude-User, Claude-SearchBot
  These crawlers have "Disallow: /" rules. If you want AI agents to access your site, change to "Allow: /" for each blocked crawler.
  https://axrush.com/guides/robots-txt#explicitly-blocked
- **WARN** 2 assistant search crawler(s) blocked — your site cannot be cited by those assistants
  PerplexityBot — Blocking it removes your site from Perplexity answers and citations. Not used for model training.
Claude-SearchBot — Blocking it removes your site from the index Claude cites when it searches the web.
  Search crawlers build the index an assistant cites from; they are separate from the training crawlers. If the intent was to opt out of training only, allow these and block the training tokens instead.
  https://axrush.com/guides/robots-txt#blocked-search-crawlers
- **PASS** 6 training crawler(s) blocked — recorded as a deliberate policy choice
  CCBot, GPTBot, ClaudeBot, Amazonbot, Applebot-Extended, Timpibot
- **WARN** 3 robots.txt rule(s) target a retired or non-existent crawler token
  anthropic-ai — Never documented by Anthropic. Use ClaudeBot.
Claude-Web — Never documented by Anthropic and absent from its current crawler page. Use ClaudeBot.
cohere-ai — Cohere states it operates no web crawlers.
  These rules have no effect. Remove them so the file reflects your actual policy.
  https://axrush.com/guides/robots-txt#legacy-tokens
- **PASS** Sitemap directive present
- **WARN** No Content-Signal directive found (optional)
  Declare how crawlers may use your content after access with the Content Signals Policy, e.g.: Content-Signal: search=yes, ai-train=no. Known signals: search, ai-input, ai-train, plus the optional use=immediate|reference|full. Generate yours at contentsignals.org.
  https://axrush.com/guides/robots-txt#missing-content-signals
- **WARN** 9/57 known AI crawlers have explicit rules
  Add explicit User-agent entries for more AI crawlers to maximize discoverability.
  https://axrush.com/guides/robots-txt#low-coverage

#### Meta Tags — 52/100

- **WARN** No AI meta tags (ai:*) found
  Add AI meta tags to your HTML <head>: <meta name="ai:summary" content="Brief description">, <meta name="ai:content_type" content="website">, <meta name="ai:author" content="Your Name">.
  https://axrush.com/guides/meta-tags#no-ai-meta
- **WARN** No rel="alternate" link to llms.txt in HTML
  Add to your <head>: <link rel="alternate" type="text/plain" href="/llms.txt" title="LLM-optimized content">
  https://axrush.com/guides/meta-tags#no-llms-alternate
- **WARN** No rel="alternate" link to the Agent Card in HTML
  Add to your <head>: <link rel="alternate" type="application/json" href="/.well-known/agent-card.json" title="Agent Card">
  https://axrush.com/guides/meta-tags#no-agent-alternate
- **WARN** No rel="me" identity links found
  Add rel="me" links to verify your identity across platforms: <link rel="me" href="https://github.com/yourname">, <link rel="me" href="https://twitter.com/yourname">.
  https://axrush.com/guides/meta-tags#no-rel-me
- **PASS** OpenGraph required tags present (og:title, og:description, og:url, og:type)
- **PASS** Twitter Card required tags present (twitter:card, twitter:title, twitter:description)
- **WARN** Twitter Card recommended tags missing: twitter:image
  Add twitter:image for richer previews.
  https://axrush.com/guides/meta-tags#twitter-recommended-missing

#### Sitemap — 95/100

- **PASS** Sitemap located: https://apnews.com/ap-sitemap.xml
- **PASS** Content-Type is XML (text/xml)
- **PASS** Sitemap-index references 230 child sitemap(s)
- **PASS** 3/3 sample child sitemap(s) reachable
- **PASS** Sample yielded 3 URL(s) across 3 child(ren)
- **WARN** Newest <lastmod> is 569 days old
  Refresh <lastmod> on URLs that have changed. A stale sitemap signals to crawlers that nothing on the site has updated.
  https://axrush.com/guides/sitemap#stale

#### HTTP Headers — 70/100

- **FAIL** Missing critical header: X-Content-Type-Options
  Add the X-Content-Type-Options response header to your server configuration. This is a critical security header.
  https://axrush.com/guides/http-headers#missing-critical-header
- **WARN** Only 2/7 security headers present
  Add security headers like Strict-Transport-Security, X-Content-Type-Options, X-Frame-Options, and Referrer-Policy to your server response.
  https://axrush.com/guides/http-headers#low-security-headers
- **WARN** No Link header for AI discovery (llms.txt, Agent Card)
  Add a Link response header pointing to your AI discovery files: Link: </llms.txt>; rel="alternate"; type="text/plain", </.well-known/agent-card.json>; rel="alternate"; type="application/json"
  https://axrush.com/guides/http-headers#no-link-header
- **WARN** No machine-readable discovery relations beyond llms.txt and the Agent Card
  describedby — llms.txt v2 uses this relation to point a page at the llms.txt that covers it.
api-catalog — RFC 9727: the catalog of APIs this publisher offers.
service-desc — RFC 8631: a machine-readable API description.
service-doc — RFC 8631: human documentation for the API.
ai-catalog — Draft: the AI catalog listing agent cards and MCP server cards.
c2pa-manifest — C2PA 2.4: content provenance for media on the page.
license — RSL and other machine-readable licensing terms.
  Advertise what you publish with Link relations so agents stop guessing paths. Add the ones that apply, for example: Link: </llms.txt>; rel="describedby", </.well-known/api-catalog>; rel="api-catalog". Informational in 3.x: this does not affect your score.
  https://axrush.com/guides/http-headers#discovery-relations

### Protocolli — 0/100

#### Agent Card (A2A) — 0/100

- **FAIL** /.well-known/agent-card.json not found
  Site is agent-facing (an API surface (navigation into an API or developer area)). Also tried the pre-0.3 path /.well-known/agent.json. IANA-registered well-known URI.
  This site offers something an agent could call, but nothing tells an agent what. Publish an A2A Agent Card at /.well-known/agent-card.json. A minimal 1.0 card needs name, description, version, capabilities, supportedInterfaces, defaultInputModes, defaultOutputModes and skills. Spec: https://a2a-protocol.org/latest/specification/
  https://axrush.com/guides/agent-card#not-found

#### OpenAPI Spec — 0/100

- **FAIL** API surface present but no machine-readable description found
  Evidence of an API: navigation into an API or developer area. Checked /.well-known/api-catalog, Link and <link> rel="service-desc", and /.well-known/openapi.json, /openapi.json, /openapi.yaml, /.well-known/openapi.yaml, /api/openapi.json, /v1/openapi.json, /swagger.json, /api-docs, /asyncapi.json, /arazzo.json
  The API exists; nothing describes it in a form an agent can read, so using it requires a human to read your documentation first. Serve an OpenAPI description at /openapi.json and advertise it with Link: </openapi.json>; rel="service-desc". For several APIs, publish an RFC 9727 catalog.
  https://axrush.com/guides/api-discovery#not-found

#### MCP (Model Context Protocol) — N/D

- **PASS** No MCP server — MCP discovery does not apply to this site
  A server card describes an MCP server so agents can find it. A site that runs none has nothing to advertise. Run with --profile mcp to audit as though it did.

#### Catalogo IA — 0/100 (informa, non assegna un punteggio)

- **WARN** No agent resource catalog found
  Checked robots.txt Agentmap: directive, Link header rel="ai-catalog", <link rel="ai-catalog">, well-known path and /.well-known/ai-catalog.json, /.well-known/ard.json. Both specifications are drafts.
  A catalog is one document listing everything an agent can call here — agent cards, MCP servers, APIs, skills — so a client stops probing four conventions to find out. Worth publishing once you have more than one of those. Informational: both ai-catalog.json and ard.json are still drafts, so this never affects your score.
  https://axrush.com/guides/ai-catalog#not-found

#### Competenze agente — N/D

- **PASS** No developer-facing surface — skills do not apply to this site
  No documentation links, llms.txt, or API description found. Skills describe procedures an agent follows; a site with no procedures to teach has nothing to publish.

#### WebMCP — 100/100 (informa, non assegna un punteggio)

- **WARN** 4 form(s) on the page, none declared as agent tools
  WebMCP is a W3C Community Group draft in a Chrome origin trial. It is not a standard and adoption is minimal, so this is a forward-looking note, not a defect.
  A declared form is called by an agent rather than driven pixel by pixel. Add toolname and tooldescription to the forms worth automating — search, filter, subscribe — and toolparamdescription to each field.
  https://axrush.com/guides/webmcp#no-annotations

#### Scoperta commercio — 0/100 (informa, non assegna un punteggio)

- **WARN** Site sells something but publishes no agent-readable commerce profile
  Commerce signals found: Product or storefront structured data. Checked /.well-known/ucp, /.well-known/ucp.json.
  Publish a Universal Commerce Protocol profile at /.well-known/ucp so an agent can find your catalog, cart and checkout without a human. Note that the alternatives are not discoverable by design: the OpenAI and Stripe Agentic Commerce Protocol defines no manifest, and AP2 advertises itself through an A2A card extension.
  https://axrush.com/guides/commerce-discovery#not-found

#### Scoperta autenticazione — N/D

- **PASS** Nothing on this site requires authorization — auth discovery does not apply
  No API description, API catalog, MCP server card or commerce profile found.

### Criteri — 18/100

#### Security.txt — 0/100

- **FAIL** /.well-known/security.txt not found
  HTTP 404
  Create a /.well-known/security.txt file per RFC 9116. At minimum, include Contact: and Expires: fields. See https://securitytxt.org/ for a generator.
  https://axrush.com/guides/security-txt#not-found

#### RSL License — 0/100

- **FAIL** No RSL license discovery found
  Checked robots.txt License directive, Link header, and <link rel="license" type="application/rsl+xml">
  Declare machine-readable licensing terms for your content with Really Simple Licensing. Add to robots.txt: License: https://your-site.com/license.xml — then publish the RSL document. See https://rslstandard.org.
  https://axrush.com/guides/rsl#not-found

#### Politica d'uso — 40/100

- **WARN** No machine-readable usage policy declared
  Checked robots.txt Content-Signal and Content-Usage, the Content-Usage and content-signal response headers, an RSL licence, TDMRep, and the noai meta directive.
  State your terms where they can be read without a lawyer. The lowest-effort option is a Content-Signal line in robots.txt: Content-Signal: search=yes, ai-input=yes, ai-train=no. Absence is neutral, not permission — but it also gives you nothing to point at.
  https://axrush.com/guides/usage-policy#no-policy

---

Rappresentazione Markdown di https://axrush.com/report/apnews.com
