AX Rush

Préparation de nytimes.com à l'IA

Rapport du site : 1 page mesurée. Note globale : Moyen.

Score du site · nytimes.com

Combine l'origine et 1 page mesurée. Le rapport ci-dessous détaille uniquement la page d'origine.

nytimes.com

Moyen · 56 vérifications réussies, 35 avertissements, 6 échecs · 2,6 s

Partagez ce rapport avec votre équipe technique

Transmettez les constats à ceux qui peuvent les résoudre. Copiez un résumé prêt pour Slack, Teams ou votre prochaine réunion.

E-mail

Ce rapport est public. Votre équipe peut l’ouvrir sans compte.

Contenu

76/100

Y a-t-il de la substance qu'un agent puisse lire ?

  • Homepage does not serve Markdown via content negotiation

    Serve a Markdown representation of your pages when agents request "Accept: text/markdown". Agents like Claude Code and Cursor ask for it, and Markdown cuts token usage by roughly 80% against HTML. Cloudflare ("Markdown for Agents") and Vercel can enable this without code changes.

  • No <link rel="alternate" type="text/markdown"> fallback found on the homepage

    If you cannot enable content negotiation, advertise a Markdown version with <link rel="alternate" type="text/markdown" href="/index.md"> so agents can discover it.

  • 136/138 interactive elements have an accessible name

  • 1/1 form controls are labelled

  • 25 link(s) lead nowhere without JavaScript

    A link with no destination cannot be followed by a fetch-only agent and cannot be opened in a new tab by a browsing one. Give it a real href, or make it a <button> if it is an action rather than a destination.

  • 1/1 iframes have no title

    An untitled frame is an opaque region. A title tells an agent whether it is worth entering.

  • 59/60 images, frames or videos have no declared dimensions

    Undeclared dimensions shift the layout as media loads. An agent working from a screenshot clicks where the button was a moment ago. Set width and height, or aspect-ratio.

  • Method note: this reads markup, not a rendered accessibility tree

  • Server-rendered content detected (1151 words, 6851 chars of visible text)

  • Low text-to-markup ratio (0.5%)

    A very low text-to-markup ratio is a typical SPA-shell symptom. Inline more content directly into the HTML response.

  • Semantic landmarks present (main, article, section, header, footer, nav)

  • Single <h1> heading: "New York Times - Top Stories"

  • Only 52/58 <img> tags have alt attributes

    Add descriptive alt="" to every <img>. Agents use alt text to understand images they cannot process visually.

  • 2 JSON-LD block(s) found

  • @context references schema.org

  • No @graph array (single-entity only)

    Use an @graph array to define multiple entities in one JSON-LD block: { "@context": "https://schema.org", "@graph": [...] }

  • Key types found: Organization, WebSite

  • No BreadcrumbList found

    Add a BreadcrumbList entity to help AI agents understand your site navigation hierarchy.

  • No author declared in structured data

    An assistant deciding whether to cite a page weighs where it came from. Add author as a Person or Organization, not a bare string.

  • 22 sameAs link(s) tie your entity to external identifiers

  • No dateModified or datePublished in structured data

    Assistants weigh recency when they answer time-sensitive questions, and a page with no date cannot be weighed at all. Add dateModified to anything that changes.

  • Structured-data headings appear in the visible text

  • <title> length 66 chars: "The New York Times - Breaking News, US News, World News and Videos"

  • Meta description is long (267 chars)

    Trim to 70-160 characters; longer descriptions get truncated.

  • Canonical URL: https://www.nytimes.com

  • <html lang="en">

  • UTF-8 charset declared

  • Viewport meta present: "width=device-width, initial-scale=1"

Accès

96/100

Un agent peut-il seulement la récupérer ?

  • A missing page redirects (301) instead of returning 404

    A redirect on a nonexistent path hides the error. Return 404 or 410 so a client can tell the difference.

  • Homepage answers after 1 redirect

  • HEAD requests are supported

  • Content-Type: text/html with charset

  • Response compressed with gzip

  • Cache validator present (Last-Modified)

  • Conditional request returns 304 Not Modified

  • Homepage is on the large side (1.4 MB decompressed)

    Consider trimming inlined payloads to reduce crawl cost.

  • Roughly 1,713 tokens of content in 372,449 tokens of response (100% markup)

    An agent pays to receive the markup and then discards it. Serving Markdown on Accept negotiation is the direct fix.

  • Homepage responded in 268ms

  • GPTBot blocked at the server — consistent with its robots.txt Disallow

  • ClaudeBot blocked at the server — consistent with its robots.txt Disallow

  • Meta-ExternalAgent blocked at the server — consistent with its robots.txt Disallow

  • Amazonbot is refused by Fastly while robots.txt permits it

    General crawler whose content may be used to train Amazon AI models. This edge verifies crawlers by IP range or signature, so the refusal may be correct anti-spoofing rather than a policy block. ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.

  • Bytespider blocked at the server — consistent with its robots.txt Disallow

  • CCBot blocked at the server — consistent with its robots.txt Disallow

  • OAI-SearchBot blocked at the server — consistent with its robots.txt Disallow

  • Claude-SearchBot blocked at the server — consistent with its robots.txt Disallow

  • PerplexityBot blocked at the server — consistent with its robots.txt Disallow

  • ChatGPT-User blocked at the server — consistent with its robots.txt Disallow

  • 1 crawler probe(s) could not be settled from outside

    ax-audit sends this user agent from its own network without a Web Bot Auth signature, so an edge that verifies crawlers by IP range or signature will reject the probe while admitting the real crawler. Confirm against your WAF logs before changing any rule.

  • Site is served over HTTPS

  • HTTP requests redirect to HTTPS

  • HSTS max-age=63072000

  • HSTS includes subdomains

  • HSTS preload-eligible

  • Homepage is indexable

  • Google-Extended is disallowed, but this page is still eligible for AI Overviews

    Google-Extended governs Gemini training and grounding in Gemini Apps and Vertex AI. AI Overviews and AI Mode follow Googlebot and the snippet directives instead. If the intent was to stay out of AI Overviews, use nosnippet or max-snippet. If the intent was to opt out of training, this is already correct.

Découverte

38/100

Un agent trouve-t-il ce que vous publiez ?

  • /llms.txt not found

    Create a /llms.txt file at your site root following the llmstxt.org specification. It should be a Markdown file starting with "# Your Site Name" and include a description, sections, and links.

  • /robots.txt exists

  • All 12 core AI crawlers explicitly configured

  • 28 AI crawler(s) explicitly blocked

    These crawlers have "Disallow: /" rules. If you want AI agents to access your site, change to "Allow: /" for each blocked crawler.

  • 7 assistant search crawler(s) blocked — your site cannot be cited by those assistants

    Search crawlers build the index an assistant cites from; they are separate from the training crawlers. If the intent was to opt out of training only, allow these and block the training tokens instead.

  • 15 training crawler(s) blocked — recorded as a deliberate policy choice

  • 4 user-triggered fetcher(s) blocked in robots.txt that may ignore it

    These clients fetch a page because a person asked for that URL, and their vendors document that robots.txt may not apply. Enforce at the edge if the block must hold.

  • 6 robots.txt rule(s) target a retired or non-existent crawler token

    These rules have no effect. Remove them so the file reflects your actual policy.

  • 1 AI crawler(s) have partial path restrictions

    These crawlers have Disallow rules on specific paths. For full AI access, use only "Allow: /" and let the wildcard User-agent: * handle path restrictions.

  • Sitemap directive present

  • No Content-Signal directive found (optional)

    Declare how crawlers may use your content after access with the Content Signals Policy, e.g.: Content-Signal: search=yes, ai-train=no. Known signals: search, ai-input, ai-train, plus the optional use=immediate|reference|full. Generate yours at contentsignals.org.

  • 27/57 known AI crawlers have explicit rules

  • No AI meta tags (ai:*) found

    Add AI meta tags to your HTML <head>: <meta name="ai:summary" content="Brief description">, <meta name="ai:content_type" content="website">, <meta name="ai:author" content="Your Name">.

  • No rel="alternate" link to llms.txt in HTML

    Add to your <head>: <link rel="alternate" type="text/plain" href="/llms.txt" title="LLM-optimized content">

  • No rel="alternate" link to the Agent Card in HTML

    Add to your <head>: <link rel="alternate" type="application/json" href="/.well-known/agent-card.json" title="Agent Card">

  • No rel="me" identity links found

    Add rel="me" links to verify your identity across platforms: <link rel="me" href="https://github.com/yourname">, <link rel="me" href="https://twitter.com/yourname">.

  • OpenGraph required tags present (og:title, og:description, og:url, og:type)

  • OpenGraph recommended tags missing: og:site_name

    Add these for richer previews: og:site_name.

  • Twitter Card required tags missing: twitter:card, twitter:title, twitter:description

    Add these meta tags: <meta name="twitter:card" content="...">, <meta name="twitter:title" content="...">, <meta name="twitter:description" content="...">.

  • 6/7 security headers present

  • No Link header for AI discovery (llms.txt, Agent Card)

    Add a Link response header pointing to your AI discovery files: Link: </llms.txt>; rel="alternate"; type="text/plain", </.well-known/agent-card.json>; rel="alternate"; type="application/json"

  • No machine-readable discovery relations beyond llms.txt and the Agent Card

    Advertise what you publish with Link relations so agents stop guessing paths. Add the ones that apply, for example: Link: </llms.txt>; rel="describedby", </.well-known/api-catalog>; rel="api-catalog". Informational in 3.x: this does not affect your score.

  • Sitemap located: https://www.nytimes.com/sitemaps/new/news.xml.gz

  • Content-Type is XML (text/xml)

  • 728 URL(s) declared

  • 100% of URLs have <lastmod>

  • Newest <lastmod> is recent (0 day(s) ago)

Protocoles

0/100

Qu'un agent peut-il appeler ?

  • /.well-known/agent-card.json not found

    This site offers something an agent could call, but nothing tells an agent what. Publish an A2A Agent Card at /.well-known/agent-card.json. A minimal 1.0 card needs name, description, version, capabilities, supportedInterfaces, defaultInputModes, defaultOutputModes and skills. Spec: https://a2a-protocol.org/latest/specification/

  • API surface present but no machine-readable description found

    The API exists; nothing describes it in a form an agent can read, so using it requires a human to read your documentation first. Serve an OpenAPI description at /openapi.json and advertise it with Link: </openapi.json>; rel="service-desc". For several APIs, publish an RFC 9727 catalog.

  • No agent resource catalog found

    A catalog is one document listing everything an agent can call here — agent cards, MCP servers, APIs, skills — so a client stops probing four conventions to find out. Worth publishing once you have more than one of those. Informational: both ai-catalog.json and ard.json are still drafts, so this never affects your score.

  • Site sells something but publishes no agent-readable commerce profile

    Publish a Universal Commerce Protocol profile at /.well-known/ucp so an agent can find your catalog, cart and checkout without a human. Note that the alternatives are not discoverable by design: the OpenAI and Stripe Agentic Commerce Protocol defines no manifest, and AP2 advertises itself through an A2A card extension.

  • 1 form(s) on the page, none declared as agent tools

    A declared form is called by an agent rather than driven pixel by pixel. Add toolname and tooldescription to the forms worth automating — search, filter, subscribe — and toolparamdescription to each field.

  • No MCP server — MCP discovery does not apply to this site

  • No developer-facing surface — skills do not apply to this site

  • Nothing on this site requires authorization — auth discovery does not apply

Politique

43/100

Quels droits d'usage sont déclarés ?

  • No RSL license discovery found

    Declare machine-readable licensing terms for your content with Really Simple Licensing. Add to robots.txt: License: https://your-site.com/license.xml — then publish the RSL document. See https://rslstandard.org.

  • No machine-readable usage policy declared

    State your terms where they can be read without a lawyer. The lowest-effort option is a Content-Signal line in robots.txt: Content-Signal: search=yes, ai-input=yes, ai-train=no. Absence is neutral, not permission — but it also gives you nothing to point at.

  • /.well-known/security.txt exists

  • Required field "Contact" present

  • Required field "Expires" missing (RFC 9116)

    Add "Expires:" to your security.txt. Use an ISO 8601 date, e.g., Expires: 2026-12-31T23:59:59.000Z

  • 1/5 optional fields present

Détails techniques

Moteur
ax-audit@4.0.0
Analysé
4 sept. 2026, 21:50:34 UTC
Durée
2 589 ms
Checks exécutés
26

Transformez ce rapport en améliorations continues

Offrez à votre équipe des rapports complets, des corrections prioritaires et un suivi pour préserver les progrès après chaque livraison.

Lire ce rapport via l'API ou le serveur MCP