Measured on the home page. Overall grade: Poor.
/llms.txt not found
Create a /llms.txt file at your site root following the llmstxt.org specification. It should be a Markdown file starting with "# Your Site Name" and include a description, sections, and links.
/.well-known/agent.json not found
Create a /.well-known/agent.json file following the A2A (Agent-to-Agent) protocol. It should include name, description, url, and skills fields describing your site's capabilities.
/.well-known/openapi.json not found
Create a /.well-known/openapi.json file with your API specification following the OpenAPI 3.x standard. See https://swagger.io/specification/ for the spec.
/.well-known/mcp.json not found
Create a /.well-known/mcp.json file describing your MCP server configuration. Include name, description, tools, and version fields. See https://modelcontextprotocol.io for the spec.
0/5 emerging AI discovery files published
ai.txt not found
Publish /.well-known/ai.txt declaring opt-in/opt-out signals for AI training. See https://site.spawning.ai/spawning-ai-txt for the format.
genai.txt not found
Publish /.well-known/genai.txt declaring your generative-AI usage policy.
ai-plugin.json not found
Publish /ai-plugin.json (legacy ChatGPT plugin manifest). Schema: name_for_model, description_for_model, api.url. Still consumed by some agents.
agents.json not found
Publish /agents.json describing your site as a callable agent (OpenAgents / Wildcard emerging spec). Includes name, description, and operations[].
nlweb.json not found
Publish /.well-known/nlweb.json (Microsoft NLWeb) so agents can interact with the site through a natural-language interface.
Homepage does not serve Markdown via content negotiation
Serve a Markdown representation of your pages when agents request "Accept: text/markdown". Agents like Claude Code and Cursor ask for it, and Markdown cuts token usage by ~80% vs HTML. Cloudflare ("Markdown for Agents") and Vercel can enable this without code changes.
No <link rel="alternate" type="text/markdown"> fallback found on the homepage
If you cannot enable content negotiation, advertise a Markdown version with <link rel="alternate" type="text/markdown" href="/index.md"> so agents can discover it.
No RSL license discovery found
Declare machine-readable licensing terms for your content with Really Simple Licensing. Add to robots.txt: License: https://your-site.com/license.xml — then publish the RSL document. See https://rslstandard.org.
/robots.txt exists
7/8 core AI crawlers configured
Add explicit User-agent entries for the missing crawlers with Allow: / for each one.
22 AI crawler(s) explicitly blocked
These crawlers have "Disallow: /" rules. If you want AI agents to access your site, change to "Allow: /" for each blocked crawler.
Sitemap directive present
No Content-Signal directive found (optional)
Declare how crawlers may use your content after access with the Content Signals Policy, e.g.: Content-Signal: search=yes, ai-train=no. Known signals: search, ai-input, ai-train. Generate yours at contentsignals.org.
22/48 known AI crawlers have explicit rules
No AI meta tags (ai:*) found
Add AI meta tags to your HTML <head>: <meta name="ai:summary" content="Brief description">, <meta name="ai:content_type" content="website">, <meta name="ai:author" content="Your Name">.
No rel="alternate" link to llms.txt in HTML
Add to your <head>: <link rel="alternate" type="text/plain" href="/llms.txt" title="LLM-optimized content">
No rel="alternate" link to agent.json in HTML
Add to your <head>: <link rel="alternate" type="application/json" href="/.well-known/agent.json" title="Agent Card">
No rel="me" identity links found
Add rel="me" links to verify your identity across platforms: <link rel="me" href="https://github.com/yourname">, <link rel="me" href="https://twitter.com/yourname">.
OpenGraph required tags present (og:title, og:description, og:url, og:type)
OpenGraph recommended tags missing: og:image, og:site_name
Add these for richer previews: og:image, og:site_name.
Twitter Card required tags missing: twitter:card
Add these meta tags: <meta name="twitter:card" content="...">.
Server-rendered content detected (2599 words, 16181 chars of visible text)
Low text-to-markup ratio (2.5%)
A very low text-to-markup ratio is a typical SPA-shell symptom. Inline more content directly into the HTML response.
Semantic landmarks present (main, article, section, header, footer, nav)
No <h1> heading found
Add a single <h1> describing the page. Agents and search engines treat the H1 as the primary topic indicator.
Only 67/131 <img> tags have alt attributes
Add descriptive alt="" to every <img>. Agents use alt text to understand images they cannot process visually.
1 JSON-LD block(s) found
@context references schema.org
No @graph array (single-entity only)
Use an @graph array to define multiple entities in one JSON-LD block: { "@context": "https://schema.org", "@graph": [...] }
Only 1 key type found: WebPage
Add more entity types to your @graph. AI agents use these to understand site structure. Common types: Person, Organization, WebSite, WebPage.
No BreadcrumbList found
Add a BreadcrumbList entity to help AI agents understand your site navigation hierarchy.
Response compressed with gzip
Cache validator present (ETag)
Conditional request returned 200 instead of 304 Not Modified
The server advertises a cache validator but does not honor If-None-Match / If-Modified-Since. Configure it to return 304 when the validator matches, so crawlers avoid re-downloading unchanged pages.
Homepage is on the large side (630.8 KB decompressed)
Consider trimming inlined payloads to reduce crawl cost.
4/7 security headers present
No Link header for AI discovery (llms.txt, agent.json)
Add a Link response header pointing to your AI discovery files: Link: </llms.txt>; rel="alternate"; type="text/plain", </.well-known/agent.json>; rel="alternate"; type="application/json"
Site is served over HTTPS
HTTP requests redirect to HTTPS
HSTS max-age=31536000
HSTS does not include subdomains
Add includeSubDomains to apply HSTS across api., docs., etc. Required for preload list submission.
HSTS has preload directive but does not satisfy preload-list requirements
Preload requires max-age >= 31536000 and includeSubDomains. See https://hstspreload.org.
<title> is too long (120 chars)
Shorten the title to 20-70 characters. Many agents and search engines truncate beyond ~70.
Meta description length 126 chars
2 <link rel="canonical"> tags (must be exactly 1)
Keep a single canonical link per page. Multiple canonical hints are ignored by agents and search engines.
<html lang="en-GB">
UTF-8 charset declared
Viewport meta present: "width=device-width, initial-scale=1"
Sitemap located: https://www.bbc.com/afrique/sitemap.xml
Content-Type is XML (application/xml)
100 URL(s) declared
100% of URLs have <lastmod>
Newest <lastmod> is 4363 days old
Refresh <lastmod> on URLs that have changed. A stale sitemap signals to crawlers that nothing on the site has updated.
/.well-known/security.txt exists
Required field "Contact" present
Required field "Expires" present
Expires date is in the future (2038-01-19)
2/5 optional fields present
All 8 core AI crawler user-agents receive equivalent responses