Skip to content
Agentic Web readiness · public methodology

What do AI agents need to find, understand and act on your business?

ARC uses a versioned test registry to examine whether AI systems can find, understand, trust and act on the information and journeys a business makes available.

The agent journeyWebsite → store → transaction context
  1. 01
    Discover

    Find the business.

  2. 02
    Understand

    Interpret the offer.

  3. 03
    Trust

    Establish credible signals.

  4. 04
    Recommend

    Support an informed choice.

  5. 05
    Act

    Reach a useful next step.

  6. 06
    Transact

    Establish the applicable transaction path.

Conceptual framework. Each stage needs evidence; the diagram does not mean every stage was tested or passed.

How to read ARC

Business meaning first. Evidence second.

ARC does not treat missing evidence as failure. Each result keeps its evidence, applicability and assessment boundary visible.

PASS

The check found evidence the condition holds.

NEEDS ATTENTION

ARC found something that could weaken this part of the journey.

FAIL

The check found evidence the condition does not hold.

NOT VERIFIED

The public evidence wasn't enough to decide. Not a failure.

NOT APPLICABLE

The check doesn't apply to this business.

Canonical test registry

Reach and Capability, organised into the current ARC categories.

The category names, test families and maturity labels below load from the canonical ARC registry rather than a marketing checklist.

Registry version 2.1.0

Reach

Can AI systems discover, access, understand and recommend the business?

Discovery & Crawl Access

Can AI systems find and access the information you want them to see?

11 tests
homepage_fetchable

Your site loads for AI systems

The homepage returns a successful (2xx-3xx) HTTP response.

Why it matters: An agent that cannot fetch the homepage cannot discover, evaluate, or act on the business at all.

Evidence
http_status
Availability
BOTH
Maturity
CORE
robots_txt_available

You have a crawl policy

/robots.txt returns a successful HTTP response.

Why it matters: robots.txt is the first file most crawlers, including AI crawlers, check to learn what they may access.

Evidence
file_presence
Availability
BOTH
Maturity
CORE
site_not_globally_blocked

Your site isn't hidden from every crawler

robots.txt does not contain a blanket Disallow: / for User-agent: *.

Why it matters: A blanket disallow tells every crawler, including AI agents, not to access any page on the site.

Evidence
content_parse
Availability
BOTH
Maturity
CORE
named_ai_crawlers_not_blocked

AI crawlers specifically aren't blocked

robots.txt does not disallow any of the named AI crawlers this check recognizes (GPTBot, ClaudeBot, Google-Extended, CCBot, PerplexityBot, Amazonbot, Applebot-Extended).

Why it matters: A Disallow rule for one of these is an explicit opt-out from that specific agent's discovery of the site. This is reported as the site owner's policy, not automatically penalised as wrong — a deliberate opt-out is a valid choice.

Evidence
content_parse
Availability
BOTH
Maturity
CORE
redirect_chain_resolves_cleanly

Your site doesn't bounce crawlers around

The homepage URL resolves in a small number of redirect hops to a single stable final host, with no redirect loop.

Why it matters: Long or unstable redirect chains slow or break crawl access and can cause an agent to index the wrong host.

Evidence
http_status
Availability
BOTH
Maturity
CORE
xml_sitemap_available

You publish a sitemap

/sitemap.xml (or the Sitemap: URL declared in robots.txt) returns a successful response.

Why it matters: Gives crawlers a complete, authoritative list of public URLs instead of relying on link-following alone.

Evidence
file_presence
Availability
BOTH
Maturity
CORE
xml_sitemap_valid

Your sitemap is well-formed

The fetched sitemap parses as well-formed XML matching the sitemap namespace, with at least one <url> entry.

Why it matters: A sitemap that doesn't parse is worse than no sitemap — crawlers may abandon it entirely rather than partially trust it.

Evidence
content_parse
Availability
PAID
Maturity
CORE
meta_robots_restrictions_documented

Indexing restrictions are shown, not hidden

Records any noindex, nofollow, nosnippet, or X-Robots-Tag directive found on the homepage, and whether it blocks AI-search visibility.

Why it matters: A noindex or nosnippet directive can silently remove a page from AI-search results even though the page itself is otherwise healthy.

Evidence
content_parse
Availability
BOTH
Maturity
CORE
waf_challenge_blocks_legitimate_access

Bot protection isn't blocking AI crawlers too

The homepage response is not a bot-challenge/interstitial page (e.g. a JS-challenge or CAPTCHA wall) to a plain, well-identified HTTP request.

Why it matters: Overly aggressive bot protection can block legitimate AI crawlers the same way it blocks malicious bots, with no visibility into the loss.

Evidence
content_parse
Availability
PAID
Maturity
CONDITIONAL
content_available_without_js

Your key content is readable without running JavaScript

Compares statically-fetched homepage content against a headless-render pass (where available); flags when meaningful content only appears after JS execution.

Why it matters: Many crawlers, including some AI crawlers, do not execute JavaScript — content that only renders client-side is invisible to them.

Evidence
content_parse
Availability
PAID
Maturity
CORE
meta_robots_ai_directive_detected

Explicit AI-training opt-out (if set)

Records whether a non-standard meta directive such as noai/noimageai is present. Purely informational — absence is never penalized, since this is not an established standard.

Why it matters: Surfaces an explicit AI-training opt-out signal where a site owner has chosen to set one, without treating its absence as a defect.

Evidence
content_parse
Availability
PAID
Maturity
EMERGING_INFORMATIONAL
Machine Readability & Content Structure

Can machines reliably read and interpret the important content on your site?

13 tests
llms_txt_available

You publish an llms.txt summary

/llms.txt is fetchable and looks like the llms.txt convention (Markdown with at least one heading).

Why it matters: An emerging, optional convention giving AI agents a curated summary instead of raw HTML. Not a web standard — and Google has stated Search does not use llms.txt for its generative AI features — so absence is informational, not a critical failure.

Evidence
file_presence
Availability
BOTH
Maturity
EMERGING_INFORMATIONAL
llms_full_txt_available

You publish an expanded llms-full.txt

/llms-full.txt is fetchable and non-trivial in length.

Why it matters: Where present, gives agents that consume the llms.txt convention a fuller reference than the summary file alone. Same non-standard, informational status as llms.txt.

Evidence
file_presence
Availability
PAID
Maturity
EMERGING_INFORMATIONAL
markdown_content_negotiation

You offer a Markdown version of your pages

Requesting the homepage with Accept: text/markdown returns Markdown, or an /index.md fallback exists.

Why it matters: Some agent tooling prefers Markdown over HTML for lower-noise parsing; support is emerging, not required.

Evidence
http_status
Availability
PAID
Maturity
EMERGING_INFORMATIONAL
descriptive_title_present

Your homepage has a clear title

The homepage's <title> element is at least 10 characters long.

Why it matters: The <title> is usually the first signal an agent reads to understand what the page/business is.

Evidence
content_parse
Availability
BOTH
Maturity
CORE
meta_description_present

Your homepage has a summary description

A meta or og:description tag with non-empty content is present.

Why it matters: A concise, structured summary agents can quote directly instead of extracting one from unstructured body copy.

Evidence
content_parse
Availability
BOTH
Maturity
CORE
canonical_url_tag_present

Your preferred page URL is declared

The homepage declares a <link rel="canonical"> pointing to an absolute URL.

Why it matters: Tells an agent (or search engine) which URL is authoritative when the same content is reachable at multiple addresses.

Evidence
content_parse
Availability
PAID
Maturity
CONDITIONAL
valid_json_ld_exists

Your site has machine-readable structured data

At least one schema.org @type is present in a JSON-LD block on the homepage (including client-rendered content where a render fallback succeeded).

Why it matters: JSON-LD is a primary machine-readable format for structured facts. Structured data is valuable for search features and agent parsing, but its presence is not itself a guarantee of inclusion in Google AI Overviews or any specific AI product.

Evidence
content_parse
Availability
BOTH
Maturity
CORE
schema_types_relevant_to_business

Your structured data matches your business type

Where an entity/commercial schema type is present, checks it against an expected type family (e.g. LocalBusiness for a storefront, Organization generically) rather than crediting any JSON-LD at all.

Why it matters: Generic or mismatched schema types (e.g. only WebSite/WebPage boilerplate) give agents little more than they'd get from prose.

Evidence
content_parse
Availability
PAID
Maturity
CONDITIONAL
service_product_offer_faq_schema

Your offerings are marked up in structured data

JSON-LD includes at least one of Service, Product, Offer, or FAQPage.

Why it matters: Lets an agent resolve exactly what is being sold and its attributes, not just that the page mentions a product.

Evidence
content_parse
Availability
PAID
Maturity
CONDITIONAL
structured_data_matches_visible_content

Your structured data matches what's on the page

Where both JSON-LD and visible text state a comparable fact (e.g. a price), checks they do not directly contradict each other.

Why it matters: Contradictory structured-vs-visible facts erode agent trust in both, and can cause an agent to cite a wrong price or detail.

Evidence
content_parse
Availability
PAID
Maturity
CONDITIONAL
duplicate_canonical_ambiguity

Your site isn't creating duplicate-content confusion

Flags when sampled pages appear to duplicate the homepage content without a canonical tag resolving the ambiguity.

Why it matters: Ambiguous duplicate content makes it harder for an agent to know which version to cite or trust.

Evidence
content_parse
Availability
PAID
Maturity
CONDITIONAL
freshness_signals_present

Time-sensitive content shows a date

Where a datePublished/dateModified schema field or a visible last-updated date is expected (e.g. pricing, offers), checks one is present.

Why it matters: An agent citing stale pricing or availability as current misleads the buyer it's assisting.

Evidence
content_parse
Availability
PAID
Maturity
CONDITIONAL
agentic_web_signals_present

You show at least one emerging AI-agent convention

Detects any publicly linked llms.txt, OpenAPI spec, or MCP-manifest-looking resource as a bundle of emerging-convention signals.

Why it matters: These are emerging, non-standard conventions; their presence is a positive signal, but absence is explicitly not treated as a critical failure.

Evidence
descriptor_only
Availability
PAID
Maturity
EMERGING_INFORMATIONAL
Business Understanding & Answerability

Can AI correctly explain what your business offers and answer the questions customers actually ask?

6 tests
content_substantial_enough_to_quote

You have enough real content to be quoted

The homepage's visible text is at least 150 words (partial credit at 50+).

Why it matters: Agents answering buyer questions by quoting or citing a page need enough real content to extract an accurate, attributable answer.

Evidence
text_heuristic
Availability
BOTH
Maturity
CORE
entity_name_clear

It's clear who you are

The homepage has a non-empty <title>. This checks only for a page-title identity signal, not a verified business-registry identity.

Why it matters: An agent needs an unambiguous entity name before it will treat a site as a legitimate business to cite or transact with.

Evidence
content_parse
Availability
BOTH
Maturity
CORE
entity_schema_exists

Your business identity is machine-readable

JSON-LD includes an Organization, LocalBusiness, Corporation, Person, or ProfessionalService type.

Why it matters: Gives an agent a structured, machine-verifiable identity record instead of an inferred one.

Evidence
content_parse
Availability
PAID
Maturity
CONDITIONAL
contact_identity_visible

Customers (and AI) can see how to reach you

The homepage's text or links contain a contact/email/phone/address signal word.

Why it matters: A visible contact path is a baseline trust and provenance signal agents use before recommending a business for a real transaction.

Evidence
text_heuristic
Availability
BOTH
Maturity
CORE
about_policies_proof_signals

You show proof and policy information

The homepage's text contains about/privacy/terms/testimonial/case-study/review language.

Why it matters: Gives an agent secondary evidence of legitimacy beyond the homepage's own claims.

Evidence
text_heuristic
Availability
PAID
Maturity
CONDITIONAL
business_answerability_extractive_coverage

AI can find answers to the questions buyers actually ask

For each applicable buyer question (what/who/where/cost/how-to-buy/terms/cancellation/contact/differentiation), checks whether retrieved page content contains extractable supporting text, and records the answer, evidence URL, and confidence. Never fabricates an answer when no supporting text was found. This measures whether information is present, clear, and extractable — it is not a live test of whether an AI answer engine actually answers correctly; that is assessed separately during managed paid fulfilment.

Why it matters: A technically valid website that cannot communicate pricing, services or conversion information in extractable form gives an agent nothing reliable to work with, regardless of technical SEO health.

Evidence
content_parse
Availability
PAID
Maturity
CORE
AI Visibility & Recommendation

Does your business appear when customers use AI to research and compare providers?

1 test
ai_visibility_brand_mention_citation

Do AI systems mention and recommend you?

Benchmarks realistic non-branded buyer-intent prompts for this business's category against major AI answer engines (ChatGPT/OpenAI, Gemini, Claude, Grok, Perplexity where appropriate), recording brand mention rate, citation rate, prominence, category recommendation presence, and competitor share of voice. Performed manually by Hoshiro during paid-report fulfilment for this launch phase — not scanner-automated. Scored Not verified until a ManualAuditFinding is recorded for the audit.

Why it matters: Technical readiness does not guarantee an AI system actually surfaces or recommends the business — this is the direct, model-observed signal of whether it does.

Evidence
manual
Availability
PAID
Maturity
CONDITIONAL
Capability

Can AI agents interact with the business and successfully complete useful tasks?

Agent Interfaces & Actionability

Does your site expose reliable ways for AI agents to use its functions instead of guessing through the UI?

10 tests
form_or_contact_path_exists

There's a working way to reach you

The homepage contains an HTML <form>, or /contact returns a successful HTTP response.

Why it matters: The concrete mechanism an agent (or a human it is assisting) uses to actually start a transaction.

Evidence
html_probe
Availability
PAID
Maturity
CORE
openapi_documentation_valid

Your API is documented in a machine-readable way

A probed OpenAPI/Swagger document is served and actually parses as a valid OpenAPI/Swagger document (not merely a 200 response from an SPA catch-all).

Why it matters: Only relevant where a real commercial API use case exists; this scanner never invents an API as a requirement.

Evidence
content_parse
Availability
PAID
Maturity
CONDITIONAL
api_catalog_available

Your APIs are listed in a standard catalog

/.well-known/api-catalog is fetchable and parses per RFC 9727.

Why it matters: A standard catalog location lets an agent discover all of a site's APIs without guessing paths.

Evidence
file_presence
Availability
PAID
Maturity
EMERGING_INFORMATIONAL
mcp_descriptor_discoverable

AI agents can discover your MCP tools

A documented/well-known MCP endpoint or server manifest is publicly discoverable and its manifest parses as valid JSON with the expected MCP shape. This checks discoverability only — a real MCP handshake and tool-quality review is performed manually during paid fulfilment and recorded separately, never assumed from descriptor presence alone.

Why it matters: Descriptor discoverability is a prerequisite for any agent to even attempt a connection, but presence alone does not prove the server actually works.

Evidence
descriptor_only
Availability
PAID
Maturity
CONDITIONAL
mcp_real_handshake_verified

Your MCP server actually works

A real MCP connection is attempted (initialize + tools/list JSON-RPC over the Streamable HTTP transport) against the referenced manifest's endpoint, and returned tool names/descriptions/input schemas are checked for clarity and non-overlap. No tool is ever invoked. Scanner-executed as of methodology 2.1.0 whenever a manifest reference exists; Not verified only when no manifest reference was discoverable at all (see mcp_descriptor_discoverable).

Why it matters: A published MCP descriptor that fails a real handshake gives agents nothing — real behaviour matters more than file presence.

Evidence
handshake
Availability
PAID
Maturity
CONDITIONAL
agent_skills_discoverable

Your agent skills are published correctly

Where an Agent Skills manifest is published at its documented location, checks it parses and includes usable capability descriptions.

Why it matters: Malformed or vague skill descriptions are effectively invisible to an agent trying to decide when to use them.

Evidence
descriptor_only
Availability
PAID
Maturity
EMERGING_INFORMATIONAL
a2a_agent_card_valid

Your agent-to-agent identity card is correct

/.well-known/agent-card.json is fetchable and validates against the expected A2A shape (identity, endpoints, skills, security metadata). Only applicable where an agent-to-agent service surface exists; absent for an ordinary business site, this is Not applicable, never a failure.

Why it matters: A malformed agent card breaks discovery for any A2A-capable agent trying to interoperate with this service.

Evidence
content_parse
Availability
PAID
Maturity
EMERGING_INFORMATIONAL
a2a_behavioral_verification

Your agent-to-agent service actually responds correctly

Where a valid agent card's `url` is present, a real A2A `message/send` task exchange is attempted with a short, generic capability-discovery message (never a transactional one). Scanner-executed as of methodology 2.1.0; Not verified only when no agent card with a usable endpoint URL was found.

Why it matters: A valid-looking agent card does not guarantee the underlying service behaves correctly under a real task exchange.

Evidence
real_call
Availability
PAID
Maturity
EMERGING_INFORMATIONAL
webmcp_real_interaction_verified

Your in-page AI tools actually work

Where WebMCP tool registration is detected on the page, real tool registration, naming, schema usability, and a safe task execution are verified manually during paid fulfilment. Scored Not verified until a ManualAuditFinding is recorded.

Why it matters: A page that merely references WebMCP without correctly registered, usable tools gives an agent nothing more than a descriptor would.

Evidence
manual
Availability
PAID
Maturity
EMERGING_INFORMATIONAL
public_transaction_interface_documented

Your transaction API is documented

Same OpenAPI/API-docs signal as openapi_documentation_valid, presented as a transaction-readiness rather than a discovery signal.

Why it matters: Where a business genuinely supports agent-executed transactions, a documented public interface lets agents act safely within defined boundaries.

Evidence
content_parse
Availability
PAID
Maturity
CONDITIONAL
Identity, Authentication & Agent Access

Can legitimate agents identify themselves and securely access the capabilities they are authorised to use?

3 tests
oauth_authorization_server_metadata

Delegated agent access is discoverable

/.well-known/oauth-authorization-server is fetchable and validates per RFC 8414. Not applicable when the site has no delegated/private agent-access surface at all.

Why it matters: Standard OAuth metadata lets an agent discover authorization/token endpoints and scopes without out-of-band instructions.

Evidence
content_parse
Availability
PAID
Maturity
EMERGING_INFORMATIONAL
oauth_protected_resource_metadata

Protected agent resources are discoverable

/.well-known/oauth-protected-resource is fetchable and validates per RFC 9728. Not applicable when the site has no protected agent-access surface.

Why it matters: Lets an agent locate the authorization server responsible for a protected resource without prior knowledge.

Evidence
content_parse
Availability
PAID
Maturity
EMERGING_INFORMATIONAL
authenticated_agent_access_verification

Agent sign-in actually works as documented

Where OAuth/agent-login metadata is published, a real authenticated flow (PKCE S256 support, usable scopes, agreement between advertised metadata and the actual protected service) is verified manually during paid fulfilment. Never attempted automatically without credentials.

Why it matters: Advertised auth metadata that doesn't match real service behaviour breaks every agent that trusts it.

Evidence
manual
Availability
PAID
Maturity
CONDITIONAL
Conversion & Agentic Commerce

Can an AI agent progress toward a real business outcome such as an enquiry, booking, signup or purchase?

4 tests
clear_conversion_action_visible

It's obvious what to do next

The homepage's text or links contain contact/book/buy/order/quote/signup/schedule language.

Why it matters: An agent completing a task needs an unambiguous next step rather than having to infer intent from navigation.

Evidence
text_heuristic
Availability
BOTH
Maturity
CORE
pricing_offer_context_visible

Pricing or offer information is visible

The homepage's text contains 'pricing', 'price', or 'offer' language.

Why it matters: Helps an agent compare options quickly; a 'contact sales' model is a legitimate choice, not a defect by itself.

Evidence
text_heuristic
Availability
PAID
Maturity
CONDITIONAL
emerging_commerce_protocol_signals

Emerging agent-payment protocols (if relevant)

Detects publicly documented support for x402, Universal Commerce Protocol, Agentic Commerce Protocol, AP2, or Machine Payments Protocol. Purely informational for a typical business; never penalizes absence.

Why it matters: Only material for businesses whose model actually requires agent-executed payments; premature for most SME sites.

Evidence
descriptor_only
Availability
PAID
Maturity
EMERGING_INFORMATIONAL
agent_task_completion_verified

An AI agent can actually complete your primary action

A safe, non-destructive attempt to progress through the site's actual primary conversion action (enquiry/booking/signup/purchase) as an agent would, capturing where the task succeeds or breaks. Performed manually during paid fulfilment; never submits a destructive or chargeable action during an ordinary scan.

Why it matters: This is the direct test of whether an agent can actually get a customer to the outcome the business needs, not just whether the page looks navigable.

Evidence
manual
Availability
PAID
Maturity
CORE

Tests apply according to the assessment scope, maturity and available evidence. A Free ARC Scan is a bounded public-web assessment; deeper testing may require customer context, authorised access or an agreed ARC Report scope.

From result to closure

A finding is useful only when its next decision is clear.

01Find

ARC records an applicable condition and its evidence.

02Understand

See the likely business consequence and limitation.

03Agree scope

Only purchased and approved remediation is implemented.

04Retest

Re-run the same relevant test against the target outcome.

05Record proof

Keep before-and-after evidence, including what remains open.

Start with your public website

Get the bounded starting point.

Run an ARC Scan to see the checks performed, their result states and the next available action. Use ARC Report when deeper evidence and interpretation are needed.

Start Free ARC Scan