PASS
The check found evidence the condition holds.
ARC uses a versioned test registry to examine whether AI systems can find, understand, trust and act on the information and journeys a business makes available.
Find the business.
Interpret the offer.
Establish credible signals.
Support an informed choice.
Reach a useful next step.
Establish the applicable transaction path.
Conceptual framework. Each stage needs evidence; the diagram does not mean every stage was tested or passed.
ARC does not treat missing evidence as failure. Each result keeps its evidence, applicability and assessment boundary visible.
The check found evidence the condition holds.
ARC found something that could weaken this part of the journey.
The check found evidence the condition does not hold.
The public evidence wasn't enough to decide. Not a failure.
The check doesn't apply to this business.
The category names, test families and maturity labels below load from the canonical ARC registry rather than a marketing checklist.
Registry version 2.1.0
Can AI systems find and access the information you want them to see?
homepage_fetchableThe homepage returns a successful (2xx-3xx) HTTP response.
Why it matters: An agent that cannot fetch the homepage cannot discover, evaluate, or act on the business at all.
robots_txt_available/robots.txt returns a successful HTTP response.
Why it matters: robots.txt is the first file most crawlers, including AI crawlers, check to learn what they may access.
site_not_globally_blockedrobots.txt does not contain a blanket Disallow: / for User-agent: *.
Why it matters: A blanket disallow tells every crawler, including AI agents, not to access any page on the site.
named_ai_crawlers_not_blockedrobots.txt does not disallow any of the named AI crawlers this check recognizes (GPTBot, ClaudeBot, Google-Extended, CCBot, PerplexityBot, Amazonbot, Applebot-Extended).
Why it matters: A Disallow rule for one of these is an explicit opt-out from that specific agent's discovery of the site. This is reported as the site owner's policy, not automatically penalised as wrong — a deliberate opt-out is a valid choice.
redirect_chain_resolves_cleanlyThe homepage URL resolves in a small number of redirect hops to a single stable final host, with no redirect loop.
Why it matters: Long or unstable redirect chains slow or break crawl access and can cause an agent to index the wrong host.
xml_sitemap_available/sitemap.xml (or the Sitemap: URL declared in robots.txt) returns a successful response.
Why it matters: Gives crawlers a complete, authoritative list of public URLs instead of relying on link-following alone.
xml_sitemap_validThe fetched sitemap parses as well-formed XML matching the sitemap namespace, with at least one <url> entry.
Why it matters: A sitemap that doesn't parse is worse than no sitemap — crawlers may abandon it entirely rather than partially trust it.
meta_robots_restrictions_documentedRecords any noindex, nofollow, nosnippet, or X-Robots-Tag directive found on the homepage, and whether it blocks AI-search visibility.
Why it matters: A noindex or nosnippet directive can silently remove a page from AI-search results even though the page itself is otherwise healthy.
waf_challenge_blocks_legitimate_accessThe homepage response is not a bot-challenge/interstitial page (e.g. a JS-challenge or CAPTCHA wall) to a plain, well-identified HTTP request.
Why it matters: Overly aggressive bot protection can block legitimate AI crawlers the same way it blocks malicious bots, with no visibility into the loss.
content_available_without_jsCompares statically-fetched homepage content against a headless-render pass (where available); flags when meaningful content only appears after JS execution.
Why it matters: Many crawlers, including some AI crawlers, do not execute JavaScript — content that only renders client-side is invisible to them.
meta_robots_ai_directive_detectedRecords whether a non-standard meta directive such as noai/noimageai is present. Purely informational — absence is never penalized, since this is not an established standard.
Why it matters: Surfaces an explicit AI-training opt-out signal where a site owner has chosen to set one, without treating its absence as a defect.
Can machines reliably read and interpret the important content on your site?
llms_txt_available/llms.txt is fetchable and looks like the llms.txt convention (Markdown with at least one heading).
Why it matters: An emerging, optional convention giving AI agents a curated summary instead of raw HTML. Not a web standard — and Google has stated Search does not use llms.txt for its generative AI features — so absence is informational, not a critical failure.
llms_full_txt_available/llms-full.txt is fetchable and non-trivial in length.
Why it matters: Where present, gives agents that consume the llms.txt convention a fuller reference than the summary file alone. Same non-standard, informational status as llms.txt.
markdown_content_negotiationRequesting the homepage with Accept: text/markdown returns Markdown, or an /index.md fallback exists.
Why it matters: Some agent tooling prefers Markdown over HTML for lower-noise parsing; support is emerging, not required.
descriptive_title_presentThe homepage's <title> element is at least 10 characters long.
Why it matters: The <title> is usually the first signal an agent reads to understand what the page/business is.
meta_description_presentA meta or og:description tag with non-empty content is present.
Why it matters: A concise, structured summary agents can quote directly instead of extracting one from unstructured body copy.
canonical_url_tag_presentThe homepage declares a <link rel="canonical"> pointing to an absolute URL.
Why it matters: Tells an agent (or search engine) which URL is authoritative when the same content is reachable at multiple addresses.
valid_json_ld_existsAt least one schema.org @type is present in a JSON-LD block on the homepage (including client-rendered content where a render fallback succeeded).
Why it matters: JSON-LD is a primary machine-readable format for structured facts. Structured data is valuable for search features and agent parsing, but its presence is not itself a guarantee of inclusion in Google AI Overviews or any specific AI product.
schema_types_relevant_to_businessWhere an entity/commercial schema type is present, checks it against an expected type family (e.g. LocalBusiness for a storefront, Organization generically) rather than crediting any JSON-LD at all.
Why it matters: Generic or mismatched schema types (e.g. only WebSite/WebPage boilerplate) give agents little more than they'd get from prose.
service_product_offer_faq_schemaJSON-LD includes at least one of Service, Product, Offer, or FAQPage.
Why it matters: Lets an agent resolve exactly what is being sold and its attributes, not just that the page mentions a product.
structured_data_matches_visible_contentWhere both JSON-LD and visible text state a comparable fact (e.g. a price), checks they do not directly contradict each other.
Why it matters: Contradictory structured-vs-visible facts erode agent trust in both, and can cause an agent to cite a wrong price or detail.
duplicate_canonical_ambiguityFlags when sampled pages appear to duplicate the homepage content without a canonical tag resolving the ambiguity.
Why it matters: Ambiguous duplicate content makes it harder for an agent to know which version to cite or trust.
freshness_signals_presentWhere a datePublished/dateModified schema field or a visible last-updated date is expected (e.g. pricing, offers), checks one is present.
Why it matters: An agent citing stale pricing or availability as current misleads the buyer it's assisting.
agentic_web_signals_presentDetects any publicly linked llms.txt, OpenAPI spec, or MCP-manifest-looking resource as a bundle of emerging-convention signals.
Why it matters: These are emerging, non-standard conventions; their presence is a positive signal, but absence is explicitly not treated as a critical failure.
Can AI correctly explain what your business offers and answer the questions customers actually ask?
content_substantial_enough_to_quoteThe homepage's visible text is at least 150 words (partial credit at 50+).
Why it matters: Agents answering buyer questions by quoting or citing a page need enough real content to extract an accurate, attributable answer.
entity_name_clearThe homepage has a non-empty <title>. This checks only for a page-title identity signal, not a verified business-registry identity.
Why it matters: An agent needs an unambiguous entity name before it will treat a site as a legitimate business to cite or transact with.
entity_schema_existsJSON-LD includes an Organization, LocalBusiness, Corporation, Person, or ProfessionalService type.
Why it matters: Gives an agent a structured, machine-verifiable identity record instead of an inferred one.
contact_identity_visibleThe homepage's text or links contain a contact/email/phone/address signal word.
Why it matters: A visible contact path is a baseline trust and provenance signal agents use before recommending a business for a real transaction.
about_policies_proof_signalsThe homepage's text contains about/privacy/terms/testimonial/case-study/review language.
Why it matters: Gives an agent secondary evidence of legitimacy beyond the homepage's own claims.
business_answerability_extractive_coverageFor each applicable buyer question (what/who/where/cost/how-to-buy/terms/cancellation/contact/differentiation), checks whether retrieved page content contains extractable supporting text, and records the answer, evidence URL, and confidence. Never fabricates an answer when no supporting text was found. This measures whether information is present, clear, and extractable — it is not a live test of whether an AI answer engine actually answers correctly; that is assessed separately during managed paid fulfilment.
Why it matters: A technically valid website that cannot communicate pricing, services or conversion information in extractable form gives an agent nothing reliable to work with, regardless of technical SEO health.
Does your business appear when customers use AI to research and compare providers?
ai_visibility_brand_mention_citationBenchmarks realistic non-branded buyer-intent prompts for this business's category against major AI answer engines (ChatGPT/OpenAI, Gemini, Claude, Grok, Perplexity where appropriate), recording brand mention rate, citation rate, prominence, category recommendation presence, and competitor share of voice. Performed manually by Hoshiro during paid-report fulfilment for this launch phase — not scanner-automated. Scored Not verified until a ManualAuditFinding is recorded for the audit.
Why it matters: Technical readiness does not guarantee an AI system actually surfaces or recommends the business — this is the direct, model-observed signal of whether it does.
Does your site expose reliable ways for AI agents to use its functions instead of guessing through the UI?
form_or_contact_path_existsThe homepage contains an HTML <form>, or /contact returns a successful HTTP response.
Why it matters: The concrete mechanism an agent (or a human it is assisting) uses to actually start a transaction.
openapi_documentation_validA probed OpenAPI/Swagger document is served and actually parses as a valid OpenAPI/Swagger document (not merely a 200 response from an SPA catch-all).
Why it matters: Only relevant where a real commercial API use case exists; this scanner never invents an API as a requirement.
api_catalog_available/.well-known/api-catalog is fetchable and parses per RFC 9727.
Why it matters: A standard catalog location lets an agent discover all of a site's APIs without guessing paths.
mcp_descriptor_discoverableA documented/well-known MCP endpoint or server manifest is publicly discoverable and its manifest parses as valid JSON with the expected MCP shape. This checks discoverability only — a real MCP handshake and tool-quality review is performed manually during paid fulfilment and recorded separately, never assumed from descriptor presence alone.
Why it matters: Descriptor discoverability is a prerequisite for any agent to even attempt a connection, but presence alone does not prove the server actually works.
mcp_real_handshake_verifiedA real MCP connection is attempted (initialize + tools/list JSON-RPC over the Streamable HTTP transport) against the referenced manifest's endpoint, and returned tool names/descriptions/input schemas are checked for clarity and non-overlap. No tool is ever invoked. Scanner-executed as of methodology 2.1.0 whenever a manifest reference exists; Not verified only when no manifest reference was discoverable at all (see mcp_descriptor_discoverable).
Why it matters: A published MCP descriptor that fails a real handshake gives agents nothing — real behaviour matters more than file presence.
agent_skills_discoverableWhere an Agent Skills manifest is published at its documented location, checks it parses and includes usable capability descriptions.
Why it matters: Malformed or vague skill descriptions are effectively invisible to an agent trying to decide when to use them.
a2a_agent_card_valid/.well-known/agent-card.json is fetchable and validates against the expected A2A shape (identity, endpoints, skills, security metadata). Only applicable where an agent-to-agent service surface exists; absent for an ordinary business site, this is Not applicable, never a failure.
Why it matters: A malformed agent card breaks discovery for any A2A-capable agent trying to interoperate with this service.
a2a_behavioral_verificationWhere a valid agent card's `url` is present, a real A2A `message/send` task exchange is attempted with a short, generic capability-discovery message (never a transactional one). Scanner-executed as of methodology 2.1.0; Not verified only when no agent card with a usable endpoint URL was found.
Why it matters: A valid-looking agent card does not guarantee the underlying service behaves correctly under a real task exchange.
webmcp_real_interaction_verifiedWhere WebMCP tool registration is detected on the page, real tool registration, naming, schema usability, and a safe task execution are verified manually during paid fulfilment. Scored Not verified until a ManualAuditFinding is recorded.
Why it matters: A page that merely references WebMCP without correctly registered, usable tools gives an agent nothing more than a descriptor would.
public_transaction_interface_documentedSame OpenAPI/API-docs signal as openapi_documentation_valid, presented as a transaction-readiness rather than a discovery signal.
Why it matters: Where a business genuinely supports agent-executed transactions, a documented public interface lets agents act safely within defined boundaries.
Can legitimate agents identify themselves and securely access the capabilities they are authorised to use?
oauth_authorization_server_metadata/.well-known/oauth-authorization-server is fetchable and validates per RFC 8414. Not applicable when the site has no delegated/private agent-access surface at all.
Why it matters: Standard OAuth metadata lets an agent discover authorization/token endpoints and scopes without out-of-band instructions.
oauth_protected_resource_metadata/.well-known/oauth-protected-resource is fetchable and validates per RFC 9728. Not applicable when the site has no protected agent-access surface.
Why it matters: Lets an agent locate the authorization server responsible for a protected resource without prior knowledge.
authenticated_agent_access_verificationWhere OAuth/agent-login metadata is published, a real authenticated flow (PKCE S256 support, usable scopes, agreement between advertised metadata and the actual protected service) is verified manually during paid fulfilment. Never attempted automatically without credentials.
Why it matters: Advertised auth metadata that doesn't match real service behaviour breaks every agent that trusts it.
Can an AI agent progress toward a real business outcome such as an enquiry, booking, signup or purchase?
clear_conversion_action_visibleThe homepage's text or links contain contact/book/buy/order/quote/signup/schedule language.
Why it matters: An agent completing a task needs an unambiguous next step rather than having to infer intent from navigation.
pricing_offer_context_visibleThe homepage's text contains 'pricing', 'price', or 'offer' language.
Why it matters: Helps an agent compare options quickly; a 'contact sales' model is a legitimate choice, not a defect by itself.
emerging_commerce_protocol_signalsDetects publicly documented support for x402, Universal Commerce Protocol, Agentic Commerce Protocol, AP2, or Machine Payments Protocol. Purely informational for a typical business; never penalizes absence.
Why it matters: Only material for businesses whose model actually requires agent-executed payments; premature for most SME sites.
agent_task_completion_verifiedA safe, non-destructive attempt to progress through the site's actual primary conversion action (enquiry/booking/signup/purchase) as an agent would, capturing where the task succeeds or breaks. Performed manually during paid fulfilment; never submits a destructive or chargeable action during an ordinary scan.
Why it matters: This is the direct test of whether an agent can actually get a customer to the outcome the business needs, not just whether the page looks navigable.
Tests apply according to the assessment scope, maturity and available evidence. A Free ARC Scan is a bounded public-web assessment; deeper testing may require customer context, authorised access or an agreed ARC Report scope.
ARC records an applicable condition and its evidence.
See the likely business consequence and limitation.
Only purchased and approved remediation is implemented.
Re-run the same relevant test against the target outcome.
Keep before-and-after evidence, including what remains open.
Run an ARC Scan to see the checks performed, their result states and the next available action. Use ARC Report when deeper evidence and interpretation are needed.