Skip to content
DocumentationIntelligenceWACI-1.0: Web & AI Crawler Intelligence
IntelligenceCurrent contractVersion DOCS-2.0

WACI-1.0: Web & AI Crawler Intelligence

Audit discovery, crawlability, indexability, semantic structure, AI readability and observable agent policies across owned or authorized sites.

How to interpret this document

This content describes technical and methodological behavior that is implemented or explicitly planned in the product. When a control depends on configuration, a provider, a secret, a contract or legal approval, that dependency must remain visible.

What WACI solves

WACI-1.0 connects technical site health to the AI Visibility lifecycle. The scanner looks for barriers that can make content harder to discover, read, interpret or use as evidence by search engines and AI systems.

Safe scope

The crawler accepts only HTTP/HTTPS, blocks local/private-network targets, limits pages and depth, validates redirects and keeps the baseline on the same origin. Owned or authorized targets must be explicitly registered.

Robots and agents

The product reads robots.txt and records observable rules for agent categories such as search, training and user-initiated access. Allowed, blocked or unspecified describes the observed directive and does not guarantee future provider behavior.

Per-page signals

Each page persists HTTP status, title, description, canonical, robots meta, H1/headings, text volume, links, JSON-LD, hreflang, Open Graph, Twitter Card, main-content presence, JavaScript-shell risk and indexability/crawlability indicators.

Crawler Readiness Score

WACI-1.0 combines discoverability, crawlability, indexability, semantic structure, AI readability, entity clarity, structured data and freshness. It is an internal technical readiness index and does not guarantee ranking, indexing or citation by any AI provider.

Findings

Issues are classified by category and severity. Examples include missing robots/sitemap, HTTP errors, missing canonical, noindex, inconsistent H1 structure, excessive JavaScript dependency and missing structured data where applicable.

Recommendation Management

High and Critical findings create or update REC-1.0 recommendations with fingerprint-based deduplication. A later re-scan provides fresh evidence to verify whether the technical condition changed.

llms.txt

llms.txt is treated as an auxiliary GEO/AEO signal, not as a universal indexing requirement or a crawler-control mechanism equivalent to robots.txt.