# https://www.machineread.ai/ --- title: "MachineRead - AI & Search Readiness Audit" description: "Audit public signals that help AI agents, retrieval systems, and search crawlers access and understand your site." source: "https://www.machineread.ai/" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` AI & agent visibility # See your site the way _machines_ do. Some of your readers are not people. They are assistants and retrieval systems fetching your pages to answer someone else's question, and they only get what your public surfaces actually hand them. MachineRead requests those surfaces the way an agent would, then reports what came back. GET/robots.txtai bot policy GET/llms.txtllm text access GET/.well-known/ai-catalog.jsonagent protocol discovery Three of the surfaces every audit requests. Thirteen check groups run on each scan. [Run a free audit](#audit-form)[What MachineRead measures](/about) public page robots.txtpresent/readable llms.txtmissing/needs attention structured datapresent/readable canonical URLmissing/needs attention **Illustrative signal map — not an audit result.** Status labels show how public evidence can be classified. - present/readable - missing/needs attention Run the audit ## Scan a URL Give it a public URL. It requests the same things an agent would: crawler policy, bot access, semantic HTML, structured data, text and Markdown access, freshness, discovery hints, and reports what came back. 13 check groups run on every scan. Anything it cannot verify is listed and scored at zero, never guessed at. 01 What MachineRead measures ## Three pillars of machine visibility Some of the traffic hitting your site has no eyes. It is a retrieval system pulling a page to answer a question someone asked somewhere else, and it succeeds or fails on what your markup and your public files actually hand it. MachineRead sorts every signal into three pillars: off-site presence, AI access, and search discovery. The full rubric caps them at 30, 40, and 30 points, for 100 in total. Those caps are allocation limits, not scores. Thirteen check groups run on every scan, worth 56 checked points between them. They cover social and entity metadata, Wikipedia and Wikidata lookups, robots.txt bot policy, bot fetch access, semantic HTML, JSON-LD structured data, LLM text and Markdown access, raw HTML readability, agent protocol discovery, crawl efficiency, canonical and HTTPS, indexing directives, and search discovery hints. Nine further rows are listed and locked at zero. They appear in the report so that what is not being measured stays visible rather than implied. Advanced coverage below itemizes each one. All of it measures a single thing: what public HTTP, DNS, and page metadata hand to a machine. Not what a private dashboard reports back to you. That distinction is the boundary on every claim in the report. **Pillar score caps.** Off-site 30, AI access 40, search 30 — 100-point total. Caps are rubric limits, not scores. 02 How the audit works ## From URL entry to scored report You hand it a URL and, if you want, a scope. FastAPI normalizes and validates that URL, and SSRF protection rejects unsafe targets and unsafe redirects before a single request leaves the server. The backend then fetches three anchors: the homepage, robots.txt, and sitemap.xml. Those documents ground the crawler-policy, bot-access, and discovery checks. If one of them fails, the failure is isolated into an actionable warning row rather than collapsing the whole report. Next the free checks run concurrently, each against a bounded slice of public HTTP, DNS, or page metadata. A check that fails stays failed on its own row and leaves the others alone. The backend also computes the strict agent-readiness summary from explicit discovery and protocol signals; if that step fails, the audit returns a degraded warning state instead of a dead request. Finally it assembles the result, locked rows at zero, scope metadata, pillar caps, benchmark comparison, and the dashboard renders the scores, caveats, and action items. The whole flow runs on free bounded public checks and one cached Wikimedia lookup. No logins, no paid APIs, no authenticated crawls. **Audit flow.** URL and scope input pass through safe bounded fetch, concurrent public checks, and scored report assembly. 03 What the scores mean ## Three scores, three denominators An audit returns three numbers, and they are not interchangeable. A site can hold a strong Essentials score and a weak agent-readiness score at the same time. That gap is usually the most useful thing in the report. The Full Rubric Score is the 100-point view across all three pillars. It counts the nine locked advanced rows at zero, because nothing has verified them. A site cannot reach 100 on Essentials alone, and that is deliberate. The Essentials Evidence Score ignores the locked rows entirely. It reads only the checks that actually ran, against a fixed 56-point maximum. Holding that maximum still is what makes the score comparable between sites and across time. The Strict Agent Readiness Score is the harshest of the three. It asks only whether explicit agent-native surfaces are present. Eight probes run by default; a fully scoped audit runs 21. The bands, Elite, Strong, Developing, At risk, are diagnostic labels, not grades. Developing means specific named signals are missing, and the report names them. **Score fields.** Each field with the denominator it is scored against. overall\_score / 100 benchmark.score / 56 checked agent\_readiness.score 8–21 probes band thresholds 85 / 70 / 50 04 What MachineRead does not claim ## The trust anchor Essentials only sees what a public request returns. It has no access to your analytics, no window into a search engine's index, and no telemetry from any model. That boundary is the point, so the report states it rather than blurring it. It does not measure ranking. Crawl access, indexing directives, and canonical tags are inputs a search engine may use. None of them tell you where a page sits on a results page. It does not measure traffic, backlinks, or conversion. Those live in private data that only you can see. It does not measure social presence. A site with no social links is not a company with no social presence, the audit reports what the page exposes, and nothing about the organization behind it. It does not measure real-user performance or model citation share. Publishing an agent surface is not evidence that a model read it. And it will not mark a site down for a protocol that does not apply to it. Scope toggles exist so a blog is not judged on checkout flows. **Out of scope.** Deliberately not measured. Listing them is the point. search ranking not measured traffic, conversion not measured backlinks not measured social traction not measured field performance not measured model citation share not measured 05 Benchmarks and peer context ## Relative position, not exposure A benchmark tells you where a site sits relative to peers audited under the same preset. It is context, not exposure data. Pick an unmodified named preset and the comparison runs against sites audited under that same preset. Choose Custom or Legacy, or override a named preset enough to change its family mix, and no like-for-like cohort exists. The report says so and falls back to disclosed cross-preset context, so you always know which pool you were placed in. The Essentials denominator stays fixed at 56 points across 13 rows. The agent-readiness denominator does not: it is derived from the families in scope. General presets run 8 probes, SaaS 17, and Ecommerce 21. A partial scope counts only the protocol, auth, and commerce probes you actually selected. You get a percentile, a median, and the peers nearest your evidence score. None of that claims comparable traffic, ranking, or citation share. It says only that these public signals, under this denominator, compare this way. Profiles are regenerated whenever scoring or applicability changes, and every report carries the snapshot date it was compared against. **Denominators.** Essentials is fixed. Agent readiness is derived from the families in scope. check groups 13 essentials max 56 pts probes · general 8 probes · saas 17 probes · ecommerce 21 06 Advanced coverage ## Nine rows, locked at zero Nine rows in every report are locked. They need logins, paid APIs, authenticated crawls, or private data, things a free audit cannot honestly perform. They score zero until something actually verifies them. They still appear in the report, and that is deliberate. A row you can see sitting at zero tells you what is not being measured. A row quietly omitted tells you nothing at all. Each one carries the Advanced state label and a badge naming the tier that would unlock it. Starter accounts for seven: earned mentions and backlinks, owned social presence, social traction and reviews, extraction fidelity, multi-engine index coverage, Core Web Vitals field data, and keyword and competitor gap analysis. Pro accounts for the remaining two, AI citation share and agent task simulation. When coverage is activated those rows unlock and start contributing to the full rubric score. The Essentials evidence score and its benchmark denominator do not move, so a report you ran last month still compares to one you run today. **Locked rows.** Shown in every report, scored at zero until verified. locked rows 9 score while locked 0 pts unlocked by Starter 7 rows unlocked by Pro 2 rows essentials denominator unchanged 07 Roadmap ## What is free, and why Essentials is free for a structural reason rather than a promotional one: nothing it does costs money to run. Bounded public HTTP, DNS, and page checks, one cached Wikimedia lookup, and a scoring pass. No account, no card, no trial clock. The nine locked rows are what the paid tiers buy, and they are itemized in Advanced coverage above rather than repeated here. Each one costs money on every audit, which is exactly why it stays at zero until it is paid for and verified. Planned and not yet built: accounts, billing, paid crawlers, paid search and model integrations, and bring-your-own-provider access. Accounts will be OAuth only, so no password is ever stored, and payment will run through Stripe and Link. There is no tracking today and there will be none on any tier later. The agent-consumable surfaces are already live. This site publishes the same documents it audits other sites for, which is the only honest way to ship a tool like this. The audit contract stays stable. New rows are additive: they unlock from the locked state, join the full rubric, and never move the Essentials denominator that past reports depend on. **This site's own surfaces.** MachineRead publishes what it audits others for. Every path is live. /llms.txt live /openapi.json live /.well-known/ai-catalog.json live /.well-known/api-catalog live accounts, billing planned Run audit ## Point it at a URL Free, no account, no card. Every locked row is shown at zero rather than left out. [Run audit](#audit-form)[Read the launch post](/blog/launching-machineread) ## Running an audit The website form works with or without JavaScript and returns an HTML report. Agents and JSON clients should call the HTTP API directly: - `POST https://api.machineread.ai/v1/audit` returns the full AuditResult with per-check evidence. - `POST https://api.machineread.ai/v1/audit/summary` returns a compact deterministic AuditSummary. The request schema for both endpoints is at https://www.machineread.ai/openapi.json. --- # https://www.machineread.ai/about --- title: "About - MachineRead" description: "MachineRead audits the public signals AI agents and search crawlers use to reach a site. Built and run by George Jieh." source: "https://www.machineread.ai/about" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` About # Measuring what machines can actually read A growing share of the traffic arriving at any website is not a person with a browser. It is a retrieval system fetching a page to answer someone's question, and unlike a human reader, it either parses the page or it does not. MachineRead exists to tell site owners which side of that line they are on. 1. 01 ## What the audit does The Essentials audit runs 13 check groups across 56 checked points, covering off-site presence, AI access and scrapability, and search discovery. It reports a separate strict agent-readiness score because a good SEO result can hide a fatal fetch failure. Nine advanced rows are listed but locked at zero, so the report shows what is not being measured instead of quietly leaving it out. 2. 02 ## What it refuses to claim The audit does not measure search ranking, traffic, backlinks, conversion, or model citation share. It measures the public signals that constrain retrieval, not the outcomes downstream of retrieval. Keeping that boundary is what makes the rest of the report trustworthy. The [methodology](/docs) documents every check and its weight. 3. 03 ## Who builds it MachineRead is built and operated by George Jieh. It runs on a deliberately small stack: a FastAPI backend, a static frontend on Cloudflare Pages, and no paid API calls during a free audit. That constraint is a design parameter, and it is why Essentials can stay free to run. 4. 04 ## Where it is going The locked rows require paid crawling and verification, so the Starter and Pro tiers that unlock them will involve accounts and payment. Accounts will be OAuth only, so no passwords are ever stored, and payment will run through Stripe and Link. There is no tracking today and there will be none later, on any tier. See [pricing](/pricing) for the current status. Questions, bug reports, or partnership inquiries: [contact@machineread.ai](mailto:contact@machineread.ai). --- # https://www.machineread.ai/blog --- title: "MachineRead Blog" description: "Deep dives into AI agent readability, machine-readable signals, and how to make your site discoverable by retrieval-augmented systems." source: "https://www.machineread.ai/blog" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` Latest entry2026-08-04 ## [Signals that help crawlers find updated public content](/blog/deep-dive-search-discovery) Search discovery checks public signals related to reachability, initial HTML content, and declared modification dates. search-discoveryaudit-pillardeep-dive 2026-08-04 ### [What a non-browser request can actually read](/blog/deep-dive-bot-access) A plain HTTP fetch can reveal whether public HTML is available before a client executes JavaScript. It cannot predict every crawler's behavior. bot-accessaudit-pillardeep-dive 2026-08-02 ### [A sitemap is a map, not an indexing guarantee](/blog/deep-dive-sitemap) A sitemap publishes URLs and optional modification dates for crawlers. It can support discovery, but it does not guarantee crawling or indexing. sitemapseodeep-dive 2026-08-02 ### [robots.txt is a public crawler policy](/blog/deep-dive-robots-txt) robots.txt expresses crawl directives for named user agents. AI-related rules should reflect a deliberate, reviewed policy. robots-txtseodeep-dive 2026-07-31 ### [Publishing page facts with Schema.org JSON-LD](/blog/deep-dive-schema-org) JSON-LD can publish page facts in structured fields. It reduces ambiguity for supporting clients but does not determine how they will use the data. schema-orgjson-lddeep-dive 2026-07-31 ### [Choosing one public URL for a page](/blog/deep-dive-canonical) A canonical tag expresses a preferred URL for duplicate or closely related pages. It can help search systems consolidate signals but does not dictate their behavior. canonicalseodeep-dive 2026-07-29 ### [How an OpenAPI document makes an API easier to inspect](/blog/deep-dive-openapi) An OpenAPI document describes an HTTP API in a machine-readable form. It can help a supporting client construct requests, but it does not authorize or validate their use. openapiagent-surfacesdeep-dive 2026-07-29 ### [A standard route to an API's public description](/blog/deep-dive-api-catalog) RFC 9727 defines a well-known URI for an API catalog expressed as a linkset. It gives supporting clients a predictable discovery path. api-catalogagent-surfacesdeep-dive 2026-07-27 ### [llms.txt is optional, not a visibility shortcut](/blog/deep-dive-llms-txt) llms.txt is a proposed discovery file for language-model tools. A MachineRead result describes the response and basic shape, not client support or use. llms-txtagent-surfacesdeep-dive 2026-07-27 ### [What an AI catalog can tell a supporting client](/blog/deep-dive-ai-catalog) An AI catalog can state what a service does and where its machine-readable surfaces live. The format is early, so the audit gives it limited weight. ai-catalogagent-surfacesdeep-dive 2026-07-23 ### [Why Agents Need to Read Your Site](/blog/why-agents-need-to-read-your-site) Agent retrieval depends on what a client can fetch and parse. A page that renders for people may expose less to a non-browser client. ai-agentsdiscoveryretrieval 2026-07-23 ### [Launching MachineRead](/blog/launching-machineread) MachineRead audits a public website for observable AI agent and search-readiness signals across 13 check groups and 56 checked points. launchproductai-readiness --- # https://www.machineread.ai/blog/deep-dive-ai-catalog --- title: "What an AI catalog can tell a supporting client" description: "An AI catalog can state what a service does and where its machine-readable surfaces live. The format is early, so the audit gives it limited weight." source: "https://www.machineread.ai/blog/deep-dive-ai-catalog" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "Article", "headline": "What an AI catalog can tell a supporting client", "description": "An AI catalog can state what a service does and where its machine-readable surfaces live. The format is early, so the audit gives it limited weight.", "datePublished": "2026-07-27", "dateModified": "2026-07-27", "author": { "@type": "Person", "name": "George Jieh", "url": "https://github.com/georgejieh" }, "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" }, "mainEntityOfPage": "https://www.machineread.ai/blog/deep-dive-ai-catalog" } ``` Blog # What an AI catalog can tell a supporting client An AI catalog can state what a service does and where its machine-readable surfaces live. The format is early, so the audit gives it limited weight. author [George Jieh](https://github.com/georgejieh) published 2026-07-27 reading 5 min For automated discovery, an AI catalog matters only to software that knows how to read it. Its practical job is to provide a typed inventory of artifacts that a publisher has listed (Agent Card Working Group, "AI Catalog"). It is not evidence that a registry indexed those artifacts, that a client selected them, or that any tool call succeeded. The specifications around `ai-catalog.json` are still moving (Agent Card Working Group, "Common AI Catalog"; Bu, Guha, and Smith). The file can be worthwhile as publisher-maintained documentation, but its presence should not be promoted as a general discovery guarantee. ## Two documents, two jobs The AI Catalog specification defines a typed, nestable JSON container. Each entry identifies an artifact, states its type, and either links to the artifact document or embeds its data. The catalog does not replace formats such as an MCP Server Card or an A2A Agent Card. It provides an envelope around those native descriptions (Agent Card Working Group, "AI Catalog"). The ARD v0.9 draft proposes a broader search and registry layer around that catalog. Its current document marks its status as "Proposal" (Bu, Guha, and Smith). The AI Catalog repository also calls itself a working repository under temporary governance. It says possible adoption votes by the MCP and A2A steering committees would follow finalization of the specification, not precede it (Agent Card Working Group, "Common AI Catalog"). Those are primary-source status statements. They support calling the work an emerging specification. They do not support claims of broad implementation or adoption. The path can cause another kind of confusion. RFC 8615 reserves `/.well-known/` as a path prefix for registered well-known locations, but it does not define the format found at every name under that prefix (Nottingham, sec. 3). The AI Catalog specification makes `/.well-known/ai-catalog.json` an optional discovery location and permits catalogs at other URLs or in other distribution forms (Agent Card Working Group, "AI Catalog"). The RFC status of the prefix does not make the AI Catalog specification an IETF standard, and the familiar path does not establish client support. ## What the document can describe In the current AI Catalog specification, the top-level document has a `specVersion` and an `entries` array. An entry has an identifier and type, plus exactly one of `url` or `data`; host information and trust metadata are separate, optional layers (Agent Card Working Group, "AI Catalog"). The ARD v0.9 proposal adds discovery-oriented fields and describes registry behavior (Bu, Guha, and Smith). The version field is important in a changing format. A publisher should validate against the version it actually implements rather than copy fields from a newer example into an older document. A small catalog with two accurate entries is easier to inspect than a long inventory assembled from stale deployment records. The core fields are publisher-supplied statements. A description says what a resource is intended to do. A URL says where its native document is supposed to live. Optional trust material provides inputs for further checks, but its presence is not verification. The specification says that even a verified signature without an independently anchored identity establishes integrity and internal consistency only, not publisher authenticity (Agent Card Working Group, "AI Catalog"). At the Trusted Catalog level, a consumer must verify the signature, anchor the identity, and confirm the subject digest before relying on a claim (Agent Card Working Group, "AI Catalog"). MachineRead does not perform those trust checks. Syntax alone cannot establish availability, authorization, safety, or fitness for a task. ## How to maintain a useful catalog Start with resources that are already public and documented. Give each one a stable identifier, identify its artifact format with the `type` field, and point to the canonical resource document. Update or remove retired entries rather than leave stale URLs and descriptions. Then test the catalog as an independent client would see it: - Request the published URL without a browser session. - Confirm that the response is JSON and declares the version you intend to support. - Validate every required field against that version of the specification. - Fetch each referenced public URL and compare its name, purpose, and version, where those fields exist, with the catalog entry. - Keep credentials, internal hostnames, and private operational details out of the public file. When protocol checks are selected, MachineRead treats the catalog as one optional discovery signal within its broader machine-surface check. The audit can report whether it received parseable JSON and whether selected fields meet MachineRead's limited structural expectations. It does not query an ARD registry, connect to a listed resource, verify catalog signatures or publisher identity, or infer that a provider consumes the file. ## What this check can and cannot establish The check can establish that the audited origin returned a candidate catalog at the checked path under the audit's request conditions. It can inspect JSON parsing, the declared version shape, selected host and entry fields, and whether a `trustManifest` key appears. Presence is reported as metadata, not verified. It cannot establish that a registry has indexed the catalog, that a named client supports the specification, or that a listed resource is safe and usable. It also cannot predict search ranking, referral traffic, model citations, resource selection, or successful execution. Those later outcomes are not properties of the catalog response. ## Works Cited Agent Card Working Group. "AI Catalog." _AI Catalog_, Linux Foundation, [https://ai-catalog.io/spec/](https://ai-catalog.io/spec/). Accessed 7 Aug. 2026. Agent Card Working Group. "Common AI Catalog and Registry Standard." _GitHub_, Linux Foundation, [https://github.com/Agent-Card/ai-catalog](https://github.com/Agent-Card/ai-catalog). Accessed 7 Aug. 2026. Bu, Junjie, R. V. Guha, and Shaun Smith. "Agentic Resource Discovery Specification." _AgenticResourceDiscovery.org_, version 0.9, 28 May 2026, [https://agenticresourcediscovery.org/spec/](https://agenticresourcediscovery.org/spec/). Accessed 7 Aug. 2026. Nottingham, Mark. "Well-Known Uniform Resource Identifiers (URIs)." _RFC 8615_, Internet Engineering Task Force, May 2019, RFC Editor, [https://www.rfc-editor.org/rfc/rfc8615.html](https://www.rfc-editor.org/rfc/rfc8615.html). Accessed 7 Aug. 2026. ## See also - [Methodology reference](/docs/methodology) - documents the ai-catalog probe and the machine\_surfaces check group - [Launch post](/blog/launching-machineread) - explains the audit's evidence boundaries - Related: [api-catalog.json](/blog/deep-dive-api-catalog) and [llms.txt](/blog/deep-dive-llms-txt) ai-catalogagent-surfacesdeep-dive [Back to all posts](/blog) --- # https://www.machineread.ai/blog/deep-dive-api-catalog --- title: "A standard route to an API's public description" description: "RFC 9727 defines a well-known URI for an API catalog expressed as a linkset. It gives supporting clients a predictable discovery path." source: "https://www.machineread.ai/blog/deep-dive-api-catalog" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "Article", "headline": "A standard route to an API's public description", "description": "RFC 9727 defines a well-known URI for an API catalog expressed as a linkset. It gives supporting clients a predictable discovery path.", "datePublished": "2026-07-29", "dateModified": "2026-07-29", "author": { "@type": "Person", "name": "George Jieh", "url": "https://github.com/georgejieh" }, "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" }, "mainEntityOfPage": "https://www.machineread.ai/blog/deep-dive-api-catalog" } ``` Blog # A standard route to an API's public description RFC 9727 defines a well-known URI for an API catalog expressed as a linkset. It gives supporting clients a predictable discovery path. author [George Jieh](https://github.com/georgejieh) published 2026-07-29 reading 4 min An API description is useful only after a client finds it. RFC 9727 addresses that first step with a stable address for an API catalog (Smith, secs. 1.1-2). Its contribution is modest and specific: a publisher can place typed links to public API resources behind a route that a supporting client already knows how to request (Smith, secs. 3-4.1). The catalog does not prove that an API works, that its description is accurate, or that any particular client will use it. For a site owner, an API catalog is an index, not an integration. ## The route is predictable; the API is not RFC 8615 reserves the `/.well-known/` path prefix so individual specifications can register stable locations for metadata. It does not define a generic resource at `/.well-known/`, and it leaves the relevant hostname and the scope of discovered metadata to the application-specific specification (Nottingham, secs. 1, 3). RFC 9727 supplies those missing details for API catalogs. A conforming publisher resolves an HTTPS GET request to `/.well-known/api-catalog` to an API catalog document. It also resolves an HTTPS HEAD request at that route with a `Link` header carrying the relations defined by the RFC (Smith, sec. 2). The underlying catalog may live elsewhere; the well-known route is the stable reference to it. The label "well-known" is easy to overread. It means the location has a registered convention, not that every crawler, agent, SDK, or search engine requests it. A client still has to implement the convention, choose the correct origin, retrieve the response, and decide what to do with the links. ## Read the catalog as a set of typed links RFC 9727 requires the catalog to include links to API endpoints. It recommends adding useful metadata such as usage policies, API versions, and links to OpenAPI descriptions, either in the catalog or at the listed endpoint resources (Smith, sec. 4.1). The required representation is a JSON Linkset served as `application/linkset+json`; other negotiated formats are optional additions, not substitutes for that representation (Smith, secs. 4.2, 6.2). In the JSON Linkset model, a top-level `linkset` array contains link-context objects. Each relation name maps to an array of link-target objects, and each target has an `href` (Wilde and Van de Sompel, secs. 4.2.1-4.2.3). Combining two patterns from RFC 9727, a compact catalog can link the catalog to an API endpoint, then link that endpoint to its machine-readable description and human documentation (Smith, apps. A.1-A.2): ```json { "linkset": [ { "anchor": "https://example.com/.well-known/api-catalog", "item": [{ "href": "https://api.example.com/v1" }] }, { "anchor": "https://api.example.com/v1", "service-desc": [{ "href": "https://api.example.com/openapi.json" }], "service-doc": [{ "href": "https://example.com/docs/api" }] } ] } ``` The relation names make those roles explicit. A client does not have to guess whether one URL is an endpoint, a contract, or a guide from its filename. Even so, the linked resources remain separate statements by the publisher. A stale `service-desc` link is still stale, and a listed endpoint is not evidence that an unauthenticated caller may use it. ## Publication is a maintenance task A catalog should change with the API portfolio. RFC 9727 recommends removing stale entries as part of the release lifecycle, checking syntax and metadata, and reviewing a public catalog so it does not expose private APIs or sensitive information (Smith, secs. 5.4, 8). Those are maintenance practices, not discoverability tricks. For a small API, the practical work is equally small: serve the required Linkset representation, expose the well-known route over HTTPS, include only intentionally public endpoints, link to current descriptions and policies, test GET and HEAD behavior, and keep the catalog in the same release process as the API documentation. ## What this check can and cannot establish When API or protocol checks are in scope, MachineRead requests `/.well-known/api-catalog`, `/.well-known/api-catalog.json`, and `/api-catalog.json`. It records an API catalog candidate when a successful response contains at least 20 characters after trimming whitespace and parses as a JSON object or array. Separately, it records a general discovery signal when the audited target page's HTTP `Link` header contains one of several configured tokens, including `api-catalog`. These are deliberately narrow discovery signals. The check does not perform a complete RFC 9727 conformance test. It does not establish that the server implements the required HEAD response, sends the required Linkset media type, uses every relation correctly, or keeps each target current. It also does not test API authentication, authorization, rate limits, runtime responses, security, or client support for the catalog. A missing result means the audit did not observe this public signal under its request conditions, not that the publisher has no API. ## Works Cited Nottingham, Mark. "Well-Known Uniform Resource Identifiers (URIs)." _RFC 8615_, Internet Engineering Task Force, May 2019, RFC Editor, [https://www.rfc-editor.org/rfc/rfc8615.html](https://www.rfc-editor.org/rfc/rfc8615.html). Accessed 7 Aug. 2026. Smith, Kevin. "api-catalog: A Well-Known URI and Link Relation to Help Discovery of APIs." _RFC 9727_, Internet Engineering Task Force, June 2025, RFC Editor, [https://www.rfc-editor.org/rfc/rfc9727.html](https://www.rfc-editor.org/rfc/rfc9727.html). Accessed 7 Aug. 2026. Wilde, Erik, and Herbert Van de Sompel. "Linkset: Media Types and a Link Relation Type for Link Sets." _RFC 9264_, Internet Engineering Task Force, July 2022, RFC Editor, [https://www.rfc-editor.org/rfc/rfc9264.html](https://www.rfc-editor.org/rfc/rfc9264.html). Accessed 7 Aug. 2026. ## See also - [Methodology reference](/docs/methodology) - documents the api-catalog probe and the machine\_surfaces check group - [Launch post](/blog/launching-machineread) - explains the audit's evidence boundaries - Related: [openapi.json](/blog/deep-dive-openapi) and [ai-catalog.json](/blog/deep-dive-ai-catalog) api-catalogagent-surfacesdeep-dive [Back to all posts](/blog) --- # https://www.machineread.ai/blog/deep-dive-bot-access --- title: "What a non-browser request can actually read" description: "A plain HTTP fetch can reveal whether public HTML is available before a client executes JavaScript. It cannot predict every crawler's behavior." source: "https://www.machineread.ai/blog/deep-dive-bot-access" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "Article", "headline": "What a non-browser request can actually read", "description": "A plain HTTP fetch can reveal whether public HTML is available before a client executes JavaScript. It cannot predict every crawler's behavior.", "datePublished": "2026-08-04", "dateModified": "2026-08-04", "author": { "@type": "Person", "name": "George Jieh", "url": "https://github.com/georgejieh" }, "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" }, "mainEntityOfPage": "https://www.machineread.ai/blog/deep-dive-bot-access" } ``` Blog # What a non-browser request can actually read A plain HTTP fetch can reveal whether public HTML is available before a client executes JavaScript. It cannot predict every crawler's behavior. author [George Jieh](https://github.com/georgejieh) published 2026-08-04 reading 4 min A page can work perfectly in a browser and still give a very different answer to a simple HTTP request. That difference matters when a site depends on client-side rendering, sends a challenge page to unfamiliar clients, or delivers an empty application shell before JavaScript runs. MachineRead's bot-access check stays deliberately narrow. It attempts a browser-identifying baseline request and otherwise reuses the initial audit fetch as its comparison baseline. It then makes a paced set of requests carrying tracked crawler user-agent strings and compares what they receive at audit time. A passing result does not certify access for an official crawler, whose network identity or treatment may differ from MachineRead's probe. ## The browser is not a useful stand-in for every client Google documents a rendering system that can process JavaScript before indexing content (Google). That capability is real, but it is specific to Google's crawler stack. It should not be used as a proxy for every system that fetches a URL. A Vercel and MERJ study published in 2024 reported no JavaScript rendering by the named OpenAI, Anthropic, Meta, ByteDance, and Perplexity crawlers in its tests (Zecchini et al.). It separately observed traffic the authors attributed to ChatGPT and Claude fetching some JavaScript resources without executing them, and noted that content included in the initial response could still be available (Zecchini et al.). This is a time-bound observation about the clients and sites in that study, not a rule for automated retrieval generally. That is why the check starts with the raw response. It can be accessible and still be unhelpful if it contains only a small root element and script references. Conversely, a page may be fully usable to a browser-rendering crawler while a plain HTTP client sees little beyond the shell. Neither observation establishes what every crawler, search engine, or model will do next. ## What the check observes The probe uses plain HTTP GET requests with crawler-identifying user-agent strings. Its comparison baseline comes from an attempted browser-identifying request or, if that request fails, the initial audit fetch. It does not execute JavaScript, run a browser renderer, or solve a challenge. It retries an HTTP 429 response once after a delay, but it does not fall back to a browser for the crawler probes. The result combines several observations: - Did a probe end in a fetch error, selected blocking status, or recognized challenge response? - Was the visible text much thinner than the comparison baseline? - Did the status, final URL, canonical URL, response size, or visible word count differ materially from that baseline? A denial, challenge page, persistent rate limit, or thin initial document gives a site owner a specific condition to inspect. It does not mean automated access is categorically wrong. Some sites intentionally restrict it. The question is whether the observed behavior matches the site's publication and discovery goals. ## What this check can and cannot establish The check can establish what MachineRead received for its comparison baseline and tracked user-agent requests at one point in time. It can flag selected response failures, challenge fingerprints, thin content, and routing differences under those conditions. It cannot authenticate an official crawler from a user-agent string, reproduce provider-specific IP reputation or request history, or show that another client will receive the same response. It also cannot establish indexing, training use, ranking, citation, referral, or model behavior. ## How to use a finding When the check reports limited content, fetch the affected URL with a non-browser user agent and inspect the response headers and opening HTML. If critical public text appears only after JavaScript runs, decide whether that is appropriate for the audience you want to reach. If a security control blocks the request, confirm whether the policy is intentional. Record whether the behavior is intentional, then repeat the request after any rendering, hosting, or security change. The evidence remains the response MachineRead received, not a forecast about a crawler's next action. ## Works Cited Google. "Understand the JavaScript SEO Basics." _Google Search Central_, updated 4 Mar. 2026, [https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics). Accessed 7 Aug. 2026. Zecchini, Giacomo, et al. "The Rise of the AI Crawler." _Vercel_, 17 Dec. 2024, [https://vercel.com/blog/the-rise-of-the-ai-crawler](https://vercel.com/blog/the-rise-of-the-ai-crawler). Accessed 7 Aug. 2026. ## See also - [Methodology reference](/docs/methodology) - documents the bot-access check group - [Launch post](/blog/launching-machineread) - explains the audit's evidence boundaries - Related: [robots.txt](/blog/deep-dive-robots-txt) and [why agents need to read your site](/blog/why-agents-need-to-read-your-site) bot-accessaudit-pillardeep-dive [Back to all posts](/blog) --- # https://www.machineread.ai/blog/deep-dive-canonical --- title: "Choosing one public URL for a page" description: "A canonical tag expresses a preferred URL for duplicate or closely related pages. It can help search systems consolidate signals but does not dictate their behavior." source: "https://www.machineread.ai/blog/deep-dive-canonical" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "Article", "headline": "Choosing one public URL for a page", "description": "A canonical tag expresses a preferred URL for duplicate or closely related pages. It can help search systems consolidate signals but does not dictate their behavior.", "datePublished": "2026-07-31", "dateModified": "2026-07-31", "author": { "@type": "Person", "name": "George Jieh", "url": "https://github.com/georgejieh" }, "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" }, "mainEntityOfPage": "https://www.machineread.ai/blog/deep-dive-canonical" } ``` Blog # Choosing one public URL for a page A canonical tag expresses a preferred URL for duplicate or closely related pages. It can help search systems consolidate signals but does not dictate their behavior. author [George Jieh](https://github.com/georgejieh) published 2026-07-31 reading 4 min A canonical tag is useful when it records a URL decision that the rest of the site already supports. On its own, it is only a publisher-controlled preference. The important question is not whether the element exists, but whether redirects, host choices, internal links, and sitemap entries tell the same story (Google, "How to Specify a Canonical URL"). URL consistency is observable. Downstream canonical selection is not. ## Canonicalization is a selection process The same or very similar content can be reachable through protocol variants, regional versions, sorting or filtering URLs, and accidental duplicates. Google defines canonicalization as choosing a representative URL from such a set and notes that some duplicate content is normal rather than a violation of its spam policies (Google, "What Is URL Canonicalization"). A site can state its preference through a redirect, a `rel="canonical"` link annotation, or sitemap inclusion. Google describes redirects and `rel="canonical"` as strong signals and sitemap inclusion as weaker, and says that the signals can reinforce one another (Google, "How to Specify a Canonical URL"). Google may still select a different canonical from the publisher's preference (Google, "What Is URL Canonicalization"). This is documented Google behavior, not a claim about every crawler or retrieval system. The difference between declaration and outcome matters. A tag can say that `https://example.com/guide` is preferred while an HTTP URL remains live, the sitemap lists a parameterized version, and internal navigation points somewhere else. The declaration is easy to observe. Whether a particular system groups those pages together, accepts the preference, or uses either URL later is a separate decision. ## One public address should have supporting evidence For HTML pages, Google documents the canonical element in the document head and recommends a self-referential element on the preferred page. It also recommends absolute URLs and consistent internal links to the canonical URL (Google, "How to Specify a Canonical URL"). A canonical element does not redirect a visitor, so duplicate URLs that should disappear still need an appropriate redirect (Google, "How to Specify a Canonical URL"). Start with the content relationship. If two URLs really represent the same page, choose the address you intend to maintain and make the surrounding signals consistent. Google's canonical methods apply to duplicate or very similar pages, not pages that present materially different content (Google, "How to Specify a Canonical URL"). ## What MachineRead observes MachineRead evaluates a limited slice of this configuration on the audited target page. It looks for a canonical link in the fetched HTML and compares it with the final target URL after normalizing query strings, fragments, and trailing slashes. It also probes the HTTP version and the corresponding `www` or non-`www` host, reporting a mismatch when the returned probe evidence shows a route that does not converge as expected. This is a consistency check on the response available during the audit. It does not crawl a site's duplicate clusters or decide which pages are similar enough to share a canonical. A finding is best used as a prompt to inspect the redirect chain, raw HTML, templates, internal links, and sitemap together. ## What this check can and cannot establish The check can establish whether the audited response exposes a canonical that matches the final target URL under MachineRead's normalization. It can also report an HTTP or alternate-host mismatch when a returned probe supplies evidence of one. It cannot establish convergence for a probe that failed or returned no usable response. It also cannot establish the canonical selected by Google or another system, validate site-wide canonical coverage, inspect non-HTML `Link` headers, decide whether two pages are equivalent, or predict indexing, ranking, traffic, citations, or later agent behavior. ## A practical review sequence Choose the intended public URL for the target page first. Then inspect the complete redirect chain for HTTP and the alternate host, confirm the canonical in the initial HTML, and compare that value with internal links and sitemap entries. Repeat the review on representative templates, especially pages with parameters or regional variants. Where a mismatch is intentional, document the content relationship instead of forcing every URL into a self-reference. When these pieces disagree, fix the public evidence that no longer reflects the intended URL: the redirect, canonical element, internal link, or sitemap entry. MachineRead can surface the contradiction, but the site owner still decides which URLs are duplicates and which address should represent them. ## Works Cited Google. "How to Specify a Canonical URL with rel='canonical' and Other Methods." _Google Search Central_, updated 10 July 2026, [https://developers.google.com/search/docs/crawling-indexing/consolidate-duplicate-urls](https://developers.google.com/search/docs/crawling-indexing/consolidate-duplicate-urls). Accessed 7 Aug. 2026. Google. "What Is URL Canonicalization." _Google Search Central_, updated 10 July 2026, [https://developers.google.com/search/docs/crawling-indexing/canonicalization](https://developers.google.com/search/docs/crawling-indexing/canonicalization). Accessed 7 Aug. 2026. ## See also - [Methodology reference](/docs/methodology) - documents the canonical check group - [Launch post](/blog/launching-machineread) - explains the audit's evidence boundaries - Related: [schema.org JSON-LD](/blog/deep-dive-schema-org) and [sitemap](/blog/deep-dive-sitemap) canonicalseodeep-dive [Back to all posts](/blog) --- # https://www.machineread.ai/blog/deep-dive-llms-txt --- title: "llms.txt is optional, not a visibility shortcut" description: "llms.txt is a proposed discovery file for language-model tools. A MachineRead result describes the response and basic shape, not client support or use." source: "https://www.machineread.ai/blog/deep-dive-llms-txt" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "Article", "headline": "llms.txt is optional, not a visibility shortcut", "description": "llms.txt is a proposed discovery file for language-model tools. A MachineRead result describes the response and basic shape, not client support or use.", "datePublished": "2026-07-27", "dateModified": "2026-07-27", "author": { "@type": "Person", "name": "George Jieh", "url": "https://github.com/georgejieh" }, "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" }, "mainEntityOfPage": "https://www.machineread.ai/blog/deep-dive-llms-txt" } ``` Blog # llms.txt is optional, not a visibility shortcut llms.txt is a proposed discovery file for language-model tools. A MachineRead result describes the response and basic shape, not client support or use. author [George Jieh](https://github.com/georgejieh) published 2026-07-27 reading 4 min `llms.txt` is a publishing proposal (Howard). It is not a visibility shortcut. A site owner can use it as a compact Markdown guide to important public material for tools that choose to read it (Howard). The file does not show that any particular crawler requested it, understood it, or used its links. The proposal was published by Jeremy Howard on 3 September 2024 and remains open to community input (Howard). Publishing a conforming file does not by itself demonstrate provider support. ## What the proposal defines The proposed root location is `/llms.txt`, although the format also permits a subpath (Howard). The document is Markdown, begins with an H1 naming the site or project, and may continue with a short blockquote summary, explanatory text, and H2 sections containing lists of links (Howard). An H2 named `Optional` marks secondary links that can be omitted when a shorter context is needed (Howard). The proposal also suggests offering clean Markdown versions of useful pages and leaves the processing of `llms.txt` to each application (Howard). The format stops there: it tells a publisher how to organize a file but does not require any remote system to fetch or act on it. A useful `llms.txt` is therefore closer to a curated documentation index than a second copy of the site. It can name the product, explain what the linked material covers, and direct a supporting client to canonical documentation. The linked pages still need to be accurate and independently accessible. ## Keep discovery separate from policy The proposal distinguishes `llms.txt` from both sitemaps and `robots.txt`: it presents the file as a curated overview, while those other files serve different discovery and crawler-policy purposes (Howard). The Robots Exclusion Protocol defines rules that crawlers are requested to honor, and it explicitly states that those rules are not access authorization (Koster et al., sec. 1). `llms.txt` is not an alternative crawler-control or security mechanism. Keeping those roles separate avoids two common mistakes. Do not put sensitive links in `llms.txt` on the theory that only selected tools will read it. Treat the file as public. Do not move essential policy or documentation into the file while leaving the ordinary site incomplete. Keep public pages and navigation complete, and maintain other discovery files for the clients that use them. ## A maintenance test is more useful than a publication claim Before publishing, decide what job the file serves. Documentation-heavy sites might use it to point to an overview, tutorials, API references, and change logs. A smaller brochure site may have no need for an extra index. For a file you do publish: - Use one clear H1 and a short factual summary. - Link to canonical public URLs rather than temporary exports or private environments. - Describe why each link is useful instead of repeating keywords. - Put secondary material in the `Optional` section. - Check links after documentation moves and remove stale entries. - Keep the file consistent with the facts readers see on the linked pages. MachineRead evaluates `llms.txt` as part of a broader LLM Text and Markdown Access check. It looks at the response and basic document shape alongside sitemap and Markdown-access evidence. The result reports which candidate text and discovery responses were returned during the audit and met its basic checks. It does not test a provider's production crawler or model. ## What this check can and cannot establish The check can establish whether the audited site returned a candidate `/llms.txt` response under the tested conditions and whether that response starts with a hash character (`#`) under its basic shape heuristic and contains a URL or Markdown link. The broader check can also report related sitemap and Markdown-access observations. It cannot establish that a named crawler supports the proposal, that the file was consumed, or that its links influenced retrieval. It cannot predict ranking, traffic, citations, model answers, or referral behavior. It also cannot turn public descriptions into access control or verify every factual statement on every linked page. ## Works Cited Howard, Jeremy. "The /llms.txt File." _llms-txt_, 3 Sept. 2024, [https://llmstxt.org/](https://llmstxt.org/). Accessed 7 Aug. 2026. Koster, Martijn, et al. "Robots Exclusion Protocol." _RFC 9309_, Internet Engineering Task Force, Sept. 2022, RFC Editor, [https://www.rfc-editor.org/rfc/rfc9309.html](https://www.rfc-editor.org/rfc/rfc9309.html). Accessed 7 Aug. 2026. ## See also - [Methodology reference](/docs/methodology) - documents the llms.txt probe and all 13 check groups - [Launch post](/blog/launching-machineread) - explains the audit's evidence boundaries - [Agent integration guide](/docs/agent-integration) - details the surface family - Related: [schema.org JSON-LD](/blog/deep-dive-schema-org) and [ai-catalog.json](/blog/deep-dive-ai-catalog) llms-txtagent-surfacesdeep-dive [Back to all posts](/blog) --- # https://www.machineread.ai/blog/deep-dive-openapi --- title: "How an OpenAPI document makes an API easier to inspect" description: "An OpenAPI document describes an HTTP API in a machine-readable form. It can help a supporting client construct requests, but it does not authorize or validate their use." source: "https://www.machineread.ai/blog/deep-dive-openapi" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "Article", "headline": "How an OpenAPI document makes an API easier to inspect", "description": "An OpenAPI document describes an HTTP API in a machine-readable form. It can help a supporting client construct requests, but it does not authorize or validate their use.", "datePublished": "2026-07-29", "dateModified": "2026-07-29", "author": { "@type": "Person", "name": "George Jieh", "url": "https://github.com/georgejieh" }, "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" }, "mainEntityOfPage": "https://www.machineread.ai/blog/deep-dive-openapi" } ``` Blog # How an OpenAPI document makes an API easier to inspect An OpenAPI document describes an HTTP API in a machine-readable form. It can help a supporting client construct requests, but it does not authorize or validate their use. author [George Jieh](https://github.com/georgejieh) published 2026-07-29 reading 4 min An OpenAPI description turns parts of an HTTP API contract into data that software can inspect (OpenAPI Initiative, sec. 2). That is its practical value. It replaces some guesswork about paths, operations, inputs, outputs, and security declarations with fields that can be parsed and compared (OpenAPI Initiative, secs. 4.8-4.10, 4.27). It does not turn the declaration into runtime truth, grant permission to call the service, or decide whether an automated action is appropriate. Finding an `openapi.json` file is only the first check. The published description also needs to be detailed, current, and clearly connected to the API it claims to describe. ## What the document actually describes The OpenAPI Specification defines a programming-language-neutral interface description for HTTP APIs. Its stated purpose is to let people and computers understand a service's capabilities without source-code access or traffic inspection (OpenAPI Initiative, sec. 2). A conforming OpenAPI document is a JSON object represented as JSON or YAML, and it identifies the specification version used to interpret its fields (OpenAPI Initiative, secs. 3, 4.1.1). The structure moves from the API down to individual operations. A `servers` array can identify base URLs. A `paths` object holds relative endpoint paths, a Path Item associates HTTP methods with those paths, and an Operation Object can describe parameters, a request body, possible responses, deprecation, and operation-level security requirements (OpenAPI Initiative, secs. 4.5, 4.8-4.10). Reusable schemas, parameters, responses, and security schemes can live under `components` (OpenAPI Initiative, sec. 4.7). Those fields make a declared contract inspectable. A reviewer can ask whether an operation has a stable identifier, whether path parameters are defined, whether error responses have schemas, and whether a deprecated operation is marked as such. Software can perform the same structural checks without trying to infer an interface from prose. ## Parseable is not the same as complete OpenAPI permits descriptions with limited visible detail. In version 3.2.0, the root must include at least one of `components`, `paths`, or `webhooks`, but a `paths` object may be empty because access controls can limit what a viewer sees (OpenAPI Initiative, secs. 4.1.1, 4.8). Many descriptive fields, including operation summaries and descriptions, are optional (OpenAPI Initiative, secs. 4.9.1, 4.10.1). A file can therefore satisfy basic structural requirements while remaining difficult to use. Accuracy is a separate question. The document reports what the publisher declares. It does not execute an operation, observe a production response, or compare the service with the schema. A `securitySchemes` entry describes an available mechanism, while a Security Requirement Object declares which schemes an operation requires; neither object supplies credentials or grants authorization (OpenAPI Initiative, secs. 4.27, 4.30). Here the description stops and runtime behavior begins. A detailed document can reduce ambiguity for documentation tools, client developers, and other supporting software. It cannot ensure that a caller chooses the right operation, provides valid real-world values, handles side effects, or follows the service's policies. ## Description and discovery are different jobs The specification recommends naming an entry document `openapi.json` or `openapi.yaml`, but that is a naming recommendation rather than a universal root URL (OpenAPI Initiative, sec. 4.1.2). A publisher still needs to link the description from documentation or expose a separate discovery mechanism. RFC 9727's API catalog is one such mechanism. It defines `/.well-known/api-catalog` and recommends that catalog entries include links to OpenAPI descriptions where relevant (Smith, secs. 2, 4.1). The catalog answers "where is the description?" OpenAPI answers "what interface is being declared?" Keeping those roles separate makes both files easier to maintain. For publishers, a sound workflow is straightforward. Generate what can be derived from code, then review the public result. Add operation summaries that distinguish similar actions, document required inputs, include realistic non-secret examples, describe expected error responses, mark deprecated elements, and remove internal-only routes. Validate the document against the OAS version it declares, and update it in the same release process as the API. ## What this check can and cannot establish When API or protocol checks are selected, MachineRead requests `/openapi.json`, `/api/openapi.json`, and `/swagger.json`. It records an API-description candidate when a successful response contains at least 20 characters after trimming whitespace and parses as a JSON object or array. Although OpenAPI permits YAML, this probe looks for JSON at those three paths and does not count a YAML-only description. This is a reachability and parseability signal from the audited origin under the audit's request conditions. The check does not validate the response against an OpenAPI schema, follow every reference, grade documentation quality, or compare declared operations with live behavior. It does not call the listed operations, test credentials or authorization, assess side effects, certify API security, or determine whether any particular client can use the description correctly. A missing signal also does not prove that no description exists; it may be published at a location outside the paths the audit checks. ## Works Cited OpenAPI Initiative. "OpenAPI Specification v3.2.0." _OpenAPI Initiative_, Linux Foundation, 19 Sept. 2025, [https://spec.openapis.org/oas/v3.2.0.html](https://spec.openapis.org/oas/v3.2.0.html). Accessed 7 Aug. 2026. Smith, Kevin. "api-catalog: A Well-Known URI and Link Relation to Help Discovery of APIs." _RFC 9727_, Internet Engineering Task Force, June 2025, RFC Editor, [https://www.rfc-editor.org/rfc/rfc9727.html](https://www.rfc-editor.org/rfc/rfc9727.html). Accessed 7 Aug. 2026. ## See also - [Methodology reference](/docs/methodology) - documents the openapi.json probe and the machine\_surfaces check group - [Launch post](/blog/launching-machineread) - explains the audit's evidence boundaries - Related: [api-catalog.json](/blog/deep-dive-api-catalog) and [ai-catalog.json](/blog/deep-dive-ai-catalog) openapiagent-surfacesdeep-dive [Back to all posts](/blog) --- # https://www.machineread.ai/blog/deep-dive-robots-txt --- title: "robots.txt is a public crawler policy" description: "robots.txt expresses crawl directives for named user agents. AI-related rules should reflect a deliberate, reviewed policy." source: "https://www.machineread.ai/blog/deep-dive-robots-txt" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "Article", "headline": "robots.txt is a public crawler policy", "description": "robots.txt expresses crawl directives for named user agents. AI-related rules should reflect a deliberate, reviewed policy.", "datePublished": "2026-08-02", "dateModified": "2026-08-02", "author": { "@type": "Person", "name": "George Jieh", "url": "https://github.com/georgejieh" }, "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" }, "mainEntityOfPage": "https://www.machineread.ai/blog/deep-dive-robots-txt" } ``` Blog # robots.txt is a public crawler policy robots.txt expresses crawl directives for named user agents. AI-related rules should reflect a deliberate, reviewed policy. author [George Jieh](https://github.com/georgejieh) published 2026-08-02 reading 4 min A useful robots.txt file makes a site's crawl policy legible. It does not have to be permissive. Its job is to tell crawlers, through named groups and URL-path rules, what they are asked to access or avoid (Koster et al., secs. 2.1-2.2.2). That makes the file a public policy signal, not a gate and not a forecast of what happens after a URL is discovered. ## Policy is not enforcement RFC 9309 describes the Robots Exclusion Protocol as rules that crawlers are requested to honor and states directly that those rules are not access authorization (Koster et al., sec. 1). A `Disallow` rule therefore should not be treated as protection for private material. The RFC warns that paths listed in robots.txt become publicly discoverable and recommends real application-layer security when access must be controlled (Koster et al., sec. 3). An `Allow` rule does not reverse that fact. Because the protocol is not authorization, the rule says nothing about whether an HTTP request will clear a login or another access control (Koster et al., secs. 1, 3). MachineRead therefore reports robots.txt policy separately from non-browser fetch results. ## Read the applicable group, not a keyword count A robots.txt group begins with one or more `User-agent` lines followed by `Allow` or `Disallow` rules. A crawler first looks for a group matching its product token; the `*` group applies when no product-token group matches. If neither a matching group nor a wildcard group exists, no rules apply (Koster et al., secs. 2.1-2.2.1). Path evaluation also depends on specificity. When more than one rule matches a URL, RFC 9309 says to use the rule with the longest matching path. If equally specific `Allow` and `Disallow` rules conflict, `Allow` should win (Koster et al., sec. 2.2.2). Reading a file by searching for a bot name or counting `Disallow` lines can therefore miss the policy that applies to the requested URL. Under RFC 9309, crawlers set their own product token, which should appear within the identification string they send to the service (Koster et al., sec. 2.2.1). A correctly matched rule therefore establishes the site's published instruction for that token. The file does not authenticate a requester, demonstrate compliance, or identify why a crawler wants the page. ## What MachineRead observes MachineRead fetches the public `/robots.txt` file and evaluates the audited target path under the URL it constructs for a maintained set of crawler product tokens. Its parser classifies each tracked token through an explicit group, a wildcard group, or no named rule, and it reports selected file-quality conditions such as malformed directive lines and non-absolute or non-HTTP(S) `Sitemap` directives. The result describes one file at audit time. It can reveal a blanket rule that catches more clients than intended, an explicit crawler rule that differs from the wildcard policy, or syntax that deserves review. A restrictive result may be exactly what the publisher chose. MachineRead surfaces the difference; it does not decide the publisher's policy. ## What this check can and cannot establish The check can establish whether MachineRead retrieved a robots.txt file, how its parser classified the audited target path for the crawler tokens it tracks, and which selected structural issues appeared in that response. It cannot establish that a requester is who it claims to be, that any crawler honored the rules, that the same client can fetch other paths, or that content was indexed, used for training, quoted, cited, or ranked. It also cannot turn robots.txt into authorization or a security control. ## Maintain the policy you mean to publish Start with the wildcard group, then inspect each crawler-specific exception and test representative public paths against the resulting rules. Remove stale groups when the policy changes. Keep sensitive route names out of the file and protect those routes with authentication or another real access control (Koster et al., sec. 3). After a hosting, CDN, or bot-management change, compare the published rules with an actual non-browser response. That pair of checks answers two useful but separate questions: what the site asks crawlers to do, and what the site serves when one arrives. ## Works Cited Koster, Martijn, et al. "Robots Exclusion Protocol." _RFC 9309_, Internet Engineering Task Force, Sept. 2022, RFC Editor, [https://www.rfc-editor.org/rfc/rfc9309.html](https://www.rfc-editor.org/rfc/rfc9309.html). Accessed 7 Aug. 2026. ## See also - [Methodology reference](/docs/methodology) - documents the robots\_txt check group - [Launch post](/blog/launching-machineread) - explains the audit's evidence boundaries - Related: [bot access](/blog/deep-dive-bot-access) and [sitemap](/blog/deep-dive-sitemap) robots-txtseodeep-dive [Back to all posts](/blog) --- # https://www.machineread.ai/blog/deep-dive-schema-org --- title: "Publishing page facts with Schema.org JSON-LD" description: "JSON-LD can publish page facts in structured fields. It reduces ambiguity for supporting clients but does not determine how they will use the data." source: "https://www.machineread.ai/blog/deep-dive-schema-org" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "Article", "headline": "Publishing page facts with Schema.org JSON-LD", "description": "JSON-LD can publish page facts in structured fields. It reduces ambiguity for supporting clients but does not determine how they will use the data.", "datePublished": "2026-07-31", "dateModified": "2026-07-31", "author": { "@type": "Person", "name": "George Jieh", "url": "https://github.com/georgejieh" }, "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" }, "mainEntityOfPage": "https://www.machineread.ai/blog/deep-dive-schema-org" } ``` Blog # Publishing page facts with Schema.org JSON-LD JSON-LD can publish page facts in structured fields. It reduces ambiguity for supporting clients but does not determine how they will use the data. author [George Jieh](https://github.com/georgejieh) published 2026-07-31 reading 4 min JSON-LD gives a page a second way to state what it is about. Types and properties can assign explicit meanings to names, dates, prices, and relationships (Schema.org, "Getting Started"). When those fields repeat facts from the visible page, they can become stale or contradict what a reader sees. Structured data is valuable as maintained publishing data, not as an invisible bundle of keywords. Parseable markup is an observed signal. Whether a search engine or another client trusts, supports, or uses it is a separate outcome. ## The vocabulary is not the consumer Schema.org publishes a shared vocabulary of types and properties for describing things and their relationships. Publishers can express that vocabulary through JSON-LD, Microdata, or RDFa (Schema.org, "Getting Started"). JSON-LD is therefore a serialization choice, not another name for Schema.org. Consumers define their own support. Google Search accepts all three formats and generally recommends JSON-LD when a site's setup allows it. Google also says its own documentation, rather than the broader Schema.org vocabulary, is authoritative for Google Search behavior (Google, "Introduction to Structured Data"). A term's presence in Schema.org does not establish that every client recognizes it or does anything with it. Valid JSON is only syntax. Markup can make a statement explicit without compelling a result, and its content can still be outdated, irrelevant, or inconsistent with the page (Google, "General Structured Data Guidelines"). ## Syntax is only the first review For Google Search eligibility, structured data should describe the page where it appears and should not introduce information hidden from readers. Google's guidelines call for current, relevant markup that truthfully represents visible page content (Google, "General Structured Data Guidelines"). This is a content-maintenance requirement, not just a linter rule. Completeness is specific to the intended use. Schema.org defines a broad vocabulary, while a consuming system may document a smaller set of required or recommended properties. For Google Search features, Google advises publishers to favor fewer complete and accurate recommended properties over a larger set of incomplete or inaccurate ones (Google, "Introduction to Structured Data"). Even correctly marked up content is not guaranteed to appear as a rich result (Google, "General Structured Data Guidelines"). Consider a product page that visibly shows one price while its `Offer` object contains another. A parser can report both values, but syntax alone cannot identify the authoritative one. The repair is to fix the publishing pipeline and decide which public fact is current, not to add more markup around the conflict. ## How MachineRead uses the signal MachineRead inspects `application/ld+json` script elements in the fetched target-page HTML. It attempts to parse top-level objects, arrays, and `@graph` entries, scores a limited set of common types, and checks whether selected fields are populated. When commerce scope is enabled, the check gives particular attention to `Product` and `Offer` fields and can report specific mismatches between selected structured fields and visible product content. This is a local structural inspection. It does not execute page JavaScript, consult a live Schema.org vocabulary, or prove that the published facts are true. A parse error, a type outside the scoring rubric, or a missing expected field is a concrete finding a publisher can inspect. ## What this check can and cannot establish The check can establish whether the audited page response contains JSON-LD that its parser can read, which type it selects under its scoring rubric, and whether selected fields are populated under the chosen audit scope. It may also expose a limited visible-content conflict for supported commerce fields. It cannot establish full Schema.org conformance, consumer-specific eligibility, markup injected only after browser rendering, the truth of a business fact, or what any search engine or other client will select and use. It does not promise a rich result, ranking, traffic, citation, or agent response. ## Treat markup as maintained content Begin with the main entity a page actually presents. Use the most specific applicable type, populate fields from the same source that renders the visible page, and remove claims the page no longer supports. When several objects belong together, nest or link them deliberately (Google, "General Structured Data Guidelines"). Use the Schema.org Markup Validator to inspect Schema.org syntax and the extracted graph; its documentation says the tool extracts JSON-LD, RDFa, and Microdata and identifies syntax mistakes (Schema.org, "Schema.org Markup Validator"). If Google Search is an intended consumer, use Google's Rich Results Test and feature-specific documentation as a separate, Google-specific review (Google, "Introduction to Structured Data"). Recheck deployed pages after template or data-feed changes. MachineRead can flag a structural condition during a broader site audit. The validator, the consumer's documentation, and the site's source data answer different questions. ## Works Cited Google. "General Structured Data Guidelines." _Google Search Central_, updated 10 July 2026, [https://developers.google.com/search/docs/appearance/structured-data/sd-policies](https://developers.google.com/search/docs/appearance/structured-data/sd-policies). Accessed 7 Aug. 2026. Google. "Introduction to Structured Data Markup in Google Search." _Google Search Central_, updated 10 Dec. 2025, [https://developers.google.com/search/docs/appearance/structured-data/intro-structured-data](https://developers.google.com/search/docs/appearance/structured-data/intro-structured-data). Accessed 7 Aug. 2026. Schema.org. "Getting Started with Schema.org Using Microdata." _Schema.org_, version 30.0, 19 Mar. 2026, [https://schema.org/docs/gs.html](https://schema.org/docs/gs.html). Accessed 7 Aug. 2026. Schema.org. "Schema.org Markup Validator." _Schema.org_, version 30.0, 19 Mar. 2026, [https://schema.org/docs/validator.html](https://schema.org/docs/validator.html). Accessed 7 Aug. 2026. ## See also - [Methodology reference](/docs/methodology) - documents the schema\_ld check group - [Launch post](/blog/launching-machineread) - explains the audit's evidence boundaries - Related: [llms.txt](/blog/deep-dive-llms-txt) and [canonical URL](/blog/deep-dive-canonical) schema-orgjson-lddeep-dive [Back to all posts](/blog) --- # https://www.machineread.ai/blog/deep-dive-search-discovery --- title: "Signals that help crawlers find updated public content" description: "Search discovery checks public signals related to reachability, initial HTML content, and declared modification dates." source: "https://www.machineread.ai/blog/deep-dive-search-discovery" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "Article", "headline": "Signals that help crawlers find updated public content", "description": "Search discovery checks public signals related to reachability, initial HTML content, and declared modification dates.", "datePublished": "2026-08-04", "dateModified": "2026-08-04", "author": { "@type": "Person", "name": "George Jieh", "url": "https://github.com/georgejieh" }, "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" }, "mainEntityOfPage": "https://www.machineread.ai/blog/deep-dive-search-discovery" } ``` Blog # Signals that help crawlers find updated public content Search discovery checks public signals related to reachability, initial HTML content, and declared modification dates. author [George Jieh](https://github.com/georgejieh) published 2026-08-04 reading 6 min Search discovery is best treated as a path audit. It asks whether a site exposes routes to public URLs, whether sampled requests retrieve useful responses, and whether it publishes freshness cues that MachineRead can parse. Those are observable conditions. Crawling, rendering, indexing, ranking, traffic, citations, and agent use all happen later, under systems the publisher does not control. MachineRead tests that public chain without turning its links into a forecast. ## Start with routes a client can see A page can be exposed through ordinary site navigation, a sitemap, or both. For Google specifically, the dependable HTML pattern for a crawlable link is an `a` element with an `href` that resolves to a web address; Google says most other link formats are not parsed and extracted reliably (Google, "Link Best Practices for Google"). A sitemap is another route into the site, but it is a declaration, not proof of an outcome. Google says that a sitemap can help its systems discover URLs while explicitly warning that listed items are not guaranteed to be crawled or indexed (Google, "Learn about Sitemaps"). This makes a sitemap useful evidence of what a publisher has exposed, but not evidence of what any search engine has done with it. The same restraint applies to `robots.txt`. The Robots Exclusion Protocol lets a service owner publish rules that crawlers are requested to honor, and the standard states that those rules are not access authorization (Koster et al., sec. 1). Reading a rule can establish the site's declared policy for a named crawler. It cannot establish whether that crawler requested the page, how it interpreted every other signal, or whether an unrelated client followed the rule. ## The response matters as much as the route A discovered URL still has to return something useful. MachineRead fetches selected pages over HTTP and examines the response without running a browser. That gives the audit the returned status, final URL, selected response headers, and initial HTML under that retrieval condition. Initial HTML and rendered content are not interchangeable. Google's own documentation describes an app-shell case in which the first response lacks the page content and Google later executes JavaScript to produce rendered HTML. The same documentation cautions that not all bots can run JavaScript (Google, "Understand the JavaScript SEO Basics"). Google's rendering capability is therefore not a sound assumption about another search crawler, archival tool, or agent. MachineRead therefore samples what is present before client-side execution. On sampled sitemap pages, it checks titles, canonical links, robots directives, parseable JSON-LD, and extractable text in the initial response. Separate heuristics inspect target-page anchors and sitemap URL strings for public trust pages, and validate `hreflang` only where it is published. These observations do not prove that another system will use the signals. An absent signal identifies a gap in the raw document, but determining whether browser rendering supplies it requires a separate test. ## Dates should report changes, not deployments Freshness cues are useful only when they describe the content. Google's sitemap guidance says it uses `lastmod` when the values are consistently and verifiably accurate, and that the field should reflect the last significant update to the page rather than a routine copyright change (Google, "Build and Submit a Sitemap"). By that standard, a build that stamps every URL with today's date may produce valid XML while weakening the meaning of the date. MachineRead inspects sitemap `lastmod` coverage, syntax, and future dates. It does not compare a declared date with the page's revision history. The fallback heuristic can also observe a linked publishing section, dated target-page metadata, or a reachable, nonempty RSS, Atom, or JSON feed. It reports feed date coverage and recency when dates are present. These alternatives describe different publishing practices. A brochure site may have no reason to maintain a feed, while a publication may use one as a compact record of recent URLs. In either case, the audit records what is exposed; it does not infer a crawl schedule from it. ## What MachineRead records The check brings several observations into one report row: - whether the published `robots.txt` rules block Googlebot or Bingbot from the audited target path under the URL constructed by the check; - whether a valid sitemap is found at the conventional path or through a `robots.txt` reference, and whether sampled entries use coherent same-site HTTPS URLs; - whether selected sitemap URLs return accessible responses whose initial HTML exposes useful text, titles, canonical links, robots directives, and parseable JSON-LD; the normal shared-evidence path also applies a local title-and-description coherence heuristic; - whether at least 80 percent of sampled sitemap entries declare `lastmod`, with no invalid or future values encountered in the parsed sitemap documents, or the target page exposes a publishing-section link, dated metadata, or a nonempty parseable feed; and - whether target-page anchors or sitemap URLs expose About, Contact, and Privacy routes, with contextual findings for other policy pages and additional checks for `hreflang` only when a site publishes it. The sample is intentionally limited. MachineRead does not crawl every URL, execute JavaScript, query a search-result page, inspect Search Console, or call a paid indexing provider. A full score means the four point-bearing conditions passed and no applicable `hreflang`, core trust-page, or commerce-policy issue capped the row. Optional findings and sample caveats still require review. A weak result identifies a response or declaration to inspect; it does not diagnose a search-performance problem. Review the chain in the same order. Request `robots.txt` and the sitemap directly. Open several listed URLs with JavaScript disabled or inspect their raw responses. Confirm that important pages have ordinary links from navigation or relevant content. Compare `lastmod` and feed dates with actual editorial changes. If those checks disagree, repair the declaration or the page rather than trying to predict how a crawler will compensate. ## What this check can and cannot establish It can establish which HTTP responses, HTML elements, links, directives, sitemap entries, and declared dates MachineRead observed in its limited sample at audit time. It can also identify contradictions, such as a sitemap URL that fails to load or a modification date that does not parse. It cannot establish that a named search engine or agent discovered the URL, executed its scripts, crawled it again, indexed it, selected it for a result, ranked it, sent traffic, cited it, or used it in a task. It also cannot prove that a syntactically valid `lastmod` value corresponds to an actual content change. Those outcomes and historical claims require evidence beyond this check. ## Works Cited Google. "Build and Submit a Sitemap." _Google Search Central_, updated 8 July 2026, [https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap](https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap). Accessed 7 Aug. 2026. Google. "Learn about Sitemaps." _Google Search Central_, updated 10 Dec. 2025, [https://developers.google.com/search/docs/crawling-indexing/sitemaps/overview](https://developers.google.com/search/docs/crawling-indexing/sitemaps/overview). Accessed 7 Aug. 2026. Google. "Link Best Practices for Google." _Google Search Central_, updated 10 Dec. 2025, [https://developers.google.com/search/docs/crawling-indexing/links-crawlable](https://developers.google.com/search/docs/crawling-indexing/links-crawlable). Accessed 7 Aug. 2026. Google. "Understand the JavaScript SEO Basics." _Google Search Central_, updated 4 Mar. 2026, [https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics). Accessed 7 Aug. 2026. Koster, Martijn, et al. "Robots Exclusion Protocol." _RFC 9309_, Internet Engineering Task Force, Sept. 2022, RFC Editor, [https://www.rfc-editor.org/rfc/rfc9309.html](https://www.rfc-editor.org/rfc/rfc9309.html). Accessed 7 Aug. 2026. ## See also - [Methodology reference](/docs/methodology) - documents the search\_discovery check group - [Launch post](/blog/launching-machineread) - explains the audit's evidence boundaries - Related: [sitemap](/blog/deep-dive-sitemap) and [bot access](/blog/deep-dive-bot-access) search-discoveryaudit-pillardeep-dive [Back to all posts](/blog) --- # https://www.machineread.ai/blog/deep-dive-sitemap --- title: "A sitemap is a map, not an indexing guarantee" description: "A sitemap publishes URLs and optional modification dates for crawlers. It can support discovery, but it does not guarantee crawling or indexing." source: "https://www.machineread.ai/blog/deep-dive-sitemap" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "Article", "headline": "A sitemap is a map, not an indexing guarantee", "description": "A sitemap publishes URLs and optional modification dates for crawlers. It can support discovery, but it does not guarantee crawling or indexing.", "datePublished": "2026-08-02", "dateModified": "2026-08-02", "author": { "@type": "Person", "name": "George Jieh", "url": "https://github.com/georgejieh" }, "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" }, "mainEntityOfPage": "https://www.machineread.ai/blog/deep-dive-sitemap" } ``` Blog # A sitemap is a map, not an indexing guarantee A sitemap publishes URLs and optional modification dates for crawlers. It can support discovery, but it does not guarantee crawling or indexing. author [George Jieh](https://github.com/georgejieh) published 2026-08-02 reading 4 min A sitemap is an inventory published by a site owner. It can make important public URLs easier to discover, but the presence of a URL in the file says nothing by itself about whether that URL was crawled or indexed (Google, "Learn about Sitemaps"). Its practical value is narrower: it publishes the URLs the site prefers as canonical (Google, "Build and Submit a Sitemap"). ## Treat the file as an inventory Google defines a sitemap as a file that describes pages and other files on a site, their relationships, and information such as modification dates. Its guidance says that a sitemap can help discovery, especially on large or complex sites, while a well-linked site may already expose most important pages through navigation (Google, "Learn about Sitemaps"). A listed URL establishes only that the publisher put it in the sitemap. Google explicitly says that sitemaps do not guarantee that every listed item will be crawled or indexed (Google, "Learn about Sitemaps"). The file should therefore be evaluated as an inventory, not as evidence of search performance. This framing also makes omissions easier to reason about. If a page matters but is missing, first ask whether the publishing pipeline left it out and whether ordinary crawlable links expose it elsewhere. If a retired or duplicate URL remains, the sitemap may be describing a site that no longer exists in that form. ## Consistency matters more than decoration Google tells publishers to use fully qualified URLs and to list the canonical versions they prefer, rather than every alternate URL that reaches the same content (Google, "Build and Submit a Sitemap"). A sitemap filled with mixed hosts, HTTP and HTTPS variants, or duplicate locations may still be parseable, but it is not a clear inventory. The optional `lastmod` value needs similar discipline. Google says it uses that value when it is consistently and verifiably accurate, and that it should reflect the last significant update to the page rather than a routine change such as a copyright-date edit (Google, "Build and Submit a Sitemap"). A build timestamp copied onto every URL may be valid XML while conveying little about which content changed. If the publishing system cannot produce meaningful dates, omitting them is clearer than manufacturing freshness. Google also ignores the XML `priority` and `changefreq` values (Google, "Build and Submit a Sitemap"). Time spent tuning those fields for Google is better spent reconciling the URL list with the site's actual canonical content. ## Submission is still a hint A publisher can make a sitemap available to Google through Search Console, its API, or a `Sitemap` line in robots.txt (Google, "Build and Submit a Sitemap"). Google describes submission as a hint and says it does not guarantee that Google will download the sitemap or use it when crawling the site's URLs (Google, "Build and Submit a Sitemap"). A robots.txt reference is useful discovery evidence, but it does not prove that a sitemap was submitted through a private Search Console property or processed by any particular crawler. ## How MachineRead reads the map There is no standalone sitemap score. MachineRead uses sitemap evidence within the LLM Text & Markdown Access and Search Discovery Hints report groups. Within those groups, MachineRead looks for `/sitemap.xml` and for sitemap locations named in robots.txt. It parses a limited set of XML sitemap and sitemap-index responses, records URL and `lastmod` conditions, and samples for issues such as duplicate entries, off-site locations, non-HTTPS URLs, and pages that do not return an accessible, indexable response. A clean sample is evidence about the files and URLs inspected during that audit, not a complete crawl of every entry. Likewise, a missing robots.txt reference means MachineRead did not observe that submission route; it does not establish whether the sitemap was submitted through Search Console or another private channel. ## What this check can and cannot establish The check can establish whether MachineRead found a parseable XML sitemap at the locations it inspected, what the sample contained, whether robots.txt named a sitemap, and whether selected URL and date conditions were present. It cannot establish complete site coverage, private submission status, crawler processing, indexing, ranking, traffic, citations, or later use by an automated client. It also cannot infer that an unlisted page is unreachable without examining the site's links and other discovery routes. ## Keep the map tied to the site Generate the sitemap from the same source that determines published canonical URLs. On each release, compare additions, removals, redirects, host changes, and meaningful modification dates with that source. If the sitemap is split into indexes, check the child files as well as the index. When MachineRead flags an entry, inspect the returned URL and the sitemap generator before treating the symptom as a search problem. Correct the public inventory first. What a crawler does with that inventory requires separate evidence. ## Works Cited Google. "Build and Submit a Sitemap." _Google Search Central_, updated 8 July 2026, [https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap](https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap). Accessed 7 Aug. 2026. Google. "Learn about Sitemaps." _Google Search Central_, updated 10 Dec. 2025, [https://developers.google.com/search/docs/crawling-indexing/sitemaps/overview](https://developers.google.com/search/docs/crawling-indexing/sitemaps/overview). Accessed 7 Aug. 2026. ## See also - [Methodology reference](/docs/methodology) - documents sitemap evidence within the public check groups - [Launch post](/blog/launching-machineread) - explains the audit's evidence boundaries - Related: [search discovery](/blog/deep-dive-search-discovery) and [canonical URL](/blog/deep-dive-canonical) sitemapseodeep-dive [Back to all posts](/blog) --- # https://www.machineread.ai/blog/launching-machineread --- title: "Launching MachineRead" description: "MachineRead audits a public website for observable AI agent and search-readiness signals across 13 check groups and 56 checked points." source: "https://www.machineread.ai/blog/launching-machineread" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "Article", "headline": "Launching MachineRead", "description": "MachineRead audits a public website for observable AI agent and search-readiness signals across 13 check groups and 56 checked points.", "datePublished": "2026-07-23", "dateModified": "2026-07-23", "author": { "@type": "Person", "name": "George Jieh", "url": "https://github.com/georgejieh" }, "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" }, "mainEntityOfPage": "https://www.machineread.ai/blog/launching-machineread" } ``` Blog # Launching MachineRead MachineRead audits a public website for observable AI agent and search-readiness signals across 13 check groups and 56 checked points. author [George Jieh](https://github.com/georgejieh) published 2026-07-23 reading 4 min The useful question in a website audit is not whether a site is "ready for AI." It is whether specific public signals can be fetched, parsed, and checked. Everything after that, including indexing, retrieval, citation, ranking, and action, belongs to a different layer of evidence. MachineRead starts with this narrower question. For a submitted public URL, the Essentials audit records technical conditions related to access, discovery, and interpretation. The report exposes individual findings because a score without its observations is easy to overread. ## Start with what the site publishes Automated clients do not all receive or process a page in the same way. A browser may execute JavaScript and display a complete document even when the first HTTP response contains little more than an application shell. Google, for example, describes crawling, rendering, and indexing as separate phases, and says some JavaScript sites require rendering before their generated content is visible to Google (Google, "Understand the JavaScript SEO Basics"). That documented Google process should not be treated as a description of every crawler. Other signals are narrower still. A `robots.txt` file states rules that crawlers are requested to honor, but the standard explicitly says those rules are not access authorization (Koster et al., sec. 1). A sitemap can present URLs for discovery, but Google calls sitemap submission a hint and does not guarantee that it will use the file to crawl those URLs (Google, "Build and Submit a Sitemap"). These files matter because they express inspectable intent, not because their presence determines what happens later. MachineRead therefore tests the representation and declarations a site actually publishes, then describes only what those tests support. ## Two views of the same public surface Essentials evaluates [13 included check groups with a current maximum of 56 checked points](/docs/methodology). The groups cover public signals such as the initial HTML response, crawler directives, canonical metadata, sitemaps, structured data, and machine-readable discovery files. MachineRead checks [llms.txt](/blog/deep-dive-llms-txt) in the general audit. When API or protocol scope is active, it also checks [ai-catalog.json](/blog/deep-dive-ai-catalog), [api-catalog.json](/blog/deep-dive-api-catalog), and [openapi.json](/blog/deep-dive-openapi). The report also includes a stricter agent-readiness view. Its 8 default probes form the default scope; the full methodology describes 21 probes when all applicable scope options are enabled. This view keeps explicit agent-facing and protocol signals legible without changing the 13-group, 56-point Essentials contract. The two summaries answer related but different questions. The Essentials score condenses the included evidence. The strict view isolates a smaller set of agent-oriented observations. Neither is a probability that a model will mention the site. Nine advanced rows are displayed as [locked coverage areas](/docs/methodology#locked-vs-unlocked-rows). They show the shape of coverage that did not run. A locked row is not a failed live test, and MachineRead does not fill missing evidence with an inferred result. ## Read findings before scores Start with findings that change what later evidence means. If a plain request is denied or the returned HTML lacks the expected public content, well-formed metadata elsewhere may have limited value to that requesting client. Check the response status, headers, and body first. Then inspect directives and discovery files. Finally, review whether metadata agrees with the visible page. That order also keeps remediation proportional. A malformed canonical has a different fix from an intentional crawler block. The audit supplies evidence and context; the site owner decides whether a finding conflicts with the site's actual publishing policy. MachineRead is deliberately not a full-site crawler, a browser task runner, or a content judge. It runs specific checks against a public target. The [methodology reference](/docs/methodology) documents the checks and scoring, while the [agent integration guide](/docs/agent-integration) describes the public machine-readable interfaces. ## What this check can and cannot establish The audit can establish that the tested URL and related public resources returned particular responses at audit time. It can report whether the inspected signals were present, parseable, internally coherent, and consistent with each check's documented criteria. It cannot establish that an untested URL behaves the same way, that every crawler will honor or interpret a signal alike, or that a search engine or model will index, retrieve, cite, rank, or act on the site. It also cannot replace a full crawl, rendered-browser testing, log analysis, or provider-specific verification. ## Works Cited Google. "Build and Submit a Sitemap." _Google Search Central_, updated 8 July 2026, [https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap](https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap). Accessed 7 Aug. 2026. Google. "Understand the JavaScript SEO Basics." _Google Search Central_, updated 4 Mar. 2026, [https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics). Accessed 7 Aug. 2026. Koster, Martijn, et al. "Robots Exclusion Protocol." _RFC 9309_, Internet Engineering Task Force, Sept. 2022, RFC Editor, [https://www.rfc-editor.org/rfc/rfc9309.html](https://www.rfc-editor.org/rfc/rfc9309.html). Accessed 7 Aug. 2026. ## See also - [Methodology reference](/docs/methodology) - definitions, weights, and evidence boundaries - [Why agents need to read your site](/blog/why-agents-need-to-read-your-site) - why access is a prerequisite rather than an outcome - [Public audit interface](/) - accepts a URL for the checks described above launchproductai-readiness [Back to all posts](/blog) --- # https://www.machineread.ai/blog/why-agents-need-to-read-your-site --- title: "Why Agents Need to Read Your Site" description: "Agent retrieval depends on what a client can fetch and parse. A page that renders for people may expose less to a non-browser client." source: "https://www.machineread.ai/blog/why-agents-need-to-read-your-site" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "Article", "headline": "Why Agents Need to Read Your Site", "description": "Agent retrieval depends on what a client can fetch and parse. A page that renders for people may expose less to a non-browser client.", "datePublished": "2026-07-23", "dateModified": "2026-07-23", "author": { "@type": "Person", "name": "George Jieh", "url": "https://github.com/georgejieh" }, "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" }, "mainEntityOfPage": "https://www.machineread.ai/blog/why-agents-need-to-read-your-site" } ``` Blog # Why Agents Need to Read Your Site Agent retrieval depends on what a client can fetch and parse. A page that renders for people may expose less to a non-browser client. author [George Jieh](https://github.com/georgejieh) published 2026-07-23 reading 5 min Before an automated system can interpret a page, it has to receive a usable representation of that page. This is a delivery requirement, not a ranking strategy. Readability can remove one technical obstacle, but it cannot decide whether a crawler visits, whether an index retains the content, or whether a model uses it later. That sequence sounds obvious until a browser hides the gap. A visitor may see a complete interface after scripts run, while a plain HTTP client receives an almost empty shell. A `200` response alone does not show whether the returned body contains the article, product details, documentation, or links the client needs. ## Reading is a chain of separate conditions It helps to split the problem into four questions: 1. Can the client discover or otherwise obtain the URL? 2. Does the server permit and complete the request? 3. Does the response contain usable content, or can that client render what is missing? 4. Can the client interpret the content well enough for its own task? These questions are related, but a pass at one step does not answer the next. Google documents its own JavaScript pipeline as crawling, rendering, and indexing, with links extracted before and again after rendering. It also notes that an app-shell page may omit its actual content from the initial HTML (Google, "Understand the JavaScript SEO Basics"). That is evidence about Google Search's documented process, not a default capability shared by every automated client. Vercel and MERJ reported a different result in 2024. Their primary data came from the Vercel network and `nextjs.org`, with two job-board sites used to compare observations across different stacks. In that sample, the listed OpenAI, Anthropic, Meta, ByteDance, and Perplexity crawlers did not execute JavaScript. The authors also reported that traffic they attributed to ChatGPT and Claude requested some JavaScript files without executing them (Zecchini et al.). The study supports inspecting initial HTML for the crawlers it measured. It does not establish permanent behavior for those services, cover every agent, or prove what any model does with fetched content. ## Access policy and content delivery are not the same thing The Robots Exclusion Protocol gives service owners a standard way to publish rules that crawlers are requested to honor. RFC 9309 is equally clear that those rules are not access authorization (Koster et al., sec. 1). A permissive `robots.txt` file therefore does not prove that a request will succeed. A firewall, authentication layer, rate limit, challenge, redirect loop, or server error can still prevent retrieval. Technical fetchability can also conflict with the published policy. A client may be able to retrieve a URL that the site's crawler policy asks it not to crawl. A readiness review should report both the observed response and the stated directive rather than collapse them into a single label such as "accessible." Discovery signals have similar limits. Google says submitting a sitemap is only a hint and does not guarantee that Google will download it or use it for crawling (Google, "Build and Submit a Sitemap"). Structured data can label facts in a standardized form, but even correct markup does not guarantee a Google rich result (Google, "General Structured Data Guidelines"). Those documents can make a site's published intent easier to inspect. They do not control later selection. ## A practical reading test Start with one public URL that matters. Fetch it without a browser and save the final status, response headers, and response body. Compare the body with the page a visitor sees after rendering. Look for the primary text, a descriptive title, ordinary crawlable links, and any metadata that the page is supposed to publish. If important content is absent, identify where it appears. It may arrive in the initial HTML, embedded application data, a later API request, or client-generated DOM. That location is an observation. Whether to change the delivery architecture depends on the clients the site intends to support, along with performance, security, and maintenance constraints. Next, inspect `robots.txt`, page-level robots directives, the canonical URL, relevant structured data, and discovery files. Check whether they agree with the visible page and with the site's actual policy. Do not add a file solely to collect a green check. An accurate absence is better than machine-readable metadata that points to stale, private, or unsupported resources. MachineRead applies this sequence through specific public checks. Its [bot-access check](/blog/deep-dive-bot-access) sends tracked bot user-agent requests and compares their status, routing, and basic content characteristics with a comparison baseline, normally an attempted browser-identifying request. The broader [methodology](/docs/methodology) keeps directives, raw-HTML readability, schema, and discovery evidence in separate findings. Remediation can then follow the observed condition instead of a broad claim about "agent visibility." ## What this check can and cannot establish The check can establish what MachineRead's tracked user-agent requests received at a specific time and whether their status, routing, or basic content characteristics differed from its comparison baseline. Related checks can identify published directives, parseable metadata, and reachable discovery files. It cannot establish that every crawler receives the same response, that a named provider rendered or retained the page, or that a model will retrieve, cite, rank, recommend, or act on the content. It also cannot infer intent from a block or prove content quality from successful parsing. ## Works Cited Google. "Build and Submit a Sitemap." _Google Search Central_, updated 8 July 2026, [https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap](https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap). Accessed 7 Aug. 2026. Google. "General Structured Data Guidelines." _Google Search Central_, updated 10 July 2026, [https://developers.google.com/search/docs/appearance/structured-data/sd-policies](https://developers.google.com/search/docs/appearance/structured-data/sd-policies). Accessed 7 Aug. 2026. Google. "Understand the JavaScript SEO Basics." _Google Search Central_, updated 4 Mar. 2026, [https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics](https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics). Accessed 7 Aug. 2026. Koster, Martijn, et al. "Robots Exclusion Protocol." _RFC 9309_, Internet Engineering Task Force, Sept. 2022, RFC Editor, [https://www.rfc-editor.org/rfc/rfc9309.html](https://www.rfc-editor.org/rfc/rfc9309.html). Accessed 7 Aug. 2026. Zecchini, Giacomo, et al. "The Rise of the AI Crawler." _Vercel_, 17 Dec. 2024, [https://vercel.com/blog/the-rise-of-the-ai-crawler](https://vercel.com/blog/the-rise-of-the-ai-crawler). Accessed 7 Aug. 2026. ## See also - [Launch post](/blog/launching-machineread) - what the audit measures and does not measure - [Methodology reference](/docs/methodology) - the 13 check groups and probe definitions - Related: [bot access](/blog/deep-dive-bot-access) and [schema.org JSON-LD](/blog/deep-dive-schema-org) ai-agentsdiscoveryretrieval [Back to all posts](/blog) --- # https://www.machineread.ai/contact --- title: "Contact - MachineRead" description: "Contact MachineRead for methodology questions, bug reports, partnership inquiries, or enterprise interest in future paid tiers." source: "https://www.machineread.ai/contact" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` Contact # Questions, bug reports, partnership inquiries Methodology questions, bug reports, partnership inquiries, or enterprise interest in future paid tiers: reach out at the email below. Email channel [contact@machineread.ai](mailto:contact@machineread.ai) Replies go to a personal inbox; expect a response within a few days. --- # https://www.machineread.ai/docs --- title: "Docs - MachineRead" description: "MachineRead documentation: methodology, API reference, and agent integration guide." source: "https://www.machineread.ai/docs" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "TechArticle", "headline": "MachineRead Documentation", "url": "https://www.machineread.ai/docs", "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" } } ``` Docs # Documentation Reference documentation for MachineRead's audit methodology, API contract, and agent integration patterns. [ ## Methodology The 13 included check groups, three pillars, nine locked rows, and scoring formula behind the MachineRead Essentials audit. ](/docs/methodology)[ ## API Reference The full and compact audit endpoints, request and response models, scope, errors, and rate limits. ](/docs/api)[ ## Agent Integration How agents call MachineRead, discover its public catalogs, and use the hosted MCP server. ](/docs/agent-integration) --- # https://www.machineread.ai/docs/agent-integration --- title: "Agent Integration - MachineRead Docs" description: "How agents call MachineRead, discover its public catalogs, and use the hosted MCP server." source: "https://www.machineread.ai/docs/agent-integration" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "TechArticle", "headline": "Agent Integration", "description": "How agents call MachineRead, discover its public catalogs, and use the hosted MCP server.", "url": "https://www.machineread.ai/docs/agent-integration", "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" } } ``` Docs Docs # Agent Integration How agents call MachineRead, discover its public catalogs, and use the hosted MCP server. ## Call the audit API The production API host is `https://api.machineread.ai`. Send JSON to the full endpoint when you need every check row: ```http POST https://api.machineread.ai/v1/audit Content-Type: application/json {"url":"https://example.com","preset":"blog"} ``` The full response includes scores, per-check findings, benchmark context, scope, and strict agent readiness. For a smaller agent-oriented projection, call `POST https://api.machineread.ai/v1/audit/summary` with the same request body. Neither endpoint uses an LLM to write the result. ### Minimal agent workflow 1. Choose the closest valid preset. 2. POST the public URL to the API host. 3. Parse `overall_score`, benchmark context, and `agent_readiness`. 4. For the full endpoint, summarize applicable `checks` and keep its caveats. 5. Treat locked rows as unverified coverage, not failures. ## Discovery surfaces The www static host publishes `llms.txt`, `llms-full.txt`, OpenAPI, an RFC 9264 API Catalog linkset, an ARD `ai-catalog.json`, an MCP Server Card, and an Agent Skills index. The API host also serves the canonical OpenAPI document and API Catalog route. These files describe or link to executable endpoints; their presence does not itself prove live uptime or agent invocation. - Product guide: `https://www.machineread.ai/llms.txt` - ARD: `https://www.machineread.ai/.well-known/ai-catalog.json` - API Catalog: `https://api.machineread.ai/.well-known/api-catalog` - OpenAPI: `https://api.machineread.ai/openapi.json` - MCP Server Card: `https://www.machineread.ai/.well-known/mcp/server-card.json` - Agent Skills: `https://www.machineread.ai/.well-known/agent-skills/index.json` ## Hosted MCP server MachineRead runs an anonymous, rate-limited Streamable HTTP MCP server at `https://api.machineread.ai/.well-known/mcp/mcp`. It exposes four read-only tools: - `run_essentials_audit(url, preset?, custom_overrides?)` - `get_audit_report(audit_id)` - `list_available_checks()` - `explain_check(check_name)` `run_essentials_audit` waits for completion and returns a compact MCP projection with audit ID, scores, check count, strict readiness, rate-limit state, and caveat. It does not return the REST endpoint's full per-check evidence array. The same FastMCP instance is also available over optional local stdio with `python -m mcp_server.server`. ## Agent-readiness scoring The separate strict score evaluates explicit agent-native discovery and protocol signals. Its denominator follows the resolved family scope: named general presets use 8 probes, SaaS uses 17, and Ecommerce uses 21. Custom and partial scopes derive their own exact denominator. Retrieval or metadata observations do not prove ranking, indexing, model citation, crawler identity, or protocol use. [Back to all docs](/docs) --- # https://www.machineread.ai/docs/api --- title: "API Reference - MachineRead Docs" description: "The full and compact audit endpoints, request and response models, scope, errors, and rate limits." source: "https://www.machineread.ai/docs/api" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "TechArticle", "headline": "API Reference", "description": "The full and compact audit endpoints, request and response models, scope, errors, and rate limits.", "url": "https://www.machineread.ai/docs/api", "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" } } ``` Docs Docs # API Reference The full and compact audit endpoints, request and response models, scope, errors, and rate limits. ## Endpoints Production base URL: `https://api.machineread.ai` ```text POST /v1/audit POST /v1/audit/summary ``` Both endpoints accept the same `AuditRequest` and run the same bounded Essentials pipeline. `/v1/audit` returns the full `AuditResult`. `/v1/audit/summary` returns a compact deterministic `AuditSummary` for agents. The deprecated `POST /audit` compatibility alias is available at runtime but is absent from OpenAPI. All audit routes share one per-client-IP REST rate-limit bucket. They accept anonymous `application/json` requests. The transport rejects request bodies larger than 65,536 bytes before JSON parsing. ## Request body ```json { "url": "https://example.com/", "preset": "blog" } ``` | Field | Type | Required | Meaning | | --- | --- | --- | --- | | `url` | string | yes | Public HTTP(S) URL. A missing scheme defaults to HTTPS. Private, loopback, link-local, and reserved targets are rejected. | | `preset` | string or null | no | Recommended scope selector: `blog`, `corporate`, `services`, `ecommerce`, `news`, `saas`, or `custom`. | | `custom_overrides` | object or null | no | Boolean family overrides applied on top of a selected preset. Rejected when `preset` is null. | | `include_protocols` | boolean | no, deprecated | Legacy protocol/API scope toggle. Ignored when a preset is selected. | | `include_account_auth` | boolean | no, deprecated | Legacy account/auth scope toggle. Ignored when a preset is selected. | | `include_ecommerce` | boolean | no, deprecated | Legacy commerce scope toggle. Ignored when a preset is selected. | Supported `custom_overrides` keys are `protocols`, `account_auth`, `ecommerce`, `feed_discovery`, `article_schema`, `localbusiness_schema`, `news_article_schema`, `claimreview_schema`, `product_offer_schema`, `commerce_fields`, `api_catalog`, `mcp`, `a2a`, `agent_skills`, `webmcp`, `oauth_oidc`, `ard_catalog`, and `auth_md`. Unknown, incoherent, or preset-inapplicable combinations return 422. ## Presets - `blog`: feed discovery and Article/BlogPosting schema. - `corporate`: universal and contextual core only. - `services`: LocalBusiness schema. - `news`: feeds, Article/BlogPosting, NewsArticle, and ClaimReview schema. - `saas`: API catalog, MCP, A2A, Agent Skills, WebMCP, ARD, OAuth/OIDC, and `auth.md`. - `ecommerce`: the SaaS protocol/auth families plus feeds, Product/Offer schema, and commerce fields. - `custom`: starts from the blog family defaults and accepts any supported override. Every preset keeps the same 13 included rows and 56-point checked denominator. Presets change applicable sub-signals, strict agent-readiness probes, scope metadata, and benchmark cohort selection. ## Full response: AuditResult `POST /v1/audit` returns these top-level fields: | Field | Meaning | | --- | --- | | `api_version` | Public contract version, currently `1.0`. | | `url` | Normalized audited URL. | | `scope` | Resolved `AuditScope`. | | `overall_score` | Points earned on the 100-point full rubric. Locked rows remain zero until verified. | | `pillar_scores` | Raw earned points for `off_site`, `scrapability`, and `seo`. | | `pillar_max` | Full-rubric caps: 30, 40, and 30. | | `agent_readiness` | Separate strict agent-native percentage plus earned/max probes and benchmark. | | `benchmark` | Singular Essentials peer comparison. `score` and `median_score` are percentages; `checked_score` and `checked_max` are raw points. | | `checks` | 13 included rows and 9 locked rows in the same `checks` array. | Do not read `benchmark.score` as raw points. For example, 42 earned of 56 checked points produces `benchmark.score: 75`. `overall_score: 42` remains 42 raw points on the full 100-point rubric. ### AuditScope keys The full response `scope` contains: - `include_protocols`, `include_account_auth`, and `include_ecommerce`: resolved coarse dimensions. - `label`: human-readable resolved scope. - `included_optional_surfaces` and `excluded_optional_surfaces`: coarse display buckets. - `preset_applied`: selected preset, or null for the legacy path. - `overrides_applied`: the accepted custom override map. - `included_families` and `excluded_families`: exact resolved family keys. - `machine_surfaces_scope`: `common-contextual` or `full`. ### CheckResult fields | Field | Meaning | | --- | --- | | `pillar` | `off_site`, `scrapability`, or `seo`. | | `check_name` | Stable machine identifier. | | `label` | Human-readable row title. | | `state` | `pass`, `partial`, `fail`, `warn`, or `locked`. | | `evidence_level` | `verified`, `inferred`, `unknown`, or `not_applicable`. | | `available_in` | `Essentials`, `Starter`, or `Pro`. | | `score` | Raw points earned for this row. | | `max_score` | Raw points available for this row. | | `finding` | Deterministic explanation of observed evidence. | | `fix` | Deterministic remediation hint. | | `effort` | `low`, `medium`, or `high`. | A locked row has `state: "locked"`, `score: 0`, `evidence_level: "not_applicable"`, and a Starter or Pro `available_in` value. A `warn` row is inconclusive, not negative verified evidence. ```json { "overall_score": 42, "benchmark": { "score": 75, "checked_score": 42, "checked_max": 56, "median_score": 75, "percentile": 100 }, "checks": [ { "pillar": "scrapability", "check_name": "robots_txt", "label": "AI Bot Policy Signals", "state": "partial", "evidence_level": "verified", "available_in": "Essentials", "score": 4, "max_score": 6, "finding": "robots.txt mentions GPTBot but does not mention ClaudeBot.", "fix": "Add an explicit ClaudeBot directive (Allow or Disallow) to robots.txt.", "effort": "low" } ] } ``` ## Compact response: AuditSummary `POST /v1/audit/summary` preserves score denominators, compact scope, benchmark percentiles and medians, row counts, up to five attention rows, and fixed limitation codes. It omits full finding/fix prose, peer entries, raw fetch context, and generated prose. The top-level keys are `api_version`, `summary_version`, `url`, `scope`, `scores`, `benchmarks`, `checks`, `attention`, and `limitations`. Note that full `AuditResult` uses singular `benchmark`; the compact projection groups two comparisons under plural `benchmarks`. ```json { "api_version": "1.0", "summary_version": "1.0", "url": "https://example.com/", "scope": { "preset": "blog", "protocols": false, "account_auth": false, "ecommerce": false, "overrides": {} }, "scores": { "overall": { "earned": 42, "max": 100 }, "pillars": { "off_site": { "earned": 10, "max": 30 }, "scrapability": { "earned": 20, "max": 40 }, "seo": { "earned": 12, "max": 30 } }, "essentials": { "percent": 75, "earned": 42, "max": 56 }, "agent_readiness": { "percent": 88, "earned": 7, "max": 8 } }, "benchmarks": { "essentials": { "percentile": 100, "median_percent": 75, "peer_count": 1, "snapshot": "2026-08-26" }, "agent_readiness": { "percentile": 100, "median_percent": 88, "peer_count": 1, "snapshot": "2026-08-26" } }, "checks": { "included": 13, "locked": 9, "pass": 2, "partial": 11, "fail": 0, "warn": 0, "attention_total": 5 }, "attention": [], "limitations": ["relative_scores", "no_live_ranking", "no_provider_ip_auth", "no_paid_crawlers"] } ``` ## Progressive HTML form adapter The public page also supports a non-JavaScript browser submission. Its standard form posts `application/x-www-form-urlencoded` data to same-origin `/audit` on the www host. A bounded Pages Function validates the form and issues a 307 to `https://api.machineread.ai/v1/audit/form`, preserving the browser POST and the real client-IP rate-limit identity. `/v1/audit/form` accepts only the browser form encoding and returns escaped, semantic HTML with `Cache-Control: no-store` and restrictive security headers. It is a progressive-enhancement adapter, not a JSON API. It is intentionally `include_in_schema=False` and must not appear in the JSON OpenAPI contract. JSON agents and SDKs must continue to use `/v1/audit` or `/v1/audit/summary`. ## Error responses Both public audit endpoints document the same error set in OpenAPI. | Status | Model | Meaning | | --- | --- | --- | | `400` | `ErrorMessage` | The URL resolves to a blocked private, loopback, link-local, or reserved target. | | `413` | `ErrorMessage` | The request body exceeds 65,536 bytes. | | `422` | `ValidationErrorMessage` | URL syntax, body, preset, or override validation failed. | | `429` | `RateLimitErrorMessage` | The shared REST audit limit was exceeded. The body includes `retry_after`. | | `500` | `ErrorMessage` | Audit context setup failed. No partial result is returned. | | `503` | `ErrorMessage` | The bounded outbound-request budget was exhausted, process capacity was unavailable, or the end-to-end audit wall deadline elapsed. No partial result is returned or cached. | ## Rate limits and retry guidance The default REST audit limit is 3 requests per minute per client IP. Successful 200 and rejected 429 responses include: - `X-RateLimit-Limit`: configured request limit. - `X-RateLimit-Remaining`: requests left in the current window. - `X-RateLimit-Reset`: absolute Unix epoch reset time. - `Retry-After`: whole seconds until reset. A 429 body also includes integer `retry_after`, which mirrors the delta-seconds `Retry-After` header. Wait at least that long before retrying. Do not treat `X-RateLimit-Reset` as a duration. ## OpenAPI The current schema is at `https://api.machineread.ai/openapi.json`. It is the source for exact models, enums, response statuses, examples, and headers. Interactive Swagger UI is at `https://api.machineread.ai/docs`. [Back to all docs](/docs) --- # https://www.machineread.ai/docs/methodology --- title: "Methodology - MachineRead Docs" description: "The 13 included check groups, three pillars, nine locked rows, and scoring formula behind the MachineRead Essentials audit." source: "https://www.machineread.ai/docs/methodology" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "TechArticle", "headline": "Methodology", "description": "The 13 included check groups, three pillars, nine locked rows, and scoring formula behind the MachineRead Essentials audit.", "url": "https://www.machineread.ai/docs/methodology", "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" } } ``` Docs Docs # Methodology The 13 included check groups, three pillars, nine locked rows, and scoring formula behind the MachineRead Essentials audit. ## Overview MachineRead audits a public website URL for AI visibility, agent accessibility, scrapability, and search discovery readiness. The free Essentials audit runs 13 check groups across 3 pillars, scoring 56 checked points total. ## The three pillars Each pillar has a maximum score cap. The total is 100 points. - **Off-site presence** (cap: 30). Signals that exist outside the website itself: social and entity metadata, Wikipedia and Wikidata entity lookup. - **AI access and scrapability** (cap: 40). Signals that determine whether AI agents and crawlers can fetch, parse, and navigate the site: robots.txt AI bot policy, bot fetch access, semantic HTML, JSON-LD structured data, llms.txt, raw HTML readability, agent protocol discovery. - **SEO** (cap: 30). Signals that affect search engine indexing and discovery: crawl efficiency, canonical and HTTPS, indexing directives, search discovery hints. ## The 13 check groups Off-site pillar (cap 30): - Social and Entity Metadata (2 points) - Wikipedia and Wikidata Entity (4 points) AI access pillar (cap 40): - AI Bot Policy Signals (6 points) - MachineRead Bot Fetch Access (6 points) - Semantic HTML and Agent Navigation (4 points) - JSON-LD Structured Data (5 points) - LLM Text and Markdown Access (5 points) - Raw HTML Readability (4 points) - Agent Protocol Discovery (3 points) SEO pillar (cap 30): - Crawl Efficiency and HTML Performance (3 points) - Canonical and HTTPS (5 points) - Indexing Directives (5 points) - Search Discovery Hints (4 points) A single check group can contain multiple underlying sub-signals. The frontend describes these as groups, not atomic checks. ## Locked vs unlocked rows The 13 check groups above are "unlocked." They are actively checked and scored. The audit also reports 9 "locked" advanced rows in the same `checks` array as the 13 included rows. These are coverage gaps: the free tier acknowledges these signals matter but does not verify them. Locked rows are scored as zero, meaning the gap is visible in the total rather than hidden. ### The 9 locked advanced rows Off-site pillar: - Earned mentions and backlinks (Starter) - Owned social presence (Starter) - Social traction and reviews (Starter) - AI citation share (Pro) AI access pillar: - Extraction fidelity (Starter) - Agent task simulation (Pro) SEO pillar: - Multi-engine index coverage (Starter) - Core Web Vitals (Starter) - Keyword competitor gap (Starter) A locked row scoring zero does not mean the underlying signal is absent. It means the free audit did not verify it. ## The scoring formula There are three separate scoring concepts: 1. **Full Rubric Score** (`overall_score`): 100 point total. Includes locked advanced rows with score 0 until verified. Pillar caps: off-site 30, AI access 40, SEO 30. 2. **Essentials Evidence Score** (`benchmark.score`): Computed only from included, non-locked Essentials rows. Current checked max: 56 points. Used for peer-relative benchmark comparison. 3. **Strict Agent Readiness Score** (`agent_readiness.score`): Measures explicit agent-native discovery and protocol signals. Default scope max: 8 probes. Full scope max: 21 probes. Intended to be stricter than general crawlability or SEO. ## Scope and applicability Presets resolve applicability before checks run. The audit always returns the same 13 included rows and fixed 56-point denominator. Scope changes contextual sub-signals inside those rows, the strict agent-readiness probe maximum, exact included/excluded family copy, and benchmark cohort. New clients should send `preset` and optional `custom_overrides`. The three legacy booleans remain for compatibility: - `include_ecommerce`: commerce schema and metadata expectations. - `include_protocols`: API and agent-protocol expectations. - `include_account_auth`: account and authorization discovery. A selected preset wins over these booleans. Exact families prevent a site from being penalized for irrelevant commerce, account, protocol, local-business, or publisher surfaces. ## Presets The audit supports website-category presets that select check families before the scan runs: - Blog / Content - Corporate / Brand - Services / Local - News / Publisher - SaaS / Product / API - Ecommerce / Catalog - Custom / Power User Custom mode lets the user add or remove supported free check families outside the default preset while keeping impossible or contradictory combinations blocked. Free preset selection is user-controlled and deterministic. [Back to all docs](/docs) --- # https://www.machineread.ai/pricing --- title: "Pricing - MachineRead" description: "Essentials is the only tier available today: it is free and anonymous. Starter and Pro are planned and not available for purchase; prices are not published." source: "https://www.machineread.ai/pricing" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "ItemList", "name": "MachineRead Tier Availability", "itemListElement": [ { "@type": "ListItem", "position": 1, "name": "Essentials", "description": "Free anonymous audit available now. 13 check groups, 56 checked points, 9 locked advanced rows." }, { "@type": "ListItem", "position": 2, "name": "Starter", "description": "Planned tier. Not available for purchase and no price is published." }, { "@type": "ListItem", "position": 3, "name": "Pro", "description": "Planned tier. Not available for purchase and no price is published." } ] } ``` Pricing # Pricing MachineRead has three tiers. Essentials is free, anonymous, and available now. Starter and Pro are planned for future release. No prices are listed for the paid tiers because they are not yet available for purchase. The pricing page does not collect email addresses, does not take payment, and does not have a sign-up button. ## Essentials Free Available now - 13 check groups, 56 checked points - 9 locked advanced rows (coverage gaps, not guesses) - No account required to run an Essentials audit - 4 agent surfaces served statically - Benchmark comparison against peer profiles ## Starter Coming soon Not yet available - Paid-crawler probes (index coverage, backlinks) - Custom audit scope configuration - Weekly monitoring with change detection - All Essentials check groups included ## Pro Coming soon Not yet available - Paid LLM probes (extraction fidelity, citation share) - Advanced agent-readiness checks (task simulation) - Multi-site comparison and benchmarking - All Starter features included The free Essentials audit is the only tier available today. It covers 13 check groups across 3 pillars (off-site presence, AI access and scrapability, and SEO) with a 56-point checked denominator. The 9 advanced rows are locked, meaning they are reported as zero-scored coverage gaps rather than guessed at. This design choice keeps the free tier honest about what it does not measure. Starter and Pro describe planned coverage, not purchasable services. Their listed features may change before launch. Essentials is free and anonymous today. [Read the methodology](/docs/methodology) --- # https://www.machineread.ai/privacy --- title: "Privacy - MachineRead" description: "MachineRead's free audit is anonymous. No account, no API key, no email. The URL is the only input and is not stored. 24-hour log retention. No tracking cookies." source: "https://www.machineread.ai/privacy" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "WebPage", "name": "MachineRead Privacy", "url": "https://www.machineread.ai/privacy", "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" } } ``` Privacy # Privacy Last updated 2026-07-22 MachineRead's free audit is anonymous. There is no account, no API key, and no email required to run a scan. This is a deliberate design choice due to the cost ceiling the project operates under (under $20 per month in compute), which makes account infrastructure and persistent storage impractical for the free tier. As a result, the audit collects the minimum data needed to function, and that data is transient. ## What the audit collects The URL you submit is the only user input. It is fetched server-side, scored against 13 check groups and 56 checked points, and the result is returned to your browser. The URL string itself is not stored in a database. It is used in memory for the duration of the request, then discarded. Meaning the audit does not build a profile of sites you have scanned, does not retain a history of your audits, and does not link your activity across sessions. It is through this server-side fetch that the audit score is computed. The backend validates the URL, blocks private IP ranges to prevent SSRF abuse, fetches the homepage, robots.txt, and sitemap, runs the checks, and returns the JSON response. In other words, the URL goes in, the score comes out, and nothing persists in between. ## Log retention The audit server keeps request logs for 24 hours. These logs contain the submitted URL string, a timestamp, and the response status. The 24-hour window exists for two reasons: rate limiting (the audit endpoint is throttled to 3 requests per minute per IP) and abuse detection (if someone is scanning hundreds of sites in a burst, the logs are how that pattern surfaces). After 24 hours, the URL string is dropped from the logs. The logs are not permanently archived and are not shared with third parties. ## Agent surfaces The 4 agent-readable surfaces (llms.txt, ai-catalog.json, api-catalog.json, and openapi.json) are static files served from the static export. They do not execute code, do not collect data, and do not track who reads them. They are public reference documents for AI agents and crawlers to discover what MachineRead does and how to call the audit API. ## Contact form The contact page uses a mailto link to contact@machineread.ai. The email is delivered through Migadu Micro to a personal Gmail inbox. It is not stored on any MachineRead server. However, once the email leaves the mailto link, it is subject to the email provider's own infrastructure, which is outside MachineRead's control. If you prefer not to use email, there is no alternative contact channel at this time. ## Ko-fi and third-party services The support link on the site points to Ko-fi, which is a third-party service. Ko-fi's own privacy policy applies to any interaction on their platform. MachineRead does not process payments, does not store payment information, and does not receive transaction data from Ko-fi beyond a notification that a contribution was made. ## Tracking and analytics There are no tracking cookies, no analytics scripts, and no fingerprinting. The static export is served as plain HTML, CSS, and JavaScript. There is no client-side telemetry. The only server-side data is the 24-hour request log described above. My recommendation is straightforward: if you want to audit a site without leaving a trace, the free Essentials audit is designed for exactly that. The URL goes in, the score comes out, and after 24 hours even the log entry is gone. Questions about this policy: contact@machineread.ai --- # https://www.machineread.ai/support --- title: "Support - MachineRead" description: "Get help with a MachineRead audit, report a bug, or ask a question about the public methodology and API contract." source: "https://www.machineread.ai/support" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` Support # Help with an audit or integration Send the audited URL, the check name, and the finding text when reporting a problem. Do not include credentials, private URLs, or personal data. MachineRead only evaluates public web surfaces. Support channel [contact@machineread.ai](mailto:contact@machineread.ai?subject=MachineRead%20support) This inbox handles audit questions, reproducible bug reports, and public API or MCP integration questions. Expect a reply within a few days. --- # https://www.machineread.ai/terms --- title: "Terms - MachineRead" description: "MachineRead's free audit is provided as-is. The score is informational, not a guarantee. Not legal, SEO, or compliance advice. User is responsible for actions taken." source: "https://www.machineread.ai/terms" --- ## Structured data copied from the HTML source ```json { "@context": "https://schema.org", "@graph": [ { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai", "logo": "https://www.machineread.ai/og-image.png", "description": "Free public-readiness audit for AI agents, retrieval systems, and search crawlers.", "sameAs": [ "https://github.com/georgejieh/MachineRead-Preview" ], "contactPoint": { "@type": "ContactPoint", "contactType": "Product support", "email": "contact@machineread.ai", "url": "https://www.machineread.ai/contact" }, "address": { "@type": "PostalAddress", "addressCountry": "US" } }, { "@type": "WebSite", "name": "MachineRead", "url": "https://www.machineread.ai" }, { "@type": "SoftwareApplication", "name": "MachineRead Essentials", "description": "Free, deterministic website audit for AI agent accessibility, search-discovery readiness, and public machine-surface visibility.", "applicationCategory": "WebApplication", "operatingSystem": "Any", "offers": { "@type": "Offer", "price": "0", "priceCurrency": "USD" } } ] } ``` ```json { "@context": "https://schema.org", "@type": "WebPage", "name": "MachineRead Terms", "url": "https://www.machineread.ai/terms", "publisher": { "@type": "Organization", "name": "MachineRead", "url": "https://www.machineread.ai" } } ``` Terms # Terms Last updated 2026-07-22 The free Essentials audit is provided as-is. The score it produces is informational, meaning it reflects what the audit's 13 check groups and 56 checked points found at the moment the URL was scanned. It is not a guarantee of any outcome: not a search ranking, not a citation rate, not a conversion lift, not an AI model's likelihood of referencing your site. The score is a snapshot of machine-readable signals, and the signals can change between the scan and any action you take based on it. ## What the audit does and does not cover MachineRead does not warrant that the audit will catch every issue on your site. The free tier checks 56 points across 3 pillars (off-site presence, AI access and scrapability, SEO). However, 9 advanced rows are locked, meaning they are reported as zero-scored coverage gaps rather than guesses at the actual values. Due to this locked-row design, a low score in an advanced row does not mean the underlying signal is absent. It means the free audit did not verify it. As a result, acting on a locked row's zero score as if it were a confirmed failure would be a misreading of what the audit reports. The audit does not constitute legal advice, SEO advice, or compliance advice. It is a technical scan of publicly visible machine-readable signals. If you need a legal opinion, an SEO strategy tailored to your competitive landscape, or a compliance audit for a regulated industry, the free audit is not a substitute for professional consultation. ## Locked rows and deep-dives The audit's locked rows are zero-scored coverage gaps. It is through these locked rows that the audit is honest about what it does not check. The 9 locked rows cover earned mentions and backlinks, social traction, AI citation share, extraction fidelity, agent task simulation, multi-engine index coverage, Core Web Vitals, and keyword competitor gaps. In other words, the free tier tells you these signals matter, tells you it does not measure them, and scores them as zero so the gap is visible rather than hidden. The deep-dive blog posts linked from the launch page explain the methodology behind individual check groups. They are not a substitute for the methodology reference at /docs/methodology, which is the canonical and up-to-date source for how scoring works, what each pillar cap is, and how the locked-versus- unlocked distinction affects your total. ## Your responsibility You are responsible for how you act on the audit's findings. The audit reports what it found. It does not prescribe specific changes, does not guarantee that fixing a flagged issue will produce a measurable result, and does not track whether you followed up. The audit is a diagnostic tool (a snapshot of machine-readable signals at a point in time), and the decision to act on any finding is yours. ## Service continuity MachineRead may modify or discontinue the free audit at any time. The audit runs on infrastructure with a monthly cost ceiling under $20, which means the service is not backed by redundant hosting or guaranteed uptime. If the service is temporarily unavailable or permanently retired, there is no Service Level Agreement and no refund obligation, because the free tier has no payment to refund. My recommendation: treat the audit as a starting point, not a final verdict. Run it, read the findings, cross-reference the methodology at /docs, and decide what to fix. The score is a signal, and signals are most useful when you understand the mechanism behind them. Questions about these terms: contact@machineread.ai