Technical website checks

Can AI crawlers access your website? A practical checking guide

A robots.txt rule, an HTTP response and a source citation describe different stages of discovery. Checking just one of them cannot establish that an AI system visited or used a page. Start with the exact URL and work through the evidence.

By the Sikwati team · Practical guide

1. Choose the page and the agent you want to investigate

Write down the full public URL, including its hostname and path. A check of your homepage does not establish access to a nested service page or a different subdomain. Name the crawler you are investigating instead of treating all ‘AI bots’ as interchangeable.

Decide your own access policy before editing anything. You may want certain public pages discoverable while limiting other uses. Do not remove privacy protections, login requirements or bot controls simply to make a checker show green.

2. Read the relevant robots.txt rules

Open /robots.txt on the same host and inspect the rules relevant to the agent and path. A specific user-agent group can matter more than a wildcard group; check the provider's documented interpretation before changing rules.

The simplified example below requests that cooperating crawlers avoid /private/. It does not protect that directory from access. Use authentication for private content. Robots rules govern cooperative crawling and are not a security boundary or a reliable way to remove a URL from search results.

User-agent: *
Disallow: /private/

3. Check what the server actually returns

Ask your developer to inspect the final response for the public page. Look for login redirects, access-denied responses, rate limits, bot challenges and missing content. A page that works in your signed-in browser may return something different to an unauthenticated request.

Review CDN and hosting rules as well as the application. A robots permission cannot override a firewall block. Likewise, changing a request's user-agent text to a bot name is only a diagnostic—it does not reproduce the provider's network, request behaviour or verified identity.

  • Does the exact public URL resolve without a login?
  • Are redirects intentional and bounded?
  • Does the response contain the actual service information rather than a challenge?
  • Do hosting logs show a denial or rate limit at the relevant time?
  • Can a narrow correction solve the issue without disabling security protections globally?

4. Inspect readable content and discovery paths

Check whether key facts are present in the returned page and whether visitors can reach it from relevant internal links. If essential service details are only visible after an interaction, discuss rendering with your developer. Different consumers may process pages differently; do not assume that one fetch proves every system can read it.

Structured data and llms.txt are separate checks, not access overrides. An llms.txt file is an optional convention, not a universal inclusion requirement. Google explicitly says its AI search features do not require special AI files or special schema. Do not extrapolate that statement into a guarantee about another provider.

5. Retest the narrow issue and retain the evidence

After an authorised change, repeat the same URL check and record the old finding, change and new response. If you need proof of a crawler visit, inspect server logs and validate the agent using its provider's current published verification method. A user-agent string alone can be spoofed.

Sikwati's free crawler checker inspects homepage robots permissions for its listed agents. Use it to spot rules worth reviewing, not to certify actual provider access, a complete site crawl or future citations. Its readiness and schema checks answer other technical questions.

Even a successful visit does not prove that the page will be selected as a source. Track actual saved answers separately and avoid repeatedly weakening security when the missing evidence is an AI mention rather than a failed request.

Put the guide into practice

Free website checks inspect technical signals; they do not measure AI mentions. Ongoing Sikwati tracking is a paid feature. Neither a technical score nor an observed citation guarantees visibility or business results.

Check homepage crawler permissions free

Explore the tracking and page-improvement workflow

Sources and further reading

The checklists and worked examples are Sikwati editorial guidance, not measured customer results. Provider-specific guidance applies to that provider, not every AI system.

Cookies

We use essential cookies to run the site. Analytics is optional and helps us improve it.

See our Privacy Policy and Cookies Policy.

Can AI crawlers access your website? A practical checking guide | Sikwati