Skip to content

Add robots.txt (allow all crawlers)#694

Open
fredericsimard wants to merge 3 commits into
mainfrom
chore/add-robots-txt/2026-07-24
Open

Add robots.txt (allow all crawlers)#694
fredericsimard wants to merge 3 commits into
mainfrom
chore/add-robots-txt/2026-07-24

Conversation

@fredericsimard

Copy link
Copy Markdown
Contributor

Summary

Adds docs/en/robots.txt, which the English build copies verbatim to the site root (generated/robots.txtgtfs.org/robots.txt).

  • Lists the major AI/search crawlers (GPTBot, OAI-SearchBot, ChatGPT-User, Google-Extended, Claude-Web, ClaudeBot, anthropic-ai, CCBot, PerplexityBot, Amazonbot, Bytespider, Applebot-Extended) each with Allow: /.
  • Catch-all User-agent: * with Content-Signal: ai-train=yes, search=yes, ai-input=no and Allow: /.

Why here (not per-language)

robots.txt is only honored at the site root. Only the English config builds to generated/ (the root); every other language builds to generated/<lang>/, where a robots.txt would be ignored by crawlers. So a single file under docs/en/ is the correct and sufficient placement.

Verification

Built the English site locally and confirmed:

  • generated/robots.txt exists at the site root and is byte-for-byte identical to the source.
  • No second or auto-generated robots.txt anywhere in the build (no collision), and no theme wrapping/alteration.

Note for reviewers

The Content-Signal was intentionally set to ai-train=yes to stay consistent with the allow-all rules (an earlier draft had ai-train=no alongside allow-all, which was contradictory). Confirm the allow-all training posture is the intended policy.

fredericsimard and others added 3 commits July 24, 2026 05:45
Adds docs/en/robots.txt, which the English build copies to the site
root (generated/robots.txt -> gtfs.org/robots.txt) — the only location
crawlers read, so it is not duplicated per language.

Blocks AI-training crawlers (GPTBot, Google-Extended, CCBot,
PerplexityBot, Amazonbot, Bytespider, Applebot-Extended, ChatGPT-User)
while allowing search/assistant bots (OAI-SearchBot, Claude-Web,
ClaudeBot, anthropic-ai), and sets a Content-Signal of
ai-train=no, search=yes, ai-input=no for all other agents.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Switch every per-bot rule from Disallow to Allow, so AI-training and
assistant/search crawlers (GPTBot, ChatGPT-User, Google-Extended,
CCBot, PerplexityBot, Amazonbot, Bytespider, Applebot-Extended, and
the already-allowed OAI-SearchBot/Claude-Web/ClaudeBot/anthropic-ai)
may all access the site.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Aligns the catch-all Content-Signal with the now allow-all crawler
rules: training was previously signalled off (ai-train=no) while every
bot was allowed, which was contradictory. Now ai-train=yes, search=yes,
ai-input=no.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
@fredericsimard fredericsimard self-assigned this Jul 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant