The Great Gpt Firewall banner
samber samber

The Great Gpt Firewall

AI community

Description

🤖 A curated list of websites that restrict access to AI Agents, AI crawlers and GPTs

Installation

This entry records only its repository, not the path inside it, so there is no exact command to give. Open the source below and copy the folder into ~/.claude/skills/, or the file into ~/.claude/agents/.

README

The Great GPT Firewall 📛

This collection is a curated list of websites that employ the `robots.txt` file to restrict access to AI Agents, AI crawlers and GPTs.

It will be updated monthly.

We need a plan!

User agents & robots.txt

The `robots.txt` file allows website owners to control and limit the access of these user agents to certain areas of their website by specifying rules and directives.

# OpenAI’s web crawler: GPT3.5, GPT4, ChatGPT
# https://platform.openai.com/docs/bots
User-agent: GPTBot

# ChatGPT plugins
# https://platform.openai.com/docs/bots
User-agent: ChatGPT-User

# OpenAI Search bot
# https://platform.openai.com/docs/bots
User-agent: OAI-SearchBot

# Google's web crawler: Bard, VertexAI, Gemini
# https://blog.google/technology/ai/an-update-on-web-publisher-controls/
User-agent: Google-Extended

# Apple's web crawler, dedicated to GenAI projects
# https://support.apple.com/en-us/119829
User-agent: Applebot-Extended

# Claude
User-agent: anthropic-ai

# Claude Bot
User-agent: ClaudeBot

# Claude web
User-agent: Claude-Web

# Amazonbot
# https://developer.amazon.com/amazonbot
User-agent: Amazonbot

# Cohere
User-agent: Cohere-ai

# Perplexity
User-agent: PerplexityBot

# You
# https://about.you.com/fr/youbot/
User-agent: YouBot

# Common Crawl
# https://commoncrawl.org/ccbot
User-agent: CCBot

# Omglibot: webz.io
# https://webz.io/blog/web-data/what-is-the-omgili-bot-and-why-is-it-crawling-your-website/
User-agent: Omgilibot
User-agent: Omgili
User-agent: Webzio-Extended

# Facebook: Llama
# https://developers.facebook.com/docs/sharing/bot/
User-agent: FacebookBot

# Facebook
# https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/
User-agent: Meta-ExternalAgent

# ByteDance: Duobao
User-agent: Bytespider

# Ai2
# https://allenai.org/crawler
User-agent: Ai2bot
User-agent: Ai2Bot-Dolma

# Diffbot
User-