# Standard Web Crawlers User-agent: * Allow: / Crawl-delay: 10 # === AI SEARCH CRAWLERS === # OpenAI ChatGPT User-agent: OAI-SearchBot Allow: / Disallow: /api/ Crawl-delay: 0.5 User-agent: ChatGPT-User Allow: / User-agent: GPTBot Allow: / Disallow: /api/ Crawl-delay: 2 # Perplexity AI User-agent: PerplexityBot Allow: / Disallow: /api/ Crawl-delay: 0.5 # Google Gemini User-agent: Google-Extended Allow: / Disallow: /api/ # Anthropic Claude User-agent: Claude-Web Allow: / Disallow: /api/ User-agent: ClaudeBot Allow: / Disallow: /api/ Crawl-delay: 1 # X/Twitter Grok User-agent: GrokBot Allow: / Disallow: /api/ Crawl-delay: 1 # You.com User-agent: YouBot Allow: / Disallow: /api/ Crawl-delay: 1 # Meta AI User-agent: FacebookBot Allow: / Disallow: /api/ Crawl-delay: 1 # === TRAINING CRAWLERS (BLOCKED) === # Meta's AI-training crawler (NOT facebookexternalhit link previews, which # stay allowed). On 2026-07-27 it swept every customer domain with a # multi-IP swarm and saturated the SSR origin — also blocked at the # Traefik edge; this entry is the polite layer. User-agent: Meta-ExternalAgent Disallow: / User-agent: meta-externalagent Disallow: / User-agent: CCBot Disallow: / User-agent: Bytespider Disallow: / User-agent: PetalBot Disallow: / User-agent: SemrushBot Disallow: / User-agent: AhrefsBot Disallow: / User-agent: AwarioBot Disallow: / # Sitemap Sitemap: https://restaurant-sahaj.de/sitemap.xml # LLMs.txt (AI optimization file) # https://restaurant-sahaj.de/llms.txt