# robots.txt for WhitelistVideo (whitelist.video) # Last Updated: 2026-05-26 User-agent: * Allow: / # Disallow internal/test pages Disallow: /api/ Disallow: /admin/ Disallow: /dashboard/ Disallow: /result Disallow: /download/success Disallow: /blog?author= # Disallow Next.js build artifacts (except static assets needed for rendering) Disallow: /_next/ Allow: /_next/static/media/ # Disallow raw markdown content from search engine indexing # (AI crawlers have explicit Allow rules below) Disallow: /content/ # Disallow help pages (client-side redirect to ChatGPT) Disallow: /help Disallow: /*/help # =========================================== # Content Signals (RFC draft-romm-aipref-contentsignals) # https://contentsignals.org/ # =========================================== # Opt in to AI search indexing, content citation, AND training # (ai-train=yes is intentional — we trade training access for AI-search visibility) Content-Signal: ai-train=yes, search=yes, ai-input=yes # Sitemap location Sitemap: https://whitelist.video/sitemap.xml # =========================================== # LLM/AI Context Files (llms.txt standard) # =========================================== # Concise context for LLMs (recommended <10KB) # https://llmstxt.org/ LLMs-Txt: https://whitelist.video/llms.txt # Extended context with full product details LLMs-Full-Txt: https://whitelist.video/llms-full.txt # Markdown content for AI crawlers # Documentation: https://whitelist.video/content/docs/ # Blog posts: https://whitelist.video/content/blog/ # Manifest: https://whitelist.video/content/manifest.json # =========================================== # Traditional Search Engines # =========================================== User-agent: Googlebot Allow: / Allow: /_next/static/media/ Disallow: /_next/ Disallow: /content/ Disallow: /help Disallow: /*/help User-agent: Bingbot Allow: / Allow: /_next/static/media/ Disallow: /_next/ Disallow: /content/ Disallow: /help Disallow: /*/help # =========================================== # AI Search/Citation Crawlers (drives visibility) # These crawlers index content for AI search results # =========================================== # OpenAI SearchBot - ChatGPT Search results User-agent: OAI-SearchBot Allow: / Allow: /content/ # Perplexity AI search User-agent: PerplexityBot Allow: / Allow: /content/ # Claude search User-agent: Claude-SearchBot Allow: / Allow: /content/ # =========================================== # User-Initiated AI Fetchers # When a user asks an AI to read a specific page # =========================================== User-agent: ChatGPT-User Allow: / Allow: /content/ User-agent: Claude-User Allow: / Allow: /content/ User-agent: Perplexity-User Allow: / Allow: /content/ # Apple Intelligence / Siri User-agent: Applebot Allow: / Allow: /content/ User-agent: Applebot-Extended Allow: / Allow: /content/ # Microsoft Copilot / Bing AI User-agent: FacebookBot Allow: / Allow: /content/ User-agent: meta-externalagent Allow: / Allow: /content/ # =========================================== # AI Training Crawlers (rate-limited) # These crawlers gather data for model training # =========================================== User-agent: GPTBot Crawl-delay: 10 Allow: / Allow: /content/ User-agent: Google-Extended Crawl-delay: 10 Allow: / Allow: /content/ User-agent: ClaudeBot Crawl-delay: 10 Allow: / Allow: /content/ User-agent: CCBot Crawl-delay: 10 Allow: / Allow: /content/ User-agent: Bytespider Crawl-delay: 10 Allow: / Allow: /content/ # Cohere AI User-agent: cohere-ai Crawl-delay: 10 Allow: / Allow: /content/ # You.com AI User-agent: YouBot Crawl-delay: 10 Allow: / Allow: /content/ # =========================================== # Additional AI Crawlers (explicit Allow) # =========================================== # Anthropic web retrieval User-agent: anthropic-ai Allow: / Allow: /content/ # xAI / Grok User-agent: GrokBot Allow: / Allow: /content/ # Diffbot AI User-agent: Diffbot Allow: / Allow: /content/ # Meta fetcher (link previews + AI) User-agent: Meta-ExternalFetcher Allow: / Allow: /content/ # Omgili / Webz.io data crawler User-agent: Omgilibot Allow: / Allow: /content/ # Google catchall (Gemini, etc.) User-agent: GoogleOther Allow: / Allow: /content/ # ImagesiftBot (image AI) User-agent: ImagesiftBot Allow: / Allow: /content/ # =========================================== # Blocked: Aggressive Scrapers # =========================================== User-agent: AhrefsBot Disallow: / User-agent: SemrushBot Disallow: / User-agent: DotBot Disallow: / User-agent: MJ12bot Disallow: /