# itcosc.com, crawler & content-usage policy # Goal: WELCOME search engines and AI answer engines (they cite & link back to us = leads). # REFUSE AI model-training scrapers (no business value; enables content theft by competitors). # Default rules, applies to all crawlers, incl. AI answer engines (ChatGPT search, # Perplexity, etc.) which are intentionally allowed so they cite us. User-agent: * Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / Disallow: /thank-you Disallow: /thank-you.html Disallow: /coming-soon Disallow: /coming-soon.html # --- AI TRAINING scrapers: blocked (take content, send no traffic) --- # Each is the training-only bot; the matching search/answer bot stays allowed above. User-agent: GPTBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: ClaudeBot Disallow: / User-agent: anthropic-ai Disallow: / User-agent: CCBot Disallow: / User-agent: Bytespider Disallow: / User-agent: meta-externalagent Disallow: / # Additional training-only crawlers (added 2026-07-02) User-agent: Amazonbot Disallow: / User-agent: cohere-ai Disallow: / User-agent: AI2Bot Disallow: / User-agent: Diffbot Disallow: / User-agent: omgili Disallow: / User-agent: omgilibot Disallow: / # NOTE: facebookexternalhit, OAI-SearchBot, ChatGPT-User, PerplexityBot, and # Claude-Web are deliberately NOT blocked, they power social link-previews and # AI answer citations (these are covered by the wildcard Allow: / above). Sitemap: https://itcosc.com/sitemap.xml