# MWM — robots.txt # # Policy (revised 2026-06): # - OPEN to all compliant crawlers, INCLUDING AI training / grounding — we # accept training use in exchange for AI visibility and citations. # - EXCEPT our proprietary IP surfaces: /apps (app intelligence) and /rankings # (ranking data), incl. localized /{lang}/ paths. These are withheld from AI # training / ingestion crawlers; search engines keep full access so the pages # stay indexed for SEO. # # Agent classification aligned with knownagents.com/insights taxonomy. # Last reviewed: 2026-09-11. # ============================================================ # ALLOW — Traditional search, AI search, citation, user-triggered, # agentic task automation, coding agents, SEO tools. # (All share the same rule set via stacked User-agent lines.) # ============================================================ User-agent: Googlebot User-agent: bingbot User-agent: Applebot User-agent: PetalBot User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: Claude-SearchBot User-agent: Claude-User User-agent: Claude-Code User-agent: PerplexityBot User-agent: Perplexity-User User-agent: meta-webindexer User-agent: meta-externalfetcher User-agent: Amzn-SearchBot User-agent: Amzn-User User-agent: AmazonBuyForMe User-agent: Google-CloudVertexBot User-agent: Google-NotebookLM User-agent: Gemini-Deep-Research User-agent: Google-Agent User-agent: GoogleAgent-Mariner User-agent: GoogleOther User-agent: DuckAssistBot User-agent: YouBot User-agent: Bravebot User-agent: kagi-fetcher User-agent: MistralAI-User User-agent: TavilyBot User-agent: ExaBot User-agent: ExaSearchBot User-agent: LinkupBot User-agent: Diffbot User-agent: Manus-User User-agent: Devin User-agent: AhrefsBot User-agent: SemrushBot Allow: / # ============================================================ # RESTRICT — AI training / ingestion crawlers. # Open to the whole site EXCEPT our proprietary IP surfaces below # (/apps + /rankings, including localized /{lang}/ paths). # ============================================================ User-agent: GPTBot User-agent: ClaudeBot User-agent: anthropic-ai User-agent: Claude-Web User-agent: CloudVertexBot User-agent: Applebot-Extended User-agent: Google-Extended User-agent: Amazonbot User-agent: meta-externalagent User-agent: FacebookBot User-agent: CCBot User-agent: Bytespider User-agent: imageSpider User-agent: PanguBot User-agent: Cohere-training-data-crawler User-agent: cohere-ai User-agent: Timpibot Disallow: /apps/ Disallow: /rankings/ Disallow: /*/apps/ Disallow: /*/rankings/ # ============================================================ # DENY - crawlers with no measured return. # Measured 2026-09-11 over 7 days: 147k (ShapBot) and 38k (SleepBot) # origin requests, zero attributed visitor over 30 days, and no # identifiable operator - ShapBot advertises no URL at all. # Also enforced at the edge (Cloud Armor rule 1005): neither agent is # assumed to honour this file. # ============================================================ User-agent: ShapBot User-agent: SleepBot Disallow: / # ============================================================ # Default — all other crawlers # ============================================================ User-agent: * Allow: / Sitemap: https://mwm.ai/sitemap.xml