# KeeperTally crawler policy # # The rule behind every line below: a crawler that SENDS READERS is welcome, a # crawler that only TAKES is not. Those are two different sets of bots, and they # are controlled separately — which is the whole reason this file is more than # three lines long. # # Search, and AI answers that cite the source and link back -> allowed. # They are how anyone finds a new site. Blocking them to protect the data # would cost the traffic the site exists to earn. # # Crawlers that collect text to train a model -> disallowed. # Nothing comes back: no link, no citation, no reader. The scores are the # work, and they are not a free input to someone else's product. # # Two things to be honest about. This file is a REQUEST, not a wall: well-run # crawlers honour it, badly-run ones ignore it, and nothing here stops a person # copying a page by hand. What it does do is state the reservation in the # machine-readable place, which is what UK and EU text and data mining law # expects of a publisher who means to keep those rights. /terms says the same # thing in words and /.well-known/tdmrep.json says it a third way. All three # must agree: change one, change all three. # # Adding a bot: put it in the group that matches what it DOES, not who owns it. # Several companies run one of each. # --- Search engines ------------------------------------------------------ # The business. Everything open, no exceptions. User-agent: Googlebot Allow: / User-agent: Bingbot Allow: / User-agent: DuckDuckBot Allow: / User-agent: Applebot Allow: / # --- AI answers that cite and link --------------------------------------- # These fetch a page to answer a question and name the source, which is a # referral like any search result. Two kinds appear here: the indexers, and the # "-User" agents that fetch a page only because a reader asked for it. Blocking # the second kind means refusing a request an actual human made. User-agent: OAI-SearchBot Allow: / User-agent: ChatGPT-User Allow: / User-agent: Claude-SearchBot Allow: / User-agent: Claude-User Allow: / User-agent: PerplexityBot Allow: / User-agent: Perplexity-User Allow: / # --- Model training ------------------------------------------------------ # Content goes into model weights. No link, no citation, no reader, ever. # Reserved here, in /terms, and under UK and EU text and data mining law. User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: CCBot Disallow: / User-agent: meta-externalagent Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Bytespider Disallow: / # Deliberately NOT blocked, against the pattern above, and the exception is # worth writing down. Google-Extended is a mixed token: blocking it stops # Gemini training AND stops the site being cited as a source inside Gemini # answers, which is referral traffic. It has no bearing on Google Search # ranking or on AI Overviews either way, so blocking it would cost readers and # buy almost nothing. User-agent: Google-Extended Allow: / # --- Everything else ----------------------------------------------------- User-agent: * Allow: / Sitemap: https://keepertally.com/sitemap.xml