BusinessDiverBot

What it is

BusinessDiverBot is the crawler of Business Diver, a company research service. It runs only when a Business Diver user asks for a report on a company, and the text it reads is used as input to the AI model that writes that report. Business Diver does not use what it collects to train or fine-tune AI models.

How it identifies itself

Every page it reads is requested with this User-Agent: BusinessDiverBot/0.1 (+https://businessdiver.com/bot; info@businessdiver.com)

What it fetches

It calls the public APIs of official registries and exchange feeds and of reference services such as Wikidata, Wikipedia and GitHub, and it reads web pages: the company's own website, press-release and market-news search results, and the pages that searches lead to. It never signs in and makes no attempt to get past login pages, paywalls or bot challenges. It fetches at most two pages at a time from one website, and starts the next page from that site only after the site's Crawl-delay, or one second when none is set.

What it honours

Before it fetches a web page, it reads the site's robots.txt and follows the group for BusinessDiverBot, or the * group when there is none. Content-Signal and Content-Usage values in robots.txt that refuse search or AI input (search=no, ai-input=no, ai-use=n) are treated as a refusal, and so is a TDMRep reservation in tdmrep.json or in a page's tdm-reservation header or meta tag. A signal that refuses only AI training, such as ai-train=no, does not stop it, because Business Diver trains no models. A 401 or 403 answer or a bot challenge on robots.txt counts as a refusal of the whole site, and when robots.txt cannot be read the site is skipped for that report. A refused page is never retried under another identity.

Search

Search queries that find candidate pages are sent to DuckDuckGo and, for news, also to Bing, through a third-party search library that does not identify as BusinessDiverBot; every page we then read ourselves is fetched as BusinessDiverBot after the checks above.

Blocking it and contacting us

To block BusinessDiverBot, add a group with the lines User-agent: BusinessDiverBot and Disallow: / to your robots.txt; a change takes effect within 24 hours. To ask us to remove material taken from your site, or with any other question, write to info@businessdiver.com. Business Diver is operated by ResourceHub Cph (CVR DK46200462), Copenhagen, Denmark, and our privacy policy explains how we handle information about people.