FlowhelpBot
FlowhelpBot is the web crawler of flowhelp. It reads websites that flowhelp customers add as knowledge sources, so that their AI assistant can answer visitors’ questions from those pages.
Which pages it reads
FlowhelpBot reads a website only after its address has been added as a knowledge source of a flowhelp assistant, normally by the flowhelp customer who runs that assistant. It does not crawl the web on its own. flowhelp does not check whether whoever added the address owns that site; robots.txt, described below, works no matter who added it.
From that address it looks for pages on the same site: first in the sitemaps listed in robots.txt and at /sitemap.xml, and if those give no pages, by following links from the starting page, at most 3 links deep. The same site means the same host name, with or without www. Other subdomains and other domains are left out, and so are links to files such as images, PDFs and archives.
Each crawl of a knowledge source covers at most as many pages as the customer’s plan allows: 25 on Trial, 200 on Starter, 2,000 on Growth and 10,000 on Scale. On Enterprise, a single crawl covers at most 10,000 pages.
It comes back when the customer retrains the assistant, by hand or automatically every week, day or hour, depending on the plan. Every visit starts over with robots.txt and the search for pages.
What it reads becomes the knowledge of that customer’s assistant, which uses it to answer visitors’ questions. flowhelp does not use it to train AI models.
How to recognise it
Requests from FlowhelpBot itself, for robots.txt, sitemaps and pages, carry this user agent:
FlowhelpBot/0.1 (+https://flowhelp.ai/bot)Each of its own requests follows at most 5 redirects. The target of a redirect is fetched even when it is on another host, and robots.txt is not checked again for it.
When a page returns an error status (for example 403, 404 or 5xx), needs more than 5 redirects, cannot be reached or does not answer in time, or has less than 200 characters of text, FlowhelpBot opens it again in a headless Chromium browser. That browser sends the user agent of headless Chrome (it contains HeadlessChrome), not that of FlowhelpBot, and like any browser it also loads the page’s scripts, styles, images and other resources, including those on other hosts. It starts only from addresses that robots.txt allows and follows redirects on its own.
robots.txt
Before every crawl, FlowhelpBot fetches robots.txt from the host of the address that was added. It follows the group for User-agent: FlowhelpBot, in any letter case, or the group for * when there is none, and it does not read pages that those rules disallow. robots.txt itself and sitemap files are fetched regardless of Disallow.
If that group sets Crawl-delay, FlowhelpBot waits that long between the pages it reads from the host, up to 30 seconds; a larger value counts as 30. robots.txt and sitemaps are fetched one after another without that pause. Without Crawl-delay it may read several pages of the site at the same time.
If robots.txt cannot be read, for example because it does not exist, the server returns an error or does not answer in time, FlowhelpBot treats the whole site as allowed.
To keep FlowhelpBot off your whole site, add this to your robots.txt:
User-agent: FlowhelpBot
Disallow: /To keep it away from part of the site only, list those paths in the same group:
User-agent: FlowhelpBot
Disallow: /account/
Disallow: /checkout/FlowhelpBot reads robots.txt again at the start of every crawl, so a new rule applies from the next visit; pages already queued in a crawl that has started may still be read. Block it in robots.txt rather than by its user agent on the server alone: a page that returns an error status, including one the server refuses because of the user agent, is opened again in the headless browser described above, which does not send the FlowhelpBot user agent.
Requests during assistant setup
When a customer enters a site address while setting up an assistant, the flowhelp application also fetches that page, and sometimes the icon it links to, to suggest a colour for the chat widget. These requests carry a different user agent:
flowhelp-onboarding/1.0 (+https://flowhelp.ai)They do not read robots.txt and follow at most 3 redirects. The icon may be on another host. The suggested colour is kept for 24 hours, so the same site is normally not fetched again in that time.
Contact
Have a question, or is FlowhelpBot causing trouble on your site? Write to hello@flowhelp.ai. Include your domain and, if you can, when the requests happened.