Consentire HootyBot nella protezione dai bot
Sei qui perché un prodotto di protezione dai bot impedisce a HootyBot di leggere un sito. Questa pagina spiega cosa cambiare, prodotto per prodotto.
Che cos’è HootyBot: il crawler che legge un sito perché il suo assistente hooty.to possa rispondere alle domande partendo da esso. Si identifica a ogni richiesta e rispetta robots.txt. Non esegue JavaScript, quindi non può superare una verifica del browser né risolvere un captcha, e non ci proverà. Vedi hooty.to/bot.
Che cosa consentire
Qualunque prodotto tu usi, quello che stai autorizzando è questo:
- Lo User-Agent contiene
HootyBot - Stringa completa:
Mozilla/5.0 (compatible; HootyBot/1.0; +https://hooty.to/bot) - Le richieste portano anche
From: bot@hooty.to
HootyBot esegue la scansione da un provider di hosting con indirizzi a rotazione, quindi non esiste un elenco fisso di IP da autorizzare. Fai invece corrispondere la stringa User-Agent. Se il tuo prodotto supporta solo elenchi di IP, scrivi a bot@hooty.to e definiremo insieme l’intervallo di uscita attuale.
Dove una sezione qui sotto indica nomi di menu esatti, li abbiamo verificati sulla documentazione del fornitore. Dove non lo fa, lo diciamo invece di tirare a indovinare e ti spieghiamo cosa chiedere al suo supporto. I nomi dei menu cambiano anche tra una versione e l’altra: considera ogni percorso come un punto di partenza.
Spesso i siti usano due prodotti insieme. Un servizio specializzato anti-bot dietro a una CDN generica è una configurazione normale, quindi il prodotto che ci blocca può non essere quello che consideri «la nostra CDN». Se una modifica su uno non basta, controlla l’altro.
Le istruzioni per prodotto qui sotto restano in inglese in tutte le lingue. Nominano menu e pulsanti esattamente come li scrive ciascun fornitore nella propria console, e un’etichetta tradotta che non corrisponde a ciò che vedi ti porterebbe all’impostazione sbagliata.
Cloudflare
You will see server: cloudflare and a cf-ray header on responses. A challenge specifically carries cf-mitigated: challenge.
Free plan, read this first. Bot Fight Mode on the free plan runs outside the rules engine, and Cloudflare’s own documentation states that it cannot be skipped or bypassed by WAF custom rules or Page Rules. If that is what is blocking HootyBot, an allow rule will not help. Your options are to turn Bot Fight Mode off under Security → Settings → Bot traffic, or to wait for HootyBot to be listed as a Cloudflare verified bot, since verified bots are excluded from Bot Fight Mode by default.
On paid plans, a custom rule works:
- Dashboard → your site → Security rules → Create rule → Custom rules.
- Expression:
http.user_agent contains "HootyBot". - Action: Skip, then tick what to skip, including Managed Rules, rate limiting rules and Super Bot Fight Mode rules.
- Deploy it above any rule that blocks.
With Super Bot Fight Mode (Pro and Business), also open Security → Settings → Bot traffic and set Verified Bots to Allow.
Cloudflare also has a separate AI Crawl Control gate. If it is on, open it in the dashboard, find HootyBot in the Crawlers list, and set its action to Allow.
Qrator Labs
Responses carry server: QRATOR. The check is answered with an unusual HTTP 401 whose whole body is a script tag pointing at /__qrator/qauth.js, and a qrator_jsr cookie is set.
Whitelisting is configured by you in your own Qrator account, not by Qrator support on our behalf, and there is no public bot registry for us to join. We were not able to verify the current menu path, so rather than guess: in your Qrator dashboard, look for the exceptions or whitelist rules for the domain, and add one matching the User-Agent header against HootyBot. If you cannot find it, ask Qrator support for exactly that.
Qrator can also stop answering a source address entirely, with no block page and no error, so a crawler simply times out. If hooty.to reports that your site did not answer at all, this is the likely cause.
KillBot
KillBot is installed at the site itself rather than as a CDN, so the Server header stays your own. It answers with HTTP 200 and a page whose title begins KillBot user verification, echoing the visitor’s address back, and it bounces links to URLs ending ?from=capt. Because the status says success, a crawler that does not inspect the body will store the verification page as if it were your content. The vendor’s site is killbot.ru.
We could not find any public documentation for allow-listing a bot in KillBot, and its own site is behind the same product, so we will not invent a procedure. The honest instruction is to ask whoever installed it, which is usually your hosting provider, agency or KillBot reseller, to permit requests whose User-Agent contains HootyBot. KillBot already ships exceptions for search-engine crawlers, so the mechanism exists.
DDoS-Guard
Responses carry server: ddos-guard and cookies beginning __ddg. Those cookies appear on ordinary traffic too, so seeing them does not by itself mean anything is being blocked.
The allow-list here matches on IP or network only. It has no User-Agent condition, so allow-listing HootyBot by name is not possible in this feature. Write to bot@hooty.to and we will give you the address range to enter.
- Личный кабинет → Домены → select the domain → Чёрный/белый список.
- Добавить правило, then enter the IP or network.
- Set the action to Пропускать and save.
Networks wider than /24 need the paid “Маска чёрного/белого списка” add-on.
StormWall
StormWall puts no product name in its headers: the Server value is a bare IP address from its own network. A block reads Access to resource was blocked. with a support ID, and the JavaScript check sets cookies named __js_p_, __jhash_ and __jua_.
We could not verify the current path in the StormWall client area, so ask their support, or look for the filtering rules for your site, and request an exception for requests whose User-Agent contains HootyBot.
Variti
Variti identifies itself plainly, with a Server: Variti/<version> header, a 307 redirect for the check, and ipp_uid and ipp_key cookies.
We could not verify the console path. Ask Variti support, or look in the filtering rules for the resource, for an exception matching the User-Agent header against HootyBot.
Servicepipe
Servicepipe leaves no trace in the response: a bare nginx server header and no documented cookie names, so there is no way to confirm it from the outside. It challenges with a cookie check, a JavaScript check, or a captcha, the last of which is off unless support enabled it.
We could not verify a self-service path. Their published support addresses are support@servicepipe.ru and cybert@servicepipe.ru; ask them to permit requests whose User-Agent contains HootyBot.
Yandex SmartCaptcha
SmartCaptcha is a widget your developers embed, plus a server-side token check in your own backend, rather than something sitting in front of the site.
You can add a rule in the Yandex Cloud console (Создать капчу or edit an existing one → Добавить правило) with a condition on an HTTP header, matching User-Agent by value, prefix or regular expression. Yandex’s own documentation uses a User-Agent example, so matching HootyBot is supported.
The catch: a rule chooses which captcha to show, and we found no documented action that skips the captcha altogether. So the reliable fix is in your own backend: skip the SmartCaptcha token validation for requests whose User-Agent contains HootyBot.
Imperva Incapsula
Responses carry x-cdn: Imperva and an x-iinfo header, plus cookies beginning visid_incap_, incap_ses_ or nlbi_.
We could not verify the current Imperva console path. In the Cloud WAF settings for the site, look for the bot access control or a delivery rule, and allow requests whose User-Agent contains HootyBot. If the site is set to block every unclassified client, that setting may need relaxing too.
DataDome
Responses carry x-datadome and x-dd-b headers and a datadome cookie. DataDome is often deployed behind another CDN, so both may need checking.
We could not verify the dashboard path. In the custom rules for your domain, add an allow rule matching the User-Agent against HootyBot, above any rule that challenges or blocks.
HUMAN Security (PerimeterX)
A block carries an x-px-blocked header and a page titled “Access to this page has been denied”. The enforcer often runs inside a CDN, so the Server header will name the CDN rather than HUMAN.
We could not verify the portal path. In the Bot Defender policy for the application, look for custom rules or an allow-list and permit requests whose User-Agent contains HootyBot.
Akamai Bot Manager
Look for an Akamai-GRN header. A block is a short “Access Denied” page citing a reference number and a link to errors.edgesuite.net.
We could not verify the Akamai Control Center path, and there are several plausible ones, so we will not guess. What you need is a custom bot definition matching the User-Agent against HootyBot, with that category’s action set to allow in the bot policy. Your Akamai account team can point you to it.
AWS WAF Bot Control
AWS WAF adds no vendor header, so there is nothing in the response that proves it is AWS WAF rather than the origin refusing us. You will simply see a 403.
We could not verify the console steps. What you need is a rule in the web ACL, at a priority above the Bot Control managed rule group, matching the User-Agent header against HootyBot with the action set to allow. Bot Control otherwise classifies any non-browser client as an HTTP library and acts on that label.
Sucuri
Responses carry Server: Sucuri/Cloudproxy and an x-sucuri-id header.
We could not verify the dashboard path, or whether the allow-list accepts a User-Agent as well as an IP. In the firewall’s access-control settings, add an allow entry for HootyBot; if only IP entries are offered, write to bot@hooty.to for the address range.
Fastly / Signal Sciences
We were not able to verify either a reliable response fingerprint or the console path for this one. In the Signal Sciences request rules for the site, add a rule where the User-Agent contains HootyBot and the action allows the request to skip further inspection.
Wallarm
Wallarm usually runs as a module on your own edge, so there is no distinctive header to look for; a block is a plain 403.
Wallarm’s allow-lists match on IP, not User-Agent. The one to use is the exception list under API Abuse Prevention, which exists precisely to stop legitimate bots and crawlers being blocked, or the broader IP Lists allow-list. Either way you need our address range, so write to bot@hooty.to. Matching on User-Agent is possible only by editing the nginx configuration on the Wallarm node itself.
Hosting-level protection
Hosting providers often enable bot filtering for you, and the control sits in the hosting panel rather than in a product you chose. If you are on Beget, Timeweb, REG.RU, Selectel or similar, look for a “protection against bots”, “anti-DDoS” or “web application firewall” toggle in the panel for your site. We have not verified the panel paths for individual providers, so the quickest route is usually a support ticket asking them to allow the crawler whose User-Agent contains HootyBot for your domain.
The same applies to rules on your own server. If you run ModSecurity with the OWASP Core Rule Set, the rules that catch a well-behaved crawler are 913101 and 913102 in REQUEST-913-SCANNER-DETECTION.conf, which match user agents against lists of HTTP clients and crawlers. Add a runtime exclusion for HootyBot in REQUEST-900-EXCLUSION-RULES-BEFORE-CRS.conf. Placement matters: ctl: exclusions must come before the CRS include, while SecRuleRemoveById must come after it, because a rule has to exist before it can be removed.
If you use nginx rate limiting, requests are rejected with a 503 by default. nginx does not account for requests whose limit key is empty, so the usual way to exempt a crawler is a map on the user agent that yields an empty key for it:
map $http_user_agent $limit_key {
default $binary_remote_addr;
"~*HootyBot" "";
}Qualcos’altro
Molti siti usano filtri propri anziché un prodotto con un nome, e alcuni rifiutano un crawler senza restituire mai un errore: rispondendo con una pagina di verifica con stato di successo, oppure reindirizzando all’infinito tra due indirizzi. Questi casi li segnaliamo come blocco senza nome, invece di indovinare un fornitore e mandarti alle impostazioni sbagliate.
Se hooty.to ha segnalato un blocco ma non riconosci il prodotto, o i passaggi qui sopra non corrispondono al tuo pannello, scrivi a bot@hooty.to indicando l’indirizzo del sito. Ti diremo esattamente che risposta stiamo ricevendo, il che di solito identifica il prodotto, e ti aiuteremo a trovare l’impostazione giusta.
Se preferisci che HootyBot non legga affatto il sito, non ti serve nulla di tutto questo. Aggiungi quanto segue al tuo robots.txt e si fermerà:
User-agent: HootyBot
Disallow: /Informazioni su HootyBot · hooty.to · Ultimo aggiornamento: 26 luglio 2026