Identity: Publishes its IP addresses.
How to recognise it
- Published IP list: https://developers.google.com/static/crawling/ipranges/common-crawlers.json
- Reverse DNS ends in
googlebot.com,google.com,googleusercontent.com(then forward-confirm) - User agent contains
Google-CloudVertexBot,CloudVertexBot,Google-Extended
Allow or block
To opt out, add to robots.txt:
User-agent: Google-Extended
Disallow: /
With Double Agent
Sessions from GoogleOther are labelled bot with agent_id: google.googleother, and tagged in your analytics. Nothing is blocked unless you choose to. Install: https://doubleagent.so/install.md
Source: https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers ยท Operator docs: https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers