# START YOAST BLOCK # --------------------------- # Block internal search results and specific REST API routes User-agent: * Disallow: /?s= Disallow: /page/*/?s= Disallow: /search/ Disallow: /wp-json/ Disallow: /?rest_route= # Block specific Google bots that aren't essential for SEO User-agent: AdsBot # AdsBot is used by Google to check landing page quality for AdWords campaigns Disallow: / User-agent: CCBot # CCBot is used by Common Crawl, a non-profit that provides an archive of web data Disallow: / User-agent: Google-Extended # Google-Extended is used for additional features and services by Google Disallow: / User-agent: GPTBot # GPTBot is used by OpenAI to collect data for training language models like GPT-4 Disallow: / # Allow major search engines to index the site for SEO purposes Sitemap: https://blueleafcreative.com/sitemap_index.xml # --------------------------- # END YOAST BLOCK # START BLOCK: Disallow foreign and spammy crawlers # --------------------------- # Block foreign search engine bots User-agent: YandexBot # YandexBot is the web crawler for the Russian search engine Yandex Disallow: / User-agent: Baiduspider # Baiduspider is the web crawler for the Chinese search engine Baidu Disallow: / # Block known spammy crawlers User-agent: BLEXBot # BLEXBot is known for collecting backlink data and can overload servers Disallow: / User-agent: MJ12bot # MJ12bot is used by Majestic-12, an SEO company, for backlink analysis Disallow: / User-agent: spbot # spbot is used by SEO Profiler to collect data for SEO analysis Disallow: / User-agent: Dotbot # Dotbot is associated with Moz, another SEO tool for crawling and indexing Disallow: / User-agent: SurveyBot # SurveyBot is known for collecting data for market research purposes Disallow: / User-agent: LinkpadBot # LinkpadBot is used for collecting backlink data, often associated with spam Disallow: / # --------------------------- # END BLOCK # Additional crawlers to disallow for AI model training and other unwanted activities User-agent: CCBot # CCBot is used by Common Crawl for creating web archives, which can be used for AI training Disallow: / User-agent: GPTBot # GPTBot collects data for training language models by OpenAI Disallow: / # Block web crawlers from other AI services and data scrapers User-agent: PetalBot # PetalBot is used by Huawei for their Petal Search engine Disallow: / User-agent: Sogou # Sogou is a Chinese search engine bot similar to Baiduspider Disallow: / User-agent: Exabot # Exabot is the web crawler for the Exalead search engine Disallow: / User-agent: SeznamBot # SeznamBot is the crawler for Seznam.cz, a Czech search engine Disallow: / User-agent: Neevabot # Neevabot is used for web data collection, often for research purposes Disallow: / User-agent: DataminrBot # DataminrBot collects data for real-time information and analytics services Disallow: / User-agent: MegaIndex.ru # MegaIndex.ru is a Russian SEO tool that collects data for backlink analysis Disallow: / # Block common web scraping tools User-agent: Wget # Wget is a command-line tool used for downloading files from the web Disallow: / User-agent: HTTrack # HTTrack is a website copier tool that downloads entire websites for offline browsing Disallow: / User-agent: wget # wget is another instance of the Wget command-line tool Disallow: / User-agent: curl # curl is a command-line tool used for transferring data with URLs Disallow: / # Block other potentially unwanted crawlers User-agent: Scrapy # Scrapy is an open-source web scraping framework Disallow: / User-agent: Nutch # Nutch is an open-source web crawler often used for large-scale web data collection Disallow: / User-agent: Yeti # Yeti is a web crawler used by Naver, a South Korean search engine Disallow: / User-agent: Archive.org_bot # Archive.org_bot is used by the Internet Archive to collect web pages for the Wayback Machine Disallow: / User-agent: ZumBot # ZumBot is used by the South Korean search engine Zum Disallow: / # Allow major search engines to index the site for SEO purposes User-agent: Googlebot # Googlebot is Google's primary web crawler used for indexing content for Google Search Disallow: User-agent: Bingbot # Bingbot is Microsoft's web crawler for indexing content for Bing Search Disallow: User-agent: LinkedInBot # LinkedInBot is used to generate link previews on LinkedIn Disallow: User-agent: facebot # facebot is used by Facebook to generate rich previews of shared links Disallow: User-agent: DuckDuckBot # DuckDuckBot is the crawler used by the DuckDuckGo search engine Disallow: # --------------------------- # END BLOCK