# # robots.txt # # This file is to prevent the crawling and indexing of certain parts # of your site by web crawlers and spiders run by sites like Yahoo! # and Google. By telling these "robots" where not to go on your site, # you save bandwidth and server resources. # # This file will be ignored unless it is at the root of your host: # Used: http://example.com/robots.txt # Ignored: http://example.com/site/robots.txt # # For more information about the robots.txt standard, see: # http://www.robotstxt.org/robotstxt.html # # --------------------------------------------------------------------------- # ICO NOTES -- read before editing # # 1. NEVER PUT A BLANK LINE INSIDE A GROUP. A blank line ends the current # user-agent record: every rule after it belongs to no group and is silently # dropped by strict parsers (Python's urllib.robotparser, and historically # Google's). Comment lines are fine -- they are stripped before the # blank-line check -- so use comments, never blank lines, to break up the # block below. Blank lines BETWEEN groups are correct and required. # # 2. This file is served from THIS ConfigMap, not from the image. nginx aliases # /robots.txt straight to the mounted copy, so Drupal core's stock robots.txt # in the docroot is shadowed and irrelevant. Editing here + an ArgoCD sync is # the whole change: no image rebuild, no digest bump, no pod restart (kubelet # refreshes the mount within ~a minute and nginx opens the file per request -- # there is no open_file_cache). # The "Drupal core defaults" section further down is core's original content, # unchanged apart from the blank lines that had to go (see 1) -- keep it that # way so a core upgrade diff stays readable. # # 3. ONE FILE SERVES EVERY DOMAIN. This is a Drupal multisite: all ~70 brand # domains share a single docroot, so they all serve this exact file. Never # add a host-specific rule here -- it would apply to every brand. # That is also why there is deliberately NO `Sitemap:` line: it would have to # name one host, and would be wrong for all the others. Sitemaps are # announced per-domain via Search Console / the sitemap module instead. # # 4. THIS FILE IS A REQUEST, NOT A CONTROL. Well-behaved crawlers honour it; # the ones that actually took the site down did not. The enforcing layer is # the nginx bot gate + rate limit on the same paths, in the infra repo: # kubernetes/apps/drupal/base/nginx-config.yaml (`$ico_bot_ua`, # `limit_req zone=booking`). Keep the two in sync: if you add an expensive # path family here, add it there too -- and vice versa. # --------------------------------------------------------------------------- User-agent: * Crawl-delay: 10 # --- Expensive routes: keep crawlers off the uncached search/booking paths --- # These run heavy uncached searches over the ~9.4M-row availability table. # Each takes seconds against the DB, and a crawler walking infinite booking-URL # parameter combinations exhausts the php-fpm workers and takes the whole # instance down -- that is the Princess incident of 2026-07-23, and the # distributed scrape of 2026-08-01. # # Measured live on www.princesscruises.de (2026-07-27), all sailing through # with HTTP 200 before the nginx gate was widened: # GET /preise?action=findByDestination&code=EUR&number_of_pages=50 bingbot # GET /booking AhrefsBot # GET /booking/node/node/2733 Amazonbot # # Drupal serves these BOTH bare and behind an i18n prefix (/de/booking, # /fr/booking, /gsw-berne/booking), so each family needs two rules: the bare # path, and one `*` wildcard form for the language prefix. # # The prefix set is deliberately NOT enumerated. It is open-ended -- de, ch, # at, fr, se, dk, fi, no, is plus locale forms like de-at, fr-ch, it-ch, # en-gb, gsw-berne -- and a new market would silently escape a hardcoded list. # `/*/booking` covers all of them, present and future. # # The `*` wildcard is not in the original 1994 standard, but it is codified in # RFC 9309 (2.2.3) and honoured by Google, Bing and Yandex -- i.e. by every # crawler that actually generated the load above. Some minimal parsers do only # literal prefix matching and will ignore these six lines (Python's # urllib.robotparser is one). That is acceptable: anything that ignores them is # by definition not a well-behaved crawler, and the nginx gate stops it anyway. # # A prefix match covers the query string and everything below the path, so # `Disallow: /booking` already covers /booking?destination=..., /booking/flight, # /booking/cabin_detail/..., /booking/itinerary/... and /booking/confirmation. Disallow: /booking Disallow: /*/booking Disallow: /prices Disallow: /*/prices Disallow: /preise Disallow: /*/preise # --- Partner data + machine endpoints: nothing here is content --- # /dailyprices/ is the daily CSV/ZIP/TXT price export served off the legacy # uploads share for partner clients. Crawling it is pure waste. The Java # partner client that polls it does not read robots.txt, so this cannot break # that sync. Disallow: /dailyprices/ # /ICRWEB is the touristic_api partner webservice dispatcher (POST only). Disallow: /ICRWEB # =========================================================================== # Drupal core (7.66) defaults below -- content unchanged, blank lines removed # =========================================================================== # CSS, JS, Images Allow: /misc/*.css$ Allow: /misc/*.css? Allow: /misc/*.js$ Allow: /misc/*.js? Allow: /misc/*.gif Allow: /misc/*.jpg Allow: /misc/*.jpeg Allow: /misc/*.png Allow: /modules/*.css$ Allow: /modules/*.css? Allow: /modules/*.js$ Allow: /modules/*.js? Allow: /modules/*.gif Allow: /modules/*.jpg Allow: /modules/*.jpeg Allow: /modules/*.png Allow: /profiles/*.css$ Allow: /profiles/*.css? Allow: /profiles/*.js$ Allow: /profiles/*.js? Allow: /profiles/*.gif Allow: /profiles/*.jpg Allow: /profiles/*.jpeg Allow: /profiles/*.png Allow: /themes/*.css$ Allow: /themes/*.css? Allow: /themes/*.js$ Allow: /themes/*.js? Allow: /themes/*.gif Allow: /themes/*.jpg Allow: /themes/*.jpeg Allow: /themes/*.png # Directories Disallow: /includes/ Disallow: /misc/ Disallow: /modules/ Disallow: /profiles/ Disallow: /scripts/ Disallow: /themes/ # Files Disallow: /CHANGELOG.txt Disallow: /cron.php Disallow: /INSTALL.mysql.txt Disallow: /INSTALL.pgsql.txt Disallow: /INSTALL.sqlite.txt Disallow: /install.php Disallow: /INSTALL.txt Disallow: /LICENSE.txt Disallow: /MAINTAINERS.txt Disallow: /update.php Disallow: /UPGRADE.txt Disallow: /xmlrpc.php # Paths (clean URLs) Disallow: /admin/ Disallow: /comment/reply/ Disallow: /filter/tips/ Disallow: /node/add/ Disallow: /search/ Disallow: /user/register/ Disallow: /user/password/ Disallow: /user/login/ Disallow: /user/logout/ # Paths (no clean URLs) Disallow: /?q=admin/ Disallow: /?q=comment/reply/ Disallow: /?q=filter/tips/ Disallow: /?q=node/add/ Disallow: /?q=search/ Disallow: /?q=user/password/ Disallow: /?q=user/register/ Disallow: /?q=user/login/ Disallow: /?q=user/logout/ # =========================================================================== # Google ad crawlers -- MUST keep access to the booking/price pages # =========================================================================== # These are the same agents exempted by `$ico_ad_crawler` in the nginx gate. # They are not indexing crawlers: they fetch a page to decide which ads to # serve (Mediapartners-Google) or to score a landing page we pay for # (AdsBot-Google). Blocking them costs ad context and Ads quality scores on # exactly the commercial pages that earn money -- observed 2026-07-27, a # /gsw-berne/booking?utm_medium=cpc landing page 403ing AdsBot-Google-Mobile. # # Each gets its own group with an empty `Disallow:` (= allow everything). # Google's ad crawlers do not fall back to the `User-agent: *` group, so # without these explicit groups their access would be undefined rather than # granted. Stating it also keeps the intent visible next to the rules above. # # Note this is a plain user-agent claim, exactly as in nginx: anyone can send # it. That is acceptable because the only thing it unlocks is the same access # a plain browser UA already gets. Do not treat it as authorization. User-agent: Mediapartners-Google Disallow: User-agent: AdsBot-Google Disallow: User-agent: AdsBot-Google-Mobile Disallow: