PATERNOGA Research · 2026 study

How do DAX 40 companies control AI crawlers?

For a clearly defined sample of public corporate websites, PATERNOGA checked which crawler policies and discovery files were reachable on 30 August 2026. The study measures published configuration — not actual bot visits, indexing or citations.

Central observation

A usable robots.txt file was available on 33 of 40 domains. None of those 33 files blocked the tested homepage for Googlebot, Bingbot, OAI-SearchBot, GPTBot or PerplexityBot. Policy remains unknown for 7 domains because of a 403 response or timeout.

usable robots.txt
33
of 40 domains
explicit OAI-SearchBot rules
2
in usable policies
usable llms.txt
6
of 40 domains
sitemaps declared in robots.txt
28
of 40 domains

What the data supports — and what it does not

Observation

No complete homepage block for the five tested user agents appeared in the readable subsample.

Interpretation

Most access came from generic rules. Only a few companies named AI crawlers explicitly. This shows configuration, not intent.

Not measured

The study does not show whether a bot actually visited the site, a CDN allowed it, content was indexed or an answer cited it.

No extrapolation

DAX 40 companies are not a representative sample of all German companies or the Mittelstand.

Method

Collected on 30 August 2026 at 22:31. Domains were processed sequentially. Exactly four public resources were requested once per domain: robots.txt, llms.txt, /sitemap.xml and the homepage.

  1. 1

    Sample

    The 40 DAX constituents observed on the collection date. One public corporate domain was defined in advance for each company.

  2. 2

    robots.txt

    Rules were evaluated by group for the / path. Specific user-agent groups take precedence over the generic wildcard group.

  3. 3

    Discovery files

    The study recorded sitemap declarations, a recognisable URL set or sitemap index at /sitemap.xml, and usable plain text at /llms.txt.

  4. 4

    Homepage

    The study only recorded whether JSON-LD syntax was present. It did not assess content quality or schema validity.

Detailed findings

  • 28 of 40 robots.txt files declared a sitemap.
  • 20 of 40 domains returned a recognisable URL set or sitemap index at /sitemap.xml. A sitemap at another undeclared path may therefore be missed.
  • 6 of 40 domains returned usable plain text at /llms.txt. No ranking or citation effect is inferred.
  • 16 of 40 homepages contained recognisable JSON-LD. Presence is not a quality assessment.
  • 3 policies named GPTBot explicitly, 2 named OAI-SearchBot and 3 named PerplexityBot.

All 40 observations

Companyrobots.txtOAI SearchGPTBotPerplexitySitemap declaredllms.txt
adidasunknownunknownunknownnono
Airbus200allowedallowedallowedyesno
Allianz403unknownunknownunknownnono
BASF200allowedallowedallowedyesno
Bayer403unknownunknownunknownnono
Beiersdorf200allowedallowedallowedyesno
BMW200allowedallowedallowedyesno
Brenntag200allowedallowedallowedyesno
Commerzbank200allowedallowedallowedyesno
Continental200allowedallowedallowedyesno
Daimler Truck200allowedallowedallowedyesno
Deutsche Bank200allowedallowedallowedyesyes
Deutsche Börse200allowedallowedallowedyesno
Deutsche Telekom200allowedallowedallowedyesno
DHL Groupunknownunknownunknownnono
E.ON403unknownunknownunknownnono
Fresenius200allowedallowedallowednono
Fresenius Medical Care200allowedallowedallowedyesno
GEA200allowedallowedallowedyesno
Hannover Rück200allowedallowedallowedyesno
Heidelberg Materials200allowedallowedallowedyesyes
Henkel200allowedallowedallowedyesno
HOCHTIEF200allowedallowedallowednono
Infineon200allowedallowedallowedyesno
Mercedes-Benz Group403unknownunknownunknownnono
Merck KGaAunknownunknownunknownnono
MTU Aero Engines200allowedallowedallowednono
Munich Re200allowedallowedallowedyesno
QIAGEN200allowedallowedallowedyesno
Rheinmetall200allowedallowedallowedyesno
RWE200allowedallowedallowedyesno
SAP200allowedallowedallowedyesyes
Scout24200allowedallowedallowedyesno
Siemens200allowedallowedallowedyesyes
Siemens Energy200allowedallowedallowedyesno
Siemens Healthineers200allowedallowedallowedyesno
Symrise200allowedallowedallowednono
Volkswagen200allowedallowedallowedyesyes
Vonovia200allowedallowedallowedyesyes
Zalando200allowedallowedallowednono

Study limitations

  • Public policy, not observed bot traffic
  • No IP verification of external crawlers
  • No WAF, CDN or rendering simulation
  • No assessment of indexing or citation
  • Point-in-time measurement; domains and rules may change
  • Not representative of German companies as a whole

Raw data and reproducibility

The complete observations, status codes, destination URLs, methodology version and technical limitations are available as JSON and CSV. The collection script remains versioned with the project so a later edition can use the same logic.

Sample and technical primary sources

Crawler access is a prerequisite, not a visibility guarantee

A credible assessment separates policy, server access, indexing, content, sources and observed answer patterns.