- Japan's "AI-blocking" hypothesis is wrong. AI exclusion rate of 2.7% (within fetched robots.txt) is the lowest of the three regions.
- The real lag is robots.txt itself. Only 19.7% of Japanese listed companies serve a fetchable robots.txt, versus 68.6% in the US and 88.2% in Europe.
- llms.txt adoption is 1/2.5 of Western peers (Japan 6.7% / US 15.8% / Europe 16.2%).
- Overall SEO equipment score (sitemap.xml, schema.org, Open Graph, canonical, etc.) is also lowest in Japan (Japan 3.77 / US 5.97 / Europe 6.94 on a 0–10 scale).
- Common pattern across all three regions: Infrastructure-sensitive sectors (Utilities, Energy) show the highest AI-blocking rates (US Utilities 25%, US Energy 20%, Europe Utilities 17.4%) — likely reflecting deliberate restriction of operational details, not regional culture.
Why this matters for global investors in Japanese SMEs
Cross-border M&A involving Japanese mid-market companies is increasingly mediated by AI-assisted research: Claude, ChatGPT, Perplexity, and Google's AI Overviews are now standard tools for initial market scoping, target identification, and regulatory analysis. When a target company is invisible to these systems (no robots.txt, no llms.txt, weak schema markup), it effectively disappears from this layer of due diligence.
This audit reveals that Japan's underexposure on AI surfaces is not a deliberate stance — it is the consequence of being one technology cycle behind on basic web infrastructure. For acquirers, this creates an information arbitrage: Japanese targets with strong fundamentals but weak web standards are systematically underrepresented in AI-driven market scans, which can translate into less competitive bidding and more favorable acquisition dynamics.
Methodology
Sample composition (N = 4,909)
| Region | N | Source |
|---|---|---|
| Japan | 3,698 | Tokyo Stock Exchange Prime-listed firms (corp_jp_listed_seo_ai.csv, May 2026) |
| United States | 507 | S&P 500 + Fortune 500 sample (corp_us_seo_ai.csv) |
| Europe | 704 | STOXX 600 + FTSE 100 + selected indices (corp_eu_seo_ai.csv) |
Measurement protocol
For each company's primary corporate domain, we measured:
- robots.txt fetch status: HTTP 200 with parseable content, or non-200 / timeout
- AI-bot blocking: presence of disallow directives for GPTBot, ClaudeBot, anthropic-ai, PerplexityBot, Google-Extended, CCBot, OAI-SearchBot, Bytespider, etc., parsed from successfully fetched robots.txt only (sites without fetchable robots.txt are excluded from the blocking-rate denominator to avoid false negatives)
- llms.txt adoption: presence of /llms.txt at the document root, parseable as the proposed llms.txt format
- SEO equipment score (0–10): weighted sum of sitemap.xml, schema.org (Organization), Open Graph tags, canonical link, and structured FAQ presence
Data quality note
corp_full_seo_ai.csv) that turned out
to be a curated set of AI-aware companies, producing a misleading AI-blocking rate of 51.3%. The figures here
use the canonical 3,698-company Japanese sample. This correction is documented as part of JFSC's
data-quality protocol.
Three-region comparison (N = 4,909)
| Region | N | robots.txt fetch | AI blocking (within fetched) | llms.txt | SEO equipment (avg, 0–10) |
|---|---|---|---|---|---|
| Japan (canonical) | 3,698 | 19.7% | 2.7% | 6.7% | 3.77 |
| United States | 507 | 68.6% | 8.9% | 15.8% | 5.97 |
| Europe | 704 | 88.2% | 5.6% | 16.2% | 6.94 |
Sector-level findings (selected, N ≥ 20)
United States by GICS Level 1
| Sector | N | robots.txt | AI blocking | llms.txt | SEO |
|---|---|---|---|---|---|
| Industrials | 79 | 67.1% | 3.8% | 21.5% | 6.76 |
| Financials | 74 | 74.3% | 9.1% | 14.9% | 6.58 |
| Information Technology | 73 | 75.3% | 10.9% | 27.4% | 4.77 |
| Health Care | 59 | 71.2% | 7.1% | 8.5% | 5.97 |
| Utilities | 31 | 51.6% | 25.0% | 6.5% | 5.81 |
| Energy | 21 | 71.4% | 20.0% | 9.5% | 6.95 |
Observation: US Information Technology shows the highest llms.txt adoption (27.4%), suggesting AI-native firms are leading the standard's diffusion. Utilities and Energy show the highest AI-blocking rates.
Europe (sector aggregates)
| Sector | N | robots.txt | AI blocking | llms.txt | SEO |
|---|---|---|---|---|---|
| Financials | 88 | 89.8% | 3.8% | 17.0% | 7.50 |
| Industrials | 78 | 93.6% | 2.7% | 21.8% | 7.65 |
| Health Care | 40 | 100.0% | 2.5% | 17.5% | 7.00 |
| Real Estate | 31 | 83.9% | 11.5% | 12.9% | 7.97 |
| Utilities | 25 | 92.0% | 17.4% | 12.0% | 7.20 |
| Banks | 22 | 90.9% | 0.0% | 22.7% | 7.55 |
Observation: European Health Care achieves 100% robots.txt fetch rate. Utilities and Real Estate exhibit elevated AI blocking similar to the US pattern.
Cross-region patterns
Infrastructure-sensitive sectors block AI consistently across regions
Utilities and Energy display elevated AI-blocking rates in both the US (Utilities 25%, Energy 20%) and Europe (Utilities 17.4%). Japan's small sample in these sectors (Utilities N=28, Energy N=14) precludes confident cross-region comparison, but the pattern itself appears region-independent — driven by content sensitivity (operational details, critical infrastructure data) rather than national culture.
Financials and Industrials are AI-friendly with strong llms.txt diffusion
Banks and large industrial firms benefit from AI distribution of corporate disclosures and product information. European Banks reach 22.7% llms.txt adoption; US Industrials reach 21.5%.
Japan's lag is structural, not strategic
Across every Japanese GICS sector measured, robots.txt fetch rates remain in the 10–25% range. Consumer Staples (food) shows a particularly low SEO equipment score of 1.95. The Japan-wide pattern is consistent with delayed adoption of basic web standards rather than deliberate AI exclusion.
Implications
For global investors evaluating Japanese targets
- AI-driven initial scoping under-counts Japanese mid-market opportunities. Manual research remains necessary.
- Sectors with significant Japanese mid-market deal flow (food, materials, industrials, construction-adjacent) tend to be the least AI-visible — creating information asymmetries that informed acquirers can exploit.
- Cross-border buyers benefit from working with Japan-based advisors who maintain direct knowledge of these "AI-invisible" targets.
For Japanese listed companies and their advisors
- Implementing robots.txt, llms.txt, and basic schema.org markup is now a low-cost, high-leverage step for global visibility.
- Blocking AI bots is not the default Japanese stance, despite earlier interpretations; firms that wish to retain editorial control can do so via standard mechanisms.
Data and reproducibility
Raw source files are maintained at _blog/research/serp_industry_2026/data/corp_*_seo_ai.csv
in JFSC's research repository. The analysis script aio_industry_compare.py performs the
cross-region aggregation reported here. Anonymized public release is under license review.
To request data access for academic or institutional research, please contact JFSC via the Japanese site contact form. A dedicated English request mechanism is in preparation.
Citation
Japan Financial Strategy Center (JFSC). (2026). 4,909-Company AI-Readiness Audit: Japan vs US vs Europe. Retrieved from https://jfsc.jp/en/research/ai-readiness-2026/