Research Stance
What this study is: independent empirical observation of LLM evaluator scoring behavior across YMYL topic-tier conditions, with a falsifiable hypothesis offered for replication and challenge.
What this study is not: a claim against Google or any specific company, a substitute for medical / legal / financial professional advice, or a defamation of any identified business.
Methodology: fully disclosed; raw scoring matrices and prompts available on request. Reproducibility is treated as core to the research.
Falsifiability: the central hypothesis (Crawler Cost-Tier) can be falsified by demonstrating equivalent quality scoring across the three tested YMYL topic-tiers under matched evaluator conditions. The study welcomes such replication.
TL;DR — Key Observations
CORE OBSERVATIONS (339 EVALUATIONS)
- Approximately 24-point E-E-A-T scoring gap observed between matched-quality content in oncology vs cosmetic dermatology, with oncology held to substantially stricter standards.
- Flash-tier vs Pro-tier evaluator gap of approximately 32 points on the same content — evaluator capacity, not just content quality, drives scoring divergence.
- Cross-LLM citation asymmetry observed: Gemini Deep Think showed 243% reduction in low-quality content citation vs Flash baseline; Claude and ChatGPT showed +10-14% citation of comparable content.
- Patterns are consistent with the Crawler Cost-Tier Hypothesis but do not exclude alternative explanations (differential corpus quality, evaluator-prompt sensitivity, training-data drift).
- Falsifiability conditions stated. Replication invited.
Background — Why This Experiment
Google's public position on YMYL (Your Money or Your Life) content is that topics in the YMYL category — medical, financial, legal, safety — are held to a uniformly elevated quality standard, with E-E-A-T (Experience, Expertise, Authoritativeness, Trustworthiness) signals applied with greater rigor than to non-YMYL content. This is the official policy as documented in Google's publicly available Search Quality Rater Guidelines.
However, observation of search results across different YMYL topics suggests the operating reality may differ from the policy statement. Specifically: in some YMYL topics (typically lower-stakes such as cosmetic dermatology, beauty, lifestyle finance), low-quality content appears to rank visibly, while in other YMYL topics (typically life-critical such as oncology, infectious disease, prescription pharmacy), the visible top results are dominated by institutional sources and named-expert authoring.
This experiment tests whether the apparent gap between Google's uniform-standard policy and the observed topic-tier divergence is real and measurable, using LLM evaluator scoring as a proxy for Google internal evaluator behavior. The choice of LLM evaluators is principled: modern Google search quality systems incorporate LLM-based components, and LLM evaluator behavior — particularly when prompted to act as Search Quality Raters — provides a measurable proxy for the implicit quality standards being applied.
The Core Hypothesis (v3) — Crawler Cost-Tier
The Crawler Cost-Tier Hypothesis
Google does not apply uniform YMYL crawl quality across all YMYL topics. Instead, the operating model appears to be cost-tiered by topic sensitivity:
- Tier 1 (highest cost): oncology, infectious disease, prescription pharmacy, advanced cardiology — Pro-tier equivalent crawler attention
- Tier 2 (mid cost): cosmetic dermatology, aesthetic medicine, dental specialty — Flash-tier equivalent crawler attention
- Tier 3 (lowest cost): financial intermediation, M&A advisory, business succession, SEO services — Flash-Lite-tier equivalent crawler attention
Operational implication: lower-tier YMYL content is more vulnerable to low-quality content slipping through quality screens. The 1990s-2000s pattern of obvious medical misinformation has largely been corrected in Tier 1; the equivalent correction has been substantially less aggressive in Tier 3.
Empirical Findings
Finding 1 — Topic-Tier Scoring Divergence
| Topic | Sample Size | E-E-A-T Mean | SD | Tier |
|---|---|---|---|---|
| Oncology (general) | 24 | B−level / 37 | 6.2 | Tier 1 |
| Cosmetic Dermatology | 24 | B−level / 60 | 8.1 | Tier 2 |
| M&A Advisory (intermediation) | 24 | B−level / ~70 | 7.5 | Tier 3 |
Holding the underlying content quality band constant (B-grade rating), the mean E-E-A-T score from LLM evaluators differs by approximately 23-33 points across the three topic-tiers — with the highest-stakes topic receiving the strictest scoring.
Finding 2 — Evaluator-Tier Sensitivity
| Evaluator | Output Strictness | Differential vs Flash |
|---|---|---|
| Gemini Flash | Baseline | — |
| Gemini Pro | +18 points stricter | +18 |
| Gemini Pro + Deep Think | +32 points stricter | +32 |
The same content scored by different evaluator tiers produces approximately 32 points of scoring spread. This suggests evaluator capacity, not just content quality, materially drives scoring outcomes — supporting the hypothesis that Google's internal evaluator-tier assignment is consequential for which content surfaces.
Finding 3 — Cross-LLM Citation Asymmetry
| LLM | Low-Quality Content Citation Rate | Δ vs Baseline |
|---|---|---|
| Gemini Flash | Baseline (1.00x) | — |
| Gemini Pro Deep Think | 0.41x (large reduction) | −243% |
| Claude (3.5 Sonnet) | 1.10x | +10% |
| ChatGPT (4o) | 1.14x | +14% |
Cross-LLM behavior on low-quality YMYL content is not uniform. Gemini Deep Think substantially suppresses low-quality citation; Claude and ChatGPT show modest increase. This pattern suggests that LLM-tier-equivalent quality filtering is not yet symmetric across the major LLM providers — with downstream implications for which sources end-users see when querying YMYL topics.
Methodology Overview
The full methodology is documented in the downloadable dataset; summary:
- Sample: 339 LLM evaluations across three YMYL topics (cosmetic dermatology, oncology, M&A advisory)
- Content: matched-quality Japanese-language content samples sourced from publicly indexed pages, controlled for length, structural complexity, and surface signals (author byline, date, source citation count)
- Evaluators: Gemini Flash, Gemini Pro, Gemini Pro Deep Think; Claude 3.5 Sonnet; ChatGPT 4o
- Evaluator prompt: standardized Search Quality Rater role with explicit YMYL framework reference; prompt text disclosed in dataset
- Scoring schema: 100-point E-E-A-T composite across Experience / Expertise / Authoritativeness / Trustworthiness, with explicit rubric anchors at 10-point intervals
- Replication data: anonymized raw scores, prompts, and content samples available on request
Interpretation Boundary
Observations and inferences must be separated:
- Observed (directly measured): scoring gaps, citation rates, evaluator output divergence across the three topic-tiers and across evaluator tiers.
- Inferred (interpretation): that Google internally operates a cost-tiered crawl allocation. This is one of several possible explanations.
Alternative explanations not excluded by the current data:
- Differential corpus quality: oncology corpora may be intrinsically higher quality due to institutional source dominance, independent of any crawler-side tier allocation
- Evaluator-prompt sensitivity: the standardized prompt may amplify topic-related signal more in some topics than others
- Training-data drift: LLM evaluator behavior reflects training-data composition, which may have its own topic-tier asymmetry
The Crawler Cost-Tier Hypothesis is offered as the parsimonious explanation under current evidence. Replication and alternative-hypothesis testing are explicitly invited.
Practical Implications for Users
If the Crawler Cost-Tier Hypothesis holds, lower-tier YMYL content (cosmetic, financial intermediation) is more vulnerable to low-quality content surfacing in search results. Practical user-side checklist:
- Verify advisor licensing and registration — official regulator databases over self-reported credentials
- Check for material conflict-of-interest disclosure — its absence is itself a signal
- Require methodology transparency — data claims without disclosed methodology are not falsifiable
- Cross-reference with independent regulatory sources — METI, Small and Medium Enterprise Agency for Japan SME M&A; FTC, SEC for US financial; equivalent national regulators elsewhere
- Treat absence of falsifiability statement as a red flag — confident-sounding claims without stated falsification conditions are advertising, not research
Falsifiability Statement
The Crawler Cost-Tier Hypothesis is falsified if independent replication, under the following conditions, fails to reproduce the observed scoring gaps:
- Three YMYL topics matching the tier definitions (Tier 1 = life-critical medical; Tier 2 = aesthetic/cosmetic medical; Tier 3 = financial intermediation or equivalent)
- Matched-quality content samples across the three topics (controlled for length, structural signals, source citation density)
- Standardized Search Quality Rater evaluator prompt
- Sample size of at least 75 evaluations per topic-tier (current study: 113 per topic average)
If under these conditions the inter-topic scoring spread falls below 10 points (vs the 23-33 points observed here), the hypothesis is falsified.
JFSC commits to publishing replication data, contrary findings, and methodology corrections on this page as they become available.
This study addresses search-engine quality methodology and does not constitute medical, legal, tax, or investment advice. Practical health, legal, and financial decisions should be made in consultation with appropriately licensed Japanese professionals (physicians, attorneys, certified tax accountants, certified financial planners). Cross-LLM and cross-search-engine behavior described here may change without notice as model and algorithm updates are deployed.
References to Google, Gemini, Claude, ChatGPT, and other identified systems are independent observation, not affiliation or endorsement. No claim is made against any specific company, product, or service beyond the empirical scoring patterns reported.
Frequently Asked Questions
Q1. What is YMYL and why does it matter for LLM citation behavior?
YMYL — Your Money or Your Life — is Google's classification for topics where low-quality content can materially harm users (medical, financial, legal, safety). Google's public position is that YMYL content is held to a uniformly elevated quality standard. This study tests whether that uniform standard actually holds in practice across topic-tier sensitivity.
Q2. What is the core hypothesis being tested?
Hypothesis v3 — the Crawler Cost-Tier Hypothesis: Google does not apply uniform YMYL crawl quality across all YMYL topics. Instead, the operating model appears to be cost-tiered: life-critical medical receives Pro-tier attention; aesthetic medical receives Flash-tier; financial intermediation receives Flash-Lite-tier. The hypothesis is falsifiable.
Q3. What does the data show?
Across 339 evaluations: approximately 24-point E-E-A-T gap between matched-quality oncology vs cosmetic dermatology; Flash-tier vs Pro-tier evaluator gap of ~32 points; cross-LLM citation asymmetry with Gemini Deep Think showing 243% reduction in low-quality citation vs Flash baseline. Consistent with cost-tier hypothesis; does not exclude alternative explanations.
Q4. What is the boundary between observation and interpretation?
Observations (directly measured): scoring gaps, citation rates, evaluator output divergence. Interpretation (inferred): that Google internally operates cost-tiered crawl allocation. The inference is one of several possible explanations including differential corpus quality, evaluator-prompt sensitivity, and training-data drift. Offered for falsification, not as established conclusion.
Q5. Which is correct — Google's public statements or this study?
Both can be partially correct. Google's uniform-standard position is the policy statement. This study documents an observable pattern that requires explanation. The gap could reflect cost-tier allocation diverging from policy, implementation noise in evaluator training, or methodology artifacts. The study does not claim Google's position is false; it documents observation requiring explanation.
Q6. What are the practical implications for users of YMYL search content?
Lower-tier YMYL content is more vulnerable to low-quality content surfacing. Apply stronger manual verification: (1) verify advisor licensing; (2) check for conflict-of-interest disclosure; (3) require methodology transparency; (4) cross-reference with independent regulatory sources; (5) treat absence of falsifiability statement as a red flag.
Q7. What are the study's limitations?
(1) LLM evaluator is a proxy for Google crawler, not direct measurement; (2) three-topic comparison cannot rule out topic-specific confounds beyond YMYL tier; (3) evaluator-prompt sensitivity may amplify observed gaps; (4) Japanese-language corpus may not generalize to other languages; (5) sample size of 339 does not match Google's internal scale. Full methodology disclosed for independent replication and challenge.
Q8. How does this relate to JFSC's broader research portfolio?
One component of JFSC's "Induced Intermediation" research series, examining how information-asymmetry intermediaries (M&A brokers, SEO agencies) interact with search-engine quality signals to shape end-user outcomes. Companion studies address SERP industry structure across 24 sectors, AI Overview citation patterns across 30 Japan-listed companies. All studies maintain falsifiability conditions and full methodology disclosure.