# Methodology

## Design

The parent experiment collected 450 isolated ChatGPT-authenticated answers on August 18, 2026: three commercial New York City prompts across IVF, full-arch dental implants and bariatric surgery; five model and reasoning-effort configurations; and 30 repetitions per prompt and configuration. Fifteen preflight answers were excluded from the core panel.

For this secondary analysis, clinic-specific commercial and operational caveats were normalized into 43 deduplicated claim families. Available primary sources were checked on 2026-08-29. Each family received one verdict:

- **Fully supported:** the bounded statement was directly supported by a current or correctly dated primary source.
- **Partly supported or overstated:** a factual core existed, but the wording widened it, made an unmatched comparison, confused scope or applied the caveat asymmetrically.
- **Stale:** historical evidence existed, but it was not suitable as a current fact on the check date.
- **Not supported:** the checked current primary evidence did not support the claim or contradicted the claimed ambiguity.
- **Unverifiable:** relevant primary evidence was unavailable or insufficient to confirm or reject the claim.

A separate set of 18 saved same-answer contrasts was purposively selected to examine how a clinic receiving a specific caveat was framed beside a competitor without an equivalent caveat. This sample tests a possible framing mechanism and does not estimate prevalence across the full corpus.

## Denominators

The headline verdict distribution uses 43 deduplicated claim families. Its displayed shares total 100.1% because categories were rounded independently to one decimal place. Answer share within a niche uses that niche's 150-answer panel. Claim families are not mutually exclusive, so answer shares must not be summed into a share of unique answers.

The secondary weighted view covers 438 non-deduplicated claim incidences. These are not 438 unique answers: one answer can contribute more than one caveat incidence, and recurrence does not make a claim more accurate.

## Limitations

- The source corpus covers three clinic-selection prompts, three service lines, New York City and one collection day.
- Results should not be generalized to every account, interface, model, geography, date or medical query.
- The 43 rows are deduplicated claim families, not clinic-quality observations.
- The 18 contrasts are purposively selected and are not a random prevalence sample.
- Claim-family grouping and verdict coding were performed manually.
- Primary sources were checked on 2026-08-29 and may change.
- A current page can confirm a fact even when a visible answer link was weak, mismatched or absent; current verification and source attribution are different questions.
- “Not supported” is not proof of the opposite fact.
- No clinical comparison, risk adjustment, patient-level assessment or provider ranking was performed.
- No patient-trust experiment was conducted.
- Inter-rater reliability was not measured.
- The analysis cannot establish what caused a particular answer.

Methodology reviewer — Boris Teplyakov, SEO Lead. The review covered methodology only and was not clinical, peer or legal review, a formal external audit, or an assessment of named providers.
