# Methodology

## Design

- Study ID: `codex-clinic-source-selection-20260817`.
- Core: 30 rounds × 3 exact prompts × 5 configurations = 450 successful answers.
- Preflight: 15 answers, excluded from all published findings.
- Collection window: 2026-08-18T13:46:41.011Z to 2026-08-18T23:13:39.706Z.
- Every observation used a fresh `codex exec --ephemeral` process; sessions were not resumed.
- Collection was sequential, concurrency 1, with deterministic within-round shuffling using seed `20260817`.
- Ordinary pacing was 5 seconds; round-boundary pacing was 60 seconds.
- Live search was available through `--search`.
- Authentication used an existing ChatGPT login, not API-key billing.
- The protocol, manifests and run plan were locally frozen before collection.

## Configurations

C1 is Luna Medium; C2 is Terra Medium; C3 is Sol Medium; C4 is Sol Low; C5 is
Sol High. Public five-column displays use Luna Medium, Terra Medium, Sol Low,
Sol Medium, Sol High to group Sol effort levels. C3 and C4 remain explicitly
identified and are never swapped.

## Provider shortlist outcome

The analysis counted the first contiguous primary provider shortlist in each
final answer. Later “also consider” lists were excluded. Aliases were mapped to
canonical entities; duplicates inside an answer were retained at their earliest
position. An explicit compound row gave both providers the same presentation
position. Bariatric programs were also rolled up to parent systems as a
separate sensitivity view.

Within a 30-answer prompt/configuration cell, all 435 answer pairs were
compared. For each prompt and pair of configurations, all 900 cross-
configuration answer pairs were compared. These computed pairs are not
statistically independent experimental observations; the experimental base is
30 executions per prompt/configuration cell.

## Visible-source outcome

Only `event_type=final_answer_markdown_link` records were eligible. Search-event
URLs, other trace URLs and duplicate plain-URL extraction rows were excluded.
URLs were normalized, then URLs and registrable domains were deduplicated
inside each answer. Publication totals are 3,757 eligible link occurrences,
3,690 unique observation–URL pairs and 3,194 unique observation–domain pairs.
All 450 core answers contained at least one visible linked domain.

## Quality controls

The source project's ten analytical tests passed. All 450 planned core
observations completed, all primary shortlists had at least three rows, all
shortlist labels were mapped and every core answer had a visible markdown link.
The source archive checksum recorded at validation was
`60125b165bece7e31ab9da66dde0cf1a9f7732070172b983f29c4b79b6024121`; the archive is not redistributed here.

## Required limitations

The study used one exact prompt per niche, three niches, one city named in the
prompts, one account, one Mac, one network context and one continuous collection
window. Physical location and IP were not controlled. The prompts combined
“best,” comparison and cost intent. Provider extraction covered the first
primary shortlist only. Visible links can support provider, pricing,
regulatory or general context and do not reveal every source retrieved or used
internally. The study did not test factual accuracy, clinical quality, patient
outcomes or source quality. Results are descriptive, not causal or statistical
significance claims.

Methodology reviewer — Boris Teplyakov, SEO Lead. This credit is not clinical
review, peer review or a formal external audit.
