Rotgar Research · Medical
Clinic Shortlists and Sources Across ChatGPT Models and Reasoning Efforts
A repeated-run New York City study of how five ChatGPT model and reasoning-effort configurations named medical providers, ordered their shortlists and linked to visible web sources.
Methodology review only; not clinical review, peer review, a formal external audit or an assessment of providers named in the responses. The study does not assess clinical accuracy, provider quality or treatment appropriateness.
Direct answer
ChatGPT models and reasoning efforts did not produce one stable clinic list, order or source set
Not reliably, and not to the same degree for every medical service. Across the tested ChatGPT configurations, mean provider-shortlist overlap within the same configuration was 80.3% for IVF, 34.7% for full-arch dental implants and 53.9% for bariatric programs. The visible source layer was less stable: domain overlap was 51.1%, 28.3% and 44.8%, respectively.
One answer should not be treated as a clinic’s stable AI visibility, stable position or stable source set. Provider inclusion, presentation order and visible citations are separate outcomes.
Study at a glance
A frozen 3 × 5 × 30 repeated-execution panel
- Core answers
- 450
- Exact prompts
- 3
- Configurations
- 5
- Runs per cell
- 30
- Collection surface
- ChatGPT-authenticated CLI
- Authentication
- ChatGPT login
- Session design
- Fresh ephemeral process
- Fieldwork
- Prompt geography
- New York City
- Visible-link coverage
- 450 of 450
- Unique answer × URL pairs
- 3,690
- Unique answer × domain pairs
- 3,194
Short answers for healthcare CMOs
What should a CMO infer from the ChatGPT configuration shifts?
Can one screenshot be a baseline?
No. It records one observation, while repeated answers showed material changes in provider sets, order and exact pages.
Did one configuration win?
No. Stability varied by service and measured layer. More visible links did not mean a more repeatable domain set.
Does first position mean “best”?
No. Position is presentation order in the generated shortlist, not a clinical-quality ranking, endorsement or outcome estimate.
Are recurring domains causal levers?
No. They are reader-visible links worth auditing; recurrence does not prove a domain caused a provider mention.
Explore the three prompts
Service line changed the normal range of variation
I’m looking for the best IVF clinic in New York City. Which clinics should I compare, and how much does IVF usually cost?
- Provider overlap, within
- 80.3%
- Domain overlap, within
- 51.1%
- Exact-URL overlap, within
- 23.8%
I’m looking for the best dental practice for full-arch dental implants in New York City. Which practices should I compare, and how much does treatment usually cost?
- Provider overlap, within
- 34.7%
- Domain overlap, within
- 28.3%
- Exact-URL overlap, within
- 20.2%
I’m looking for the best bariatric surgery clinic in New York City. Which clinics should I compare, and how much does weight-loss surgery usually cost?
- Program overlap, within
- 53.9%
- Parent-system overlap
- 84.3%
- Domain overlap, within
- 44.8%
Interpretation boundary
A descriptive ChatGPT model benchmark, not a patient recommendation
Codex CLI 0.147.0 authenticated through one ChatGPT account, with live search available. This was not consumer ChatGPT Free or Plus and not an API-key run.
New York City was written into each prompt. Physical US location, IP geolocation and device location were not controlled.
Only links visibly present in final-answer Markdown were counted. They do not expose every retrieved or internally used source.
Presentation order is an exposure signal, not a clinical ranking, provider endorsement or estimate of user attention.
Research design
Thirty isolated executions of every prompt and configuration
The protocol, manifests and run plan were locally frozen before collection. The core was a 3 × 5 × 30 repeated-execution panel: three exact commercial prompts, five requested configurations and 30 repetitions of every prompt/configuration combination.
The exact question was how much primary provider-shortlist composition, presentation order and reader-visible linked sources varied across repeated runs of the same configuration and across Luna Medium, Terra Medium, Sol Medium, Sol Low and Sol High.
Each of 30 rounds contained all 15 combinations in a deterministically shuffled order using seed 20260817. Collection was sequential, concurrency one, with ordinary five-second pacing and a 60-second round-boundary delay.
Every observation started a fresh ephemeral CLI process; sessions were not resumed. Live search was available and was invoked in all 450 core answers. Optional context surfaces, persistence, user rules, agents, shell, plugins, apps, memory, browser/computer control and multi-agent behavior were disabled where supported by the strict loader.
The 15 preflight answers were kept separate and excluded from every reported denominator. Content was not retried. One technical attempt, R04-P1-C4 attempt 1, failed; attempt 2 succeeded, and both artifacts were retained. The 450 successful core answers are the analytical panel.
Configuration manifest
Five configurations of three requested GPT-5.6 model variants
| ID | Requested model | Reasoning effort | Reader-facing label |
|---|---|---|---|
| C1 | gpt-5.6-luna | medium | Luna Medium |
| C2 | gpt-5.6-terra | medium | Terra Medium |
| C3 | gpt-5.6-sol | medium | Sol Medium |
| C4 | gpt-5.6-sol | low | Sol Low |
| C5 | gpt-5.6-sol | high | Sol High |
The requested labels are the identities recorded by the collector. The study does not estimate consumer usage share, identify a default consumer model or justify weighting one configuration more heavily than another.
Visible source stability
More links did not mean a more repeatable domain set
Sol High exposed the broadest visible source set, averaging 10.3 unique URLs and 8.1 unique domains per answer across the three prompts. Its mean domain-set overlap was 40.9%, within the narrow 40.5%–42.1% range across all five configurations. Link breadth and repeatability are distinct measures.

| Configuration | Mean unique URLs | Mean unique domains | Exact-URL overlap | Domain overlap |
|---|---|---|---|---|
| Luna Medium | 7.80 | 7.01 | 19.6% | 40.5% |
| Terra Medium | 6.41 | 5.83 | 20.7% | 42.0% |
| Sol Low | 7.66 | 6.91 | 21.7% | 42.1% |
| Sol Medium | 8.83 | 7.68 | 21.0% | 41.6% |
| Sol High | 10.30 | 8.06 | 24.1% | 40.9% |
Within versus across configurations
Provider names were generally more stable than visible source pages
IVF had a compact provider core, while dental implants were fragmented in both providers and visible domains. Bariatric results looked substantially more stable when named programs were rolled up to parent health systems.

| Service | Entity level | Provider within | Provider across | URL within | URL across | Domain within | Domain across |
|---|---|---|---|---|---|---|---|
| IVF | Clinic/program | 80.3% | 79.5% | 23.8% | 17.5% | 51.1% | 43.5% |
| Full-arch dental implants | Practice/program | 34.7% | 28.6% | 20.2% | 15.9% | 28.3% | 23.0% |
| Bariatric surgery | Specific program | 53.9% | 49.9% | 20.2% | 16.6% | 44.8% | 41.8% |
| Bariatric surgery | Parent health system | 84.3% | 82.1% | — | — | — | — |
The parent-system row regroups named bariatric providers only. Source URLs and domains were not reclassified by health-system family. The differences are descriptive; they are not statistical or causal estimates.
Provider appearance
Configuration profiles differed most visibly in dental
Several IVF providers formed a stable inclusion core, while dental providers moved in different directions across the tested configurations. In bariatric answers, the exact program sometimes changed while the parent health system remained present.

| Service | Provider | Luna Medium | Terra Medium | Sol Low | Sol Medium | Sol High |
|---|---|---|---|---|---|---|
| IVF | NYU Langone | 100% | 100% | 100% | 100% | 100% |
| IVF | Columbia | 100% | 100% | 96.7% | 100% | 100% |
| IVF | Weill Cornell | 100% | 96.7% | 100% | 100% | 100% |
| IVF | RMA of New York | 100% | 96.7% | 100% | 100% | 100% |
| IVF | CCRM New York | 90.0% | 90.0% | 66.7% | 76.7% | 86.7% |
| IVF | Spring Fertility | 20.0% | 23.3% | 40.0% | 56.7% | 43.3% |
| Dental | Manhattan Arch | 86.7% | 90.0% | 93.3% | 96.7% | 93.3% |
| Dental | Columbia dental programs | 80.0% | 23.3% | 86.7% | 90.0% | 100% |
| Dental | Central Park Oral Surgery | 90.0% | 63.3% | 40.0% | 33.3% | 3.3% |
| Dental | All-on-Four Dental Implant Centers | 50.0% | 56.7% | 23.3% | 40.0% | 56.7% |
| Dental | Manhattan Aesthetic Dentistry / Smart Arches | 13.3% | 6.7% | 43.3% | 36.7% | 53.3% |
| Dental | NYC Dental Implants Center | 36.7% | 16.7% | 13.3% | 23.3% | 30.0% |
| Dental | ClearChoice | 3.3% | 3.3% | 10.0% | 33.3% | 50.0% |
| Bariatric | NYU Langone | 100% | 100% | 100% | 100% | 100% |
| Bariatric | NYP / Weill Cornell | 100% | 96.7% | 100% | 100% | 100% |
| Bariatric | NYP / Columbia | 56.7% | 50.0% | 90.0% | 80.0% | 96.7% |
| Bariatric | Mount Sinai Morningside | 43.3% | 66.7% | 40.0% | 73.3% | 66.7% |
| Bariatric | NYC Health + Hospitals / Bellevue | 60.0% | 33.3% | 66.7% | 73.3% | 73.3% |
| Bariatric | NYC Health + Hospitals / Jacobi | 60.0% | 10.0% | 3.3% | 3.3% | 3.3% |
Presentation order
Shortlist composition and order were different stability problems
Sol High’s IVF provider set was 86.6% similar across repeat pairs, yet the same provider occupied first position in only 29.4% of pairs. A visibility monitor should not collapse “included somewhere” and “presented first” into one rank.
The full ordered shortlist rarely repeated exactly. IVF produced 16–24 unique sequences per 30 answers across configurations, dental implants produced 29–30 and bariatric surgery produced 26–29.

| Service | Configuration | Full set overlap | Top-three overlap | Same first | Same position |
|---|---|---|---|---|---|
| IVF | Luna Medium | 80.7% | 88.6% | 31.7% | 41.3% |
| IVF | Terra Medium | 81.4% | 86.1% | 35.9% | 41.0% |
| IVF | Sol Low | 74.4% | 96.7% | 93.3% | 81.8% |
| IVF | Sol Medium | 78.5% | 78.5% | 64.4% | 57.9% |
| IVF | Sol High | 86.6% | 61.2% | 29.4% | 32.6% |
| Dental implants | Luna Medium | 37.5% | 35.5% | 41.8% | 32.1% |
| Dental implants | Terra Medium | 28.8% | 28.7% | 64.1% | 57.0% |
| Dental implants | Sol Low | 32.4% | 28.5% | 75.4% | 50.3% |
| Dental implants | Sol Medium | 32.9% | 26.0% | 74.7% | 50.2% |
| Dental implants | Sol High | 41.8% | 31.9% | 63.9% | 44.4% |
| Bariatric surgery | Luna Medium | 43.6% | 63.2% | 64.8% | 49.9% |
| Bariatric surgery | Terra Medium | 47.5% | 57.3% | 74.9% | 53.0% |
| Bariatric surgery | Sol Low | 57.5% | 64.6% | 56.8% | 41.4% |
| Bariatric surgery | Sol Medium | 59.3% | 68.9% | 71.3% | 50.3% |
| Bariatric surgery | Sol High | 61.4% | 72.6% | 57.9% | 45.0% |
Visible domain appearance
The domains visible to readers also shifted by configuration
The corrected source pipeline counted only event_type=final_answer_markdown_link, normalized the URLs and deduplicated URLs and domains inside each answer. It retained 3,757 visible Markdown-link occurrences, 3,690 unique answer × URL pairs and 3,194 unique answer × registrable-domain pairs.

| Service | Domain | Luna Medium | Terra Medium | Sol Low | Sol Medium | Sol High |
|---|---|---|---|---|---|---|
| IVF | ny.gov | 93.3% | 100% | 100% | 100% | 100% |
| IVF | nyulangone.org | 90.0% | 96.7% | 96.7% | 100% | 100% |
| IVF | sartcorsonline.com | 33.3% | 43.3% | 0.0% | 26.7% | 80.0% |
| IVF | rmany.com | 86.7% | 20.0% | 80.0% | 76.7% | 66.7% |
| Dental | centralparkoralsurgery.nyc | 96.7% | 93.3% | 80.0% | 60.0% | 26.7% |
| Dental | manhattanarchcollective.com | 80.0% | 83.3% | 73.3% | 60.0% | 60.0% |
| Dental | columbia.edu | 70.0% | 13.3% | 80.0% | 90.0% | 93.3% |
| Dental | clearchoice.com | 3.3% | 6.7% | 10.0% | 33.3% | 50.0% |
| Bariatric | mountsinai.org | 100% | 96.7% | 100% | 100% | 100% |
| Bariatric | nyulangone.org | 83.3% | 100% | 96.7% | 100% | 100% |
| Bariatric | nyp.org | 66.7% | 80.0% | 60.0% | 60.0% | 76.7% |
| Bariatric | nychealthandhospitals.org | 73.3% | 33.3% | 70.0% | 80.0% | 83.3% |
| Bariatric | cms.gov | 0.0% | 0.0% | 0.0% | 36.7% | 50.0% |
These links can support provider, cost, regulatory or general guidance claims. Their appearance does not prove source authority, endorsement or the probability that editing a domain will change a future answer.
Three service-line readings
The same monitoring rule would misdescribe all three niches
IVF: a stable inclusion core, but unstable first position
NYU Langone appeared in 100% of the 150 IVF answers. Columbia, RMA of New York and Weill Cornell each appeared in at least 99.3%. The core did not imply stable order: Weill Cornell’s first-position rate ranged from 20.0% in Luna Medium to 96.7% in Sol Low.
Dental: the most fragmented provider and source profile
Central Park Oral Surgery appeared in 90.0% of Luna Medium answers and 3.3% of Sol High answers. ClearChoice moved in the opposite direction, from 3.3% in Luna Medium to 50.0% in Sol High. This is an observed profile difference, not proof that reasoning effort caused the shift.
Bariatric: entity granularity changed the conclusion
Mean within-configuration provider overlap was 53.9% at exact-program level and 84.3% after programs were grouped into parent health systems. A hospital network can therefore appear consistently while the named facility changes.
Practical framework
How a CMO should monitor ChatGPT models and reasoning efforts
- Measure inclusion, position and sources separately. Track provider appearance, top-three, first position, parent/facility identity, visible domains and exact pages as different outcomes.
- Use repeated observations. One response is an observation, not a baseline. Choose repetition based on the precision required for the decision; 30 is not a universal minimum.
- Build service-specific baselines. IVF, dental and bariatric prompts had materially different normal ranges of variation.
- Keep configurations as separate strata. The study does not reveal which tested configuration represents prospective-patient usage.
- Audit recurring domains without calling them causal levers. Check whether owned and third-party pages are current, accessible and factually accurate.
- Preserve hospital-system hierarchy. Report parent-network and exact-facility presence when both identities affect interpretation.
- Retain enough evidence to distinguish drift from noise. Compare repeated batches with exact prompt, configuration, timestamp, provider sequence and visible-link records.
Methodology
Two measured surfaces with explicit denominators
Primary provider shortlist
The analysis used the first contiguous primary provider shortlist in each saved final answer. Later “also consider” lists were excluded. Aliases were mapped to canonical clinic, practice or program entities. Duplicates inside one answer were retained at their earliest presentation position; a compound row could give two providers a tied position. Bariatric providers were also mapped to parent health systems for a separate sensitivity view.
Appearance rate uses 30 answers in one prompt/configuration cell or 150 answers for one prompt across configurations. Within each 30-answer cell, all 435 answer pairs were compared. Between two configurations, all 900 cross-configuration answer pairs were compared. Those pairings are computed comparisons, not statistically independent users; the experimental base remains 30 executions per cell.
Reader-visible links
Only final_answer_markdown_link records were eligible. Raw JSON trace URLs, search-event URLs and duplicate plain-URL regex rows were excluded. URLs were normalized and deduplicated within each answer; registrable domains were also deduplicated within each answer.
All 450 core answers contained at least one visible linked domain. The 3,757 link-occurrence total is provenance, not an exposure or authority score. Publication metrics use the 3,690 unique answer × normalized-URL pairs and 3,194 unique answer × registrable-domain pairs.
Quality checks
- All 450 planned core observations completed successfully; 15 preflight observations were excluded.
- All 450 primary shortlists contained at least three rows; 2,276 shortlist rows were extracted.
- Canonical normalization produced 2,359 provider mentions after within-answer deduplication, with zero unmapped shortlist rows.
- C3 is Sol Medium and C4 is Sol Low; the configuration identities were checked against the frozen manifest.
- Ten automated analytical tests passed for expected counts, configuration identity, metric ranges and pair denominators.
Limitations
What the benchmark cannot establish
- Surface: this was a ChatGPT-authenticated CLI collection, not browser ChatGPT Free or Plus, a mobile app, an API-key run, Gemini or Google Search.
- Geography: New York City was specified in each prompt; physical location and IP were not controlled.
- Scope: one exact prompt per niche, three niches, one city, one account, one Mac, one network context and one continuous collection window.
- Prompt wording: every prompt combined “best,” a comparison shortlist and a cost question; paraphrases and other intents were not tested.
- Time: the sequential core ran for 9 hours 26 minutes 59 seconds. Shuffling distributed configurations through the window, but time and configuration are not fully separable.
- Backend state: every observation used a fresh ephemeral local process, but server-side caching, retrieval state and platform changes were not observable or controllable.
- Inference: no confidence intervals, hypothesis tests or causal model were used. Differences cannot be attributed solely to model family or reasoning effort.
- Sources: only visible final-answer Markdown links were measured. Hidden retrieval sources and source quality were not evaluated.
- Domains: registrable-domain keys are mechanical normalization fields, not source-class or authority judgments.
- Normalization: aliases and parent systems were mapped using explicit rules; the entity level materially changes the bariatric result.
- Position: shortlist position is presentation order, not an objective rank, endorsement or estimate of user attention.
- Medical and commercial outcomes: the study did not assess factual or clinical accuracy, treatment appropriateness, impressions, clicks, leads, appointments, revenue or patient outcomes.
- Audience: it does not estimate which configuration real prospective patients use.
- Retry: one observation required a technical retry; the failed artifact was retained and the successful replacement was analyzed.
Definitions and further research
A vocabulary for repeatable AI visibility measurement
| Term | Definition |
|---|---|
| Core response | One successful observation in the frozen 450-answer analysis panel. |
| Primary provider shortlist | The first contiguous list or table of provider organizations in the final answer. |
| Presentation position | Ordinal row in that shortlist; not a quality rank. |
| Visible cited URL | A Markdown link visibly present in the saved final answer. |
| Jaccard similarity | Intersection divided by union for two sets; 100% means identical sets. |
| Within-configuration stability | Mean across all 435 pairs among 30 answers in the same prompt/configuration cell. |
| Across-configuration stability | Mean across 900 pairs for a configuration pair, summarized across ten pairs per prompt. |
Next replications should freeze new observation windows across days, add prompt paraphrases and cities, and test other consumer or API surfaces as separate datasets. The current results should remain a dated baseline rather than being silently combined with a future wave.
Downloadable research materials
Download the public v1.0 research package
The versioned package contains publication-safe aggregate tables, sanitized observation-level evidence, frozen prompts and configurations, methods, citation metadata and checksums under CC BY 4.0. It excludes raw JSONL, full answers, stderr, machine paths, authentication state, the private archive and unsafe raw provider-extraction fields.
Headline provider, exact-URL and domain stability metrics.
text/csv · 703 BDataComplete public JSONMachine-readable metadata and all publication-safe tables.
application/json · 3.4 MBDataAnalysis workbookFormatted workbook of results, manifests, definitions and limits.
application/vnd.openxmlformats-officedocument.spreadsheetml.sheet · 290 KBReproducibilityPrompt manifestThe three frozen commercial New York City prompts.
text/csv · 603 BReproducibilityConfiguration manifestRequested model labels and reasoning settings for C1–C5.
text/csv · 251 BProvider dataProvider frequencyAppearance, top-one, top-three and conditional position measures by configuration.
text/csv · 85 KBProvider dataWithin-configuration provider stabilityPairwise provider-set and presentation-order stability within each cell.
text/csv · 5.0 KBProvider dataCross-configuration provider stabilityProvider-set comparisons across configuration pairs.
text/csv · 9.2 KBSource dataVisible-source summaryVisible link breadth and repeat stability for each configuration.
text/csv · 605 BSource dataWithin-configuration source stabilityExact-URL and registrable-domain overlap within each cell.
text/csv · 1.7 KBSource dataCross-configuration source stabilityExact-URL and registrable-domain overlap across configuration pairs.
text/csv · 2.6 KBDocumentationData dictionaryField definitions, denominators and interpretation boundaries.
text/csv · 22 KBDocumentationPackage READMEScope, headline results, package map and public-data boundary.
text/markdown · 2.7 KBDocumentationMethodologyFrozen design, extraction rules, metrics, checks and limitations.
text/markdown · 3.5 KBVerificationPackage manifestVersioned inventory with file sizes, row counts and SHA-256 values.
application/json · 10 KBVerificationSHA-256 checksumsIntegrity checksums for the complete public release.
text/plain · 4.4 KBCitationCitation File FormatCFF 1.2 metadata for research and repository tools.
text/plain · 837 BCitationBibTeX citationReady-to-import BibTeX citation metadata.
text/plain · 593 BVersion 1.0 · Published August 29, 2026 · Manifest · SHA-256 checksums
Citation and research record
A traceable, versioned publication record
Yudin, E. (2026). Clinic Shortlists and Sources Across ChatGPT Models and Reasoning Efforts (Version 1.0) [Data set]. Rotgar Research. https://doi.org/10.5281/zenodo.22162725
- Author
- Evgeniy Yudin, Founder and Strategy Lead · ORCID 0009-0007-8400-9561
- Methodology reviewer
- Boris Teplyakov, SEO Lead
- Review scope
- Methodology review; not peer review, clinical review or an assessment of named providers
- Version and license
- Version 1.0; CC BY 4.0
- Version DOI
- 10.5281/zenodo.22162725
- Archive
- Zenodo record 22162725
- Fieldwork
- August 18, 2026
- Publication
- August 29, 2026
- Surface
- ChatGPT-authenticated CLI collection with live search available
- Language and scope
- US English; New York City named in each prompt
Free audit
Build a repeatable medical AI visibility baseline
The simplest free audit starts with one clinic or selected location, one priority market and one patient language.
- Dental
- Aesthetic medicine
- Physiotherapy
- Fertility & IVF
- Dermatology
- Hair transplant
- Plastic surgery
- Mental health
- Multi-specialty groups
- Your current visibility and named competitors
- The sources AI answers rely on
- A prioritized list of improvements to implement
Several clinics, locations, countries, markets or patient languages can be included without a call. Prefer to discuss the scope? Book a 30-minute call We’ll tailor the audit at no cost.