Rotgar Research · Medical

Clinic Shortlists and Sources Across ChatGPT Models and Reasoning Efforts

A repeated-run New York City study of how five ChatGPT model and reasoning-effort configurations named medical providers, ordered their shortlists and linked to visible web sources.

Methodology review only; not clinical review, peer review, a formal external audit or an assessment of providers named in the responses. The study does not assess clinical accuracy, provider quality or treatment appropriateness.

Direct answer

ChatGPT models and reasoning efforts did not produce one stable clinic list, order or source set

Not reliably, and not to the same degree for every medical service. Across the tested ChatGPT configurations, mean provider-shortlist overlap within the same configuration was 80.3% for IVF, 34.7% for full-arch dental implants and 53.9% for bariatric programs. The visible source layer was less stable: domain overlap was 51.1%, 28.3% and 44.8%, respectively.

One answer should not be treated as a clinic’s stable AI visibility, stable position or stable source set. Provider inclusion, presentation order and visible citations are separate outcomes.

80.3%IVF provider overlap within the same configuration
34.7%dental provider overlap within the same configuration
23.8%highest prompt-level exact-URL overlap within a configuration
450isolated core executions; 15 preflight responses excluded

Study at a glance

A frozen 3 × 5 × 30 repeated-execution panel

Core answers
450
Exact prompts
3
Configurations
5
Runs per cell
30
Collection surface
ChatGPT-authenticated CLI
Authentication
ChatGPT login
Session design
Fresh ephemeral process
Fieldwork
Prompt geography
New York City
Visible-link coverage
450 of 450
Unique answer × URL pairs
3,690
Unique answer × domain pairs
3,194

Short answers for healthcare CMOs

What should a CMO infer from the ChatGPT configuration shifts?

Can one screenshot be a baseline?

No. It records one observation, while repeated answers showed material changes in provider sets, order and exact pages.

Did one configuration win?

No. Stability varied by service and measured layer. More visible links did not mean a more repeatable domain set.

Does first position mean “best”?

No. Position is presentation order in the generated shortlist, not a clinical-quality ranking, endorsement or outcome estimate.

Are recurring domains causal levers?

No. They are reader-visible links worth auditing; recurrence does not prove a domain caused a provider mention.

Explore the three prompts

Service line changed the normal range of variation

P1 · IVF
I’m looking for the best IVF clinic in New York City. Which clinics should I compare, and how much does IVF usually cost?
Provider overlap, within
80.3%
Domain overlap, within
51.1%
Exact-URL overlap, within
23.8%
P2 · Full-arch dental implants
I’m looking for the best dental practice for full-arch dental implants in New York City. Which practices should I compare, and how much does treatment usually cost?
Provider overlap, within
34.7%
Domain overlap, within
28.3%
Exact-URL overlap, within
20.2%
P3 · Bariatric surgery
I’m looking for the best bariatric surgery clinic in New York City. Which clinics should I compare, and how much does weight-loss surgery usually cost?
Program overlap, within
53.9%
Parent-system overlap
84.3%
Domain overlap, within
44.8%

Interpretation boundary

A descriptive ChatGPT model benchmark, not a patient recommendation

Tested surface

Codex CLI 0.147.0 authenticated through one ChatGPT account, with live search available. This was not consumer ChatGPT Free or Plus and not an API-key run.

Prompt geography

New York City was written into each prompt. Physical US location, IP geolocation and device location were not controlled.

Measured sources

Only links visibly present in final-answer Markdown were counted. They do not expose every retrieved or internally used source.

Measured position

Presentation order is an exposure signal, not a clinical ranking, provider endorsement or estimate of user attention.

Research design

Thirty isolated executions of every prompt and configuration

The protocol, manifests and run plan were locally frozen before collection. The core was a 3 × 5 × 30 repeated-execution panel: three exact commercial prompts, five requested configurations and 30 repetitions of every prompt/configuration combination.

The exact question was how much primary provider-shortlist composition, presentation order and reader-visible linked sources varied across repeated runs of the same configuration and across Luna Medium, Terra Medium, Sol Medium, Sol Low and Sol High.

Each of 30 rounds contained all 15 combinations in a deterministically shuffled order using seed 20260817. Collection was sequential, concurrency one, with ordinary five-second pacing and a 60-second round-boundary delay.

Every observation started a fresh ephemeral CLI process; sessions were not resumed. Live search was available and was invoked in all 450 core answers. Optional context surfaces, persistence, user rules, agents, shell, plugins, apps, memory, browser/computer control and multi-agent behavior were disabled where supported by the strict loader.

The 15 preflight answers were kept separate and excluded from every reported denominator. Content was not retried. One technical attempt, R04-P1-C4 attempt 1, failed; attempt 2 succeeded, and both artifacts were retained. The 450 successful core answers are the analytical panel.

Configuration manifest

Five configurations of three requested GPT-5.6 model variants

Table 1. Exact requested configuration mapping
IDRequested modelReasoning effortReader-facing label
C1gpt-5.6-lunamediumLuna Medium
C2gpt-5.6-terramediumTerra Medium
C3gpt-5.6-solmediumSol Medium
C4gpt-5.6-sollowSol Low
C5gpt-5.6-solhighSol High

The requested labels are the identities recorded by the collector. The study does not estimate consumer usage share, identify a default consumer model or justify weighting one configuration more heavily than another.

Visible source stability

More links did not mean a more repeatable domain set

Sol High exposed the broadest visible source set, averaging 10.3 unique URLs and 8.1 unique domains per answer across the three prompts. Its mean domain-set overlap was 40.9%, within the narrow 40.5%–42.1% range across all five configurations. Link breadth and repeatability are distinct measures.

Grouped bars compare mean exact-URL and registrable-domain overlap for repeated answers under five ChatGPT model and reasoning configurations. Exact-URL overlap ranges from 19.6% to 24.1%, while domain overlap remains between 40.5% and 42.1%.
Figure 1. Across the tested ChatGPT model and reasoning-effort configurations, Sol High exposed the most links but did not have the most repeatable domain set. Overlap is the mean within-configuration pairwise Jaccard score across the three prompts. SVG · PNG
Table 2. Citation breadth and repeat stability by configuration, averaged across the three prompts
ConfigurationMean unique URLsMean unique domainsExact-URL overlapDomain overlap
Luna Medium7.807.0119.6%40.5%
Terra Medium6.415.8320.7%42.0%
Sol Low7.666.9121.7%42.1%
Sol Medium8.837.6821.0%41.6%
Sol High10.308.0624.1%40.9%

Within versus across configurations

Provider names were generally more stable than visible source pages

IVF had a compact provider core, while dental implants were fragmented in both providers and visible domains. Bariatric results looked substantially more stable when named programs were rolled up to parent health systems.

Paired bars compare provider-set and visible-domain overlap within the same configuration and across configurations for IVF, dental implants and bariatric surgery. IVF provider overlap is highest; dental is lowest.
Figure 2. Across ChatGPT models and reasoning efforts, stability depended more on the service and measured layer than on a universal configuration hierarchy. Parent-system regrouping is shown separately for bariatric providers. SVG · PNG
Table 3. Provider-set, exact-URL and domain overlap by service
ServiceEntity levelProvider withinProvider acrossURL withinURL acrossDomain withinDomain across
IVFClinic/program80.3%79.5%23.8%17.5%51.1%43.5%
Full-arch dental implantsPractice/program34.7%28.6%20.2%15.9%28.3%23.0%
Bariatric surgerySpecific program53.9%49.9%20.2%16.6%44.8%41.8%
Bariatric surgeryParent health system84.3%82.1%

The parent-system row regroups named bariatric providers only. Source URLs and domains were not reclassified by health-system family. The differences are descriptive; they are not statistical or causal estimates.

Provider appearance

Configuration profiles differed most visibly in dental

Several IVF providers formed a stable inclusion core, while dental providers moved in different directions across the tested configurations. In bariatric answers, the exact program sometimes changed while the parent health system remained present.

Heatmap of selected provider appearance rates across Luna Medium, Terra Medium, Sol Low, Sol Medium and Sol High for IVF, dental implants and bariatric surgery. Several providers show large configuration-to-configuration shifts.
Figure 3. Selected provider appearance rates across 30 answers per ChatGPT prompt/configuration cell. The labels are descriptive visibility observations, not clinical comparisons or endorsements. SVG · PNG
Table 4. Selected provider appearance rates across 30 answers per cell
ServiceProviderLuna MediumTerra MediumSol LowSol MediumSol High
IVFNYU Langone100%100%100%100%100%
IVFColumbia100%100%96.7%100%100%
IVFWeill Cornell100%96.7%100%100%100%
IVFRMA of New York100%96.7%100%100%100%
IVFCCRM New York90.0%90.0%66.7%76.7%86.7%
IVFSpring Fertility20.0%23.3%40.0%56.7%43.3%
DentalManhattan Arch86.7%90.0%93.3%96.7%93.3%
DentalColumbia dental programs80.0%23.3%86.7%90.0%100%
DentalCentral Park Oral Surgery90.0%63.3%40.0%33.3%3.3%
DentalAll-on-Four Dental Implant Centers50.0%56.7%23.3%40.0%56.7%
DentalManhattan Aesthetic Dentistry / Smart Arches13.3%6.7%43.3%36.7%53.3%
DentalNYC Dental Implants Center36.7%16.7%13.3%23.3%30.0%
DentalClearChoice3.3%3.3%10.0%33.3%50.0%
BariatricNYU Langone100%100%100%100%100%
BariatricNYP / Weill Cornell100%96.7%100%100%100%
BariatricNYP / Columbia56.7%50.0%90.0%80.0%96.7%
BariatricMount Sinai Morningside43.3%66.7%40.0%73.3%66.7%
BariatricNYC Health + Hospitals / Bellevue60.0%33.3%66.7%73.3%73.3%
BariatricNYC Health + Hospitals / Jacobi60.0%10.0%3.3%3.3%3.3%

Presentation order

Shortlist composition and order were different stability problems

Sol High’s IVF provider set was 86.6% similar across repeat pairs, yet the same provider occupied first position in only 29.4% of pairs. A visibility monitor should not collapse “included somewhere” and “presented first” into one rank.

The full ordered shortlist rarely repeated exactly. IVF produced 16–24 unique sequences per 30 answers across configurations, dental implants produced 29–30 and bariatric surgery produced 26–29.

Heatmap compares full provider-set overlap, top-three overlap, same first provider and same-position rate for all 15 prompt-configuration cells. Composition and presentation-order stability often differ.
Figure 4. Across the tested ChatGPT configurations, a stable provider set did not imply a stable first position or presentation order. Position is the order shown in the generated shortlist, not a quality rank. SVG · PNG
Table 5. Within-configuration provider composition and presentation-order stability
ServiceConfigurationFull set overlapTop-three overlapSame firstSame position
IVFLuna Medium80.7%88.6%31.7%41.3%
IVFTerra Medium81.4%86.1%35.9%41.0%
IVFSol Low74.4%96.7%93.3%81.8%
IVFSol Medium78.5%78.5%64.4%57.9%
IVFSol High86.6%61.2%29.4%32.6%
Dental implantsLuna Medium37.5%35.5%41.8%32.1%
Dental implantsTerra Medium28.8%28.7%64.1%57.0%
Dental implantsSol Low32.4%28.5%75.4%50.3%
Dental implantsSol Medium32.9%26.0%74.7%50.2%
Dental implantsSol High41.8%31.9%63.9%44.4%
Bariatric surgeryLuna Medium43.6%63.2%64.8%49.9%
Bariatric surgeryTerra Medium47.5%57.3%74.9%53.0%
Bariatric surgerySol Low57.5%64.6%56.8%41.4%
Bariatric surgerySol Medium59.3%68.9%71.3%50.3%
Bariatric surgerySol High61.4%72.6%57.9%45.0%

Visible domain appearance

The domains visible to readers also shifted by configuration

The corrected source pipeline counted only event_type=final_answer_markdown_link, normalized the URLs and deduplicated URLs and domains inside each answer. It retained 3,757 visible Markdown-link occurrences, 3,690 unique answer × URL pairs and 3,194 unique answer × registrable-domain pairs.

Heatmap of 13 selected source-domain appearance rates across the five tested configurations and three services. Several domains recur consistently while others vary sharply by configuration.
Figure 5. Selected response-level domain appearance rates across ChatGPT models and reasoning efforts. A visible link can support pricing, regulation or general context; recurrence is not proof that a domain caused a provider recommendation. SVG · PNG
Table 6. Selected response-level domain appearance rates across 30 answers per cell
ServiceDomainLuna MediumTerra MediumSol LowSol MediumSol High
IVFny.gov93.3%100%100%100%100%
IVFnyulangone.org90.0%96.7%96.7%100%100%
IVFsartcorsonline.com33.3%43.3%0.0%26.7%80.0%
IVFrmany.com86.7%20.0%80.0%76.7%66.7%
Dentalcentralparkoralsurgery.nyc96.7%93.3%80.0%60.0%26.7%
Dentalmanhattanarchcollective.com80.0%83.3%73.3%60.0%60.0%
Dentalcolumbia.edu70.0%13.3%80.0%90.0%93.3%
Dentalclearchoice.com3.3%6.7%10.0%33.3%50.0%
Bariatricmountsinai.org100%96.7%100%100%100%
Bariatricnyulangone.org83.3%100%96.7%100%100%
Bariatricnyp.org66.7%80.0%60.0%60.0%76.7%
Bariatricnychealthandhospitals.org73.3%33.3%70.0%80.0%83.3%
Bariatriccms.gov0.0%0.0%0.0%36.7%50.0%

These links can support provider, cost, regulatory or general guidance claims. Their appearance does not prove source authority, endorsement or the probability that editing a domain will change a future answer.

Three service-line readings

The same monitoring rule would misdescribe all three niches

IVF: a stable inclusion core, but unstable first position

NYU Langone appeared in 100% of the 150 IVF answers. Columbia, RMA of New York and Weill Cornell each appeared in at least 99.3%. The core did not imply stable order: Weill Cornell’s first-position rate ranged from 20.0% in Luna Medium to 96.7% in Sol Low.

Dental: the most fragmented provider and source profile

Central Park Oral Surgery appeared in 90.0% of Luna Medium answers and 3.3% of Sol High answers. ClearChoice moved in the opposite direction, from 3.3% in Luna Medium to 50.0% in Sol High. This is an observed profile difference, not proof that reasoning effort caused the shift.

Bariatric: entity granularity changed the conclusion

Mean within-configuration provider overlap was 53.9% at exact-program level and 84.3% after programs were grouped into parent health systems. A hospital network can therefore appear consistently while the named facility changes.

Practical framework

How a CMO should monitor ChatGPT models and reasoning efforts

  1. Measure inclusion, position and sources separately. Track provider appearance, top-three, first position, parent/facility identity, visible domains and exact pages as different outcomes.
  2. Use repeated observations. One response is an observation, not a baseline. Choose repetition based on the precision required for the decision; 30 is not a universal minimum.
  3. Build service-specific baselines. IVF, dental and bariatric prompts had materially different normal ranges of variation.
  4. Keep configurations as separate strata. The study does not reveal which tested configuration represents prospective-patient usage.
  5. Audit recurring domains without calling them causal levers. Check whether owned and third-party pages are current, accessible and factually accurate.
  6. Preserve hospital-system hierarchy. Report parent-network and exact-facility presence when both identities affect interpretation.
  7. Retain enough evidence to distinguish drift from noise. Compare repeated batches with exact prompt, configuration, timestamp, provider sequence and visible-link records.

Methodology

Two measured surfaces with explicit denominators

Primary provider shortlist

The analysis used the first contiguous primary provider shortlist in each saved final answer. Later “also consider” lists were excluded. Aliases were mapped to canonical clinic, practice or program entities. Duplicates inside one answer were retained at their earliest presentation position; a compound row could give two providers a tied position. Bariatric providers were also mapped to parent health systems for a separate sensitivity view.

Appearance rate uses 30 answers in one prompt/configuration cell or 150 answers for one prompt across configurations. Within each 30-answer cell, all 435 answer pairs were compared. Between two configurations, all 900 cross-configuration answer pairs were compared. Those pairings are computed comparisons, not statistically independent users; the experimental base remains 30 executions per cell.

Reader-visible links

Only final_answer_markdown_link records were eligible. Raw JSON trace URLs, search-event URLs and duplicate plain-URL regex rows were excluded. URLs were normalized and deduplicated within each answer; registrable domains were also deduplicated within each answer.

All 450 core answers contained at least one visible linked domain. The 3,757 link-occurrence total is provenance, not an exposure or authority score. Publication metrics use the 3,690 unique answer × normalized-URL pairs and 3,194 unique answer × registrable-domain pairs.

Quality checks

  • All 450 planned core observations completed successfully; 15 preflight observations were excluded.
  • All 450 primary shortlists contained at least three rows; 2,276 shortlist rows were extracted.
  • Canonical normalization produced 2,359 provider mentions after within-answer deduplication, with zero unmapped shortlist rows.
  • C3 is Sol Medium and C4 is Sol Low; the configuration identities were checked against the frozen manifest.
  • Ten automated analytical tests passed for expected counts, configuration identity, metric ranges and pair denominators.

Limitations

What the benchmark cannot establish

  • Surface: this was a ChatGPT-authenticated CLI collection, not browser ChatGPT Free or Plus, a mobile app, an API-key run, Gemini or Google Search.
  • Geography: New York City was specified in each prompt; physical location and IP were not controlled.
  • Scope: one exact prompt per niche, three niches, one city, one account, one Mac, one network context and one continuous collection window.
  • Prompt wording: every prompt combined “best,” a comparison shortlist and a cost question; paraphrases and other intents were not tested.
  • Time: the sequential core ran for 9 hours 26 minutes 59 seconds. Shuffling distributed configurations through the window, but time and configuration are not fully separable.
  • Backend state: every observation used a fresh ephemeral local process, but server-side caching, retrieval state and platform changes were not observable or controllable.
  • Inference: no confidence intervals, hypothesis tests or causal model were used. Differences cannot be attributed solely to model family or reasoning effort.
  • Sources: only visible final-answer Markdown links were measured. Hidden retrieval sources and source quality were not evaluated.
  • Domains: registrable-domain keys are mechanical normalization fields, not source-class or authority judgments.
  • Normalization: aliases and parent systems were mapped using explicit rules; the entity level materially changes the bariatric result.
  • Position: shortlist position is presentation order, not an objective rank, endorsement or estimate of user attention.
  • Medical and commercial outcomes: the study did not assess factual or clinical accuracy, treatment appropriateness, impressions, clicks, leads, appointments, revenue or patient outcomes.
  • Audience: it does not estimate which configuration real prospective patients use.
  • Retry: one observation required a technical retry; the failed artifact was retained and the successful replacement was analyzed.

Definitions and further research

A vocabulary for repeatable AI visibility measurement

Table 7. Definitions used in the study
TermDefinition
Core responseOne successful observation in the frozen 450-answer analysis panel.
Primary provider shortlistThe first contiguous list or table of provider organizations in the final answer.
Presentation positionOrdinal row in that shortlist; not a quality rank.
Visible cited URLA Markdown link visibly present in the saved final answer.
Jaccard similarityIntersection divided by union for two sets; 100% means identical sets.
Within-configuration stabilityMean across all 435 pairs among 30 answers in the same prompt/configuration cell.
Across-configuration stabilityMean across 900 pairs for a configuration pair, summarized across ten pairs per prompt.

Next replications should freeze new observation windows across days, add prompt paraphrases and cities, and test other consumer or API surfaces as separate datasets. The current results should remain a dated baseline rather than being silently combined with a future wave.

Downloadable research materials

Download the public v1.0 research package

The versioned package contains publication-safe aggregate tables, sanitized observation-level evidence, frozen prompts and configurations, methods, citation metadata and checksums under CC BY 4.0. It excludes raw JSONL, full answers, stderr, machine paths, authentication state, the private archive and unsafe raw provider-extraction fields.

DataSummary CSV

Headline provider, exact-URL and domain stability metrics.

text/csv · 703 B
DataComplete public JSON

Machine-readable metadata and all publication-safe tables.

application/json · 3.4 MB
DataAnalysis workbook

Formatted workbook of results, manifests, definitions and limits.

application/vnd.openxmlformats-officedocument.spreadsheetml.sheet · 290 KB
ReproducibilityPrompt manifest

The three frozen commercial New York City prompts.

text/csv · 603 B
ReproducibilityConfiguration manifest

Requested model labels and reasoning settings for C1–C5.

text/csv · 251 B
Provider dataProvider frequency

Appearance, top-one, top-three and conditional position measures by configuration.

text/csv · 85 KB
Provider dataWithin-configuration provider stability

Pairwise provider-set and presentation-order stability within each cell.

text/csv · 5.0 KB
Provider dataCross-configuration provider stability

Provider-set comparisons across configuration pairs.

text/csv · 9.2 KB
Source dataVisible-source summary

Visible link breadth and repeat stability for each configuration.

text/csv · 605 B
Source dataWithin-configuration source stability

Exact-URL and registrable-domain overlap within each cell.

text/csv · 1.7 KB
Source dataCross-configuration source stability

Exact-URL and registrable-domain overlap across configuration pairs.

text/csv · 2.6 KB
DocumentationData dictionary

Field definitions, denominators and interpretation boundaries.

text/csv · 22 KB
DocumentationPackage README

Scope, headline results, package map and public-data boundary.

text/markdown · 2.7 KB
DocumentationMethodology

Frozen design, extraction rules, metrics, checks and limitations.

text/markdown · 3.5 KB
VerificationPackage manifest

Versioned inventory with file sizes, row counts and SHA-256 values.

application/json · 10 KB
VerificationSHA-256 checksums

Integrity checksums for the complete public release.

text/plain · 4.4 KB
CitationCitation File Format

CFF 1.2 metadata for research and repository tools.

text/plain · 837 B
CitationBibTeX citation

Ready-to-import BibTeX citation metadata.

text/plain · 593 B

Version 1.0 · Published August 29, 2026 · Manifest · SHA-256 checksums

Citation and research record

A traceable, versioned publication record

Yudin, E. (2026). Clinic Shortlists and Sources Across ChatGPT Models and Reasoning Efforts (Version 1.0) [Data set]. Rotgar Research. https://doi.org/10.5281/zenodo.22162725

Author
Evgeniy Yudin, Founder and Strategy Lead · ORCID 0009-0007-8400-9561
Methodology reviewer
Boris Teplyakov, SEO Lead
Review scope
Methodology review; not peer review, clinical review or an assessment of named providers
Version and license
Version 1.0; CC BY 4.0
Fieldwork
August 18, 2026
Publication
August 29, 2026
Surface
ChatGPT-authenticated CLI collection with live search available
Language and scope
US English; New York City named in each prompt

Free audit

Build a repeatable medical AI visibility baseline

The simplest free audit starts with one clinic or selected location, one priority market and one patient language.

Best fitClinics, hospitals and medical groups in:
  • Dental
  • Aesthetic medicine
  • Physiotherapy
  • Fertility & IVF
  • Dermatology
  • Hair transplant
  • Plastic surgery
  • Mental health
  • Multi-specialty groups
  • Your current visibility and named competitors
  • The sources AI answers rely on
  • A prioritized list of improvements to implement

Several clinics, locations, countries, markets or patient languages can be included without a call. Prefer to discuss the scope? Book a 30-minute call We’ll tailor the audit at no cost.

Free Google + AI visibility audit

Google Search + Maps, Google AI Overviews, ChatGPT + Gemini.

  1. 1Contact
  2. 2Priorities
  3. 3Focus

Step 1 of 3: Contact

Step 1 of 3

Contact details

Enter at least one contact: email or WhatsApp.

Step 2 of 3

Audit priorities

Step 3 of 3

Search focus

What are you most interested in?

Your request is sent securely to Rotgar.

Public data is sufficient. No call or account access is required. No obligation. Paid work is quoted separately.

We will confirm the scope and that the audit is in progress within 2 business days. The initial report will be delivered within 6 business days.