Rotgar Research · Medical

ChatGPT Named Most Final-Shortlist Medical Providers in Its First Search Queries

Across 180 commercial medical-provider responses, 85% of final provider positions were already named in ChatGPT's first returned search-action queries.

Boris Teplyakov reviewed the research methodology as SEO Lead. This was not a clinical subject-matter review or an assessment of the providers named in the responses.

Direct answer

Most final provider positions were already visible in the first returned search action

Across 180 commercial medical provider-selection responses, 85.0% of final provider positions were occupied by organizations already named in the query strings of ChatGPT’s first returned web-search action. The connection was strongest for broad comparisons and weakest for cost-and-insurance questions.

85.0%612 of 720 final positions already named
179/180responses with at least one provider in the first action
36.1%cost-and-insurance positions absent from the first action

Study at a glance

A frozen US medical provider-selection panel

Responses
180
Prompt cells
72
Service lines
6
US metros
3
Patient tasks
3
Fieldwork

Explore the three patient tasks

The shortlist path changed with the commercial question

Broad comparison93.0% retained
Responses
108
Final positions absent from first action
7.4%
Final shortlist stability
77.9%
Appointment access87.0% retained
Responses
36
Final positions absent from first action
16.7%
Final shortlist stability
61.2%
Cost and insurance72.4% retained
Responses
36
Final positions absent from first action
36.1%
Final shortlist stability
50.4%

How to read this study

An observational shortlist study, not a provider rating

Observable stages only

The study compares provider names in the first returned search action with providers in the final answer. It does not expose hidden reasoning or establish where initial candidates originated.

Visibility, not clinical quality

Inclusion is a measured visibility outcome. It is not a clinical-quality assessment, a patient recommendation or evidence that one provider is better than another.

Dated API configuration

The results describe one August 18, 2026 Responses API configuration. They are not a test of every model, consumer interface or personalized account state.

The answer in brief

Most final-shortlist providers were already present in the first observable query set

Across the 180 responses, 612 of 720 final provider positions, or 85.0%, were occupied by organizations already named in the query strings of the first returned web-search action.

At least one provider organization appeared at that stage in 179 of 180 responses. Each first action contained four query strings.

The relationship between that first observable candidate set and the final answer depended on the patient’s task. Broad comparison answers retained 93.0% of provider mentions from the first search action. Only 7.4% of their final provider positions contained an organization absent from that first action. For cost-and-insurance questions, retention fell to 72.4%, while 36.1% of final provider positions contained an organization absent from the first action.

The final shortlists also became less consistent across repeated runs. Mean provider-set similarity was 77.9% for broad comparisons, 61.2% for appointment-access questions and 50.4% for cost-and-insurance questions.

For a health system CMO, this creates two measurement questions: Did the organization enter the first observable candidate set? Did it remain in, or appear only in, the final answer? A single AI visibility score cannot answer both.

What we tested

A frozen set of commercial provider-selection prompts

We created brand-free commercial questions asking ChatGPT to return exactly four local providers. The study covered:

  • six service lines: fertility, orthopedics, ophthalmology, urgent care, primary care and bariatric surgery;
  • three metros: New York, Chicago and Houston;
  • three patient tasks: broad provider comparison, appointment access and cost or insurance;
  • 180 fresh responses across 72 frozen prompt cells.

Each response began without prior conversation state. The broad comparison used two neutral prompt phrasings and three repeats per service-line and metro cell. The appointment-access and cost-and-insurance extensions used two repeats per cell.

We captured the four query strings in the first returned web-search action and the four providers in the final answer. We then compared normalized provider entities at the two stages.

First action

Provider names appeared in 179 of 180 first search actions

The user prompts contained no provider brands. Even so, at least one provider organization appeared in the first returned search action in:

  • 108 of 108 broad comparison responses;
  • 36 of 36 appointment-access responses;
  • 35 of 36 cost-and-insurance responses.

This is not a claim that 99.4% of providers were chosen before search. It is a response-level result: 179 of 180 first search actions contained at least one provider among their four query strings.

The query strings are observable API output. They show the names used at the first search step, but not where those names originated.

Patient task

The patient’s task changed the shortlist path

Table 1. Shortlist-path metrics by patient task
Patient taskInitial provider mentions retained in the final answerFinal provider positions absent from the first actionFinal shortlist stability
Broad comparison93.0% (400 of 430)7.4% (32 of 432)77.9%
Appointment access87.0% (120 of 138)16.7% (24 of 144)61.2%
Cost and insurance72.4% (92 of 127)36.1% (52 of 144)50.4%

The first two percentages in each row use different denominators. Retention starts with provider mentions in the first search action. Later addition starts with provider positions in the final answer. They are related measures, not complementary shares.

Broad comparison prompts were mostly candidate-led in this collection: nearly all initial provider mentions survived into the final answer, and relatively few final positions were filled by providers absent from the first action.

Cost-and-insurance prompts produced a larger difference between the first observable set and the final answer. More than one-third of final provider positions were occupied by organizations absent from the first action. Appointment-access prompts sat between the two.

For reporting, do not combine these tasks into one score. A health system can look established in a general comparison and still be inconsistent when a patient asks about insurance, self-pay, financing, scheduling or locations.

Two-panel chart comparing three patient tasks. Initial provider mentions retained in final answers were 93.0% for broad comparison, 87.0% for appointment access and 72.4% for cost and insurance. Final providers absent from the first action were 7.4%, 16.7% and 36.1%, respectively. The panels use different denominators.
Figure 1. The relationship between the first returned search action and the final provider list varied by commercial task. Retention uses first-action provider mentions as its denominator; first-action-absent presence uses final provider positions.

Cost and insurance

Cost answers included first-action-absent providers across all six service lines

In the cost-and-insurance sample, the share of final provider positions absent from the first action ranged from 25.0% to 50.0%:

Table 2. Final cost-and-insurance provider positions absent from the first action
Service lineFinal providers absent from the first action
Bariatric surgery50.0% (12 of 24)
Orthopedics41.7% (10 of 24)
Fertility37.5% (9 of 24)
Ophthalmology37.5% (9 of 24)
Primary care25.0% (6 of 24)
Urgent care25.0% (6 of 24)

Each service-line cut contains six responses across three prompt cells, so the ordering is exploratory. The practical point is broader: every service line tested produced final cost-and-insurance answers containing providers absent from the first action.

For a CMO, cost visibility needs its own audit. Review whether the public web gives a patient clear, current answers about accepted insurance, self-pay, financing, estimate requests and the next step for contacting the organization. Review both owned pages and external pages that answer the same commercial question.

Horizontal bars showing the share of final cost-and-insurance provider positions absent from the first search action: bariatric surgery 50.0%, orthopedics 41.7%, fertility 37.5%, ophthalmology 37.5%, primary care 25.0% and urgent care 25.0%.
Figure 3. Every tested service line produced cost-and-insurance answers containing providers absent from the first action. Each service-line cut contains six responses across three prompt cells and is exploratory.

Repeatability

One answer is not a stable benchmark

We measured stability with pairwise Jaccard similarity. A score of 100% means repeated responses returned the same provider set; 0% means they shared none.

The final provider set averaged 77.9% similarity for broad comparisons, 61.2% for appointment access and 50.4% for cost and insurance. The cost shortlist was therefore the least consistent of the three task types in this collection.

One screenshot can show whether a provider appeared once. It cannot show whether the result persists. A useful monitoring program freezes the exact question and location, repeats the run, records the configuration and reports appearance frequency alongside shortlist similarity.

Paired dot plot of mean provider-set similarity. First-action and final-answer similarity were 78.4% and 77.9% for broad comparison, 67.1% and 61.2% for appointment access, and 65.6% and 50.4% for cost and insurance.
Figure 2. Final provider lists were less stable across identical repeated prompts for appointment-access and cost-and-insurance questions than for broad comparisons. Similarity is mean within-cell pairwise Jaccard.

For health system CMOs

A measurement framework for health system CMOs

Track the stages separately:

  1. First-action presence. Did the provider appear in any query string in the first returned search action?
  2. Final inclusion. Did the provider appear in the final four-provider answer?
  3. Retention. When the provider appeared in the first action, did it remain in the final answer?
  4. Final-only presence. When the provider appeared in the final answer, was it absent from the first action?
  5. Stability. How often did it appear across repeated runs of the same frozen prompt?

Build separate views for broad comparison, appointment access and cost or insurance. Then segment by service line and metro. This keeps a strong general-comparison result from masking a weak or unstable result for a high-intent patient task.

Use the result to choose the next audit:

  • If a provider appears early and stays, document the commercial questions for which that pattern holds.
  • If it appears early but drops out, compare the final answer’s retained providers and their supporting public information.
  • If it appears only in the final answer, identify which public pages answer that specific patient task.
  • If it rarely appears across repeats, treat the result as unstable rather than as a fixed rank.

These are measurement paths, not a universal optimization recipe. The study shows where shortlist changes occurred; the page and source review explains what to investigate next.

Context

How this extends earlier work

This study was prompted by Suganthan Mohanadasan’s analysis of provider and product names in ChatGPT’s first search queries. That work examined a logged-in ChatGPT account across a mixed, largely commercial query set and explicitly treated its percentages as directional.

We tested a narrower question in health care: whether provider organizations appeared in the first returned search action for brand-free medical provider-selection prompts, whether they persisted into the final answer and how the pattern changed across commercial tasks. The two studies use different surfaces and samples, so their percentages should not be pooled.

Methodology

A dated observational snapshot

The analytical sample contains 180 completed responses from August 18, 2026. Collection used the OpenAI Responses API with the requested and returned model alias chat-latest, required web_search, low search context, fixed approximate metro location, no prior conversation state, store=false and a 1,200-token output ceiling. OpenAI describes web search as a Responses API tool that lets models access current web information and return sourced answers.

The response was the observational unit. Provider names were normalized to organizations or consumer-facing provider brands. Rule-based extraction was followed by manual adjudication of every unmatched query string and all identified ambiguous positive matches. The analytical sample had complete coverage: four first-action query strings, four final providers and four citation URLs in every response.

Retention equals retained initial provider mentions divided by all normalized provider mentions in the first action. First-action-absent final presence equals final provider entities absent from the first action divided by all final provider entities. Stability is the mean within-cell pairwise Jaccard similarity of normalized provider sets across repeated responses. Ninety-five percent prompt-cell cluster bootstrap intervals used 10,000 replications with the exact prompt cell as the cluster unit; the six fixed scenario seeds are published in the downloadable methodology and scenario data. The resulting intervals were 89.8%–95.8%, 81.2%–92.2% and 66.2%–79.4% for retention; and 4.4%–10.9%, 9.7%–24.3% and 28.5%–44.4% for first-action-absent final presence, in broad comparison, appointment access and cost or insurance respectively.

The 180 analytical responses cost an estimated $16.98. Including the excluded feasibility pilot, calibration and stopped first attempt, total estimated API spend was $19.12.

Limits

What this study does and does not show

This is a dated snapshot of one chat-latest API configuration, not a test of the consumer ChatGPT interface, every model or personalized account behavior. Service-line and metro cuts are descriptive and exploratory. Two repeats in the commercial extensions and three in the broad comparison estimate short-run variation but do not create new prompt families.

The first returned search action is observable; hidden reasoning and the origin of initial provider candidates are not. Provider inclusion is a visibility outcome, not a clinical-quality rating or a recommendation to patients.

Raw API responses remain private. The page reports reviewed aggregate results and does not publish response identifiers, full answers or encrypted tool content.

Suggested citation: Yudin, Evgeniy. “ChatGPT Named Most Final-Shortlist Medical Providers in Its First Search Queries.” Rotgar Research dataset, version 1.0, August 18, 2026. https://rotgar.com/medical/resources/chatgpt-medical-provider-shortlists

Downloadable research materials

Download the public v1.0 research package

The public package contains aggregate and anonymous response-level metrics, frozen prompts, documentation and integrity records under CC BY 4.0. It excludes raw API responses, returned query text, provider-level mention records, response IDs, credentials and internal collection identifiers.

DataPrimary CSV

Overall and scenario summary metrics in a compact tabular file.

text/csv · 1.6 KB
DataComplete public JSON

Machine-readable public tables, metadata, metrics and interpretation boundaries.

application/json · 331 KB
DataResearch workbook

Formatted workbook with ten documented research sheets.

application/vnd.openxmlformats-officedocument.spreadsheetml.sheet · 62 KB
ReproducibilityPrompt manifest

The 72 frozen, brand-free commercial prompt cells used in the study.

text/csv · 46 KB
ReproducibilityStability pairs

The 144 repeat-pair rows used to reproduce provider-set stability without exposing provider names.

text/csv · 4.3 KB
ReproducibilityData dictionary

Definitions for every field in the public package.

text/csv · 7.7 KB
DocumentationPackage README

Scope, direct result, package contents and public-data boundary.

text/markdown · 2.6 KB
DocumentationMethodology

Research design, observable stages, metrics, entity handling and interpretation limits.

text/markdown · 3.5 KB
DocumentationDataset metadata

Machine-readable identity, scope, configuration, counts and license metadata.

application/json · 4.2 KB
VerificationPackage manifest

Versioned file inventory with bytes, media types, row counts and SHA-256 values.

text/plain · 5.4 KB
VerificationSHA-256 checksums

Integrity checksums for the complete public package.

text/plain · 2.2 KB
CitationBibTeX citation

Ready-to-import BibTeX citation metadata.

text/plain · 600 B
CitationCitation File Format

CFF 1.2 metadata for research and repository tools.

text/plain · 883 B

Version 1.0 · Published August 18, 2026 · Manifest · SHA-256 checksums

Citation and research record

A traceable, versioned publication record

Yudin, Evgeniy. “ChatGPT Named Most Final-Shortlist Medical Providers in Its First Search Queries.” Rotgar Research dataset, version 1.0, August 18, 2026. https://rotgar.com/medical/resources/chatgpt-medical-provider-shortlists

Author
Evgeniy Yudin, Founder and Strategy Lead
Methodology reviewer
Boris Teplyakov, SEO Lead
Version and license
Version 1.0; CC BY 4.0
Fieldwork
August 18, 2026
Surface
OpenAI Responses API with required web_search
Language and market
US English; New York, Chicago and Houston

Boris Teplyakov reviewed the research methodology as SEO Lead. This was not a clinical subject-matter review or an assessment of the providers named in the responses.

No DOI or external archive record is assigned to this version.

Free audit

See how your organization appears in medical AI answers

The simplest free audit starts with one clinic or selected location, one priority market and one patient language.

Best fitClinics, hospitals and medical groups in:
  • Dental
  • Aesthetic medicine
  • Physiotherapy
  • Fertility & IVF
  • Dermatology
  • Hair transplant
  • Plastic surgery
  • Mental health
  • Multi-specialty groups
  • Your current visibility and named competitors
  • The sources AI answers rely on
  • A prioritized list of improvements to implement

Several clinics, locations, countries, markets or patient languages can be included without a call. Prefer to discuss the scope? Book a 30-minute call We’ll tailor the audit at no cost.

Free Google + AI visibility audit

Google Search + Maps, Google AI Overviews, ChatGPT + Gemini.

  1. 1Contact
  2. 2Priorities
  3. 3Focus

Step 1 of 3: Contact

Step 1 of 3

Contact details

Enter at least one contact: email or WhatsApp.

Step 2 of 3

Audit priorities

Step 3 of 3

Search focus

What are you most interested in?

Your request is sent securely to Rotgar.

Public data is sufficient. No call or account access is required. No obligation. Paid work is quoted separately.

We will confirm the scope and that the audit is in progress within 2 business days. The initial report will be delivered within 6 business days.