Methodology
How DeepContour measures AI visibility.
AI answers are probabilistic. The same question, asked twice, can name different surgeons. So we do not report a rank. We measure presence, share and persistence across repeated observations, we state the sample behind every number, and we say plainly what cannot be measured.
Updated · See it applied in our Houston study of 117 practices
The principle: observation over assertion
A patient asking an AI system which surgeon to see receives an answer assembled from many sources at that moment. DeepContour treats that answer as something to observe, repeatedly and under recorded conditions, rather than something to predict or promise.
Every number we report carries its context: the engine, the market, the procedure, the date range and the number of runs behind it.
Question sets: the way patients actually ask
For each engagement we agree a set of patient-intent questions per procedure and market, written the way patients phrase them rather than the way marketers do. For example:
- “Who are the best facelift surgeons in Houston?”
- “Which surgeon should I see for a deep plane facelift near me?”
- “Is Dr. [name] board certified?”
The set includes recommendation questions (“whom should I see”), association questions (“who performs”) and recognition questions about the practice by name, which together separate being known, being associated and being recommended.
Geography
Recommendations are local. Each question carries its location explicitly, and where an engine supports it, the location context is set as well. A market is defined with the practice: a metro area, a set of neighbourhoods, or a radius around each office.
Procedures, measured one at a time
Authority in one procedure does not transfer to another. A strong facelift reputation says nothing about whether AI recommends the same practice for rhinoplasty. We measure each procedure separately:
Facelift · Deep Plane Facelift · Neck Lift · Rhinoplasty · Revision Rhinoplasty · Blepharoplasty · Breast Augmentation · Breast Lift · Breast Reduction · Mommy Makeover · Tummy Tuck · Liposuction · Body Contouring · Post-GLP-1 Body Contouring · Hair Restoration · Non-Surgical Facial Rejuvenation
Engines
Engines in scope are agreed per engagement from: ChatGPT, Google AI, Gemini, Perplexity, Claude, Copilot. Each observation records the engine, the access surface (for example an app, a search results page or an API) and the date, because the same engine can answer differently on different surfaces.
Repeated observations
Each question is asked repeatedly on each engine within the same period. Ten runs per question per engine is our standard sample. One answer is an anecdote; ten are a measurement.
This is a deliberate response to the evidence. Independent research found less than a 1-in-100 chance that AI tools return the same list of recommendations twice, and that measuring visibility across many prompts and runs is reasonable where tracking a ranking position is not. Technical work on language-model inference shows that even with randomness settings fixed, hosted models can return different answers to identical requests.
The six measurements
Each measurement answers one question a surgeon would ask. Definitions are fixed and applied the same way to every client and competitor.
- AAnswer PresenceDeepScan
Does AI name your practice at all?
Each patient-intent question in the agreed set is asked of each engine in scope. A question counts as present when your practice or one of your surgeons is named in the answer.
Questions where you are named ÷ questions asked
- BShare of AnswerDeepScan
When AI recommends surgeons, how many of those recommendations are yours?
Every surgeon or practice named as a recommendation is counted across all questions, runs and engines in the period. Your share is your count divided by the total.
Your recommendations ÷ all surgeon and practice recommendations observed
- CRecommendation PersistenceDeepScan
Does AI recommend you consistently, or did you appear once?
Each question is asked repeatedly on each engine within the same period. Persistence is the share of those runs in which you are recommended, reported with the run count, engine, market, procedure and dates.
Runs in which you are recommended ÷ total runs
- DEntity AccuracyDeepShield
When AI talks about you, are the facts right?
Every factual claim an answer makes about the practice is compared with the verified record the practice supplies and with public registries. A fact stated inconsistently across runs counts as unverified.
Facts verified ÷ facts checked
- ECitation AuthorityDeepCite
What evidence is AI relying on when it talks about you?
Sources cited or evidently drawn on in each answer are classified by type and weighted by authority, relevance to the procedure, and consistency with your verified record, producing a 0–100 index with the source map behind it.
- FProcedure OwnershipDeepGraph
Which procedures does AI connect to you strongly enough to recommend you?
Presence, share, persistence and citation strength are computed separately for each procedure and classified as Strong, Emerging, Weak, Not established, or Opportunity where no competitor holds the answer.
Share of Answer is not market share. It is the observed share of recommendations within the monitored question set, engines and period, and it is reported with that scope attached.
Citation and source mapping
For each answer we record the sources it cites, and where an engine does not cite, the sources whose content it evidently reflects. Sources are classified by type: the practice’s own site, physician profiles, medical directories, professional societies, review platforms, publications, media and local sources. The result is a map of the evidence path behind each recommendation, for you and for the practices recommended instead of you.
Engines do not always disclose their sources, so the map distinguishes cited sources from inferred ones.
Entity validation
At the start of an engagement the practice supplies a verified record: surgeon and practice names, locations, phone numbers, website, board certifications, procedures, specialties, affiliations, financing and consultation options. Credentials are additionally checked against public registries such as the American Board of Plastic Surgery.
Every factual claim an answer makes is compared with that record. A fact stated correctly in some runs and incorrectly in others counts as unverified, because a patient sees one run, not the average.
Competitor comparison
Competitors are not chosen by us. They are the practices the engines name in the same answers. Each is measured with the same question set and the same definitions, which is what makes Share of Answer and displacement comparable.
DeepContour accepts a limited number of practices and does not represent directly competing clients within the same defined market.
Temporal tracking
Measurements repeat on a fixed cycle so change can be separated from noise. Reports read four horizons: today (what AI says now), 30 days (what changed), 90 days (whether recommendation presence is building or eroding) and 180 days (the structural trend). A practice appearing once is not the goal; durable recommendation presence is.
Limitations
- AI outputs are probabilistic and can change between runs, even for identical questions.
- Engines update models, retrieval and interfaces without notice; a change in a number can reflect the engine, not the practice.
- Personalisation, signed-in state, conversation history and device can change what an individual patient sees. Our observations use controlled, non-personalised conditions and are a sample, not a census.
- Engines do not publish how they choose recommendations. Source maps show what answers cite or reflect, not the engines’ internal logic.
- Every reported figure is bounded by its sample: the questions, engines, markets, runs and dates it was measured on.
Claims we don’t make
Correlation is not causation. Studies of AI answers report patterns, for instance in the kinds of sources that recommended practices tend to have. We report such patterns as observed, associated or frequently present among recommended practices.
We do not claim that financing, video, word count, reviews, structured data, booking software or any other single website characteristic has been proven to make ChatGPT or Google AI recommend a surgeon, and we do not guarantee rankings, recommendations or patient outcomes.
Research informing our methodology
Each study is summarised from its primary source. We note what it informs in our method, not what it proves about any practice.
Industry research
2026-01
AIs are highly inconsistent when recommending brands or productsRand Fishkin (SparkToro) and Patrick O’Donnell (Gumshoe.ai) · SparkToro
600 volunteers ran 12 prompts through ChatGPT, Claude and Google’s AI 2,961 times. There was less than a 1-in-100 chance of the same list of recommendations appearing twice, and about 1 in 1,000 of the same list in the same order. The authors conclude that visibility measured across many prompts and runs is a reasonable metric, while tracking ranking position is not.
Informs Why DeepContour reports Recommendation Persistence and Share of Answer across repeated runs, and never a single “rank”.
Technical report
2025-09
Defeating Nondeterminism in LLM InferenceHorace He and Thinking Machines Lab · Thinking Machines Lab
Sampling 1,000 completions of the same prompt at temperature 0 produced 80 unique completions. The authors attribute most run-to-run variation in hosted inference endpoints to server load changing batch sizes.
Informs Why a single observation is not evidence, even with settings fixed, and why every number carries its run count.
Peer-reviewed
2026-07
Identification of Board-Certified Plastic Surgeons Using Artificial Intelligence: An Accuracy AssessmentGupta R, Bhagwat AM, Hainline M, Lund H, Mailey BA, Kaswan S · Aesthetic Surgery Journal Open Forum, 2026;8:ojag123 (doi:10.1093/asjof/ojag123)
ChatGPT, Perplexity, Gemini and Claude were asked to name board-certified plastic surgeons across all 50 states. Across 1,000 results, 94.1% were plastic surgeons; most errors were otolaryngologists and general surgeons. The authors advise patients to verify AI answers independently.
Informs Entity Accuracy: credentials are a fact AI can get wrong, so DeepShield checks them every cycle.
Peer-reviewed
2025-10
When does ChatGPT refer someone to a plastic surgeon?Dagi, Jones, Bogue et al. · Journal of Plastic, Reconstructive & Aesthetic Surgery, 2025;109:20–24
A peer-reviewed examination of when ChatGPT directs a user with a plastic-surgery question to a plastic surgeon.
Informs Answer Presence: whether the answer routes a patient to a surgeon at all is measured before who is named.
Survey research
2025-07
Google users are less likely to click on links when an AI summary appears in the resultsPew Research Center · Pew Research Center
Using March 2025 browsing data from 900 U.S. adults, users clicked a traditional search result in 8% of visits to pages with an AI summary, against 15% without one, and clicked a link inside the summary in 1% of visits.
Informs Why the answer itself is measured: when the summary satisfies the question, the click to your website may never happen.
Peer-reviewed
2024-08
GEO: Generative Engine OptimizationPranjal Aggarwal et al. · Proceedings of KDD 2024 (arXiv:2311.09735)
Introduces a benchmark and visibility metrics for how often and how prominently sources appear in generative engine answers, reporting visibility changes of up to 40% under the paper’s experimental conditions.
Informs Citation Authority measures source visibility inside answers. The reported gains are benchmark results, not an expected outcome for any practice.
Questions surgeons ask
- Is DeepContour an SEO service?
- No. DeepContour measures the AI answer itself: whether a practice is named when patients ask, how consistently, which competitors appear instead, whether the facts are right, and what evidence the answer relies on.
- Can DeepContour guarantee that AI will recommend my practice?
- No. AI answers are probabilistic and engines change without notice. DeepContour provides measurement, intelligence and evidence-backed improvement pathways. It does not guarantee rankings, recommendations or patient outcomes.
- Why don’t you report my “rank” in ChatGPT?
- Because a rank does not survive a second run. Published research found less than a 1-in-100 chance of AI tools returning the same list of recommendations twice. Presence, share and persistence across repeated runs are stable enough to act on; a single position is not.
- Which factors make AI recommend a surgeon?
- Nobody outside the AI companies can say with certainty. Studies report patterns and correlations, such as which source types recommended practices tend to have. DeepContour reports what it observes and labels it as observation, never as a proven ranking factor.
See what AI says about your practice, measured this way.
Request a Private Visibility Audit