Guide · Measurement
How to track whether ChatGPT recommends your plastic surgery practice.
To track whether ChatGPT recommends your plastic surgery practice, ask it the questions patients ask, in a clean session with your city named, ten times each, and record every surgeon it names. Your recommendation rate is the share of runs that name you. One answer proves nothing, because the next run can name someone else.
By Shayne Beavan, founder of Deep AI Solutions · Published
The nine steps
- Write the questions patients ask. List 12 to 20 questions per procedure and city, phrased the way patients phrase them. Mix recommendation questions (“Who is the best facelift surgeon in Houston?”), association questions (“Who performs deep plane facelifts in Houston?”) and questions about you by name (“Is Dr. [name] board certified?”).
- Use a clean session. Sign out of ChatGPT or use a Temporary Chat with memory off. Your own account history personalises answers, so checking from your usual login measures your history, not what a new patient sees. Keep the city in the question.
- Ask each question ten times. Start a new chat for every run. Ten runs per question is enough to see whether a recommendation repeats; a single answer is an anecdote.
- Record every practice named. For each run, log the date, the engine, the question, every surgeon or practice named, whether you were named, and the sources the answer cites.
- Calculate presence, share and persistence. Answer Presence is questions where you were named divided by questions asked. Share of Answer is your mentions divided by all practice mentions. Recommendation Persistence is runs in which you were named divided by total runs, reported per question.
- Check what it says about you. Compare every fact the answers state about your practice (names, locations, phone, board certification, procedures, financing, consultation options) with your verified record. A fact that is right in some runs and wrong in others counts as wrong.
- Map the sources. Tally the domains cited across all runs. They show which directories, review platforms, publications and pages the answers draw on, for you and for the practices named instead of you.
- Repeat on the other engines. Run the same questions on Google AI Overviews and AI Mode, Gemini, Perplexity, Claude and Copilot. Each engine answers differently, so report each one separately.
- Repeat monthly, the same way. Keep the questions, wording, run count and conditions identical so changes in the numbers reflect changes in the answers. Compare today with 30, 90 and 180 days ago.
Tracking template
One row per run. Ten rows per question per engine.
| Date | Engine | Question | Run | Named (in order) | You? | Sources cited |
|---|---|---|---|---|---|---|
| 2026-10-05 | ChatGPT | Best facelift surgeon in Houston? | 1 | Practice A, Practice B, Practice C | No | realself.com, healthgrades.com |
| 2026-10-05 | ChatGPT | Best facelift surgeon in Houston? | 2 | Practice B, You, Practice D | Yes | realself.com, your site |
From the finished sheet, persistence for a question is the count of “Yes” rows divided by its runs. In the example, one of two runs named you: 50%, on a sample too small to act on.
Common mistakes
- Checking from your own login. Your history personalises the answer.
- Trusting one screenshot. In DeepContour’s Houston study, one recommendation in six appeared in only one of three runs.
- Rewording questions between months. A new wording is a new question, and the comparison breaks.
- Reporting a rank. Positions in AI answers rarely survive a second run; shares of runs do.
- Merging engines. ChatGPT and Perplexity can disagree completely. Report them separately.
- Counting mentions without checking facts. Being named with the wrong phone number or a missing office still loses patients.
DeepContour runs this method across engines, markets and procedures and reports it as six measurements. See the methodology or request a private snapshot.
Frequently asked questions
- Does ChatGPT give the same recommendations every time?
- No. In research by SparkToro and Gumshoe covering 2,961 runs of 12 prompts across ChatGPT, Claude and Google’s AI, there was less than a 1-in-100 chance of the same list of recommendations appearing twice. Separate technical work found that a language model sampled 1,000 times on the same prompt with randomness set to zero still produced 80 different completions.
- How many times should I ask each question?
- Ten runs per question per engine is a practical standard. It is enough to separate a practice that is recommended reliably from one that appeared once by chance. Close calls need more runs.
- Why not just check my ranking in ChatGPT?
- Because a position in an AI answer rarely survives a second run. The share of runs in which you are named is stable enough to track; your place in one list is not.
- Can I see which sources ChatGPT used?
- When ChatGPT searches the web it usually shows citations with the answer, and those are worth logging on every run. Engines do not always disclose every source they draw on, so treat the citations as a partial map.
- How much work is manual tracking?
- Fifteen questions asked ten times on six engines is 900 answers a month, each read and logged by hand. That volume is why practices automate it or use a service such as DeepContour.
Sources
- Fishkin & O’Donnell, SparkToro (2026). AIs are highly inconsistent when recommending brands or products.
- He & Thinking Machines Lab (2025). Defeating nondeterminism in LLM inference.
- DeepContour (2026). Which Houston plastic surgeons does AI recommend? A study of 117 practices.
- DeepContour methodology.
See what AI says about your practice.
A private snapshot of Answer Presence, Share of Answer and Recommendation Persistence for your procedures and market, read by a partner.
Request My Private Snapshot