Filling the information gap: a validated comparison of ChatGPT, Gemini, and Claude for penile cancer patient education
Original Article

Filling the information gap: a validated comparison of ChatGPT, Gemini, and Claude for penile cancer patient education

Jeffery Yang1,2#, Ymer Bushati1,3#, Sarah Lorger1, Jeremy Saad1,4, Prem Rathore1,4

1Department of Urology, Campbelltown Hospital, Campbelltown, NSW, Australia; 2School of Medicine, Western Sydney University, Sydney, NSW, Australia; 3Faculty of Medicine and Health, University of Sydney, Sydney, NSW, Australia; 4Department of Urology, Liverpool Hospital, Sydney, NSW, Australia

Contributions: (I) Conception and design: All authors; (II) Administrative support: All authors; (III) Provision of study materials or patients: All authors; (IV) Collection and assembly of data: All authors; (V) Data analysis and interpretation: All authors; (VI) Manuscript writing: All authors; (VII) Final approval of manuscript: All authors.

#These authors contributed equally to this work as co-first authors.

Correspondence to: Dr. Jeffery Yang, MD. Department of Urology, Campbelltown Hospital, Therry Rd., Campbelltown, NSW 2560, Australia; School of Medicine, Western Sydney University, Sydney, NSW, Australia. Email: jeffrey.yang1@health.nsw.gov.au.

Background: ChatGPT, Gemini, and Claude are three large language models (LLMs) increasingly consulted by patients researching a new cancer diagnosis. Penile cancer is rare, stigmatised, and dominated by decisions around organ preservation and sexual function, which makes the quality of patient-facing information particularly consequential. The aim of this study was to assess the quality of responses generated by these LLMs to patient-oriented questions on penile cancer.

Methods: Fourteen questions reflecting common patient information needs were adapted from established educational resources and refined by senior urology consultant review. Each query was presented to ChatGPT, Gemini, and Claude using an identical patient-centred prompt. Responses (n=42) were anonymised and randomly ordered before being independently reviewed in a consensus meeting by a urology panel using the Quality Analysis of Medical Artificial Intelligence (QAMAI) framework. Comparative analysis used the Friedman test with Bonferroni-corrected pairwise Wilcoxon signed-rank testing.

Results: Mean total QAMAI scores were 22.7 for ChatGPT (median, 23; range, 17 to 25), 22.6 for Claude (median, 23; range, 18 to 24), and 17.0 for Gemini (median, 17.5; range, 14 to 20), corresponding to moderate quality for ChatGPT and Claude and the boundary of poor and moderate quality for Gemini. The difference across models was significant (Friedman χ2=18.78; P<0.001). ChatGPT and Claude each significantly outperformed Gemini (both adjusted P=0.004) but did not differ from one another (adjusted P>0.99). ChatGPT produced the clearest responses and Claude the most complete, while Gemini scored lowest in every domain, most strikingly for surgery and organ preservation (14.0 vs. 23.7 and 23.0). No response provided citations or referred readers to external sources of information.

Conclusions: ChatGPT and Claude delivered broadly equivalent, moderate-to-good-quality information on patient-derived penile cancer questions, and both clearly outperformed Gemini, which frequently produced incomplete answers in the surgical and organ preservation domain. The universal absence of source citation is a clinically relevant limitation. These tools should be regarded as a supplement to, rather than a substitute for, specialist consultation.

Keywords: Penile cancer; large language models (LLMs); patient education; artificial intelligence (AI); organ preservation


Submitted Jun 28, 2026. Accepted for publication Jul 29, 2026. Published online Aug 11, 2026.

doi: 10.21037/tau-2026-0603


Highlight box

Key findings

• In a blinded, Quality Analysis of Medical Artificial Intelligence (QAMAI) based comparison, ChatGPT and Claude produced moderate-to-good-quality answers to patient-derived penile cancer questions and significantly outperformed Gemini, which scored poorly, most markedly on surgery and organ preservation. No model cited any source.

What is known and what is new?

• Large language models are increasingly used by patients seeking cancer information, and prior work in other urological cancers describes fluent but variable output; no previous study had evaluated or compared ChatGPT, Gemini, and Claude for penile cancer specifically.

• This is the first validated comparison of these three models in penile cancer. It identifies a clinically important, domain specific weakness in the surgery and organ preservation questions central to this disease, and the universal absence of source citation.

What is the implication, and what should change now?

• Clinicians counselling newly diagnosed men should anticipate prior large language model use, recognize that quality varies substantially between platforms, and actively direct patients towards reliable, sourced information. These tools supplement, but do not replace, specialist consultation.


Introduction

Background

Penile cancer is an uncommon malignancy in high-income settings, with a global age-standardised incidence of approximately 0.8 per 100,000 men and markedly higher rates across parts of South America, sub-Saharan Africa and South Asia (1,2). More than 95% of tumours are squamous cell carcinomas, and incidence climbs with age to peak in the sixth and seventh decades (1,3). Between one third and one half of cases are attributable to high-risk human papillomavirus (HPV), whereas HPV independent disease develops along chronic inflammatory pathways associated with phimosis, lichen sclerosus and retained smegma (1,3). Despite its rarity, the disease carries considerable morbidity: inguinal lymph node involvement is the single most important determinant of survival, and treatment of the primary lesion entails surgery to the genitalia (3).

These features give penile cancer a distinctive position among urological malignancies. The overriding therapeutic aim, which is to achieve complete oncological clearance while preserving as much of the organ as is oncologically safe, must be balanced against far-reaching consequences for urinary function, sexual function, body image and masculine identity (3,4). Substantial unmet psychosexual and supportive care needs are documented in this population, and the stigma surrounding genital disease often delays presentation and discourages candid discussion (4). Information delivered at the point of diagnosis must therefore be not only accurate but clear and psychosocially attuned.

Rationale and knowledge gap

Patients increasingly turn to the internet for such information, and artificial intelligence (AI), in particular large language models (LLMs) such as ChatGPT (OpenAI), Gemini (Google), and Claude (Anthropic), has rapidly become one of the first resources they consult (5-8). These models generate fluent, conversational answers to clinical questions (9). Because they operate by predicting the most probable next token rather than sourcing from verified evidence, their output can be incomplete, outdated, or confidently incorrect (10,11). The stakes are heightened when an LLM is the first source a patient reaches before any clinical encounter, shaping understanding and expectation before the diagnosis has even been discussed with a clinician (11,12).

Recent literature has begun to appraise LLM use across urological oncology, consistently describing fluent but variable output with weak source attribution and limited actionability (13-20). Comparisons between models have repeatedly exposed meaningful performance differences between platforms that appear superficially similar (14,15,17). To our knowledge, however, no study has yet evaluated LLM-generated patient information specifically in penile cancer, nor compared ChatGPT, Gemini, and Claude in this disease, nor examined the surgical and organ preservation questions that dominate counselling of this cohort.

Objective

The aim of the present study was therefore to evaluate and compare the quality of responses generated by ChatGPT, Gemini, and Claude to patient-derived frequently asked questions on penile cancer, with particular attention to the surgical and organ preservation domain.


Methods

Study design

This was a blinded comparison of how three consumers (LLMs, ChatGPT, Gemini, and Claude) answer the questions a man might ask after being told he has penile cancer. Each model was given the same fixed set of standardised queries, and the resulting answers were graded by a panel of urology clinicians across two consensus sessions using the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool (21), the validated instrument for appraising AI-generated health information.

AI platform selection

The three platforms were chosen because they are the most widely used systems a member of the public can reach without payment, so the free tiers were the ones tested, these being what patients are most likely to open. Each platform was accessed through its free public web interface rather than an application programming interface, and every query was entered under default consumer settings, with no extended reasoning or chain-of-thought mode manually enabled. For each platform, a single account was created specifically for this study, with no prior interaction history, so that account-level personalisation would not affect the output. The models served on the query date were ChatGPT (GPT-5.3; knowledge cutoff August 2025), Gemini (2.5 Flash; knowledge cutoff January 2025), and Claude (Sonnet 4.6; training data cutoff January 2026, reliable knowledge cutoff August 2025), each queried in July 2025.

Deriving the question set

The questions were assembled to mirror the information a man actually looks for after a penile cancer diagnosis. Common queries were gathered from established patient education resources (European Association of Urology patient information, National Health Service resources, the American Cancer Society, Cancer Research UK and National Comprehensive Cancer Network patient guidelines) and grouped by recurring theme, and a small number of further questions were added where clinically important areas, in particular organ preservation, inguinal node surgery and recovery after treatment, tend to be thinly covered in general material. A senior consultant urologist then reviewed and refined the list before any data were collected. The final instrument held 14 questions spanning five domains: basics and risk (n=3), symptoms and diagnosis (n=2), treatment options (n=4), surgery and organ preservation (n=3), and prognosis and survivorship (n=2). Table 1 sets out the complete set.

Table 1

Full list of final questions grouped by thematic domain

Thematic domain Patient-derived questions
Overview and risk Q1. What is penile cancer and what causes it?
Q2. What are the risk factors for penile cancer?
Q3. How quickly does penile cancer grow and spread, and how urgently do I need treatment?
Symptoms and diagnosis Q4. What are the symptoms of penile cancer, and how is it diagnosed?
Q5. Has my cancer spread, and where does penile cancer usually spread to?
Treatment options Q6. How is penile cancer treated, and what treatment options will I have?
Q10. Can I be cured if I don’t have surgery?
Q11. What nonsurgical options are there?
Q12. When is chemotherapy used, and can I have multiple or combined treatments?
Surgery and organ preservation Q7. What are my surgical options, and can my penis be preserved?
Q8. What is lymph node surgery? Will I need it, and what are its risks?
Q9. How will my life change after surgery, and what could go wrong?
Prognosis and survivorship Q13. How successful is penile cancer treatment, and what is my prognosis?
Q14. Can penile cancer come back, and what follow-up will I need?

Question numbers reflect the order in which the questions were presented to the models; the thematic grouping shown is for analysis only and does not follow the numerical sequence.

Response capture

Every question was put to the models behind the same opening line, written in the patient’s own voice [for instance, “I am a patient newly diagnosed with penile cancer. Please tell me (insert question)”], so that the framing held constant across platforms. The queries were run on 1 day, each in a fresh incognito conversation to rule out carryover from earlier answers, with the prompt and question typed exactly as written and only the first reply kept; no clarifications, edits, or follow-up prompts were allowed. The 42 replies were collated into one scoring document and stripped of any wording that revealed which platform had produced them, leaving the clinical content untouched. Within each question, the three answers were shuffled and labelled A, B, or C, and the key tying labels back to models stayed with the lead author until every score had been entered.

Response quality appraisal

Grading used the six QAMAI domains, namely accuracy, clarity, relevance, completeness, provision of sources, and usefulness, each marked from 1 (strongly disagree) to 5 (strongly agree) to give a total of 6 to 30, in which higher is better. The validation study’s five bands were applied: very poor [6 to 11], poor [12 to 17], moderate [18 to 23], good [24 to 29], and excellent [30] (21). Because none of the responses provided any source or citation, as reported below, the provision of sources item scored 0 for every response and was excluded from all comparative analyses; the reported total scores therefore carry an effective maximum of 25 rather than 30. The QAMAI validation bands were derived for the complete six-item 30-point scale, so with this item structurally absent, the bands are applied only as an approximate interpretive guide, a limitation revisited below. Following the approach of Pompili et al. (14), a panel of one consultant urologist and three registrars at differing stages of training scored every answer blinded to its source. To keep seniority from steering the result, each rater first committed an independent mark across all six domains. These independent scores were retained and not subsequently altered. The panel then discussed each response, and where the independent scores diverged, the discussion continued until a single consensus score was agreed for the record, with the consultant urologist’s score given no greater weight than those of the registrars. Only the agreed consensus score was used as the final value in the analysis, although individual scores were retained for inter-rater reliability assessment. Scoring was completed across two sessions of approximately 2 hours each, identifying clinically important points, misconceptions, how well each answer met the needs of a newly diagnosed man, whether chatbots gave sensible safety advice, and whether they acknowledged the psychosocial side of care.

Organ preservation appraisal

A focused secondary analysis was performed for the surgery and organ preservation domain because of its clinical significance and the anticipated challenges LLMs may face in communicating the balance between organ preservation and oncological control. For the penile preservation question (Q7), responses were evaluated against four predefined criteria: whether the response acknowledged that penile preservation is often feasible in early-stage or superficial disease, whether it described the available organ-sparing surgical approaches, whether it discussed partial and total penectomy, including their functional implications, and whether it explicitly recognised that oncological control should take precedence over preservation when necessary.

Statistical analysis

QAMAI performance was summarised descriptively, continuous measures as means and median with interquartile range (IQR) and the quality grades as counts and percentages; domain means, grade distributions, and mean totals by theme were derived for each model. To compare the three platforms, the Friedman test for related samples treated each question as a matched observation, and where it flagged an overall difference, Bonferroni-adjusted Wilcoxon signed-rank tests were applied pairwise. Two-sided P below 0.05 denoted significance. For each pairwise comparison, a paired effect size was calculated as Cohen’s dz, the mean of the paired differences divided by their standard deviation, with the Hedges’ g small-sample correction, and interpreted using conventional thresholds of 0.2, 0.5, and 0.8 for small, medium, and large effects. Inter-rater reliability among the four raters’ independent scores was quantified with the two-way intraclass correlation coefficient (absolute agreement) for the total QAMAI scores, reported for a single rater and for the four-rater average, and with Fleiss’ kappa for the item-level ratings. Performance across the five domains was a prespecified secondary question. All computations were carried out in RStudio (v2024.12.1+563; Posit Software, PBC, 2025).


Results

Overall performance

All three platforms answered each of the 14 questions (Table S1). On the QAMAI total, ChatGPT and Claude were near level at 22.7 and 22.6, respectively (ChatGPT: median, 23; range, 17 to 25; Claude: median, 23; range, 18 to 24), while Gemini trailed at 17.0 (median, 17.5; range, 14 to 20); in grade terms, the first two were moderate on average and Gemini sat on the poor to moderate border. The Friedman test confirmed a marked separation across models (χ2=18.78; P<0.001), and pairwise testing placed both ChatGPT and Claude above Gemini (each adjusted P=0.004) while leaving the two leaders statistically indistinguishable (adjusted P>0.99). The magnitude of the difference was large for both comparisons against Gemini (ChatGPT vs. Gemini, Cohen’s dz =1.62; Claude vs. Gemini, dz =1.94) and negligible between ChatGPT and Claude (dz =0.05). The grade counts told the same story: ChatGPT returned 7 good, 6 moderate, and 1 poor answer; Claude 5 good and 9 moderate; and Gemini 7 moderate and 7 poor (Table 2), with nothing reaching excellent and nothing falling to very poor.

Table 2

QAMAI quality grades by model

QAMAI grade Score range ChatGPT, n (%) Gemini, n (%) Claude, n (%) Overall, n (%)
Very poor 6 to 11 0 (0.0) 0 (0.0) 0 (0.0) 0 (0.0)
Poor 12 to 17 1 (7.1) 7 (50.0) 0 (0.0) 8 (19.0)
Moderate 18 to 23 6 (42.9) 7 (50.0) 9 (64.3) 22 (52.4)
Good 24 to 29 7 (50.0) 0 (0.0) 5 (35.7) 12 (28.6)
Excellent 30 0 (0.0) 0 (0.0) 0 (0.0) 0 (0.0)

Because no response provided a source, the provision of sources item scored 0 for all responses; total scores therefore carry an effective maximum of 25, and the grade bands shown derive from the original 30-point QAMAI scale. N, number of responses; QAMAI, Quality Analysis of Medical Artificial Intelligence.

Comparative testing

Not one model cited a reference or pointed the reader to an authoritative external source, so every one of the 42 answers scored the floor on provision of sources, and that item was set aside for the comparative tests. The remaining domain means appear in Table 3. ChatGPT was clearest of the three (mean, 4.86), ahead of both Gemini (mean, 4.07) and Claude (mean, 4.07) by a significant margin (Friedman P<0.001; Bonferroni adjusted P=0.003 and 0.007), whereas Gemini and Claude did not differ on clarity. Completeness ran the other way, highest for Claude (4.71) and lowest for Gemini (2.71), and here every pairwise contrast was significant (Friedman P<0.001). Gemini also lagged both competitors on accuracy, relevance and usefulness. The upshot is that ChatGPT and Claude behaved alike across most domains, parting company only on clarity, which favoured ChatGPT, and completeness, which favoured Claude.

Table 3

Mean QAMAI per item scores (provision of sources item excluded; scored 0 across all responses)

QAMAI item ChatGPT Gemini Claude Friedman P
Accuracy 4.43 3.71 4.64 0.001
Clarity 4.86 4.07 4.07 <0.001
Relevance 4.50 3.29 4.43 <0.001
Completeness 4.43 2.71 4.71 <0.001
Usefulness 4.50 3.21 4.71 <0.001

Two-sided P values from the Friedman test across the three models. QAMAI, Quality Analysis of Medical Artificial Intelligence.

Thematic subanalysis

Table 4 breaks the totals down by theme. ChatGPT and Claude moved together across the domains, ChatGPT peaking in symptoms and diagnosis (25.00) and Claude in prognosis and survivorship (24.00), while Gemini sat lowest in every domain without exception. The models pulled furthest apart in surgery and organ preservation: ChatGPT (23.67) and Claude (23.00) held the moderate to good range, but Gemini dropped to 14.00, squarely poor. That gap came mostly from weak completeness and usefulness on the questions dealing with penile preservation, inguinal node surgery, and life after an operation.

Table 4

Mean total QAMAI score by thematic domain

Thematic domain ChatGPT Gemini Claude Overall
Basics and risk 24.00 18.67 23.00 21.89
Symptoms and diagnosis 25.00 18.50 20.00 21.17
Treatment options 20.75 17.25 22.50 20.17
Surgery and organ preservation 23.67 14.00 23.00 20.22
Prognosis and survivorship 21.00 17.00 24.00 20.67

QAMAI, Quality Analysis of Medical Artificial Intelligence.

Surgery and organ preservation questions

On the focused penile preservation question (Q7), all three models listed the organ-sparing options (circumcision, wide local excision, Mohs micrographic surgery, laser ablation, glans resurfacing, and glansectomy) and set out partial and total penectomy with perineal urethrostomy. What separated them was the framing of the governing principle, that clearing the cancer must come before sparing the organ. ChatGPT (25, good quality) and Claude (24, good quality) both conveyed it, ChatGPT in plain terms and Claude through the balance between clear margins and local recurrence, with a note on later reconstruction. Gemini (14, poor quality) ran through the same list of operations but scored far lower on completeness and usefulness, leaving the preservation versus control balance implicit and omitting the guidance a newly diagnosed man would need to weigh the choice (Table 5). Q9, on life after surgery, followed suit: ChatGPT (21) and Claude (22) took in urinary and sexual changes alongside the psychological and relationship toll, whereas Gemini (14) gave a thinner account.

Table 5

Focused analysis of the penile preservation question (Q7) against four prespecified clinical elements

Model QAMAI (Q7) Stated preservation often achievable Enumerated organ-sparing options Described partial/total penectomy and functional impact Framed oncological control as priority
ChatGPT 25 Yes Yes Yes Yes (explicit)
Gemini 14 Partial Yes Yes Partial (implied, not explicit)
Claude 24 Yes Yes Yes Yes (via margin/recurrence tradeoff)

QAMAI, Quality Analysis of Medical Artificial Intelligence.


Discussion

Key findings

To our knowledge, this is the first study to evaluate AI-generated patient information in penile cancer and the first to directly compare ChatGPT, Gemini, and Claude in this disease setting. Generative AI has moved into routine clinical and consumer use at a pace few anticipated (12), and patients increasingly access these systems before engaging with healthcare professionals. The principal finding of this study was a clear separation between models. ChatGPT and Claude produced moderate-to-good-quality responses and did not differ significantly from one another, whereas Gemini scored approximately five to six QAMAI points lower overall and generated poor-quality responses for half of the questions assessed. Given the rarity of penile cancer (2) and the fact that existing publicly available patient information is often of only moderate quality and written at a reading level that exceeds that of the average reader (22), this variation has important practical implications. Specifically, the platform through which an individual first seeks information may substantially influence whether their initial understanding of the disease is accurate and comprehensive or incomplete and potentially misleading.

Strengths and limitations

These findings should be interpreted in light of several limitations. First, each response represents a single interaction captured on 1 day from systems that are not deterministic and subject to ongoing retraining. Consequently, the reported scores reflect performance at a specific point in time and will not be reproducible in future model versions. Second, only free tier versions were evaluated, which may underestimate the capabilities of paid or reasoning-enhanced models. Third, although platform-identifying wording was removed from every response before assessment, no formal blinding assessment was undertaken, for example, asking raters to identify the source model, so the effectiveness of blinding was not quantified; stylistic characteristics may have remained recognisable and could have influenced scoring. The four raters scored every response independently before discussion, and agreement on the total QAMAI score was good to excellent (two way intraclass correlation coefficient 0.82 for a single rater, 95% confidence interval: 0.64 to 0.91; and 0.95 for the four rater average, 95% confidence interval: 0.88 to 0.97); item level agreement was fair when the ordinal ratings were treated as nominal categories (Fleiss’ kappa =0.32), reflecting their skewed distribution, although pairwise agreement between raters was exact for 54% of item ratings and within one point for 94%, and any residual divergence was settled through the consensus process, which together with the blinded independent scoring limited individual and seniority bias. Fourth, the fourteen questions, while derived from established patient information resources, cannot encompass the full spectrum of patient concerns. In addition, the rarity of penile cancer means that the quantity of disease-specific training data available to any model is likely to be limited. Readability was likewise not formally assessed in the present study, and where the discussion refers to readability, it summarises prior literature rather than measurements made here. Finally, the study was restricted to English language outputs and to the domains assessed by QAMAI. Future work should formally assess readability and actionability using dedicated instruments, and examine multilingual performance and longitudinal changes in model behaviour as these systems continue to evolve.

Comparison with similar research

The marked divergence between contemporary LLMs is consistent with findings from the wider literature. Comparative studies in prostate cancer have demonstrated differences in comprehensiveness, accuracy, and readability between models rather than convergence in performance (15). Similarly, a multiplatform evaluation in bladder cancer reported stylistically distinct outputs despite an absence of hallucinations (16), while a comparison of AI-generated urology patient leaflets identified meaningful quality differences beneath superficially similar content (14). The ability to detect such differences in the present study may partly reflect the assessment tool employed. Unlike readability and consumer information indices used in earlier work, QAMAI was specifically developed and validated to evaluate the accuracy, completeness, and usefulness of AI-generated health information (21). These dimensions revealed meaningful distinctions between models. Notably, differences were also evident between the two highest-performing systems. ChatGPT achieved the highest scores for clarity, whereas Claude scored highest for completeness, suggesting a balance between accessibility and comprehensiveness. No individual platform consistently maximised both attributes.

Explanations of findings

Importantly, Gemini performance was not uniformly weak across topics but was concentrated within questions relating to surgery and organ preservation. This domain is particularly important because treatment decisions frequently require patients to weigh oncological outcomes against functional preservation. The lower scores were primarily attributable to deficiencies in completeness and usefulness rather than outright inaccuracies. Contemporary penile cancer management relies on a spectrum of organ-preserving approaches for appropriately selected early-stage disease, with partial or total penectomy and inguinal nodal staging reserved for more advanced or higher-risk tumours (3). Responses that list available procedures without adequately explaining the rationale for organ preservation, the indications for more radical surgery, or the implications for urinary and sexual function may provide technically correct information while failing to support informed understanding. In effect, such responses describe the available operations without fully addressing the decisions that surround them. Given the well-documented psychosexual burden associated with penile cancer (4), omissions in these areas may contribute to uncertainty and unmet informational needs.

Implications and actions needed

The clinical significance of these findings depends partly on how patients use these technologies. Many individuals seek information about symptoms and diagnoses online before consulting a healthcare professional (6), and this behaviour may be particularly relevant in penile cancer, where embarrassment and stigma are recognised contributors to delayed presentation (7). The anonymity and immediate accessibility of conversational AI may therefore increase its appeal for patients reluctant to seek medical advice. However, accessibility should not be equated with completeness or clinical reliability. Studies evaluating AI-generated information in prostate cancer radiotherapy (17), simulated bladder cancer consultations (18), and common questions across urological malignancies (13,19) have consistently found that while factual content is often broadly accurate, readability frequently exceeds recommended levels and actionable guidance remains limited. A recent systematic review of LLM-based healthcare chatbots reached similar conclusions across multiple specialties (23). Correspondingly, patients continue to report substantially greater trust in urologists than in LLMs for treatment decisions and reassurance (20). In the present study, the strongest performing models provided useful information, but none consistently delivered the contextualised guidance required for patient counselling.

One finding was universal across all three platforms. None of the 42 responses cited information sources or directed users towards additional resources. From a patient safety perspective, this is important because unsourced clinical information provides no mechanism for independent verification and no pathway to authoritative educational materials following a new diagnosis (10,11). In addition, no response carried an explicit statement that it did not constitute medical advice, although several did appropriately defer patient-specific determinations, such as whether the cancer had spread or the individual prognosis, to the treating clinician; this safety netting was applied inconsistently across questions and models. Previous QAMAI-based evaluations have reported highly variable performance in source attribution across specialties, ranging from near absence to routine inclusion (24,25), suggesting that citation behaviour may be influenced by both model design and prompting strategy. The patient-oriented prompts used in the present study may therefore have reduced the likelihood of source provision. Nevertheless, reliable referencing and signposting are widely regarded as essential prerequisites for the safe deployment of medical chatbots (9). Their complete absence in this study highlights an important limitation of current systems, consistent with direct evaluations in which ChatGPT has been shown to provide inaccurate urological healthcare advice (26). This concern is further reinforced by evidence that LLMs may generate inaccurate or fabricated references when prompted for citations (27) and may elaborate upon incorrect information contained within user prompts rather than challenge it (28). Clinicians should therefore assume that information obtained from these platforms is unsourced and should actively direct patients towards evidence-based resources, including professional society guidelines and survivorship support materials.


Conclusions

ChatGPT and Claude delivered broadly equivalent, moderate-to-good-quality information on patient-derived penile cancer questions, and both significantly outperformed Gemini, which frequently produced incomplete answers, most notably in the surgical and organ preservation domain so central to this disease. ChatGPT was the clearest and Claude the most complete, while no model provided a single source or onward signpost. These tools are a useful supplement to, but not a substitute for, specialist consultation, and clinicians counselling newly diagnosed men should anticipate prior LLM use, recognise that quality varies considerably between platforms, and actively direct patients towards reliable information and support.


Acknowledgments

During the preparation of this work, the authors used generative AI only to generate responses for the dataset. No AI input was used for language editing. All output was subsequently reviewed and edited by the authors, who take full responsibility for the content of this publication.


Footnote

Data Sharing Statement: Available at https://tau.amegroups.com/article/view/10.21037/tau-2026-0603/dss

Peer Review File: Available at https://tau.amegroups.com/article/view/10.21037/tau-2026-0603/prf

Funding: None.

Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://tau.amegroups.com/article/view/10.21037/tau-2026-0603/coif). The authors have no conflicts of interest to declare.

Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.

Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.


References

  1. Mannam G, Miller JW, Johnson JS, et al. HPV and Penile Cancer: Epidemiology, Risk Factors, and Clinical Insights. Pathogens 2024;13:809. [Crossref] [PubMed]
  2. Huang J, Chan SC, Pang WS, et al. Incidence, risk factors, and temporal trends of penile cancer: a global population-based study. BJU Int 2024;133:314-23. [Crossref] [PubMed]
  3. Brouwer OR, Albersen M, Parnham A, et al. European Association of Urology-American Society of Clinical Oncology Collaborative Guideline on Penile Cancer: 2023 Update. Eur Urol 2023;83:548-60. [Crossref] [PubMed]
  4. Maddineni SB, Lau MM, Sangar VK. Identifying the needs of penile cancer sufferers: a systematic review of the quality of life, psychosexual and psychosocial literature in penile cancer. BMC Urol 2009;9:8. [Crossref] [PubMed]
  5. Lim MSC, Molenaar A, Brennan L, et al. Young Adults' Use of Different Social Media Platforms for Health Information: Insights From Web-Based Conversations. J Med Internet Res 2022;24:e23656. [Crossref] [PubMed]
  6. Maon SN, Hassan NM, Seman SAA. Online health information seeking behavior pattern. Adv Sci Lett 2017;23:10582-5.
  7. Cieślikowski WA, Kasperczak M, Milecki T, et al. Reasons behind the Delayed Diagnosis of Testicular Cancer: A Retrospective Analysis. Int J Environ Res Public Health 2023;20:4752. [Crossref] [PubMed]
  8. Van Booven D, Chen CB. Chapter 6 - ChatGPT and healthcare—current and future prospects. In: Arora H, editor. Artificial intelligence in urologic malignancies. London: Academic Press; 2025:173-93.
  9. Chow JCL, Wong V, Li K. Generative pre-trained transformer-empowered healthcare conversations: current trends, challenges, and future directions in large language model-enabled medical chatbots. BioMedInformatics 2024;4:837-52.
  10. Bélisle-Pipon JC. Why we need to be careful with LLMs in medicine. Front Med (Lausanne) 2024;11:1495582. [Crossref] [PubMed]
  11. Sauerbrei A, Kerasidou A, Lucivero F, et al. The impact of artificial intelligence on the person-centred, doctor-patient relationship: some problems and solutions. BMC Med Inform Decis Mak 2023;23:73. [Crossref] [PubMed]
  12. Teo ZL, Thirunavukarasu AJ, Elangovan K, et al. Generative artificial intelligence in medicine. Nat Med 2025;31:3270-82. [Crossref] [PubMed]
  13. Musheyev D, Pan A, Loeb S, et al. How Well Do Artificial Intelligence Chatbots Respond to the Top Search Queries About Urological Malignancies?. Eur Urol 2024;85:13-6. [Crossref] [PubMed]
  14. Pompili D, Richa Y, Collins P, et al. Using artificial intelligence to generate medical literature for urology patients: a comparison of three different large language models. World J Urol 2024;42:455. [Crossref] [PubMed]
  15. Alasker A, Alsalamah S, Alshathri N, et al. Performance of large language models (LLMs) in providing prostate cancer information. BMC Urol 2024;24:177. [Crossref] [PubMed]
  16. Patel K, Radcliffe R. Evaluating the Readability and Quality of Bladder Cancer Information from AI Chatbots: A Comparative Study Between ChatGPT, Google Gemini, Grok, Claude and DeepSeek. J Clin Med 2025;14:7804. [Crossref] [PubMed]
  17. Trapp C, Schmidt-Hegemann N, Keilholz M, et al. Patient- and clinician-based evaluation of large language models for patient education in prostate cancer radiotherapy. Strahlenther Onkol 2025;201:333-42. [Crossref] [PubMed]
  18. Guo AA, Razi B, Kim P, et al. The role of artificial intelligence in patient education: a bladder cancer consultation with ChatGPT. Soc Int Urol J 2024;5:214-24.
  19. Thia I, Saluja M. ChatGPT: Is This Patient Education Tool for Urological Malignancies Readable for the General Population?. Res Rep Urol 2024;16:31-7. [Crossref] [PubMed]
  20. Carl N, Nguyen L, Haggenmüller S, et al. Comparing Patient's Confidence in Clinical Capabilities in Urology: Large Language Models Versus Urologists. Eur Urol Open Sci 2024;70:91-8. [Crossref] [PubMed]
  21. Vaira LA, Lechien JR, Abbate V, et al. Validation of the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool: a new tool to assess the quality of health information provided by AI platforms. Eur Arch Otorhinolaryngol 2024;281:6123-31. [Crossref] [PubMed]
  22. Ma J, Wang M, Liu Q, et al. Quality and readability of online penile-cancer information: a Cross-Sectional Web Study. Int J Impot Res 2026; [Epub ahead of print]. [Crossref] [PubMed]
  23. Huo B, Boyle A, Marfo N, et al. Large Language Models for Chatbot Health Advice Studies: A Systematic Review. JAMA Netw Open 2025;8:e2457879. [Crossref] [PubMed]
  24. Furukawa E, Okuhara T, Okada H, et al. Reliability of AI-generated responses on frequently-posed questions by patients with chronic kidney disease. Health Informatics J 2025;31:14604582251381996. [Crossref] [PubMed]
  25. Bozhöyük MS, Yücel L. A comprehensive evaluation of artificial intelligence-provided information on common ENT surgical procedures using the QAMAI tool. J Laryngol Otol 2025;139:1109-14. [Crossref] [PubMed]
  26. Whiles BB, Bird VG, Canales BK, et al. Caution! AI Bot Has Entered the Patient Chat: ChatGPT Has Limitations in Providing Accurate Urologic Healthcare Advice. Urology 2023;180:278-84. [Crossref] [PubMed]
  27. Chelli M, Descamps J, Lavoué V, et al. Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis. J Med Internet Res 2024;26:e53164. [Crossref] [PubMed]
  28. Omar M, Sorin V, Collins JD, et al. Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support. Commun Med (Lond) 2025;5:330. [Crossref] [PubMed]
Cite this article as: Yang J, Bushati Y, Lorger S, Saad J, Rathore P. Filling the information gap: a validated comparison of ChatGPT, Gemini, and Claude for penile cancer patient education. Transl Androl Urol 2026;15(9):340. doi: 10.21037/tau-2026-0603

Download Citation