Filling the information gap: a validated comparison of ChatGPT, Gemini, and Claude for penile cancer patient education
Highlight box
Key findings
• In a blinded, Quality Analysis of Medical Artificial Intelligence (QAMAI) based comparison, ChatGPT and Claude produced moderate-to-good-quality answers to patient-derived penile cancer questions and significantly outperformed Gemini, which scored poorly, most markedly on surgery and organ preservation. No model cited any source.
What is known and what is new?
• Large language models are increasingly used by patients seeking cancer information, and prior work in other urological cancers describes fluent but variable output; no previous study had evaluated or compared ChatGPT, Gemini, and Claude for penile cancer specifically.
• This is the first validated comparison of these three models in penile cancer. It identifies a clinically important, domain specific weakness in the surgery and organ preservation questions central to this disease, and the universal absence of source citation.
What is the implication, and what should change now?
• Clinicians counselling newly diagnosed men should anticipate prior large language model use, recognize that quality varies substantially between platforms, and actively direct patients towards reliable, sourced information. These tools supplement, but do not replace, specialist consultation.
Introduction
Background
Penile cancer is an uncommon malignancy in high-income settings, with a global age-standardised incidence of approximately 0.8 per 100,000 men and markedly higher rates across parts of South America, sub-Saharan Africa and South Asia (1,2). More than 95% of tumours are squamous cell carcinomas, and incidence climbs with age to peak in the sixth and seventh decades (1,3). Between one third and one half of cases are attributable to high-risk human papillomavirus (HPV), whereas HPV independent disease develops along chronic inflammatory pathways associated with phimosis, lichen sclerosus and retained smegma (1,3). Despite its rarity, the disease carries considerable morbidity: inguinal lymph node involvement is the single most important determinant of survival, and treatment of the primary lesion entails surgery to the genitalia (3).
These features give penile cancer a distinctive position among urological malignancies. The overriding therapeutic aim, which is to achieve complete oncological clearance while preserving as much of the organ as is oncologically safe, must be balanced against far-reaching consequences for urinary function, sexual function, body image and masculine identity (3,4). Substantial unmet psychosexual and supportive care needs are documented in this population, and the stigma surrounding genital disease often delays presentation and discourages candid discussion (4). Information delivered at the point of diagnosis must therefore be not only accurate but clear and psychosocially attuned.
Rationale and knowledge gap
Patients increasingly turn to the internet for such information, and artificial intelligence (AI), in particular large language models (LLMs) such as ChatGPT (OpenAI), Gemini (Google), and Claude (Anthropic), has rapidly become one of the first resources they consult (5-8). These models generate fluent, conversational answers to clinical questions (9). Because they operate by predicting the most probable next token rather than sourcing from verified evidence, their output can be incomplete, outdated, or confidently incorrect (10,11). The stakes are heightened when an LLM is the first source a patient reaches before any clinical encounter, shaping understanding and expectation before the diagnosis has even been discussed with a clinician (11,12).
Recent literature has begun to appraise LLM use across urological oncology, consistently describing fluent but variable output with weak source attribution and limited actionability (13-20). Comparisons between models have repeatedly exposed meaningful performance differences between platforms that appear superficially similar (14,15,17). To our knowledge, however, no study has yet evaluated LLM-generated patient information specifically in penile cancer, nor compared ChatGPT, Gemini, and Claude in this disease, nor examined the surgical and organ preservation questions that dominate counselling of this cohort.
Objective
The aim of the present study was therefore to evaluate and compare the quality of responses generated by ChatGPT, Gemini, and Claude to patient-derived frequently asked questions on penile cancer, with particular attention to the surgical and organ preservation domain.
Methods
Study design
This was a blinded comparison of how three consumers (LLMs, ChatGPT, Gemini, and Claude) answer the questions a man might ask after being told he has penile cancer. Each model was given the same fixed set of standardised queries, and the resulting answers were graded by a panel of urology clinicians across two consensus sessions using the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool (21), the validated instrument for appraising AI-generated health information.
AI platform selection
The three platforms were chosen because they are the most widely used systems a member of the public can reach without payment, so the free tiers were the ones tested, these being what patients are most likely to open. Each platform was accessed through its free public web interface rather than an application programming interface, and every query was entered under default consumer settings, with no extended reasoning or chain-of-thought mode manually enabled. For each platform, a single account was created specifically for this study, with no prior interaction history, so that account-level personalisation would not affect the output. The models served on the query date were ChatGPT (GPT-5.3; knowledge cutoff August 2025), Gemini (2.5 Flash; knowledge cutoff January 2025), and Claude (Sonnet 4.6; training data cutoff January 2026, reliable knowledge cutoff August 2025), each queried in July 2025.
Deriving the question set
The questions were assembled to mirror the information a man actually looks for after a penile cancer diagnosis. Common queries were gathered from established patient education resources (European Association of Urology patient information, National Health Service resources, the American Cancer Society, Cancer Research UK and National Comprehensive Cancer Network patient guidelines) and grouped by recurring theme, and a small number of further questions were added where clinically important areas, in particular organ preservation, inguinal node surgery and recovery after treatment, tend to be thinly covered in general material. A senior consultant urologist then reviewed and refined the list before any data were collected. The final instrument held 14 questions spanning five domains: basics and risk (n=3), symptoms and diagnosis (n=2), treatment options (n=4), surgery and organ preservation (n=3), and prognosis and survivorship (n=2). Table 1 sets out the complete set.
Table 1
| Thematic domain | Patient-derived questions |
|---|---|
| Overview and risk | Q1. What is penile cancer and what causes it? |
| Q2. What are the risk factors for penile cancer? | |
| Q3. How quickly does penile cancer grow and spread, and how urgently do I need treatment? | |
| Symptoms and diagnosis | Q4. What are the symptoms of penile cancer, and how is it diagnosed? |
| Q5. Has my cancer spread, and where does penile cancer usually spread to? | |
| Treatment options | Q6. How is penile cancer treated, and what treatment options will I have? |
| Q10. Can I be cured if I don’t have surgery? | |
| Q11. What nonsurgical options are there? | |
| Q12. When is chemotherapy used, and can I have multiple or combined treatments? | |
| Surgery and organ preservation | Q7. What are my surgical options, and can my penis be preserved? |
| Q8. What is lymph node surgery? Will I need it, and what are its risks? | |
| Q9. How will my life change after surgery, and what could go wrong? | |
| Prognosis and survivorship | Q13. How successful is penile cancer treatment, and what is my prognosis? |
| Q14. Can penile cancer come back, and what follow-up will I need? |
Question numbers reflect the order in which the questions were presented to the models; the thematic grouping shown is for analysis only and does not follow the numerical sequence.
Response capture
Every question was put to the models behind the same opening line, written in the patient’s own voice [for instance, “I am a patient newly diagnosed with penile cancer. Please tell me (insert question)”], so that the framing held constant across platforms. The queries were run on 1 day, each in a fresh incognito conversation to rule out carryover from earlier answers, with the prompt and question typed exactly as written and only the first reply kept; no clarifications, edits, or follow-up prompts were allowed. The 42 replies were collated into one scoring document and stripped of any wording that revealed which platform had produced them, leaving the clinical content untouched. Within each question, the three answers were shuffled and labelled A, B, or C, and the key tying labels back to models stayed with the lead author until every score had been entered.
Response quality appraisal
Grading used the six QAMAI domains, namely accuracy, clarity, relevance, completeness, provision of sources, and usefulness, each marked from 1 (strongly disagree) to 5 (strongly agree) to give a total of 6 to 30, in which higher is better. The validation study’s five bands were applied: very poor [6 to 11], poor [12 to 17], moderate [18 to 23], good [24 to 29], and excellent [30] (21). Because none of the responses provided any source or citation, as reported below, the provision of sources item scored 0 for every response and was excluded from all comparative analyses; the reported total scores therefore carry an effective maximum of 25 rather than 30. The QAMAI validation bands were derived for the complete six-item 30-point scale, so with this item structurally absent, the bands are applied only as an approximate interpretive guide, a limitation revisited below. Following the approach of Pompili et al. (14), a panel of one consultant urologist and three registrars at differing stages of training scored every answer blinded to its source. To keep seniority from steering the result, each rater first committed an independent mark across all six domains. These independent scores were retained and not subsequently altered. The panel then discussed each response, and where the independent scores diverged, the discussion continued until a single consensus score was agreed for the record, with the consultant urologist’s score given no greater weight than those of the registrars. Only the agreed consensus score was used as the final value in the analysis, although individual scores were retained for inter-rater reliability assessment. Scoring was completed across two sessions of approximately 2 hours each, identifying clinically important points, misconceptions, how well each answer met the needs of a newly diagnosed man, whether chatbots gave sensible safety advice, and whether they acknowledged the psychosocial side of care.
Organ preservation appraisal
A focused secondary analysis was performed for the surgery and organ preservation domain because of its clinical significance and the anticipated challenges LLMs may face in communicating the balance between organ preservation and oncological control. For the penile preservation question (Q7), responses were evaluated against four predefined criteria: whether the response acknowledged that penile preservation is often feasible in early-stage or superficial disease, whether it described the available organ-sparing surgical approaches, whether it discussed partial and total penectomy, including their functional implications, and whether it explicitly recognised that oncological control should take precedence over preservation when necessary.
Statistical analysis
QAMAI performance was summarised descriptively, continuous measures as means and median with interquartile range (IQR) and the quality grades as counts and percentages; domain means, grade distributions, and mean totals by theme were derived for each model. To compare the three platforms, the Friedman test for related samples treated each question as a matched observation, and where it flagged an overall difference, Bonferroni-adjusted Wilcoxon signed-rank tests were applied pairwise. Two-sided P below 0.05 denoted significance. For each pairwise comparison, a paired effect size was calculated as Cohen’s dz, the mean of the paired differences divided by their standard deviation, with the Hedges’ g small-sample correction, and interpreted using conventional thresholds of 0.2, 0.5, and 0.8 for small, medium, and large effects. Inter-rater reliability among the four raters’ independent scores was quantified with the two-way intraclass correlation coefficient (absolute agreement) for the total QAMAI scores, reported for a single rater and for the four-rater average, and with Fleiss’ kappa for the item-level ratings. Performance across the five domains was a prespecified secondary question. All computations were carried out in RStudio (v2024.12.1+563; Posit Software, PBC, 2025).
Results
Overall performance
All three platforms answered each of the 14 questions (Table S1). On the QAMAI total, ChatGPT and Claude were near level at 22.7 and 22.6, respectively (ChatGPT: median, 23; range, 17 to 25; Claude: median, 23; range, 18 to 24), while Gemini trailed at 17.0 (median, 17.5; range, 14 to 20); in grade terms, the first two were moderate on average and Gemini sat on the poor to moderate border. The Friedman test confirmed a marked separation across models (χ2=18.78; P<0.001), and pairwise testing placed both ChatGPT and Claude above Gemini (each adjusted P=0.004) while leaving the two leaders statistically indistinguishable (adjusted P>0.99). The magnitude of the difference was large for both comparisons against Gemini (ChatGPT vs. Gemini, Cohen’s dz =1.62; Claude vs. Gemini, dz =1.94) and negligible between ChatGPT and Claude (dz =0.05). The grade counts told the same story: ChatGPT returned 7 good, 6 moderate, and 1 poor answer; Claude 5 good and 9 moderate; and Gemini 7 moderate and 7 poor (Table 2), with nothing reaching excellent and nothing falling to very poor.
Table 2
| QAMAI grade | Score range | ChatGPT, n (%) | Gemini, n (%) | Claude, n (%) | Overall, n (%) |
|---|---|---|---|---|---|
| Very poor | 6 to 11 | 0 (0.0) | 0 (0.0) | 0 (0.0) | 0 (0.0) |
| Poor | 12 to 17 | 1 (7.1) | 7 (50.0) | 0 (0.0) | 8 (19.0) |
| Moderate | 18 to 23 | 6 (42.9) | 7 (50.0) | 9 (64.3) | 22 (52.4) |
| Good | 24 to 29 | 7 (50.0) | 0 (0.0) | 5 (35.7) | 12 (28.6) |
| Excellent | 30 | 0 (0.0) | 0 (0.0) | 0 (0.0) | 0 (0.0) |
Because no response provided a source, the provision of sources item scored 0 for all responses; total scores therefore carry an effective maximum of 25, and the grade bands shown derive from the original 30-point QAMAI scale. N, number of responses; QAMAI, Quality Analysis of Medical Artificial Intelligence.
Comparative testing
Not one model cited a reference or pointed the reader to an authoritative external source, so every one of the 42 answers scored the floor on provision of sources, and that item was set aside for the comparative tests. The remaining domain means appear in Table 3. ChatGPT was clearest of the three (mean, 4.86), ahead of both Gemini (mean, 4.07) and Claude (mean, 4.07) by a significant margin (Friedman P<0.001; Bonferroni adjusted P=0.003 and 0.007), whereas Gemini and Claude did not differ on clarity. Completeness ran the other way, highest for Claude (4.71) and lowest for Gemini (2.71), and here every pairwise contrast was significant (Friedman P<0.001). Gemini also lagged both competitors on accuracy, relevance and usefulness. The upshot is that ChatGPT and Claude behaved alike across most domains, parting company only on clarity, which favoured ChatGPT, and completeness, which favoured Claude.
Table 3
| QAMAI item | ChatGPT | Gemini | Claude | Friedman P |
|---|---|---|---|---|
| Accuracy | 4.43 | 3.71 | 4.64 | 0.001 |
| Clarity | 4.86 | 4.07 | 4.07 | <0.001 |
| Relevance | 4.50 | 3.29 | 4.43 | <0.001 |
| Completeness | 4.43 | 2.71 | 4.71 | <0.001 |
| Usefulness | 4.50 | 3.21 | 4.71 | <0.001 |
Two-sided P values from the Friedman test across the three models. QAMAI, Quality Analysis of Medical Artificial Intelligence.
Thematic subanalysis
Table 4 breaks the totals down by theme. ChatGPT and Claude moved together across the domains, ChatGPT peaking in symptoms and diagnosis (25.00) and Claude in prognosis and survivorship (24.00), while Gemini sat lowest in every domain without exception. The models pulled furthest apart in surgery and organ preservation: ChatGPT (23.67) and Claude (23.00) held the moderate to good range, but Gemini dropped to 14.00, squarely poor. That gap came mostly from weak completeness and usefulness on the questions dealing with penile preservation, inguinal node surgery, and life after an operation.
Table 4
| Thematic domain | ChatGPT | Gemini | Claude | Overall |
|---|---|---|---|---|
| Basics and risk | 24.00 | 18.67 | 23.00 | 21.89 |
| Symptoms and diagnosis | 25.00 | 18.50 | 20.00 | 21.17 |
| Treatment options | 20.75 | 17.25 | 22.50 | 20.17 |
| Surgery and organ preservation | 23.67 | 14.00 | 23.00 | 20.22 |
| Prognosis and survivorship | 21.00 | 17.00 | 24.00 | 20.67 |
QAMAI, Quality Analysis of Medical Artificial Intelligence.
Surgery and organ preservation questions
On the focused penile preservation question (Q7), all three models listed the organ-sparing options (circumcision, wide local excision, Mohs micrographic surgery, laser ablation, glans resurfacing, and glansectomy) and set out partial and total penectomy with perineal urethrostomy. What separated them was the framing of the governing principle, that clearing the cancer must come before sparing the organ. ChatGPT (25, good quality) and Claude (24, good quality) both conveyed it, ChatGPT in plain terms and Claude through the balance between clear margins and local recurrence, with a note on later reconstruction. Gemini (14, poor quality) ran through the same list of operations but scored far lower on completeness and usefulness, leaving the preservation versus control balance implicit and omitting the guidance a newly diagnosed man would need to weigh the choice (Table 5). Q9, on life after surgery, followed suit: ChatGPT (21) and Claude (22) took in urinary and sexual changes alongside the psychological and relationship toll, whereas Gemini (14) gave a thinner account.
Table 5
| Model | QAMAI (Q7) | Stated preservation often achievable | Enumerated organ-sparing options | Described partial/total penectomy and functional impact | Framed oncological control as priority |
|---|---|---|---|---|---|
| ChatGPT | 25 | Yes | Yes | Yes | Yes (explicit) |
| Gemini | 14 | Partial | Yes | Yes | Partial (implied, not explicit) |
| Claude | 24 | Yes | Yes | Yes | Yes (via margin/recurrence tradeoff) |
QAMAI, Quality Analysis of Medical Artificial Intelligence.
Discussion
Key findings
To our knowledge, this is the first study to evaluate AI-generated patient information in penile cancer and the first to directly compare ChatGPT, Gemini, and Claude in this disease setting. Generative AI has moved into routine clinical and consumer use at a pace few anticipated (12), and patients increasingly access these systems before engaging with healthcare professionals. The principal finding of this study was a clear separation between models. ChatGPT and Claude produced moderate-to-good-quality responses and did not differ significantly from one another, whereas Gemini scored approximately five to six QAMAI points lower overall and generated poor-quality responses for half of the questions assessed. Given the rarity of penile cancer (2) and the fact that existing publicly available patient information is often of only moderate quality and written at a reading level that exceeds that of the average reader (22), this variation has important practical implications. Specifically, the platform through which an individual first seeks information may substantially influence whether their initial understanding of the disease is accurate and comprehensive or incomplete and potentially misleading.
Strengths and limitations
These findings should be interpreted in light of several limitations. First, each response represents a single interaction captured on 1 day from systems that are not deterministic and subject to ongoing retraining. Consequently, the reported scores reflect performance at a specific point in time and will not be reproducible in future model versions. Second, only free tier versions were evaluated, which may underestimate the capabilities of paid or reasoning-enhanced models. Third, although platform-identifying wording was removed from every response before assessment, no formal blinding assessment was undertaken, for example, asking raters to identify the source model, so the effectiveness of blinding was not quantified; stylistic characteristics may have remained recognisable and could have influenced scoring. The four raters scored every response independently before discussion, and agreement on the total QAMAI score was good to excellent (two way intraclass correlation coefficient 0.82 for a single rater, 95% confidence interval: 0.64 to 0.91; and 0.95 for the four rater average, 95% confidence interval: 0.88 to 0.97); item level agreement was fair when the ordinal ratings were treated as nominal categories (Fleiss’ kappa =0.32), reflecting their skewed distribution, although pairwise agreement between raters was exact for 54% of item ratings and within one point for 94%, and any residual divergence was settled through the consensus process, which together with the blinded independent scoring limited individual and seniority bias. Fourth, the fourteen questions, while derived from established patient information resources, cannot encompass the full spectrum of patient concerns. In addition, the rarity of penile cancer means that the quantity of disease-specific training data available to any model is likely to be limited. Readability was likewise not formally assessed in the present study, and where the discussion refers to readability, it summarises prior literature rather than measurements made here. Finally, the study was restricted to English language outputs and to the domains assessed by QAMAI. Future work should formally assess readability and actionability using dedicated instruments, and examine multilingual performance and longitudinal changes in model behaviour as these systems continue to evolve.
Comparison with similar research
The marked divergence between contemporary LLMs is consistent with findings from the wider literature. Comparative studies in prostate cancer have demonstrated differences in comprehensiveness, accuracy, and readability between models rather than convergence in performance (15). Similarly, a multiplatform evaluation in bladder cancer reported stylistically distinct outputs despite an absence of hallucinations (16), while a comparison of AI-generated urology patient leaflets identified meaningful quality differences beneath superficially similar content (14). The ability to detect such differences in the present study may partly reflect the assessment tool employed. Unlike readability and consumer information indices used in earlier work, QAMAI was specifically developed and validated to evaluate the accuracy, completeness, and usefulness of AI-generated health information (21). These dimensions revealed meaningful distinctions between models. Notably, differences were also evident between the two highest-performing systems. ChatGPT achieved the highest scores for clarity, whereas Claude scored highest for completeness, suggesting a balance between accessibility and comprehensiveness. No individual platform consistently maximised both attributes.
Explanations of findings
Importantly, Gemini performance was not uniformly weak across topics but was concentrated within questions relating to surgery and organ preservation. This domain is particularly important because treatment decisions frequently require patients to weigh oncological outcomes against functional preservation. The lower scores were primarily attributable to deficiencies in completeness and usefulness rather than outright inaccuracies. Contemporary penile cancer management relies on a spectrum of organ-preserving approaches for appropriately selected early-stage disease, with partial or total penectomy and inguinal nodal staging reserved for more advanced or higher-risk tumours (3). Responses that list available procedures without adequately explaining the rationale for organ preservation, the indications for more radical surgery, or the implications for urinary and sexual function may provide technically correct information while failing to support informed understanding. In effect, such responses describe the available operations without fully addressing the decisions that surround them. Given the well-documented psychosexual burden associated with penile cancer (4), omissions in these areas may contribute to uncertainty and unmet informational needs.
Implications and actions needed
The clinical significance of these findings depends partly on how patients use these technologies. Many individuals seek information about symptoms and diagnoses online before consulting a healthcare professional (6), and this behaviour may be particularly relevant in penile cancer, where embarrassment and stigma are recognised contributors to delayed presentation (7). The anonymity and immediate accessibility of conversational AI may therefore increase its appeal for patients reluctant to seek medical advice. However, accessibility should not be equated with completeness or clinical reliability. Studies evaluating AI-generated information in prostate cancer radiotherapy (17), simulated bladder cancer consultations (18), and common questions across urological malignancies (13,19) have consistently found that while factual content is often broadly accurate, readability frequently exceeds recommended levels and actionable guidance remains limited. A recent systematic review of LLM-based healthcare chatbots reached similar conclusions across multiple specialties (23). Correspondingly, patients continue to report substantially greater trust in urologists than in LLMs for treatment decisions and reassurance (20). In the present study, the strongest performing models provided useful information, but none consistently delivered the contextualised guidance required for patient counselling.
One finding was universal across all three platforms. None of the 42 responses cited information sources or directed users towards additional resources. From a patient safety perspective, this is important because unsourced clinical information provides no mechanism for independent verification and no pathway to authoritative educational materials following a new diagnosis (10,11). In addition, no response carried an explicit statement that it did not constitute medical advice, although several did appropriately defer patient-specific determinations, such as whether the cancer had spread or the individual prognosis, to the treating clinician; this safety netting was applied inconsistently across questions and models. Previous QAMAI-based evaluations have reported highly variable performance in source attribution across specialties, ranging from near absence to routine inclusion (24,25), suggesting that citation behaviour may be influenced by both model design and prompting strategy. The patient-oriented prompts used in the present study may therefore have reduced the likelihood of source provision. Nevertheless, reliable referencing and signposting are widely regarded as essential prerequisites for the safe deployment of medical chatbots (9). Their complete absence in this study highlights an important limitation of current systems, consistent with direct evaluations in which ChatGPT has been shown to provide inaccurate urological healthcare advice (26). This concern is further reinforced by evidence that LLMs may generate inaccurate or fabricated references when prompted for citations (27) and may elaborate upon incorrect information contained within user prompts rather than challenge it (28). Clinicians should therefore assume that information obtained from these platforms is unsourced and should actively direct patients towards evidence-based resources, including professional society guidelines and survivorship support materials.
Conclusions
ChatGPT and Claude delivered broadly equivalent, moderate-to-good-quality information on patient-derived penile cancer questions, and both significantly outperformed Gemini, which frequently produced incomplete answers, most notably in the surgical and organ preservation domain so central to this disease. ChatGPT was the clearest and Claude the most complete, while no model provided a single source or onward signpost. These tools are a useful supplement to, but not a substitute for, specialist consultation, and clinicians counselling newly diagnosed men should anticipate prior LLM use, recognise that quality varies considerably between platforms, and actively direct patients towards reliable information and support.
Acknowledgments
During the preparation of this work, the authors used generative AI only to generate responses for the dataset. No AI input was used for language editing. All output was subsequently reviewed and edited by the authors, who take full responsibility for the content of this publication.
Footnote
Data Sharing Statement: Available at https://tau.amegroups.com/article/view/10.21037/tau-2026-0603/dss
Peer Review File: Available at https://tau.amegroups.com/article/view/10.21037/tau-2026-0603/prf
Funding: None.
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://tau.amegroups.com/article/view/10.21037/tau-2026-0603/coif). The authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Mannam G, Miller JW, Johnson JS, et al. HPV and Penile Cancer: Epidemiology, Risk Factors, and Clinical Insights. Pathogens 2024;13:809. [Crossref] [PubMed]
- Huang J, Chan SC, Pang WS, et al. Incidence, risk factors, and temporal trends of penile cancer: a global population-based study. BJU Int 2024;133:314-23. [Crossref] [PubMed]
- Brouwer OR, Albersen M, Parnham A, et al. European Association of Urology-American Society of Clinical Oncology Collaborative Guideline on Penile Cancer: 2023 Update. Eur Urol 2023;83:548-60. [Crossref] [PubMed]
- Maddineni SB, Lau MM, Sangar VK. Identifying the needs of penile cancer sufferers: a systematic review of the quality of life, psychosexual and psychosocial literature in penile cancer. BMC Urol 2009;9:8. [Crossref] [PubMed]
- Lim MSC, Molenaar A, Brennan L, et al. Young Adults' Use of Different Social Media Platforms for Health Information: Insights From Web-Based Conversations. J Med Internet Res 2022;24:e23656. [Crossref] [PubMed]
- Maon SN, Hassan NM, Seman SAA. Online health information seeking behavior pattern. Adv Sci Lett 2017;23:10582-5.
- Cieślikowski WA, Kasperczak M, Milecki T, et al. Reasons behind the Delayed Diagnosis of Testicular Cancer: A Retrospective Analysis. Int J Environ Res Public Health 2023;20:4752. [Crossref] [PubMed]
- Van Booven D, Chen CB. Chapter 6 - ChatGPT and healthcare—current and future prospects. In: Arora H, editor. Artificial intelligence in urologic malignancies. London: Academic Press; 2025:173-93.
- Chow JCL, Wong V, Li K. Generative pre-trained transformer-empowered healthcare conversations: current trends, challenges, and future directions in large language model-enabled medical chatbots. BioMedInformatics 2024;4:837-52.
- Bélisle-Pipon JC. Why we need to be careful with LLMs in medicine. Front Med (Lausanne) 2024;11:1495582. [Crossref] [PubMed]
- Sauerbrei A, Kerasidou A, Lucivero F, et al. The impact of artificial intelligence on the person-centred, doctor-patient relationship: some problems and solutions. BMC Med Inform Decis Mak 2023;23:73. [Crossref] [PubMed]
- Teo ZL, Thirunavukarasu AJ, Elangovan K, et al. Generative artificial intelligence in medicine. Nat Med 2025;31:3270-82. [Crossref] [PubMed]
- Musheyev D, Pan A, Loeb S, et al. How Well Do Artificial Intelligence Chatbots Respond to the Top Search Queries About Urological Malignancies?. Eur Urol 2024;85:13-6. [Crossref] [PubMed]
- Pompili D, Richa Y, Collins P, et al. Using artificial intelligence to generate medical literature for urology patients: a comparison of three different large language models. World J Urol 2024;42:455. [Crossref] [PubMed]
- Alasker A, Alsalamah S, Alshathri N, et al. Performance of large language models (LLMs) in providing prostate cancer information. BMC Urol 2024;24:177. [Crossref] [PubMed]
- Patel K, Radcliffe R. Evaluating the Readability and Quality of Bladder Cancer Information from AI Chatbots: A Comparative Study Between ChatGPT, Google Gemini, Grok, Claude and DeepSeek. J Clin Med 2025;14:7804. [Crossref] [PubMed]
- Trapp C, Schmidt-Hegemann N, Keilholz M, et al. Patient- and clinician-based evaluation of large language models for patient education in prostate cancer radiotherapy. Strahlenther Onkol 2025;201:333-42. [Crossref] [PubMed]
- Guo AA, Razi B, Kim P, et al. The role of artificial intelligence in patient education: a bladder cancer consultation with ChatGPT. Soc Int Urol J 2024;5:214-24.
- Thia I, Saluja M. ChatGPT: Is This Patient Education Tool for Urological Malignancies Readable for the General Population?. Res Rep Urol 2024;16:31-7. [Crossref] [PubMed]
- Carl N, Nguyen L, Haggenmüller S, et al. Comparing Patient's Confidence in Clinical Capabilities in Urology: Large Language Models Versus Urologists. Eur Urol Open Sci 2024;70:91-8. [Crossref] [PubMed]
- Vaira LA, Lechien JR, Abbate V, et al. Validation of the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool: a new tool to assess the quality of health information provided by AI platforms. Eur Arch Otorhinolaryngol 2024;281:6123-31. [Crossref] [PubMed]
- Ma J, Wang M, Liu Q, et al. Quality and readability of online penile-cancer information: a Cross-Sectional Web Study. Int J Impot Res 2026; [Epub ahead of print]. [Crossref] [PubMed]
- Huo B, Boyle A, Marfo N, et al. Large Language Models for Chatbot Health Advice Studies: A Systematic Review. JAMA Netw Open 2025;8:e2457879. [Crossref] [PubMed]
- Furukawa E, Okuhara T, Okada H, et al. Reliability of AI-generated responses on frequently-posed questions by patients with chronic kidney disease. Health Informatics J 2025;31:14604582251381996. [Crossref] [PubMed]
- Bozhöyük MS, Yücel L. A comprehensive evaluation of artificial intelligence-provided information on common ENT surgical procedures using the QAMAI tool. J Laryngol Otol 2025;139:1109-14. [Crossref] [PubMed]
- Whiles BB, Bird VG, Canales BK, et al. Caution! AI Bot Has Entered the Patient Chat: ChatGPT Has Limitations in Providing Accurate Urologic Healthcare Advice. Urology 2023;180:278-84. [Crossref] [PubMed]
- Chelli M, Descamps J, Lavoué V, et al. Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative Analysis. J Med Internet Res 2024;26:e53164. [Crossref] [PubMed]
- Omar M, Sorin V, Collins JD, et al. Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support. Commun Med (Lond) 2025;5:330. [Crossref] [PubMed]

