Original Article
Completeness and Guideline Concordance of Artificial Intelligence Chatbot Responses to Patient Questions on Regenerative Therapies for Erectile Dysfunction: A Multi-platform Analysis
Abstract
Background: Interest in regenerative therapies for erectile dysfunction (ED), including stem cell therapy (SCT), low-intensity shockwave therapy (LiSWT), and platelet-rich plasma (PRP), has increased despite limited evidence and investigational designations from professional societies. Concurrently, patients increasingly use artificial intelligence (AI) large language model (LLM) chatbots to obtain health information. We evaluated the completeness and guideline-concordance of AI-generated responses to patient-style questions about regenerative ED therapies, and their consistency across repeated queries.
Methods: We performed a cross-sectional evaluation of four general-purpose LLM platforms: ChatGPT-4o, Gemini 2.5, Claude 3.7, and Copilot. Six patient-style prompts addressing general inquiry, evidence appraisal, safety, and treatment-seeking behavior were submitted for each therapy (SCT, LiSWT, PRP) in two independent sessions on separate days, yielding 144 responses. The primary outcome was the presence of nine prespecified counseling elements using a binary checklist. The secondary outcome was concordance with the Sexual Medicine Society of North America (SMSNA) position statement, assessed across efficacy framing, recommendation against routine use, and regulatory-status accuracy. Within-platform consistency across two sessions was assessed using Cohen’s kappa.
Results: Most responses identified investigational or non-FDA-approved status (80.6%), recommended physician consultation (76.4%) and acknowledged limitations in current research (75%). References to society guidelines (14.6%), safety profiles (38.9%), alternative therapies (23.6%), and cost (31.9%) were uncommon. Peer-reviewed citations were present in fewer than half of responses (47.9%), despite frequent references to “research”. Among the 130 responses that characterized efficacy, only 36% conveyed the guideline-concordant position that clinical efficacy has not been established for regenerative therapies. The majority of responses instead framed efficacy optimistically despite the lack of robust efficacy evidence. An explicit recommendation against routine clinical use appeared in only 18.8% of responses. Within-platform consistency was substantial for Copilot, Claude, and Gemini (κ=0.61-0.67) and moderate for ChatGPT (κ=0.54).
Conclusions: General-purpose AI chatbots provide incomplete counseling regarding regenerative ED therapies and inconsistently reflect the core elements of the SMSNA position statement. Chatbot responses specifically contain gaps in disclosing safety profiles, approved alternatives, and cost. Deliberate integration of evidence-based guidance into consumer-facing LLMs is needed to prevent AI-driven misconceptions. Clinicians should assess AI use and address AI-derived expectations during clinical encounters.

