Feasibility analysis of generative artificial intelligence tools as a means of medical science popularization on urological diseases
Highlight box
Key findings
• Generative artificial intelligence (AI) models combined with manual review can assist in the medical science popularization of urological diseases.
What is known and what is new?
• There are few reports on the application of AI in medical science popularization among patients in China.
• The generative AI tool, ChatGPT 4.0, could be used to assist in the popularization of medical science in urology among Chinese population.
What is the implication, and what should change now?
• When using the generative AI for popularizing knowledge of urological diseases, the generated content needs to be manually reviewed to ensure the accuracy and readability of the content.
Introduction
Medical science popularization is an activity that popularizes medical concepts, knowledge and technologies to the public to improve the public’s health literacy and health self-management ability. In 2021, the health literacy level of Chinese residents was only 25.40% (1), while the search share of the health and medical topic in 2018 was as high as 66.83% in China. Carrying out medical science popularization is an important means to meet the public’s growing health needs and improve the health literacy level of residents (2). Due to the heavy medical workload of medical staff and the complicated procedure of medical science popularization, there are certain difficulties and resistance in medical science popularization.
Artificial intelligence (AI) is a technology based on statistical analysis of big data and machine learning to generate new and meaningful insights, and it can potentially contribute to the medical field. ChatGPT is one of the most popular AI tools today. It is a type of large language model trained on a vast amount of internet resources, including research papers, website content, videos, podcasts, articles, and forum conversations (3). Through continuous updates and iterations, ChatGPT 4.0 can generate coherent, accurate, and context-aware responses. The expansion of its knowledge base and its multilingual comprehension ability have further broadened the scope of its medical applications worldwide (4). Leveraging its enhanced memory capacity, ChatGPT 4.0 is capable of conducting more comprehensive and in-depth patient interactions. This breakthrough may play a significant role in patient education (3). There have been some attempts and breakthroughs in clinical work and medical education, but there are few reports on the application of AI in medical science popularization among patients in China (5-8). Based on the advantages and practical performance of AI in medical education, we envisage that it may also be used for “patient education”. The study aimed to evaluate the feasibility of generative AI tool, ChatGPT 4.0, to assist in the popularization of medical science in urology among Chinese population.
Methods
Generating the content of urological diseases using ChatGPT 4.0
This study intended to evaluate the quality of urological diseases-related content generated by ChatGPT 4.0. We selected common urological diseases as representatives for research, including kidney diseases (renal cyst and renal cell carcinoma), prostate diseases (benign prostate hyperplasia and prostate cancer), adrenal diseases (adrenal aldosteronoma and pheochromocytoma), metabolic diseases and urinary tract infection (ureteral stone and cystitis). For each disease, ChatGPT 4.0 was used to generate four aspects of content, including pathogenesis, clinical manifestations, diagnosis and treatment. We asked ChatGPT 4.0 the following questions in Chinese: “What is the pathogenesis of?”, “What are the clinical manifestations of?”, “How to diagnose?”, and “How to treat?”. The Chinese content of ChatGPT 4.0 was generated on May 15th, 2025. IRB approval and informed consent are not required as no human subjects were involved.
Evaluation of ChatGPT 4.0 generated content
Evaluation was conducted on the eight diseases and four aspects of each disease. Two attending physicians (W.W. and Y.Z.) from the Department of Urology of Peking Union Medical College Hospital (PUMCH) were selected as evaluators l and 2 to score each item. We evaluated the content from the following six dimensions: scientificity (is the content scientific and accurate?), recency (are the viewpoints and concepts the latest?), comprehensiveness (is the content relatively comprehensive and complete?), understandability (is the content easy to understand?), conciseness (is the content concise without repetition?), and interest (is the content lively and not too serious?). Each of the above dimensions was set with a score from 1 to 5 points, and the higher the score, the more it met the relevant requirements: Score 1, fully does not meet the requirements; Score 2, mostly does not meet the requirements; Score 3, half meets the requirements; Score 4, mostly meets the requirements; Score 5, fully meets the requirements. Evaluators 1 and 2 marked unreasonable points during the evaluation process.
After the initial evaluation, another two attending physicians (G.Z. and J.D.) from the Department of Urology of PUMCH were selected as reviser 1 and 2 to revise the unreasonable points marked by the evaluators. The two evaluators would evaluate the content of ChatGPT 4.0 and point out unreasonable contents, including inaccurate knowledge (such as errors, outdated knowledge, overly simple description, etc.) and unclear logic (such as repeated content, inconsistent content with the topic, etc.). The two revisers would then revise the unreasonable contents. Afterwards, evaluators l and 2 evaluated all the revised content again from the six dimensions and compared them with the initial evaluation.
Statistical analysis
SPSS 27.0 was used to process and analyse the data. The scores of each item generated by ChatGPT 4.0 were presented in the form of mean ± standard deviation (SD), The comparison between different groups was performed by analysis of variance (ANOVA), and the comparison of scores before and after revision was performed by paired sample t-test. The linear weighted Kappa coefficient was used to evaluate the mutual consistency between different evaluators. The weighted Kappa values were classified as follows: Kappa <0.20 indicated poor consistency, 0.21–0.4 indicated general consistency, 0.41–0.60 indicated moderate consistency, 0.61–0.80 indicated strong consistency, and 0.81–1.00 indicated very strong consistency. P value <0.05 was considered statistically significant.
Results
Consistency analysis of evaluation of ChatGPT 4.0 generated content
The scores of each part of the content generated by ChatGPT 4.0 were shown in Table 1, which shows the mean ± SD of the scores of the two evaluators on the six dimensions of scientificity, recency, comprehensiveness, understandability, conciseness and interest for each content of various diseases.
Table 1
| Items | Diseases | Contents | Scores by evaluator 1 | Scores by evaluator 2 | Weighted Kappa coefficient | P |
|---|---|---|---|---|---|---|
| Kidney diseases | Renal cyst | Pathogenesis | 3.67±0.52 | 3.50±1.38 | 0.591 (0.321 to 0.861) | 0.01 |
| Clinical manifestations | 4.00±0.63 | 3.50±1.05 | 0.471 (0.081 to 0.860) | 0.03 | ||
| Diagnosis | 3.83±0.75 | 3.33±1.03 | 0.500 (0.073 to 0.927) | 0.03 | ||
| Treatment | 3.83±1.17 | 3.67±1.21 | 0.864 (0.605 to 1.123) | 0.003 | ||
| Renal cell carcinoma | Pathogenesis | 4.00±0.63 | 3.83±0.98 | 0.438 (−0.010 to 0.885) | 0.08 | |
| Clinical manifestations | 4.00±0.63 | 3.67±0.82 | 0.571 (0.079 to 1.064) | 0.03 | ||
| Diagnosis | 3.67±0.82 | 3.50±0.55 | 0.304 (−0.229 to 0.896) | 0.27 | ||
| Treatment | 4.17±0.75 | 4.17±0.98 | 0.625 (0.201 to 1.049) | 0.04 | ||
| Prostate diseases | Benign prostate hyperplasia | Pathogenesis | 4.33±0.82 | 3.83±1.17 | 0.526 (0.222 to 0.830) | 0.04 |
| Clinical manifestations | 4.00±0.63 | 3.50±1.05 | 0.471 (0.081 to 0.860) | 0.03 | ||
| Diagnosis | 3.83±0.98 | 4.00±0.63 | 0.769 (0.520 to 1.019) | 0.001 | ||
| Treatment | 3.67±0.52 | 3.50±0.55 | 0.667 (0.104 to 1.229) | 0.08 | ||
| Prostate cancer | Pathogenesis | 3.17±0.98 | 3.33±0.82 | 0.813 (0.500 to 1.125) | 0.01 | |
| Clinical manifestations | 3.50±0.55 | 3.17±0.41 | 0.333 (−0.229 to 0.896) | 0.27 | ||
| Diagnosis | 4.00±0.89 | 3.83±0.98 | 0.824 (0.502 to 1.145) | 0.01 | ||
| Treatment | 3.67±1.21 | 3.50±1.05 | 0.857 (0.603 to 1.111) | 0.002 | ||
| Adrenal diseases | Adrenal aldosteronoma | Pathogenesis | 3.67±0.52 | 3.17±0.75 | 0.308 (−0.073 to 0.689) | 0.22 |
| Clinical manifestations | 3.83±0.41 | 3.50±0.55 | 0.333 (−0.229 to 0.896) | 0.27 | ||
| Diagnosis | 4.00±0.89 | 3.50±0.84 | 0.471 (0.042 to 0.899) | 0.09 | ||
| Treatment | 3.83±0.41 | 3.67±0.82 | 0.625 (0.541 to 0.709) | 0.01 | ||
| Pheochromocytoma | Pathogenesis | 3.17±0.75 | 3.00±0.63 | 0.750 (0.296 to 1.204) | 0.008 | |
| Clinical manifestations | 3.83±0.41 | 4.00±0.63 | 0.625 (0.038 to 1.212) | 0.01 | ||
| Diagnosis | 3.83±0.41 | 3.67±0.52 | 0.571 (−0.121 to 1.264) | 0.12 | ||
| Treatment | 4.00±0.63 | 3.83±0.75 | 0.750 (0.296 to 1.204) | 0.008 | ||
| Metabolic diseases and urinary tract infection | Ureteral stone | Pathogenesis | 4.33±0.82 | 4.00±1.10 | 0.647 (0.353 to 0.941) | 0.01 |
| Clinical manifestations | 3.50±0.55 | 3.33±0.82 | 0.750 (0.415 to 1.085) | 0.01 | ||
| Diagnosis | 4.00±0.63 | 3.67±1.03 | 0.600 (0.242 to 0.958) | 0.008 | ||
| Treatment | 4.33±0.82 | 4.00±1.10 | 0.647 (0.353 to 0.941) | 0.01 | ||
| Cystitis | Pathogenesis | 4.33±0.82 | 3.83±0.98 | 0.471 (0.136 to 0.805) | 0.03 | |
| Clinical manifestations | 4.00±0.63 | 4 .00±1.10 | 0.600 (0.278 to 0.922) | 0.008 | ||
| Diagnosis | 3.83±0.41 | 3.67±0.52 | 0.571 (−0.121 to 1.264) | 0.12 | ||
| Treatment | 4.00±0.63 | 3.83±0.75 | 0.750 (0.296 to 1.204) | 0.008 |
Data are presented as mean ± standard deviation or (95% CI). CI, confidence interval.
Quality analysis of ChatGPT 4.0 generated content
Two evaluators evaluated the content of ChatGPT 4.0 and found 54 unreasonable contents, including inaccurate knowledge (25 in total, 11 of errors, 6 of outdated knowledge, 8 of overly simple description) and unclear logic (29 in total, 13 of repeated content, 16 of inconsistent content with the topic) (Table 2).
Table 2
| Unreasonable types | Typical cases | Modification suggestions |
|---|---|---|
| Inaccurate knowledge | ||
| Errors | Renal cyst. Diagnosis—“Pathological examination: for suspected malignant lesions or complex cysts, a puncture biopsy may be required to obtain cells or tissues inside the cyst for pathological examination. This method is usually used when the results of imaging examinations are unclear or further diagnosis is required.” | When malignant lesions of renal cysts are suspected, puncture is avoided to prevent tumor spread, and complete surgical resection should be performed |
| Outdated knowledge | Adrenal aldosteronoma. Diagnosis—“The 68Ga-Pentixafor PET/CT imaging which targets CXCR4 is not mentioned in this section.” | In recent years, 68Ga-Pentixafor PET/CT imaging which targets CXCR4 has been widely used in the diagnosis of adrenal aldosteronoma and in the determination of the functional dominant side |
| Overly simple description |
Renal cell carcinoma. Treatment—“Laparoscopic or robotic-assisted laparoscopic surgery: these minimally invasive surgical methods are often used for earlier kidney carcinoma and can reduce postoperative recovery time and pain.” |
Currently, minimally invasive surgeries including laparoscopic surgery and robotic-assisted laparoscopic surgery are the main methods of partial and radical nephrectomy. In addition to reducing recovery time and pain, it also has important advantages in surgical flexibility and precision, which should be emphasized |
| Unclear logic | ||
| Repeated content | Benign prostate hyperplasia. Clinical manifestations—The previous content “obstructive symptoms” described “urinary retention: because the bladder cannot completely empty urine, there is still residual urine after urination, and the patient may feel incomplete urination”, and the following text mentions “urinary retention symptoms: as prostatic hyperplasia progresses, some patients may experience acute or chronic urinary retention.” | Repeated content should be integrated |
| Inconsistent content with the topic |
Prostate cancer. Clinical manifestations—“Elevated PSA: PSA is a common biomarker for prostate cancer, and its increase is an important indicator of prostate cancer. Many patients with early prostate cancer will find elevated PSA levels during examination, although elevated PSA levels may also be caused by benign prostate hyperplasia or prostatitis.” | PSA test is part of the “diagnosis” section and should be deleted in “clinical manifestations” section |
CXCR4, C-X-C chemokine receptor type 4; PET/CT, positron emission tomography-computed tomography; PSA, prostate-specific antigen.
From the perspective of different contents, the scores of “pathogenesis”, “clinical manifestations”, “diagnosis”, and “treatment” were compared with each other (Table 3). For example, the six dimensions of the “pathogenesis” of a disease were first added up to calculate the total, and then the mean ± SD of the scores of the “pathogenesis” of eight diseases was calculated and then compared with the scores of “clinical manifestations”, “diagnosis”, and “treatment”.
Table 3
| Evaluators | Pathogenesis | Clinical manifestations | Diagnosis | Treatment | F | P |
|---|---|---|---|---|---|---|
| Evaluator 1 | 23.00±2.98 | 23.00±1.31 | 23.25±0.71 | 23.63±1.41 | 0.214 | 0.89 |
| Evaluator 2 | 21.38±2.20 | 21.50±1.77 | 21.88±1.25 | 22.63±1.41 | 0.879 | 0.46 |
Data are presented as mean ± standard deviation.
From the perspective of evaluation dimensions, the scores of different dimensions in various diseases were compared (Table 4). For example, the sum of the “scientificity” scores of the four contents of “pathogenesis”, “clinical manifestations”, “diagnosis”, and “treatment” of a certain disease was first calculated, and then the mean ± SD of the “scientificity” scores of eight different diseases were calculated, and compared with the scores of “recency”, “comprehensiveness”, “understandability”, “conciseness” and “interest”.
Table 4
| Evaluators | Scientificity | Recency | Comprehensiveness | Understandability | Conciseness | Interest | F | P |
|---|---|---|---|---|---|---|---|---|
| Evaluator 1 | 16.00±1.31 | 17.25±0.89 | 18.25±0.89 | 15.00±1.07 | 15.13±1.73 | 11.25±0.71 | 35.588 | <0.001 |
| Evaluator 2 | 15.38±0.92 | 16.63±1.30 | 18.38±0.92 | 13.25±1.28 | 13.88±1.46 | 9.88±1.36 | 46.617 | <0.001 |
Data are presented as mean ± standard deviation.
From the perspective of diseases, the scores of the generated content of different diseases were compared (Table 5). For example, the sum of the six dimensions of the “pathogenesis” in “renal cyst” was calculated first, and the mean ± SD of the scores of the “pathogenesis”, “clinical manifestations”, “diagnosis”, and “treatment” was calculated, and then compared with other diseases.
Table 5
| Evaluators | Renal cyst | Renal cell carcinoma | Benign prostate hyperplasia | Prostate cancer | Adrenal aldosteronoma | Pheochromocytoma | Ureteral stone | Cystitis | F | P |
|---|---|---|---|---|---|---|---|---|---|---|
| Evaluator 1 | 23.00±0.82 | 23.75±1.26 | 23.75±1.71 | 21.50±2.08 | 23.00±0.82 | 22.25±2.22 | 24.25±2.36 | 24.25±1.26 | 1.372 | 0.26 |
| Evaluator 2 | 21.00±0.82 | 22.75±1.71 | 22.25±1.50 | 20.75±1.71 | 20.75±1.26 | 21.75±2.63 | 22.50±1.91 | 23.00±0.82 | 1.243 | 0.32 |
Data are presented as mean ± standard deviation.
Comparison of content quality before and after manual review
Manual revisions were made for the unreasonable contents raised by the two evaluators in the initial evaluation. The comparison of the secondary evaluation scores after revision with the initial evaluation scores of each dimension was shown in Table 6.
Table 6
| Dimensions | Evaluator 1 | Evaluator 2 | |||
|---|---|---|---|---|---|
| After revision | P (with that before revision) | After revision | P (with that before revision) | ||
| Scientificity | 17.25±0.71 | 0.06 | 17.00±1.20 | <0.001 | |
| Recency | 17.88±0.64 | 0.10 | 17.63±0.74 | 0.10 | |
| Comprehensiveness | 18.63±0.92 | 0.35 | 18.63±0.74 | 0.52 | |
| Understandability | 17.25±0.89 | <0.001 | 15.63±0.92 | <0.001 | |
| Conciseness | 17.00±1.69 | 0.006 | 16.63±1.19 | <0.001 | |
| interest | 11.50±1.41 | 0.67 | 11.13±0.99 | 0.08 | |
Data are presented as mean ± standard deviation.
Discussion
Generative AI models such as ChatGPT are large language models that have acquired powerful ability to understand and generate text through deep learning. Compared with previous AI, it can understand the meaning behind the text more deeply and make inferences and logical judgments. For example, large-scale parameters help ChatGPT better understand and generate natural language, thereby improving its performance in anthropomorphic interaction and technical support, and providing more logical and readable answers (9,10). In addition, ChatGPT has memory, and users can continuously ask and communicate with ChatGPT to achieve more natural and realistic conversations.
Because medical staff have mastered the knowledge of medical science popularization in their professional fields, they can understand the medical blind spots of patients and their families and provide health guidance when carrying out medical activities. However, the heavy medical work of medical staff and the lack of effective incentive mechanism have seriously affected the energy and motivation of medical staff to carry out medical and health science popularization. According to a regional survey report in China on medical personnel, only 31.1% (172/553) of the respondents were willing to devote more than 1 hour a week to medical science popularization work (11). For the respondents who were unwilling to devote themselves to medical science popularization, the top three reasons were being too busy at work [75.0% (48/64)], insufficient personal science popularization ability [42.2% (27/64)], and lack of incentive mechanism [21.9% (14/64)]. Therefore, in the context of the increasing demand for medical science popularization, generative AI models may be able to solve the current dilemma.
At present, there are more and more researches report on generative AI tools for the dissemination of medical knowledge among the general population. The topics covered in the literature include multidisciplinary content such as tumors, nutrition, pain, orthopedics, ophthalmology, etc. (12-18). In these studies, the generative AI models answer and generate content based on questions retrieved by patients or questions listed by researchers. Most studies show that the content generated by the generative AI model is accurate, but performs poorly in readability. Pushpanathan et al. tried to evaluate the self-checking and correction capabilities of the generative AI model, but the performance was not satisfactory. For the answers initially judged as “poor”, the improvement was not obvious (13). In addition, the study by Ponzo et al. showed that when there were different comorbidities and overall nutritional advice was needed, there might be some contradictory and inappropriate advice (15). This indicated that for some complex conditions, the processing ability of the generative AI model might be worse than that of relatively simple conditions. Some studies have also clearly stated that the content produced by the generative AI model ultimately needed to be reviewed by a clinical professional team and could not completely replace manual work (14,15,18).
In the field of urology, literature has reported the evaluation of generative AI models on urinary stones, prostate cancer, urothelial carcinoma, erectile dysfunction, urinary incontinence, etc. (19-25). Similarly, the accuracy of these contents was acceptable, but there were still some inaccuracies, and readability was also an aspect that needed to be improved. The study by Song et al. also showed that the generative Al model was insufficient in empathy and humanistic care (23). Our study tried to evaluate the ability to conduct medical science popularization of generative Al for common diseases in urology.
In this study, ChatGPT 4.0 was used to generate multi-faceted knowledge of different types of diseases in urology, including kidney, prostate, adrenal gland, metabolic and infection-related diseases, with specific content on pathogenesis, clinical manifestations, diagnosis, and treatment. The above generated content was evaluated in six dimensions: scientificity, recency, comprehensiveness, understandability, conciseness, and interest. The linear weighted Kappa coefficient consistency analysis showed that 4 items had very strong consistency, 11 items had strong consistency, 13 items had moderate consistency, 4 items had general consistency, and no item had poor consistency between the two evaluators. The results of the analysis showed that there was no significant statistical difference in the scores of generated content of diseases. In terms of pathogenesis, clinical manifestations, diagnosis, and treatment, there was no significant statistical difference in the scores of the content generated by ChatGPT 4.0. Overall, the quality of the content generated by ChatGPT 4.0 was relatively satisfactory, but there were still some problems of inaccurate knowledge and unclear logic. This result is consistent with the results reported previously (19-25). Among these, the inaccurate knowledge may have a more critical and significant impact on readers. Errors or misleading results produced by AI are referred to as “hallucinations”. For patient groups without a medical professional background, it is difficult to distinguish between hallucinations and correct contents. Such misinformation may spread incorrect health-related information, thereby exerting negative impacts on readers’ safety and awareness (26,27). Therefore, at the current stage, the model of AI generation combined with manual review may correct hallucinations, improve information accuracy, and enhance readers’ safety. Compared with scientificity, recency and comprehensiveness, the scores of understandability, conciseness, and interest were significantly reduced. The higher scores of scientificity, recency, and comprehensiveness might benefit from the big data foundation of ChatGpT 4.0. The understandability, conciseness, and interest were mainly related to the writing style of the texts, which needed to be further improved. This is consistent with the result reported in previous study regarding the “humanistic” aspects of generative AI content (23). In light of the issues identified in both previous studies and this current research regarding the generated AI content, we have also proposed a practical solution that can be implemented at present, namely manual review and revision. In this study, a manual review process was introduced for the content generated by ChatGPT 4.0, and the results showed that manual revision could significantly improve the understandability and conciseness of the content. This might be related to the fact that the revisers further clarify the logic and remove duplicate content. It illustrated the importance of human review in the current application of generative AI models. However, there was no significant improvement in terms of interest after revision, indicating that this feature needed to be improved in a more targeted manner, and attempts such as content style setting and AI images might be applied in future researches. At present, the medical science popularization of generative AI models assisted by manual review may be able to improve the quality of medical science popularization content and alleviate the pressure of medical staff.
This study is only a preliminary attempt of generative AI model in urology-related medical science popularization, and has certain limitations. First of all, this was a small sample research which at a single institution and might introduce institutional bias. Secondly, only ChatGPT 4.0 was selected in this study, and with the continuous development of generative Al technology, more and more models can be used for further trial and verification. And the big data foundation of generative AI models is constantly updated over time, the content generated by the same topics will change. Thirdly, although in this study the revised and un-revised AI contents were compared, high-quality human-generated contents could be compared with the generated AI contents in the future. Fourth, the evaluation of AI-generated content by professionals in this study can be extended to the general population for further verification. In addition, we should also recognize that generative AI is a technology that has just emerged in recent years, and there is still a great deal of room for further advancement in AI technology. For example, the reliability of the source of generated content in generative AI cannot be fully guaranteed, and there will be some inaccurate and biased information in the generated content, and the standardization of its application needs to be improved (28-30). In terms of personal privacy and data security, the training of generative AI involves the collection of users’ personal information, and there are information security risks such as excessive collection and trafficking of health privacy generated in the process of dialogue, and anonymization may not completely eliminate the identifiability, so that sensitive information such as users’ health status and medical records may be exposed (31-33). Therefore, it is necessary to establish and improve an ethical review and supervision mechanism in the future to ensure the security of personal information and maintain the privacy security of users. With the continuous advancement of generative Al technology, it may be able to provide more assistance for medical science popularization in the future.
Conclusions
Based on the results of this study, generative AI tools such as ChatGPT may assist urologists in popularizing knowledge of urological diseases, and the generated content needs to be manually reviewed to ensure the accuracy and readability of the content. With the continuous development of AI technology, it is believed that generative AI models can be more widely and directly used in medical science popularization.
Acknowledgments
None.
Footnote
Data Sharing Statement: Available at https://tau.amegroups.com/article/view/10.21037/tau-2025-500/dss
Peer Review File: Available at https://tau.amegroups.com/article/view/10.21037/tau-2025-500/prf
Funding: This work was supported by
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://tau.amegroups.com/article/view/10.21037/tau-2025-500/coif). The authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Xiao L, Min H, Wu Y, et al. Public’s preferences for health science popularization short videos in China: a discrete choice experiment. Front Public Health 2023;11:1160629. [Crossref] [PubMed]
- Li Y, Lv X, Liang J, et al. The development and progress of health literacy in China. Front Public Health 2022;10:1034907. [Crossref] [PubMed]
- Tzelves L, Kapriniotis K, Feretzakis G, et al. ChatGPT in Clinical Medicine, Urology and Academia: A Review. Arch Esp Urol 2024;77:708-17. [Crossref] [PubMed]
- Lee P, Bubeck S, Petro J. Benefits, Limits, and Risks of GPT-4 as an AI Chatbot for Medicine. N Engl J Med 2023;388:1233-9. [Crossref] [PubMed]
- Bindra S, Jain R. Artificial intelligence in medical science: a review. Ir J Med Sci 2024;193:1419-29. [Crossref] [PubMed]
- Nagi F, Salih R, Alzubaidi M, et al. Applications of Artificial Intelligence (AI) in Medical Education: A Scoping Review. Stud Health Technol Inform 2023;305:648-51. [Crossref] [PubMed]
- Dave T, Athaluri SA, Singh S. ChatGPT in medicine: an overview of its applications, advantages, limitations, future prospects, and ethical considerations. Front Artif Intell 2023;6:1169595. [Crossref] [PubMed]
- Belge Bilgin G, Bilgin C, Burkett BJ, et al. Theranostics and artificial intelligence: new frontiers in personalized medicine. Theranostics 2024;14:2367-78. [Crossref] [PubMed]
- Telenti A, Auli M, Hie BL, et al. Large language models for science and medicine. Eur J Clin Invest 2024;54:e14183. [Crossref] [PubMed]
- Sarumi OA, Heider D. Large language models and their applications in bioinformatics. Comput Struct Biotechnol J 2024;23:3498-505. [Crossref] [PubMed]
- Qin R, Liu B, Zhang M, et al. Investigation and analysis on health science popularization cognition of medical staff in Tianjin. Chin J Med Edu 2022;42:783-6.
- Zhang C, Xu J, Tang R, et al. Novel research and future prospects of artificial intelligence in cancer diagnosis and treatment. J Hematol Oncol 2023;16:114. [Crossref] [PubMed]
- Pushpanathan K, Lim ZW, Er Yew SM, et al. Popular large language model chatbots’ accuracy, comprehensiveness, and self-awareness in answering ocular symptom queries. iScience 2023;26:108163. [Crossref] [PubMed]
- Ozduran E, Hancı V, Erkin Y, et al. Assessing the readability, quality and reliability of responses produced by ChatGPT, Gemini, and Perplexity regarding most frequently asked keywords about low back pain. PeerJ 2025;13:e18847. [Crossref] [PubMed]
- Ponzo V, Goitre I, Favaro E, et al. Is ChatGPT an Effective Tool for Providing Dietary Advice? Nutrients 2024;16:469. [Crossref] [PubMed]
- Aguirre A, Hilsabeck R, Smith T, et al. Assessing the Quality of ChatGPT Responses to Dementia Caregivers’ Questions: Qualitative Analysis. JMIR Aging 2024;7:e53019. [Crossref] [PubMed]
- Carlson JA, Cheng RZ, Lange A, et al. Accuracy and Readability of Artificial Intelligence Chatbot Responses to Vasectomy-Related Questions: Public Beware. Cureus 2024;16:e67996. [Crossref] [PubMed]
- Troian M, Lovadina S, Ravasin A, et al. An Assessment of ChatGPT’s Responses to Common Patient Questions About Lung Cancer Surgery: A Preliminary Clinical Evaluation of Accuracy and Relevance. J Clin Med 2025;14:1676. [Crossref] [PubMed]
- Razdan S, Siegal AR, Brewer Y, et al. Assessing ChatGPT’s ability to answer questions pertaining to erectile dysfunction: can our patients trust it? Int J Impot Res 2024;36:734-40. [Crossref] [PubMed]
- Rotem R, Zamstein O, Rottenstreich M, et al. The future of patient education: A study on AI-driven responses to urinary incontinence inquiries. Int J Gynaecol Obstet 2024;167:1004-9. [Crossref] [PubMed]
- Thia I, Saluja M. ChatGPT: Is This Patient Education Tool for Urological Malignancies Readable for the General Population? Res Rep Urol 2024;16:31-7. [Crossref] [PubMed]
- Cakir H, Caglar U, Yildiz O, et al. Evaluating the performance of ChatGPT in answering questions related to urolithiasis. Int Urol Nephrol 2024;56:17-21. [Crossref] [PubMed]
- Song H, Xia Y, Luo Z, et al. Evaluating the Performance of Different Large Language Models on Health Consultation and Patient Education in Urolithiasis. J Med Syst 2023;47:125. [Crossref] [PubMed]
- Gibson D, Jackson S, Shanmugasundaram R, et al. Evaluating the Efficacy of ChatGPT as a Patient Education Tool in Prostate Cancer: Multimetric Assessment. J Med Internet Res 2024;26:e55939. [Crossref] [PubMed]
- Belge Bilgin G, Bilgin C, Childs DS, et al. Performance of ChatGPT-4 and Bard chatbots in responding to common patient questions on prostate cancer (177)Lu-PSMA-617 therapy. Front Oncol 2024;14:1386718. [Crossref] [PubMed]
- Sonmez SC, Sevgi M, Antaki F, et al. Generative artificial intelligence in ophthalmology: current innovations, future applications and challenges. Br J Ophthalmol 2024;108:1335-40. [Crossref] [PubMed]
- Helvacioglu-Yigit D, Demirturk H, Ali K, et al. Evaluating artificial intelligence chatbots for patient education in oral and maxillofacial radiology. Oral Surg Oral Med Oral Pathol Oral Radiol 2025;139:750-9. [Crossref] [PubMed]
- Sallam M. ChatGPT Utility in Healthcare Education, Research, and Practice: Systematic Review on the Promising Perspectives and Valid Concerns. Healthcare (Basel) 2023;11:887. [Crossref] [PubMed]
- Thapa S, Adhikari S. ChatGPT, Bard, and Large Language Models for Biomedical Research: Opportunities and Pitfalls. Ann Biomed Eng 2023;51:2647-51. [Crossref] [PubMed]
- Benítez TM, Xu Y, Boudreau JD, et al. Harnessing the potential of large language models in medical education: promise and pitfalls. J Am Med Inform Assoc 2024;31:776-83. [Crossref] [PubMed]
- Jeyaraman M, Balaji S, Jeyaraman N, et al. Unraveling the Ethical Enigma: Artificial Intelligence in Healthcare. Cureus 2023;15:e43262. [Crossref] [PubMed]
- Ong JCL, Chang SY, William W, et al. Ethical and regulatory challenges of large language models in medicine. Lancet Digit Health 2024;6:e428-32. [Crossref] [PubMed]
- Akinci D’Antonoli T, Stanzione A, Bluethgen C, et al. Large language models in radiology: fundamentals, applications, ethical considerations, risks, and future directions. Diagn Interv Radiol 2024;30:80-90. [Crossref] [PubMed]

