Apparent diffusion coefficient-based single-sequence radiomics integrated with clinical variables and Prostate Imaging Reporting and Data System for predicting clinically significant prostate cancer
Highlight box
Key findings
• A parsimonious five-feature apparent diffusion coefficient (ADC) radiomics model achieved a test-cohort area under the curve (AUC) of 0.916 for predicting clinically significant prostate cancer (csPCa).
• Adding the Rad-score to clinical variables and Prostate Imaging Reporting and Data System (PI-RADS) increased the AUC only from 0.944 to 0.951, without a statistically significant difference (P=0.71).
What is known and what is new?
• PI-RADS and clinical variables are widely used for csPCa risk assessment, while magnetic resonance imaging radiomics has emerged as a promising quantitative imaging biomarker. The true incremental value of radiomics beyond established clinicoradiological assessment remains uncertain.
• This study used five hierarchical models to distinguish the standalone performance of ADC radiomics from its incremental value beyond established clinicoradiological predictors.
What is the implication, and what should change now?
• High standalone radiomics performance should not automatically be interpreted as meaningful incremental clinical benefit.
• Future studies should identify patient subgroups or simplified diagnostic settings in which ADC radiomics may provide clinically relevant information.
Introduction
Prostate cancer remains one of the most common malignancies in men, with an estimated 400,000 deaths annually worldwide (1). However, prostate-specific antigen (PSA) testing has limited specificity, and transrectal ultrasound-guided biopsy carries procedural risks and may miss clinically relevant lesions (2,3). Accordingly, the modern diagnostic goal is to accurately identify patients with clinically significant prostate cancer (csPCa) while limiting overdiagnosis and unnecessary biopsy (4). Current international guidelines reflect this shift, recommending magnetic resonance imaging (MRI)-based assessment before biopsy in men with suspected prostate cancer (5). Within this diagnostic pathway, the Prostate Imaging Reporting and Data System (PI-RADS) has become the dominant framework for standardized interpretation of prostate MRI. PI-RADS categorizes prostate lesions from 1 to 5, with higher scores indicating a greater likelihood of csPCa (6). Nevertheless, a systematic review and meta-analysis including 36,366 patients suggested that PI-RADS may still miss a proportion of cases (7).
Radiomics, which extracts high-dimensional quantitative imaging features, has emerged as a potential adjunct to conventional radiological assessment (8). In this context, the present study focuses on apparent diffusion coefficient (ADC)-based single-sequence radiomics. ADC reflects tissue diffusion restriction and is closely related to tumor cellularity (9). As a widely used quantitative biomarker in prostate MRI, ADC has demonstrated considerable diagnostic value (10,11).
Therefore, this study aimed to develop and internally validate a model integrating clinical variables, PI-RADS score, and ADC-based single-sequence radiomics for predicting csPCa. By comparing five hierarchical models, we further aimed to distinguish the standalone predictive performance of ADC radiomics from its incremental contribution beyond clinical and PI-RADS assessment. We present this article in accordance with the TRIPOD+AI reporting checklist (12) (available at https://tau.amegroups.com/article/view/10.21037/tau-2026-0546/rc).
Methods
Study design and patient population
This retrospective single-center study included consecutive patients who underwent prostate MRI at a tertiary hospital (The Sixth Hospital of Wuhan, Affiliated Hospital of Jianghan University) between January 2020 and February 2026 and subsequently underwent histopathological assessment. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Institutional Review Board of The Sixth Hospital of Wuhan, Affiliated Hospital of Jianghan University (No. WHSHIRB-K-2026015). The requirement for written informed consent was waived because of the retrospective nature of the study. All data were anonymized and de-identified before analysis.
Patients were included if they met the following criteria: (I) underwent prostate MRI during the study period; (II) had ADC maps available for radiomics analysis; and (III) had complete clinical, MRI, and histopathological data for endpoint determination and model construction. Patients were excluded for the following reasons: prior prostate cancer-related treatment before MRI (n=24), unavailable original ADC images from the picture archiving and communication system (PACS) (n=9), poor image quality precluding lesion segmentation (n=13), or indeterminate pathological diagnosis (n=6).
Among 353 initially screened patients, 52 were excluded, and 301 were finally included. Patient selection and cohort allocation are shown in Figure 1A.
Clinical, radiological, and pathological assessment
Clinical and laboratory data were retrospectively extracted from the electronic medical record system and laboratory information system. All PSA-related variables were obtained from the most recent laboratory test performed before MRI examination and before any prostate cancer-related treatment. Candidate clinical predictors included age, body mass index, total prostate-specific antigen, free prostate-specific antigen, FPSA/TPSA ratio, smoking history, alcohol use, hypertension, diabetes mellitus, and family history of prostate cancer. Radiological and radiomics predictors included PI-RADS score and the ADC-derived Rad-score.
PI-RADS scores were reassessed independently by two radiologists with more than 5 years of experience in prostate MRI. Both radiologists were blinded to clinical and pathological information, and scoring was performed according to PI-RADS version 2.1. The same PI-RADS version 2.1 criteria were applied consistently to all patients irrespective of the examination date. The index lesion was identified by integrating information from available MRI sequences. Disagreements were resolved by consensus with a third senior radiologist.
Histopathological findings served as the reference standard. The primary endpoint was csPCa, defined as pathologically confirmed prostate adenocarcinoma with a Gleason score ≥7, corresponding to International Society of Urological Pathology (ISUP) Grade Group ≥2. Patients with negative pathological findings or Gleason score 3+3=6 prostate cancer (ISUP Grade Group 1) were classified as non-csPCa (13). According to this definition, the final cohort included 82 patients with csPCa and 219 patients with non-csPCa.
After endpoint determination, patients were randomly divided into training and held-out internal test cohorts at a 7:3 ratio using stratified random sampling to maintain similar csPCa and non-csPCa distributions. Baseline clinical and MRI characteristics of the csPCa and non-csPCa groups are summarized in Table 1.
Table 1
| Characteristic | Overall (n=301) | Non-csPCa (n=219) | csPCa (n=82) | P value |
|---|---|---|---|---|
| Age (years) | 73.0±8.2 | 71.9±8.2 | 75.9±7.7 | <0.001 |
| BMI (kg/m2) | 23.6 [21.9–25.2] | 23.7 [22.0–25.4] | 23.1 [21.4–24.2] | 0.04 |
| TPSA (ng/mL) | 12.76 [8.80–30.59] | 10.84 [7.81–17.73] | 69.80 [14.51–124.10] | <0.001 |
| FPSA (ng/mL) | 2.18 [1.51–5.00] | 2.01 [1.40–3.27] | 7.34 [2.13–19.84] | <0.001 |
| FPSA/TPSA ratio | 0.171 [0.131–0.218] | 0.184 [0.143–0.233] | 0.137 [0.109–0.184] | <0.001 |
| Smoking history | 13 (4.3) | 12 (5.5) | 1 (1.2) | 0.20 |
| Alcohol use | 8 (2.7) | 7 (3.2) | 1 (1.2) | 0.69 |
| Hypertension | 141 (46.8) | 101 (46.1) | 40 (48.8) | 0.68 |
| Diabetes mellitus | 169 (56.1) | 123 (56.2) | 46 (56.1) | 0.99 |
| Family history of prostate cancer | 21 (7.0) | 9 (4.1) | 12 (14.6) | 0.001 |
| PI-RADS score | <0.001 | |||
| PI-RADS 1–2 | 67 (22.3) | 63 (28.8) | 4 (4.9) | |
| PI-RADS 3 | 115 (38.2) | 105 (47.9) | 10 (12.2) | |
| PI-RADS 4 | 53 (17.6) | 33 (15.1) | 20 (24.4) | |
| PI-RADS 5 | 66 (21.9) | 18 (8.2) | 48 (58.5) |
Data are presented as mean ± standard deviation, median [interquartile range], or number (percentage), as appropriate. csPCa was defined as pathologically confirmed prostate adenocarcinoma with a Gleason score ≥7 (ISUP grade group ≥2). The P value for PI-RADS score indicates the overall distributional difference between groups. BMI, body mass index; csPCa, clinically significant prostate cancer; FPSA, free prostate-specific antigen; ISUP, International Society of Urological Pathology; MRI, magnetic resonance imaging; PI-RADS, Prostate Imaging Reporting and Data System; TPSA, total prostate-specific antigen.
MRI acquisition, ADC lesion segmentation, and radiomics feature extraction
All patients underwent MRI using a standard prostate MRI protocol on a MAGNETOM Skyra 3.0T (Siemens, Germany) with a body phased-array coil. The MRI hardware remained unchanged throughout the study period. Thus, all imaging data were obtained at a single institution using one scanner model from a single vendor. The scanner software version remained unchanged throughout the study period. The core ADC acquisition protocol was generally consistent, although minor parameter variations occurred over time. ComBat harmonization was considered but was not applied because the dataset contained no distinct scanner-, vendor-, or institution-based batches. ADC was selected as the only radiomics sequence because it is a routinely available quantitative diffusion parameter in prostate MRI and may reduce methodological heterogeneity associated with multiparametric MRI variability (11). Detailed MRI acquisition parameters are provided in Table S1.
Manual region of interest (ROI) segmentation was performed slice by slice on ADC images using 3D Slicer software (version 5.8.1) by a radiologist blinded to pathological results. The ROI was defined as the MRI-visible index lesion, namely the most suspicious intraprostatic lesion. If multiple suspicious lesions were present, the lesion with the highest PI-RADS score was selected; if scores were identical, the lesion with the largest diameter was selected. Because both csPCa assessment and PI-RADS are lesion-centered, restricting radiomics extraction to the index lesion allowed the Rad-score and PI-RADS to evaluate the same target. Representative ADC lesion ROI delineation and the radiomics workflow are shown in Figure 1B.
Image preprocessing and radiomics feature extraction were performed in accordance with the recommendations of the Image Biomarker Standardisation Initiative (IBSI) (14). Before feature extraction, ADC images and ROI masks were checked for spatial alignment and resampled to 1.0×1.0×1.0 mm3 using B-spline interpolation for images and nearest-neighbor interpolation for masks. Intensity normalization was performed with a scale of 100, and gray-level discretization used a fixed bin width of 5. Pre-cropping was performed with a pad distance of 5 voxels; Laplacian of Gaussian (LoG) filters were applied with sigma values of 1.0, 2.0, 3.0, 4.0, and 5.0; and the texture matrix distance parameter was set to 1. All preprocessing settings were identical in both cohorts.
Radiomics features were extracted using PyRadiomics in Python. A total of 2,153 ADC-based features were extracted, including first-order, shape, and texture features derived from the gray-level co-occurrence matrix, gray-level dependence matrix, gray-level run-length matrix, gray-level size-zone matrix, and neighboring gray-tone difference matrix.
Feature selection and model development
For segmentation reproducibility assessment, 50 patients were selected by stratified random sampling. The first radiologist repeated segmentation after at least 4 weeks to assess intraobserver agreement, and a second radiologist independently segmented the same cases to assess interobserver agreement. Radiomics features were extracted from repeated segmentations, and intraclass correlation coefficients (ICCs) were calculated. Features with both intraobserver and interobserver ICCs ≥0.75 were retained for model development (15). Manual lesion segmentation required approximately 5 minutes per patient.
Radiomics feature selection was performed only in the training cohort to avoid information leakage. Features with high missing rates or near-zero variance were first removed. Highly correlated features were then filtered; when two features had |r| >0.90, the feature more strongly associated with csPCa was retained. LASSO logistic regression was subsequently performed, with the regularization parameter selected using the one-standard-error rule to favor a more parsimonious model. Feature-selection stability was assessed through inner resampling, and a conservative complexity constraint was applied to obtain the final five-feature radiomics signature.
The Rad-score was calculated using the selected features and their coefficients:
where β₀ is the intercept, X₁ to Xₖ are the final selected ADC radiomics features, β₁ to βₖ are the corresponding LASSO-derived coefficients, and k is the number of selected features.
Model construction and validation
Five prespecified logistic regression models were constructed using csPCa as the primary outcome: a clinical model based on selected clinical variables, a PI-RADS model based on PI-RADS score alone, an ADC radiomics model based on Rad-score alone, a clinicoradiological model combining clinical variables and PI-RADS score, and a combined model integrating clinical variables, PI-RADS score, and Rad-score.
All variable selection, radiomics feature selection, Rad-score construction, and model development were performed only in the training cohort. The test cohort was not involved in model training or selection and was used only for held-out internal evaluation of the locked models. Continuous clinical predictors were retained as continuous variables when entered into logistic regression models. PI-RADS score was treated as an ordinal variable. Binary clinical variables were coded as 0/1. Radiomics features were standardized using parameters derived from the training cohort before LASSO logistic regression.
To evaluate optimism arising from radiomics feature selection and model tuning, full radiomics-pipeline repeated nested cross-validation was performed within the training cohort. The outer validation procedure consisted of two repetitions of five-fold stratified cross-validation, with three-fold inner cross-validation for model tuning. Within each outer training fold, radiomics preprocessing, near-zero variance filtering, correlation filtering, LASSO tuning, feature selection, Rad-score construction, and model fitting were repeated independently. Each outer validation fold was used only for performance evaluation and was not involved in the corresponding feature-selection or model-tuning procedures.
Logistic regression was selected because of its interpretability and suitability for a moderate-sized prediction modeling study. A nomogram was constructed for the final combined model to provide individualized csPCa risk estimation. Model performance was evaluated using the area under the curve (AUC), sensitivity, specificity, accuracy, positive predictive value, negative predictive value, and F1 score. For each model, a classification threshold was selected by maximizing the Youden index on nested-cross-validation out-of-fold predictions generated exclusively within the training cohort. The threshold was then locked before application to the test cohort, and no test-cohort information was used for threshold determination.
Statistical analysis
Baseline clinical and MRI characteristics were compared between the non-csPCa and csPCa groups. Normally distributed continuous variables are presented as mean ± standard deviation and were compared using the independent-samples t-test. Non-normally distributed continuous variables are presented as median (interquartile range) and were compared using the Mann–Whitney U-test. Categorical variables are presented as frequencies (percentages) and were compared using the χ2 test or Fisher’s exact test, as appropriate. The P value for PI-RADS score represents the overall distributional difference between groups. Patients with incomplete clinical, MRI, or pathological data required for endpoint determination and model construction were excluded before model development. No missing values were present in the final clinical predictors, PI-RADS score, outcome, or Rad-score used for model development and validation; therefore, no imputation was performed.
Model discrimination was evaluated using AUCs with 95% confidence intervals (CIs), and AUCs were compared using the DeLong test. Calibration was assessed using calibration curves and the Brier score. Clinical utility was evaluated using decision curve analysis by comparing net benefit across threshold probabilities. AUC, Brier score, calibration, and decision curve analysis were used as the primary model-evaluation measures because they did not depend on a single classification cutoff. Threshold-dependent measures, including sensitivity, specificity, accuracy, balanced accuracy, positive predictive value, negative predictive value, F1 score, and Matthews correlation coefficient (MCC), were reported as secondary descriptive results at the locked training-derived threshold. The corresponding 95% confidence intervals were calculated for sensitivity, specificity, positive predictive value, and negative predictive value. To facilitate clinical interpretation, the numbers of biopsies potentially avoided relative to a biopsy-all strategy and csPCa cases missed were calculated from the test-cohort confusion matrix at the locked combined-model threshold. Paired bootstrap resampling was additionally used to estimate 95% confidence intervals for differences in AUC and Brier score between the combined model and each comparator. Because the test cohort contained only 25 csPCa events, pairwise model comparisons were considered exploratory and were interpreted according to the magnitude and confidence interval of each difference rather than P values alone.
All statistical tests were two-sided, and P<0.05 was considered statistically significant. Analyses were performed using R software (version 4.4.2) and Python software (version 3.13.5).
Results
Patient characteristics
After screening 353 eligible patients, 301 were included in the final analysis (Figure 1). Of these, 82 patients were classified as having csPCa and 219 as having non-csPCa. Using stratified random sampling, 210 patients were assigned to the training cohort, including 57 with csPCa and 153 with non-csPCa, and 91 patients were assigned to the test cohort, including 25 with csPCa and 66 with non-csPCa.
Baseline clinical and MRI characteristics are shown in Table 1. Compared with the non-csPCa group, patients in the csPCa group were older (75.9±7.7 vs. 71.9±8.2 years, P<0.001), had higher TPSA levels [69.80 (14.51–124.10) vs. 10.84 (7.81–17.73) ng/mL, P<0.001], higher FPSA levels [7.34 (2.13–19.84) vs. 2.01 (1.40–3.27) ng/mL, P<0.001], and a lower FPSA/TPSA ratio [0.137 (0.109–0.184) vs. 0.184 (0.143–0.233), P<0.001]. The proportion of patients with a family history of prostate cancer was also higher in the csPCa group (14.6% vs. 4.1%, P=0.001). The distribution of PI-RADS scores differed significantly between the two groups (P<0.001), with a markedly higher proportion of PI-RADS 5 lesions in the csPCa group (58.5% vs. 8.2%). No significant differences were observed between the two groups in smoking history, alcohol use, hypertension, or diabetes mellitus. Baseline clinical and MRI characteristics of the training and test cohorts are presented in Table S2.
Radiomics feature selection and Rad-score construction
A total of 2,153 radiomics features were extracted from ADC images. All feature selection procedures were performed only in the training cohort to avoid information leakage. Among all 2,153 extracted radiomics features, the mean intraobserver ICC was 0.92 (range, 0.71–0.99), and the mean interobserver ICC was 0.90 (range, 0.69–0.98). A total of 1,657 features met both the intraobserver and interobserver ICC thresholds of ≥0.75 and were retained, whereas 496 features (23.0%) were discarded because at least one of the two ICC values was below 0.75. In the training cohort, 1,657 features remained after missing-rate filtering, 1,649 remained after near-zero variance filtering, and 511 remained after correlation-based redundancy filtering. Application of the one-standard-error LASSO rule and inner-cross-validation stability assessment retained 11 candidate features. After applying the sample-size-based complexity constraint, five ADC radiomics features were retained for construction of the final Rad-score (Figure 2A,2B).
The Rad-score was calculated using the LASSO-selected features and their corresponding coefficients. The specific intercept, selected feature names, and regression coefficients are provided in Table S3. The complete feature filtering process is shown in Table S4.
The Rad-score demonstrated good discrimination between groups in both the training and test cohorts, with overall higher Rad-scores in patients with csPCa than in those with non-csPCa (Figure 2C). These findings supported the use of Rad-score as a radiomics predictor in subsequent model construction.
Predictive performance of the five models
The discriminatory performance of the five prediction models in the training and test cohorts is shown in Figure 3 and Table 2. Overall, the combined model achieved the numerically highest AUC in both cohorts. In the training cohort, the combined model achieved an AUC of 0.946 (95% CI, 0.919–0.973), followed by the ADC radiomics model [AUC, 0.913 (95% CI, 0.875–0.947)], clinicoradiological model [AUC, 0.853 (95% CI, 0.789–0.916)], PI-RADS model [AUC, 0.811 (95% CI, 0.741–0.873)], and clinical model [AUC, 0.794 (95% CI, 0.718–0.867)].
Table 2
| Model | Cohort | AUC (95% CI) | Sensitivity (95% CI) | Specificity (95% CI) | Accuracy | PPV (95% CI) | NPV (95% CI) | F1 score | Balanced accuracy | MCC | Brier score | Calibration intercept | Calibration slope | Threshold |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Clinical model | Training cohort | 0.794 (0.718, 0.867) | 0.702 (0.573, 0.805) | 0.850 (0.785, 0.898) | 0.810 | 0.635 (0.511, 0.743) | 0.884 (0.823, 0.927) | 0.667 | 0.776 | 0.535 | 0.143 | 0.033 | 1.037 | 0.323 |
| Clinical model | Test cohort | 0.904 (0.832, 0.963) | 0.760 (0.566, 0.885) | 0.939 (0.854, 0.976) | 0.890 | 0.826 (0.629, 0.930) | 0.912 (0.821, 0.959) | 0.792 | 0.850 | 0.718 | 0.115 | 0.680 | 1.880 | 0.323 |
| PI-RADS model | Training cohort | 0.811 (0.741, 0.873) | 0.825 (0.706, 0.902) | 0.739 (0.664, 0.802) | 0.762 | 0.540 (0.436, 0.641) | 0.919 (0.857, 0.955) | 0.653 | 0.782 | 0.508 | 0.138 | 0.030 | 1.043 | 0.344 |
| PI-RADS model | Test cohort | 0.894 (0.819, 0.953) | 0.840 (0.653, 0.936) | 0.833 (0.726, 0.904) | 0.835 | 0.656 (0.483, 0.796) | 0.932 (0.838, 0.973) | 0.737 | 0.837 | 0.629 | 0.112 | 0.535 | 1.476 | 0.344 |
| ADC radiomics model | Training cohort | 0.913 (0.875, 0.947) | 0.772 (0.648, 0.862) | 0.850 (0.785, 0.898) | 0.829 | 0.657 (0.537, 0.759) | 0.909 (0.851, 0.946) | 0.710 | 0.811 | 0.593 | 0.108 | 0.057 | 1.092 | 0.334 |
| ADC radiomics model | Test cohort | 0.916 (0.856, 0.965) | 0.800 (0.609, 0.911) | 0.803 (0.692, 0.881) | 0.802 | 0.606 (0.437, 0.753) | 0.914 (0.814, 0.963) | 0.690 | 0.802 | 0.560 | 0.113 | −0.037 | 1.164 | 0.334 |
| Clinicoradiological model | Training cohort | 0.853 (0.789, 0.916) | 0.632 (0.502, 0.745) | 0.895 (0.837, 0.935) | 0.824 | 0.692 (0.557, 0.801) | 0.867 (0.805, 0.911) | 0.661 | 0.764 | 0.543 | 0.124 | 0.026 | 1.038 | 0.442 |
| Clinicoradiological model | Test cohort | 0.944 (0.892, 0.982) | 0.720 (0.524, 0.857) | 0.924 (0.835, 0.967) | 0.868 | 0.783 (0.581, 0.903) | 0.897 (0.802, 0.949) | 0.750 | 0.822 | 0.662 | 0.096 | 0.444 | 1.522 | 0.442 |
| Combined model | Training cohort | 0.946 (0.919, 0.973) | 0.947 (0.856, 0.982) | 0.830 (0.763, 0.881) | 0.862 | 0.675 (0.566, 0.768) | 0.977 (0.934, 0.992) | 0.788 | 0.889 | 0.712 | 0.086 | 0.049 | 1.114 | 0.201 |
| Combined model | Test cohort | 0.951 (0.905, 0.986) | 0.960 (0.805, 0.993) | 0.848 (0.743, 0.916) | 0.879 | 0.706 (0.538, 0.832) | 0.982 (0.907, 0.997) | 0.814 | 0.904 | 0.746 | 0.078 | −0.113 | 1.008 | 0.201 |
Thresholds were selected using the Youden index from nested cross-validation out-of-fold predictions in the training cohort and locked before application to the test cohort. Threshold-dependent metrics are reported as secondary results. AUC, area under the receiver operating characteristic curve; CI, confidence interval; F1 score, harmonic mean of precision and recall; MCC, Matthews correlation coefficient; NPV, negative predictive value; PPV, positive predictive value.
In the test cohort, the combined model also achieved the highest AUC [0.951 (95% CI, 0.905–0.986)], followed by the clinicoradiological model [AUC, 0.944 (95% CI, 0.892–0.982)], ADC radiomics model [AUC, 0.916 (95% CI, 0.856–0.965)], clinical model [AUC, 0.904 (95% CI, 0.832–0.963)], and PI-RADS model [AUC, 0.894 (95% CI, 0.819–0.953)]. Using the classification threshold derived from out-of-fold predictions in the training cohort and locked before test-cohort evaluation, the sensitivity, specificity, accuracy, positive predictive value, negative predictive value, and F1 score of the combined model in the test cohort were 0.960, 0.848, 0.879, 0.706, 0.982, and 0.814, respectively. These threshold-dependent measures were reported as secondary operating-point estimates and were not used as the primary basis for model comparison.
Full radiomics-pipeline repeated nested cross-validation yielded more conservative internal-validation estimates. The mean AUCs were 0.764 for the clinical model, 0.790 for the PI-RADS model, 0.837 for the ADC radiomics model, 0.818 for the clinicoradiological model, and 0.890 for the combined model. Detailed AUC and Brier score results are provided in Table S5. The lower nested-cross-validation AUCs compared with the corresponding apparent training-cohort AUCs indicate that some model-development optimism remained despite the reduction in radiomics model complexity.
Pairwise comparisons of AUC and Brier score between the combined model and the comparator models are presented in Table 3. In the test cohort, the combined model showed numerically higher AUCs than the clinical, PI-RADS, ADC radiomics, and clinicoradiological models, with AUC differences of 0.047, 0.057, 0.035, and 0.007, respectively. However, none of these differences reached statistical significance in the DeLong tests (P=0.15, P=0.08, P=0.052, and P=0.71, respectively). The paired bootstrap confidence intervals for the AUC differences included zero for comparisons with the clinical, PI-RADS, and clinicoradiological models. For comparison with the ADC radiomics model, the paired bootstrap 95% confidence interval was 0.002 to 0.075, whereas the DeLong test yielded a borderline nonsignificant result (P=0.052). Given the small number of outcome events and the exploratory nature of these comparisons, this finding was interpreted conservatively and was not considered evidence of superiority. The combined model also had a lower Brier score than the ADC radiomics model, with a difference of −0.035 (95% CI, −0.057 to −0.013), whereas the confidence intervals for the other Brier score differences included zero. Because only 25 csPCa events were included in the test cohort, all pairwise comparisons had limited precision and statistical power. Confusion matrix results for all models at the locked threshold are provided in Table S6.
Table 3
| Comparison | Combined model AUC | Comparator AUC | AUC difference (95% CI) | DeLong P value | Brier-score difference (95% CI) |
|---|---|---|---|---|---|
| Combined model vs. clinical model | 0.951 | 0.904 | 0.047 (−0.009, 0.112) | 0.15 | −0.037 (−0.076, 0.003) |
| Combined model vs. PI-RADS model | 0.951 | 0.894 | 0.057 (−0.001, 0.128) | 0.08 | −0.034 (−0.073, 0.006) |
| Combined model vs. ADC radiomics model | 0.951 | 0.916 | 0.035 (0.002, 0.075) | 0.052 | −0.035 (−0.057, −0.013) |
| Combined model vs. clinicoradiological model | 0.951 | 0.944 | 0.007 (−0.027, 0.043) | 0.71 | −0.017 (−0.046, 0.012) |
AUC and Brier score differences were calculated as the combined model minus the comparator model. DeLong tests were used for pairwise AUC comparisons, whereas 95% confidence intervals for AUC and Brier score differences were obtained by paired bootstrap resampling of fixed test-cohort predictions. Because only 25 csPCa events were included in the test cohort, all pairwise comparisons were considered exploratory. A nonsignificant result does not demonstrate equivalence. DeLong tests compare the AUC of the combined model with each comparator in the test cohort. The comparison with the clinicoradiological model assesses the incremental value of Rad-score beyond clinical variables and PI-RADS. ADC, apparent diffusion coefficient; AUC, area under the curve; CI, confidence interval; csPCa, clinically significant prostate cancer; PI-RADS, Prostate Imaging Reporting and Data System; Rad-score, radiomics score.
Nomogram and calibration performance
A nomogram based on the combined model was constructed using Rad-score, PI-RADS score, age, log10-transformed TPSA, FPSA/TPSA ratio, and family history of prostate cancer to provide individualized prediction of csPCa risk (Figure 4A). Given the markedly right-skewed distribution of TPSA, log10(TPSA) was used in the nomogram to reduce the influence of extreme values and improve graphical readability. The Rad-score had the widest point range in the nomogram. However, nomogram point allocation is determined jointly by the regression coefficient and the observed range of each predictor and should not be interpreted as a direct measure of incremental predictive value. By summing the points assigned to each predictor, the total score for each patient could be obtained and mapped to an individualized predicted probability of csPCa.
Calibration curves showed generally acceptable agreement between predicted and observed probabilities across the models (Figure 4B,4C). In the training cohort, the combined model had the lowest Brier score (0.086). In the test cohort, the Brier score of the combined model was 0.078, which was lower than that of the clinicoradiological model (0.096). These results indicated lower overall probabilistic prediction error for the combined model in both cohorts, although validation in an external cohort is still required. Overall, the nomogram provided an intuitive tool for individualized csPCa risk estimation.
Clinical utility based on decision curve analysis
Decision curve analysis was performed to further evaluate the clinical net benefit of each model across different threshold probabilities (Figure 5). In the training cohort, the combined model achieved the highest overall net benefit across a broad range of threshold probabilities. In particular, within the approximate threshold range of 0.20–0.80, its curve consistently remained above those of the clinical model, PI-RADS model, ADC radiomics model, and clinicoradiological model, suggesting improved clinical decision-making benefit after integrating clinical variables, PI-RADS score, and Rad-score.
In the test cohort, differences in net benefit among the models were smaller than those in the training cohort, but the combined model still showed favorable clinical utility in the low-to-intermediate threshold probability range. Specifically, within the approximate threshold range of 0.10–0.60, the combined model was generally above or close to the better-performing models and outperformed both the “treat all” and “treat none” strategies. As the threshold probability increased further, the curves of different models crossed and fluctuated, indicating that the net-benefit advantage of the combined model was not stable in the high-threshold range. At the locked combined-model threshold of 0.201, using the model as a biopsy-triage tool would have recommended biopsy for 34 of 91 patients and avoided biopsy in 57 patients (62.6%) relative to biopsying all patients. This strategy would have detected 24 of 25 csPCa cases and missed one case (4.0%). These estimates are descriptive and should not be interpreted as validation of 0.201 as a clinically recommended threshold.
Discussion
This study developed and internally validated a csPCa prediction model integrating clinical variables, PI-RADS score, and ADC-based single-sequence radiomics features. The combined model had the numerically highest AUC in both cohorts, with a test-cohort AUC of 0.951. At the locked training-derived threshold, the combined model achieved a sensitivity, specificity, and accuracy of 0.960, 0.848, and 0.879, respectively, in the test cohort. These threshold-dependent estimates describe model performance at one prespecified operating point, whereas the primary interpretation was based on AUC, calibration, Brier score, and decision curve analysis. However, its improvement over the clinicoradiological model (AUC, 0.944; ΔAUC, 0.007; P=0.71) was modest and not statistically significant. Given the limited number of csPCa events in the test cohort, this numerical difference should not be interpreted as evidence that the combined model was superior. The ADC radiomics model nevertheless showed strong standalone discrimination; however, its incremental contribution beyond clinical variables and PI-RADS was not demonstrated in this cohort.
Given the limited number of csPCa events available for model development, control of radiomics model complexity was essential. A parsimonious five-feature Rad-score was therefore constructed using reproducibility filtering, redundancy reduction, the one-standard-error LASSO rule, stability assessment, and conservative complexity control. The five prediction models were fitted separately rather than as a single high-dimensional model; therefore, constructing five models did not increase the predictor count within any individual model. With 57 csPCa events in the training cohort, the crude event-to-feature ratio was 11.4 for the five-feature radiomics signature. The combined model included six predictors—four clinical predictors, PI-RADS, and Rad-score—corresponding to 9.5 events per predictor. These ratios are descriptive and do not exclude overfitting because radiomics feature selection was data-driven. Although penalized regression can reduce model complexity, it does not by itself quantify the optimism introduced by high-dimensional feature selection and hyperparameter tuning. Full radiomics-pipeline repeated nested cross-validation was consequently performed, with all data-dependent radiomics procedures repeated within each outer training fold. The nested-cross-validation AUCs were lower than the apparent training-cohort AUCs, indicating that some model-development optimism remained despite the reduced signature. Accordingly, the reported performance should be interpreted as internal validation within a single-center dataset. Multicenter external validation is required to assess model transportability across different populations, scanners, and acquisition protocols.
The current focus of prostate cancer diagnosis and management has shifted from the simple detection of prostate cancer toward the selective identification of csPCa (16). Modern diagnostic pathways increasingly emphasize prebiopsy MRI to reduce unnecessary biopsy and overdiagnosis while maintaining sensitivity for aggressive disease (5,17). PI-RADS has become the dominant framework for standardized prostate MRI interpretation, and PI-RADS v2.1 has demonstrated strong diagnostic value in prospective studies and meta-analyses (6). Nevertheless, PI-RADS still has qualitative elements and remains reader-dependent, particularly in equivocal lesions and transition-zone abnormalities (18). In the present study, the PI-RADS-only model achieved a test-cohort AUC of 0.894, whereas the model combining clinical and radiological variables achieved an AUC of 0.944, indicating that PI-RADS combined with clinical variables already provided strong predictive information. This finding also explains why the additional improvement in overall AUC after adding Rad-score was relatively limited once clinical variables and PI-RADS had already been incorporated into the model.
ADC was selected as the only radiomics sequence in this study because it is a routinely available and biologically meaningful quantitative MRI parameter in prostate imaging. It reflects the degree of water diffusion restriction and is closely associated with tissue cellularity, glandular architecture, stromal composition, and tumor aggressiveness (10,19). Compared with single-value ADC measurements, ADC-based radiomics can quantify intralesional heterogeneity through intensity, shape, and texture features (20). In the present study, the ADC radiomics model achieved a test-cohort AUC of 0.916, outperforming the PI-RADS-only model (AUC =0.894). This supports the standalone discriminatory ability of ADC-derived quantitative features for csPCa prediction. Notably, the test-cohort AUC of the ADC radiomics model was close to that of the clinicoradiological model (0.944), suggesting that the ADC map alone can provide strong quantitative diagnostic information even without relying on additional sequences such as T2-weighted imaging (T2WI), dynamic contrast-enhanced imaging (DCE), or multisequence radiomics inputs. Compared with multiparametric radiomics, which may be affected by intersequence registration, sequence selection, acquisition variability, segmentation differences, and protocol heterogeneity, ADC-only radiomics provides a simpler and more transparent workflow (21,22). This design is consistent with the growing interest in simplified prostate MRI strategies and may facilitate implementation in retrospective clinical datasets (23).
Compared with previous similar studies, the combined model in the present study showed relatively high performance. For example, Hectors et al. developed an MRI radiomics machine-learning model for equivocal PI-RADS 3 lesions and reported a test-cohort AUC of only 0.76 (24). Zhao et al. investigated PI-RADS 3 transition-zone lesions and emphasized the value of quantitative ADC measurement and radiomics for managing equivocal lesions (25). These studies mainly focused on PI-RADS 3 lesions, whereas the present study included a consecutive clinical cohort covering PI-RADS categories 1–5 and defined non-csPCa as negative pathology or Gleason score 3+3=6. In addition, recent studies have supported the potential of MRI radiomics as an auxiliary tool for csPCa risk stratification and for improving radiologists’ PI-RADS-based diagnostic performance (26,27). Antolin et al. reported in a systematic review that radiomics showed good performance for predicting csPCa, but in the limited number of studies directly comparing radiomics with radiologist PI-RADS assessment, radiomics did not demonstrate a consistently significant advantage (26). Our findings are consistent with this evidence framework: the ADC radiomics model itself showed strong discriminatory ability (AUC =0.916), but its additional AUC gain became smaller once PI-RADS and clinical variables were already included in the model (AUC =0.951). These findings demonstrate that high standalone radiomics performance does not necessarily translate into significant incremental discrimination when strong clinicoradiological predictors are already available.
Deep learning and multicenter radiomics represent important developments in MRI-based prostate cancer assessment. Deep learning can learn multilevel imaging representations directly from biparametric or multiparametric MRI and may support automated lesion detection and segmentation. In a 12-center study of 8,786 men, Rodrigues et al. reported a prospective-validation AUC of 0.91 for a model integrating automatically segmented biparametric MRI (bpMRI) radiomics, PI-RADS, and clinical variables, compared with 0.85 for PI-RADS alone (28). Samaras et al. systematically benchmarked multicenter ADC radiomics pipelines and found that combining radiomics with the lesion-to-normal ADC ratio supported external generalizability, although ComBat harmonization improved calibration rather than external discrimination (22). More recently, Ding et al. integrated automated segmentation, radiomics, deep-learning, and clinical features and achieved an external-test AUC of 0.902 (29). Nevertheless, prospective evidence has not consistently shown that deep learning significantly outperforms radiologist-based PI-RADS assessment, and methodological heterogeneity remains substantial (30,31). Compared with these larger multimodal and automated approaches, the present five-feature ADC-only model prioritizes parsimony and interpretability, but multicenter external validation is required to establish its comparative clinical value. The hierarchical modeling strategy compared the clinical, PI-RADS, ADC radiomics, clinicoradiological, and combined models rather than reporting radiomics in isolation. This stepwise modeling design better reflects real-world clinical decision-making and more directly addresses the “added value” question of this study. If only the AUC of the ADC radiomics model (0.916) had been reported, the independent value of radiomics could have been overestimated. By contrast, adding the Rad-score to the clinical + PI-RADS model increased the test-cohort AUC only from 0.944 to 0.951 [ΔAUC, 0.007; 95% CI, (−0.027 to 0.043); P=0.71]. The difference between the combined and ADC radiomics models was also not statistically significant (ΔAUC, 0.035; DeLong P=0.052). Because the test cohort included only 25 csPCa events, the precision and statistical power of these comparisons were limited. Therefore, the numerical differences among the high-performing models should be interpreted as exploratory rather than as evidence of superiority. The principal contribution of this hierarchical analysis is therefore to distinguish strong standalone radiomics performance from incremental value beyond routine clinicoradiological assessment. In the present cohort, statistically significant incremental discrimination from adding the Rad-score was not demonstrated. Direct comparison with established tools such as the European Randomized Study of Screening for Prostate Cancer (ERSPC) and Prostate Biopsy Collaborative Group (PBCG) risk calculators or published MRI-based nomograms was not feasible because the retrospective dataset did not contain all required input variables. Accordingly, the present analysis does not establish superiority over existing risk calculators. Future prospective studies should collect the complete required variables and perform direct head-to-head comparisons. From an implementation perspective, this small gain must be weighed against the additional requirements for manual lesion segmentation, radiomics extraction, quality control, and software support. Although the single-sequence five-feature design reduces complexity relative to multisequence radiomics, the present evidence does not justify routine adoption of the radiomics workflow over the simpler clinicoradiological model. Future studies should determine whether automated segmentation or use in selected clinical scenarios can provide a more favorable balance between predictive benefit and workflow burden.
The nomogram provided a visual representation of the combined model. The final nomogram incorporated Rad-score, PI-RADS score, age, log10(TPSA), FPSA/TPSA ratio, and family history of prostate cancer. Although BMI differed significantly at baseline, it was not retained in the final nomogram because variable inclusion was based on model contribution, clinical relevance, and interpretability rather than univariable P values alone. Because TPSA showed a markedly right-skewed distribution, log10(TPSA) was used to reduce the influence of extreme values on model estimation and nomogram scaling. The relatively wide point range assigned to the Rad-score reflects its regression coefficient and observed distribution within the combined model, rather than its incremental value beyond the other predictors. A predictor may occupy a wide range on a nomogram while producing little improvement in AUC if much of its discriminatory information overlaps with that already captured by PI-RADS and clinical variables. Moreover, AUC primarily measures changes in patient ranking and may remain nearly unchanged even when a predictor modifies individual predicted probabilities. Thus, the nomogram point range and the nonsignificant DeLong comparison address different aspects of model behavior and are not contradictory. Calibration and DCA provided complementary evaluation. The combined model had the lowest Brier score in both cohorts (0.086 and 0.078) and showed favorable net benefit mainly within the low-to-intermediate threshold range. However, these findings describe the combined model as a whole and do not establish significant incremental benefit from the Rad-score. As an illustrative biopsy triage scenario, the locked threshold of 0.201 would have avoided 57 of 91 biopsies while missing one of 25 csPCa cases. However, this threshold was selected using the Youden index rather than a prespecified clinical utility criterion and should not be considered a recommended clinical cutoff. The appropriate threshold depends on the acceptable balance between biopsy reduction and missed csPCa and must be established through prospective external validation.
This study also has several limitations. First, this was a single-center retrospective study with only 57 csPCa events available in the training cohort for model development. Although the Rad-score was reduced to five features and full radiomics-pipeline repeated nested cross-validation was performed, residual overfitting and feature-selection instability cannot be completely excluded. In addition, only 25 csPCa events were included in the test cohort, limiting the precision of the AUC estimates and the statistical power of the pairwise DeLong tests. Consequently, the numerical differences among the high-performing models cannot be considered definitive. Moreover, the held-out test cohort originated from the same institution and therefore represented internal testing rather than external validation. Multicenter external validation remains necessary. Accordingly, the present findings represent internal validation only and do not yet support clinical implementation. Second, ROIs were manually delineated slice by slice. Although repeated segmentation and ICC assessment were performed, manual segmentation may still be subject to observer dependence. Manual lesion segmentation required approximately 5 minutes per patient; however, the total end-to-end radiomics workflow time and implementation costs were not formally evaluated. Therefore, whether the limited incremental performance justifies the additional workflow burden remains uncertain. Third, all images were acquired at a single institution using a single-vendor, single-scanner platform with an unchanged software version. Although the core acquisition protocol was generally consistent, minor parameter variations occurred over time and may have influenced some radiomics features. Moreover, the single-scanner design prevented assessment of cross-scanner robustness, and multicenter, multi-vendor external validation remains necessary. Fourth, established clinical and MRI-based risk calculators could not be reconstructed because complete required inputs were unavailable; therefore, direct benchmarking against currently available tools was not performed.
Conclusions
In conclusion, the parsimonious five-feature ADC radiomics model showed strong standalone discrimination for csPCa using a single routinely available MRI sequence. However, adding the Rad-score to clinical variables and PI-RADS increased the test-cohort AUC only from 0.944 to 0.951, and the difference was small and statistically nonsignificant (ΔAUC =0.007; DeLong P=0.71). Thus, significant incremental discrimination beyond the clinicoradiological model was not demonstrated. The principal contribution of this study is the direct comparison of five hierarchical models, which distinguishes standalone radiomics performance from its added value beyond routine clinicoradiological assessment. Given the manual segmentation requirements, limited incremental discrimination, and absence of independent external validation, the current workflow cannot yet be recommended for clinical implementation. Multicenter, multi-vendor external validation is required to establish its generalizability and clinical value.
Acknowledgments
None.
Footnote
Reporting Checklist: The authors have completed the TRIPOD+AI reporting checklist. Available at https://tau.amegroups.com/article/view/10.21037/tau-2026-0546/rc
Data Sharing Statement: Available at https://tau.amegroups.com/article/view/10.21037/tau-2026-0546/dss
Peer Review File: Available at https://tau.amegroups.com/article/view/10.21037/tau-2026-0546/prf
Funding: None.
Conflicts of Interest: All authors have completed the ICMJE uniform disclosure form (available at https://tau.amegroups.com/article/view/10.21037/tau-2026-0546/coif). The authors have no conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved. This study was conducted in accordance with the Declaration of Helsinki and its subsequent amendments. The study was approved by the Institutional Review Board of The Sixth Hospital of Wuhan, Affiliated Hospital of Jianghan University (No. WHSHIRB-K-2026015). The requirement for written informed consent was waived because of the retrospective nature of the study.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Fonteyne V, Tree A, Castro E, et al. Prostate cancer. Lancet 2026;407:622-36. [Crossref] [PubMed]
- Nikolaevich PV, Samuel AO, Fayazovich UM, et al. Infection risks and biopsy-associated complications in prostate cancer diagnosis: a review of recent literatures. Prostate Cancer Prostatic Dis 2026;29:503-15. [Crossref] [PubMed]
- Assani KD, Pierce JM, Galloway LA, et al. Blood- and urine-based biomarkers for the detection of clinically significant prostate cancer: a contemporary review. Curr Opin Urol 2025;35:590-6. [Crossref] [PubMed]
- Schoots IG, Ahmed HU, Albers P, et al. Magnetic Resonance Imaging-based Biopsy Strategies in Prostate Cancer Screening: A Systematic Review. Eur Urol 2025;88:247-60. [Crossref] [PubMed]
- Cornford P, van den Bergh RCN, Briers E, et al. EAU-EANM-ESTRO-ESUR-ISUP-SIOG Guidelines on Prostate Cancer-2024 Update. Part I: Screening, Diagnosis, and Local Treatment with Curative Intent. Eur Urol 2024;86:148-63.
- Yilmaz EC, Shih JH, Belue MJ, et al. Prospective Evaluation of PI-RADS Version 2.1 for Prostate Cancer Detection and Investigation of Multiparametric MRI-derived Markers. Radiology 2023;307:e221309. [Crossref] [PubMed]
- Haj-Mirzaian A, Burk KS, Lacson R, et al. Magnetic Resonance Imaging, Clinical, and Biopsy Findings in Suspected Prostate Cancer: A Systematic Review and Meta-Analysis. JAMA Netw Open 2024;7:e244258. [Crossref] [PubMed]
- Maniaci A, Lavalle S, Gagliano C, et al. The Integration of Radiomics and Artificial Intelligence in Modern Medicine. Life (Basel) 2024;14:1248. [Crossref] [PubMed]
- Surov A, Eger KI, Potratz J, et al. Apparent diffusion coefficient correlates with different histopathological features in several intrahepatic tumors. Eur Radiol 2023;33:5955-64. [Crossref] [PubMed]
- Yan X, Ma K, Zhu L, et al. The value of apparent diffusion coefficient values in predicting Gleason grading of low to intermediate-risk prostate cancer. Insights Imaging 2024;15:137. [Crossref] [PubMed]
- Girometti R, Peruzzi V, Clauser P, et al. Diffusion levels for quantitative assessment of the apparent diffusion coefficient value in prostate MRI: a proof-of-concept bicentric study. Eur Radiol 2025;35:6171-82. [Crossref] [PubMed]
- Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ 2024;385:e078378. [Crossref] [PubMed]
- Hugosson J, Månsson M, Wallström J, et al. Prostate Cancer Screening with PSA and MRI Followed by Targeted Biopsy Only. N Engl J Med 2022;387:2126-37. [Crossref] [PubMed]
- IBSI. The image biomarker standardisation initiative. Accessed 2026.4.12. Available online: https://ibsi.readthedocs.io/en/latest/
- Jung S, Kim JS. Radiomics Reproducibility in Prostate Cancer Diagnosis Based on PROSTATEx. Int Neurourol J 2025;29:S95-S100. [Crossref] [PubMed]
- Lin DW, Carlsson S, Filson CP, et al. Updates to Early Detection of Prostate Cancer: AUA/SUO Guideline (2026). J Urol 2026;215:491-501. [Crossref] [PubMed]
- Margolis DJA, Chatterjee A, deSouza NM, et al. Quantitative Prostate MRI, From the AJR Special Series on Quantitative Imaging. AJR Am J Roentgenol 2025;225:e2431715. [Crossref] [PubMed]
- Taya M, Behr SC, Westphalen AC. Perspectives on technology: Prostate Imaging-Reporting and Data System (PI-RADS) interobserver variability. BJU Int 2024;134:510-8. [Crossref] [PubMed]
- Agrotis G, Pooch E, Abdelatty M, et al. Diagnostic performance of ADC and ADCratio in MRI-based prostate cancer assessment: A systematic review and meta-analysis. Eur Radiol 2025;35:404-16. [Crossref] [PubMed]
- Mylona E, Zaridis DI, Kalantzopoulos CN, et al. Optimizing radiomics for prostate cancer diagnosis: feature selection strategies, machine learning classifiers, and MRI sequences. Insights Imaging 2024;15:265. [Crossref] [PubMed]
- Zhang KS, Neelsen CJO, Wennmann M, et al. In vivo variability of MRI radiomics features in prostate lesions assessed by a test-retest study with repositioning. Sci Rep 2025;15:29703. [Crossref] [PubMed]
- Samaras D, Agrotis G, Vamvakas A, et al. Beyond Radiomics Alone: Enhancing Prostate Cancer Classification with ADC Ratio in a Multicenter Benchmarking Study. Diagnostics (Basel) 2025;15:2546. [Crossref] [PubMed]
- Ng ABCD, Asif A, Agarwal R, et al. Biparametric vs. Multiparametric MRI for Prostate Cancer Diagnosis: The PRIME Diagnostic Clinical Trial. JAMA 2025;334:1170-9.
- Hectors SJ, Chen C, Chen J, et al. Magnetic Resonance Imaging Radiomics-Based Machine Learning Prediction of Clinically Significant Prostate Cancer in Equivocal PI-RADS 3 Lesions. J Magn Reson Imaging 2021;54:1466-73. [Crossref] [PubMed]
- Zhao YY, Xiong ML, Liu YF, et al. Magnetic resonance imaging radiomics-based prediction of clinically significant prostate cancer in equivocal PI-RADS 3 lesions in the transitional zone. Front Oncol 2023;13:1247682. [Crossref] [PubMed]
- Antolin A, Roson N, Mast R, et al. The Role of Radiomics in the Prediction of Clinically Significant Prostate Cancer in the PI-RADS v2 and v2.1 Era: A Systematic Review. Cancers (Basel) 2024;16:2951. [Crossref] [PubMed]
- Bao J, Qiao X, Song Y, et al. Prediction of clinically significant prostate cancer using radiomics models in real-world clinical practice: a retrospective multicenter study. Insights Imaging 2024;15:68. [Crossref] [PubMed]
- Rodrigues AC, de Almeida JG, Rodrigues N, et al. Improving Clinically Significant Prostate Cancer Detection with a Multimodal Machine Learning Approach: A Large-Scale Multicenter Study. Radiol Imaging Cancer 2025;7:e240507. [Crossref] [PubMed]
- Ding N, Jin L, Yin S, et al. A multicenter study of automatic segmentation-based multimodal fusion integrating radiomics, deep learning, and clinical parameters for prostate cancer detection. Abdom Radiol (NY) 2026; Epub ahead of print. [Crossref]
- Lee YJ, Moon HW, Choi MH, et al. MRI-based Deep Learning Algorithm for Assisting Clinically Significant Prostate Cancer Detection: A Bicenter Prospective Study. Radiology 2025;314:e232788. [Crossref] [PubMed]
- Molière S, Hamzaoui D, Ploussard G, et al. A Systematic Review of the Diagnostic Accuracy of Deep Learning Models for the Automatic Detection, Localization, and Characterization of Clinically Significant Prostate Cancer on Magnetic Resonance Imaging. Eur Urol Oncol 2025;8:1182-202. [Crossref] [PubMed]


