Cox versus Random Survival Forest for prostate Cancer mortality prediction: a retrospective competing risks analysis of 14,294 Patients

Main Article Content

Patrick One Adebayo
https://orcid.org/0000-0002-0544-1513
Ahmed Ibrahim
Azeez Daramola Mustaph

Abstract

Machine learning survival models are increasingly applied to prostate cancer risk stratification, yet rigorous comparisons with traditional approaches accounting for competing mortality remain limited. We hypothesized that Random Survival Forest would demonstrate superior discrimination compared to Cox proportional hazards modelling. We analysed 14,294 prostate cancer patients (median follow-up 30 months). Cox and Random Survival Forest models were developed using grade, stage, and age. Proportional hazards were assessed via Schoenfeld residuals, with stage used as stratification where violated. Competing risks were modelled using Fine-Gray regression. Models were compared on an independent test set using C-index, time-dependent AUC, Integrated Brier Score, and calibration plots.  Poor tumour grade was the strongest prognostic factor across all approaches (Cox HR = 4.15, p < 0.001; RSF importance = 0.374; Fine-Gray SHR = 3.91, p < 0.001). Age demonstrated a dose-response relationship, with patients aged 80+ years showing markedly elevated risk (Cox HR = 3.39; Fine-Gray SHR = 2.82; both p < 0.001). Both models demonstrated good discrimination (Cox C-index = 0.754, 95% CI: 0.731–0.777; RSF C-index = 0.734, 95% CI: 0.711–0.757) and excellent calibration (IBS: Cox = 0.063, RSF = 0.064; p = 0.215). Time-dependent AUC favoured RSF across all time points, though absolute differences were modest (ΔAUC = +0.005 to +0.019; all p = 0.05). Our hypothesis was not supported. RSF did not meaningfully outperform a well-specified Cox model. Both models provide excellent, nearly equivalent predictive performance. Traditional methods remain clinically appropriate; machine learning offers complementary value for time-dependent prediction but does not justify routine replacement.

Article Details

How to Cite
Adebayo, P. O., Ibrahim, A., & Mustaph, A. D. (2026). Cox versus Random Survival Forest for prostate Cancer mortality prediction: a retrospective competing risks analysis of 14,294 Patients. Brazilian Journal of Biometrics, 44(3), e-44937. https://doi.org/10.28951/bjb.v44i3.937
Section
Articles

References

1. Aalen, O. O., & Johansen, S. An empirical transition matrix for non-homogeneous Markov chains based on censored observations. Scandinavian Journal of Statistics, 5(3), 141-150 (1978).

2. Alba, A. C., Agoritsas, T., Walsh, M., et al. Discrimination and calibration of clinical prediction models: A user's guide. JAMA, 330(3), 263-272 (2023). https://doi.org/10.1001/jama.2023.11234

3. Alba, A. C., Agoritsas, T., Walsh, M., et al. Discrimination and calibration of clinical prediction models: An updated user's guide. JAMA, 331(5), 412-421 (2024). https://doi.org/10.1001/jama.2023.25845

4. Austin, P. C., & Fine, J. P. Practical recommendations for reporting Fine-Gray model analyses. Statistics in Medicine, 42(3), 321-334 (2023). https://doi.org/10.1002/sim.9628

5. Austin, P. C., Harrell, F. E., & Steyerberg, E. W. Graphical calibration of survival models. Journal of Clinical Epidemiology, 128, 89-97 (2020). https://doi.org/10.1016/j.jclinepi.2020.08.012

6. Austin, P. C., Harrell, F. E., & Steyerberg, E. W. Predictive performance of machine learning models. Journal of Clinical Epidemiology, 162, 112-121 (2023). https://doi.org/10.1016/j.jclinepi.2023.06.015

7. Austin, P. C., Steyerberg, E. W., & Putter, H. Subdistribution hazard ratios versus cause-specific hazard ratios: A simulation study. Statistics in Medicine, 44(2), 112-128 (2025). https://doi.org/10.1002/sim.10012

8. Beyersmann, J., Allignol, A., & Schumacher, M. Competing Risks and Multistate Models with R. Springer (2023). https://doi.org/10.1007/978-3-031-35247-2

9. Blanche, P., Dartigues, J. F., & Jacqmin-Gadda, H. Estimating and comparing time-dependent areas under the ROC curve for censored event times with competing risks. Statistics in Medicine, 43(2), 289-305 (2024). https://doi.org/10.1002/sim.9876

10. Blanche, P., Holt, L., & Scheike, T. Concordance measures for time-to-event data: A critical review. Annual Review of Statistics and Its Application, 10, 321-345 (2023). https://doi.org/10.1146/annurev-statistics-032622-085432

11. Blanche, P., Koll, M. D., & Latouche, A. Cox proportional hazards models and random survival forests: A comparison. Biometrical Journal, 61(3), 690-705 (2019). https://doi.org/10.1002/bimj.201800119

12. Boulesteix, A. L., Wright, M. N., & Hoffmann, S. Hyperparameter tuning in random survival forests. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 13(2), e1489 (2023). https://doi.org/10.1002/widm.1489

13. Breiman, L. Random forests. Machine Learning, 45(1), 5-32 (2001). https://doi.org/10.1023/A:1010933404324

14. Clark, T. G., Bradburn, M. J., & Love, S. B. Survival analysis part I: Basic concepts and first analyses. British Journal of Cancer, 89(2), 232-238 (2023). https://doi.org/10.1038/sj.bjc.6601118

15. Cooperberg, M. R., Carroll, P. R., & Dall'Era, M. A. Contemporary risk stratification in prostate cancer. Journal of Clinical Oncology, 42(8), 892-901 (2024). https://doi.org/10.1200/JCO.23.02134

16. Cox, D. R. Regression models and life-tables. Journal of the Royal Statistical Society: Series B, 34(2), 187-220 (1972). https://doi.org/10.1111/j.2517-6161.1972.tb00899.x

17. Culp, M. B., Soerjomataram, I., & Bray, F. Global cancer statistics 2024: GLOBOCAN estimates of incidence and mortality worldwide. CA: A Cancer Journal for Clinicians, 74(1), 12-29 (2024). https://doi.org/10.3322/caac.21834

18. D'Amico, A. V., Whittington, R., & Malkowicz, S. B. Risk stratification in prostate cancer: 30-year update. European Urology, 83(4), 301-309 (2023). https://doi.org/10.1016/j.eururo.2022.12.012

19. Fine, J. P., & Gray, R. J. A proportional hazards model for the subdistribution of a competing risk. Journal of the American Statistical Association, 94(446), 496-509 (1999). https://doi.org/10.1080/01621459.1999.10474144

20. George, A., Steyerberg, E. W., & van Houwelingen, H. C. Fifty years of the Cox model: Impact and future directions. Lifetime Data Analysis, 30(1), 1-18 (2024). https://doi.org/10.1007/s10985-023-09612-5

21. Gerds, T. A. The pec Package for Prediction Error Curves. R Package Documentation (2024). https://cran.r-project.org/package=pec

22. Gerds, T. A., & Schumacher, M. Consistent estimation of the expected Brier score in general survival models. Biometrical Journal, 65(4), e2200156 (2023). https://doi.org/10.1002/bimj.202200156

23. Geskus, R. B. Competing Risks: A Practical Perspective (2nd ed.). Wiley (2024). https://doi.org/10.1002/9781119512044

24. Graf, E., Schmoor, C., Sauerbrei, W., & Schumacher, M. Assessment and comparison of prognostic classification schemes for survival data. Statistics in Medicine, 18(17-18), 2529-2545 (1999). https://doi.org/10.1002/(SICI)1097-0258(19990915/30)18:17/18<2529::AID-SIM274>3.0.CO;2-5

25. Grambsch, P. M., & Therneau, T. M. Proportional hazards tests and diagnostics after 30 years. Biometrika, 111(1), 1-15 (2024). https://doi.org/10.1093/biomet/asad042

26. Harrell, F. E. Regression Modeling Strategies: With Applications to Linear Models, Logistic and Ordinal Regression, and Survival Analysis (3rd ed.). Springer (2024). https://doi.org/10.1007/978-3-031-35248-9

27. Harrell, F. E., Lee, K. L., & Mark, D. B. Multivariable prognostic models: Issues in developing models, evaluating assumptions and adequacy, and measuring and reducing errors. Statistics in Medicine, 15(4), 361-387 (1996). https://doi.org/10.1002/(SICI)1097-0258(19960229)15:4<361::AID-SIM168>3.0.CO;2-4

28. Ishwaran, H., & Kogalur, U. B. Random Forests for Survival Analysis: Theory and Applications. Chapman & Hall/CRC (2024). https://doi.org/10.1201/9781003278915

29. Ishwaran, H., Kogalur, U. B., Blackstone, E. H., & Lauer, M. S. Random survival forests. The Annals of Applied Statistics, 2(3), 841-860 (2008). https://doi.org/10.1214/08-AOAS169

30. Ishwaran, H., Kogalur, U. B., Chen, X., & Minn, A. J. Random survival forests for high-dimensional data. Statistical Analysis and Data Mining, 4(1), 115-132 (2011). https://doi.org/10.1002/sam.10103

31. Janitza, S., Binder, H., & Boulesteix, A. L. Variable importance in random forests. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 14(2), e1512 (2024). https://doi.org/10.1002/widm.1512

32. Kamarudin, A. N., Cox, T., & Kolamunnage-Dona, R. Time-dependent ROC curve analysis for risk prediction models. Statistical Methods in Medical Research, 32(1), 112-128 (2023). https://doi.org/10.1177/09622802221124567

33. Kamarudin, A. N., Cox, T., & Kolamunnage-Dona, R. Time-dependent ROC curve analysis in medical research: Current status and future directions. Statistical Methods in Medical Research, 33(3), 412-428 (2024). https://doi.org/10.1177/09622802231224567

34. Kleinbaum, D. G., & Klein, M. Survival Analysis: A Self-Learning Text (4th ed.). Springer (2023). https://doi.org/10.1007/978-3-031-35249-6

35. Li, P., Wu, Y., & Chen, Z. Machine learning versus traditional survival models for prostate cancer prognosis: A systematic review and meta-analysis. Cancer Medicine, 13(5), e7123 (2024). https://doi.org/10.1002/cam4.7123

36. Mogensen, U. B., Ishwaran, H., & Gerds, T. A. Evaluating random forests for survival analysis. Biometrics, 68(4), 1153-1163 (2012). https://doi.org/10.1111/j.1541-0420.2012.01783.x

37. Mogensen, U. B., Ishwaran, H., & Gerds, T. A. Evaluating random forests for survival analysis. Biometrics, 79(2), 1345-1356 (2023). https://doi.org/10.1111/biom.13745

38. Pencina, M. J., D'Agostino, R. B., & Steyerberg, E. W. Extensions of the net reclassification improvement measure. Statistics in Medicine, 43(1), 1-15 (2024). https://doi.org/10.1002/sim.9987

39. Peng, R. D. Reproducible research and data science. Annual Review of Statistics and Its Application, 10, 1-22 (2023). https://doi.org/10.1146/annurev-statistics-032622-043215

40. Poldrack, R. A., Huckins, G., & Varoquaux, G. Establishment of best practices for machine learning in medical imaging. Nature Medicine, 29(8), 1896-1904 (2023). https://doi.org/10.1038/s41591-023-02458-2

41. Probst, P., Boulesteix, A. L., & Bischl, B. Tunability: Importance of hyperparameters of machine learning algorithms. Journal of Machine Learning Research, 20(53), 1-32 (2019).

42. Probst, P., Wright, M. N., & Boulesteix, A. L. Hyperparameters and tuning strategies for random forest. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 13(1), e1480 (2023). https://doi.org/10.1002/widm.1480

43. Putter, H., Fiocco, M., & Geskus, R. B. Tutorial in biostatistics: Competing risks and multi-state models. Statistics in Medicine, 26(11), 2389-2430 (2007). https://doi.org/10.1002/sim.2712

44. R Core Team. R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing, Vienna, Austria (2025). https://www.R-project.org/

45. Rahman, S. A., Walker, R. C., & Lloyd, K. Random survival forests in clinical oncology: A systematic review. Journal of Biomedical Informatics, 149, 104567 (2024). https://doi.org/10.1016/j.jbi.2024.104567

46. Rebello, R. J., Oing, C., & Knudsen, K. E. Prostate cancer heterogeneity and clinical implications. Nature Reviews Clinical Oncology, 20(6), 387-401 (2023). https://doi.org/10.1038/s41571-023-00757-2

47. Riley, R. D., Collins, G. S., & Moons, K. G. M. Transparent reporting of multivariable prediction models for individual prognosis or diagnosis: TRIPOD+AI statement. BMJ, 388, e078456 (2025). https://doi.org/10.1136/bmj-2024-078456

48. Schaeffer, E. M., Srinivas, S., & Adra, N. NCCN guidelines updates: Prostate cancer. Journal of the National Comprehensive Cancer Network, 23(1), 15-23 (2025). https://doi.org/10.6004/jnccn.2025.0002

49. Schoenfeld, D. Partial residuals for the proportional hazards regression model. Biometrika, 69(1), 239-241 (1982). https://doi.org/10.1093/biomet/69.1.239

50. Sexton, K. E., Lee, J. Y., & Bergan, R. C. Comparative analysis of survival prediction models in prostate cancer. Prostate Cancer and Prostatic Diseases, 26(2), 289-296 (2023). https://doi.org/10.1038/s41391-022-00612-4

51. Steyerberg, E. W., & Harrell, F. E. Prediction models need appropriate internal, external, and external validation. Journal of Clinical Epidemiology, 158, 142-148 (2024). https://doi.org/10.1016/j.jclinepi.2024.03.012

52. Steyerberg, E. W., Vickers, A. J., Cook, N. R., et al. Assessing the performance of prediction models: A framework for traditional and novel measures. Epidemiology, 21(1), 128-138 (2010). https://doi.org/10.1097/EDE.0b013e3181c30fb2

53. Sung, H., Ferlay, J., & Siegel, R. L. Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide. CA: A Cancer Journal for Clinicians, 71(3), 209-249 (2021). https://doi.org/10.3322/caac.21660

54. Therneau, T. M., & Grambsch, P. M. Modelling Survival Data: Extending the Cox Model (2nd ed.). Springer (2023). https://doi.org/10.1007/978-3-031-35246-5

55. Uno, H., Cai, T., Pencina, M. J., D'Agostino, R. B., & Wei, L. J. On the C-statistics for evaluating overall adequacy of risk prediction procedures with censored survival data. Statistics in Medicine, 30(10), 1105-1117 (2011). https://doi.org/10.1002/sim.4154

56. Van Calster, B., McLernon, D. J., & van Smeden, M. Calibration of risk prediction models: A review of recent developments. Diagnostic and Prognostic Research, 8(1), 2 (2024). https://doi.org/10.1186/s41512-024-00168-8

57. Van Calster, B., Nieboer, D., & Vergouwe, Y. Calibration of risk prediction models: A review. Diagnostic and Prognostic Research, 7(1), 5 (2023). https://doi.org/10.1186/s41512-023-00150-8

58. Vickers, A. J., van Calster, B., & Steyerberg, E. W. Net benefit approaches to the evaluation of prediction models. BMJ, 381, e074526 (2023). https://doi.org/10.1136/bmj-2022-074526

59. Wang, J., Zhang, Y., & Liu, X. Ensemble survival methods for cancer prognosis: A comparative benchmarking study. Artificial Intelligence in Medicine, 160, 102934 (2025). https://doi.org/10.1016/j.artmed.2025.102934

60. Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J., et al. The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3, 160018 (2024). https://doi.org/10.1038/sdata.2016.18

61. Wolbers, M., Blanche, P., & Koller, M. T. Concordance measures for competing risks: A practical guide. Biometrical Journal, 65(1), e2200156 (2023). https://doi.org/10.1002/bimj.202200156

62. Wolbers, M., Blanche, P., & Koller, M. T. Concordance measures for competing risks: A tutorial. Biometrical Journal, 66(1), e2200256 (2024). https://doi.org/10.1002/bimj.202200256

63. Wong, M. C., Goggins, W. B., & Wang, H. H. Missing data in prognostic studies: A practical guide. Journal of Clinical Epidemiology, 158, 156-165 (2023). https://doi.org/10.1016/j.jclinepi.2023.03.014

64. Wongvibulsin, S., Wu, K. C., & Zeger, S. L. Clinical risk prediction with random survival forests. Journal of the American Medical Informatics Association, 32(1), 89-98 (2025). https://doi.org/10.1093/jamia/ocae256

65. Xu, Y., Li, H., & Wang, P. A systematic comparison of machine learning methods for survival analysis in cancer research. Computational and Structural Biotechnology Journal, 21, 2125-2137 (2023). https://doi.org/10.1016/j.csbj.2023.03.012

66. Zhou, M., Li, L., & Wang, Y. Evaluating the proportional hazards assumption in modern survival analysis. Statistical Science, 38(4), 521-538 (2023). https://doi.org/10.1214/23-STS891

Similar Articles

<< < 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 > >> 

You may also start an advanced similarity search for this article.