Evolving Perspectives on Breiman’s Two Cultures: From Statistical Modeling to Contemporary Machine Learning
DOI:
https://doi.org/10.22370/pe.2026.20.5263Keywords:
Statistical modeling, Data Science, Machine learning, Two cultures debate, Model interpretability, data analysis evolution, data-driven methods, Interdisciplinary collaborationAbstract
This article presents a structured critical review of the debate between statistical modeling and algorithmic modeling, following the SALSA framework (Search, Appraisal, Synthesis, and Analysis). The review is anchored in Breiman’s (2001a) “Two Cultures” and Athey and Imbens’ (2019) perspective on machine learning for economists, with literature searches documented through Scopus, Web of Science, and OpenAlex and evaluated using explicit inclusion and exclusion criteria. The synthesis examines the evolving relationship between predictive accuracy and interpretability, algorithmic flexibility and causal inference, and their implications for economic research and policy evaluation. Recent developments in machine learning-based science, explainable and trustworthy artificial intelligence, and causal machine learning are discussed as extensions of the original debate. The review argues that future progress will depend not on the dominance of a single methodological culture, but on the integration of statistical reasoning, algorithmic approaches, and scientific judgment through transparent validation and problem-oriented methodological choices.
Downloads
References
Angrist, J. D., & Pischke, J.-S. (2009). *Mostly harmless econometrics: An empiricist’s companion*. Princeton University Press.
Athey, S., Agrawal, A., Gans, J., & Goldfarb, A. (2018). The impact of machine learning on economics. In *The economics of artificial intelligence: An agenda* (pp. 507–547). University of Chicago Press.
Athey, S., & Imbens, G. W. (2019). Machine learning methods that economists should know about. *Annual Review of Economics, 11*(1), 685–725. https://doi.org/10.1146/annurev-economics-080217-053950
Barocas, S., Hardt, M., & Narayanan, A. (2023). *Fairness and machine learning: Limitations and opportunities*. MIT Press.
Bhadra, A., Datta, J., Polson, N., Sokolov, V., & Xu, J. (2021). Merging two cultures: Deep and statistical learning. *arXiv preprint arXiv:2110.11561*. https://doi.org/10.48550/arXiv.2110.11561
Booth, A., St. James, M., Clowes, M., Sutton, A., et al. (2021). *Systematic approaches to a successful literature review*. SAGE Publications Ltd.
Borrellas, P., & Unceta, I. (2021). The challenges of machine learning and their economic implications. *Entropy, 23*(3), 275. https://doi.org/10.3390/e23030275
Breiman, L. (2001). Random forests. *Machine Learning, 45*, 5–32. https://doi.org/10.1023/A:1010933404324
Breiman, L. (2001). Statistical modeling: The two cultures (with comments and a rejoinder by the author). *Statistical Science, 16*(3), 199–231. https://doi.org/10.1214/ss/1009213729
Calin-Jageman, R. J., & Cumming, G. (2019). The new statistics for better science: Ask how much, how uncertain, and what else is known. *The American Statistician, 73*(sup1), 271–280. https://doi.org/10.1080/00031305.2019.1616221
Castelvecchi, D. (2024). The AI-quantum computing mash-up: Will it revolutionize science? *Nature*. https://doi.org/10.1038/s41586-024-08888-7
Chen, J. C., Dunn, A., Hood, K., Driessen, A., & Batch, A. (2019). Off to the races: A comparison of machine learning and alternative data for predicting economic indicators. In *Big data for 21st century economic statistics* (pp. 123–145). University of Chicago Press.
de Mast, J., Steiner, S. H., Nuijten, W. P. M., & Kapitan, D. (2023). Analytical problem solving based on causal, correlational, and deductive models. *The American Statistician, 77*(1), 51–61. https://doi.org/10.1080/00031305.2023.2100701
Easton, P. D., Kapons, M. M., Monahan, S. J., Schütt, H. H., & Weisbrod, E. H. (2024). Forecasting earnings using k-nearest neighbors. *The Accounting Review, 99*(3), 115–140. https://doi.org/10.2308/accr-2024-0029
Fayyad, U., Piatetsky-Shapiro, G., & Smyth, P. (1996). From data mining to knowledge discovery in databases. *AI Magazine, 17*(3), 37–37. https://doi.org/10.1609/aimag.v17i3.1464
García-Holgado, A., Marcos-Pablos, S., & García-Peñalvo, F. (2020). Guidelines for performing systematic research projects reviews. *International Journal of Interactive Multimedia and Artificial Intelligence, 7*(1), 19–26. https://doi.org/10.9781/ijimai.2020.02.002
Grammarly, Inc. (2024). *Grammarly* [Online tool]. https://www.grammarly.com/
Hassija, V., Chamola, V., Mahapatra, A., Singal, A., Goel, D., Huang, K., Scarda-pane, S., Spinelli, I., Mahmud, M., & Hussain, A. (2024). Interpreting black-box models: A review on explainable artificial intelligence. *Cognitive Computation, 16*(1), 45–74. https://doi.org/10.1007/s12559-024-09803-w
Hume, D. (1907). *An Enquiry Concerning Human Understanding and Selections from A Treatise of Human Nature: With Hume’s Autobiography and a Letter from Adam Smith* (Vol. 45). Open Court Publishing Company.
Kahneman, D., Sibony, O., & Sunstein, C. R. (2021). *Noise: A flaw in human judgment*. Hachette UK.
Kapoor, S., Cantrell, E. M., Peng, K., Pham, T. H., Bail, C. A., Gundersen, O. E., Hofman, J. M., Hullman, J., Lones, M. A., Malik, M. M., et al. (2024). Reforms: Consensus-based recommendations for machine-learning-based science. *Science Advances, 10*(18), eadk3452. https://doi.org/10.1126/sciadv.adk3452
Ludwig, J., & Mullainathan, S. (2024). Machine learning as a tool for hypothesis generation. *The Quarterly Journal of Economics, 139*(2), 751–827. https://doi.org/10.1093/qje/qjz016
Maccarrone, G., Morelli, G., & Spadaccini, S. (2021). GDP forecasting: Machine learning, linear or autoregression? *Frontiers in Artificial Intelligence, 4*, 757864. https://doi.org/10.3389/frai.2021.757864
Majone, G., & Quade, E. S. (1980). *Pitfalls of analysis* (Vol. 8). John Wiley & Sons.
Messeri, L., & Crockett, M. J. (2024). Artificial intelligence and illusions of understanding in scientific research. *Nature, 627*(8002), 49–58. https://doi.org/10.1038/s41586-024-08900-w
Ooi, K.-B., Tan, G. W.-H., Al-Emran, M., Al-Sharafi, M. A., Capatina, A., Chakraborty, A., Dwivedi, Y. K., Huang, T.-L., Kar, A. K., Lee, V.-H., et al. (2023). The potential of generative artificial intelligence across disciplines: Perspectives and future directions. *Journal of Computer Information Systems*, 1–32. https://doi.org/10.1080/08874417.2023.2174927
OpenAI. (2024, August 25). *ChatGPT: A large language model* [Model version]. https://chat.openai.com/
Sarker, I. H. (2021). Machine learning: Algorithms, real-world applications, and research directions. *SN Computer Science, 2*(3), 160. https://doi.org/10.1007/s42979-021-00439-x
Silva, T. C., Wilhelm, P. V. B., & Amancio, D. R. (2024). Machine learning and economic forecasting: The role of international trade networks. *Physica A: Statistical Mechanics and its Applications, 649*, 129977. https://doi.org/10.1016/j.physa.2024.129977
Tukey, J. W. (1962). The future of data analysis. *The Annals of Mathematical Statistics, 33*(1), 1–67. https://doi.org/10.1214/aoms/1177704711
van der Zant, T., Kouw, M., & Schomaker, L. (2013). *Generative artificial intelligence*. Springer.
Woloszko, N. (2017). Making better economic forecasts with machine learning. *Finance, Machine Learning, 2017. Visited on 01-11-2024
Downloads
Published
Issue
Section
License
Those authors who have publications with this journal accept the following terms:
1.- Authors will retain their copyright and grant the journal the right of first publication of their work, which will simultaneously be subject to the Creative Commons Attribution License (CC BY-NC-ND 4.0 International) https://creativecommons.org/licenses/by-nc-nd/4.0/deed.es which allows third parties to share, copy and redistribute the material in any medium or format.
- Attribution: credit must be given appropriately, provide a link to the license, and indicate if changes have been made. It may be done in any reasonable manner, but not in such a way as to suggest that the use is supported by the licensor.
- Non-Commercial: No use of the material may be made for commercial purposes.
- No Derivatives: Any remix, transformation or creation from the material, the modified material may not be distributed.
- No Additional Restrictions: No legal terms or technological measures may be applied that legally restrict others from making any use permitted by the license.
2.- Authors may adopt other non-exclusive license agreements for distribution of the published version of the work (e.g., depositing it in an institutional telematic archive or publishing it in a monographic volume) as long as the initial publication in this journal is indicated.
3.- Authors are allowed and encouraged to disseminate their work through the Internet (e.g., in institutional telematic archives or on their web page) before and during the submission process, which can produce interesting exchanges and increase citations of the published work. (See The Open Access Effect).







