Cross-validation failure: Small sample sizes lead to large error bars

Gaël Varoquaux

doi:10.1016/j.neuroimage.2017.06.061

Cross-validation failure: Small sample sizes lead to large error bars

Neuroimage. 2018 Oct 15;180(Pt A):68-77. doi: 10.1016/j.neuroimage.2017.06.061. Epub 2017 Jun 24.

Author

Gaël Varoquaux¹

Affiliation

¹ Parietal Project-team, INRIA Saclay-île de France, France; CEA/Neurospin bât 145, 91191 Gif-Sur-Yvette, France; Université Paris-Saclay, Saclay, France. Electronic address: gael.varoquaux@inria.fr.

PMID: 28655633
DOI: 10.1016/j.neuroimage.2017.06.061

Abstract

Predictive models ground many state-of-the-art developments in statistical brain image analysis: decoding, MVPA, searchlight, or extraction of biomarkers. The principled approach to establish their validity and usefulness is cross-validation, testing prediction on unseen data. Here, I would like to raise awareness on error bars of cross-validation, which are often underestimated. Simple experiments show that sample sizes of many neuroimaging studies inherently lead to large error bars, eg±10% for 100 samples. The standard error across folds strongly underestimates them. These large error bars compromise the reliability of conclusions drawn with predictive models, such as biomarkers or methods developments where, unlike with cognitive neuroimaging MVPA approaches, more samples cannot be acquired by repeating the experiment across many subjects. Solutions to increase sample size must be investigated, tackling possible increases in heterogeneity of the data.

Keywords: Biomarkers; Cross-validation; Decoding; MVPA; Model selection; Statistics; fMRI.

Publication types

Research Support, Non-U.S. Gov't

MeSH terms

Brain Mapping / methods*
Brain Mapping / standards*
Humans
Image Processing, Computer-Assisted / methods*
Image Processing, Computer-Assisted / standards*
Magnetic Resonance Imaging / methods
Magnetic Resonance Imaging / standards
Reproducibility of Results
Sample Size*