Big data, observational research and p-value: a recipe for false positive findings? A study of simulated and real prospective cohorts

IRIS - Institutional Research Information System
IRIS è il sistema di gestione integrata dei dati della ricerca (persone, progetti, pubblicazioni, attività) adottato dall'Università degli Studi dell’Insubria.

IRInSubria - Institutional Repository Insubria
IRInSubria raccoglie, conserva, documenta e dissemina le informazioni sulla produzione scientifica dell'Università degli Studi dell’Insubria anche ai fini della valutazione della ricerca.

Background: An increasing number of observational studies combine large sample sizes with low participation rates, which could lead to standard inference failing to control the false discovery rate. We investigated if the “empirical calibration of p-value” method (EPCV), reliant on negative controls, can preserve type-I error in the context of survival analysis. Methods. Simulated cohort studies with 50% participation rate and two different selection bias mechanisms, and a real-life application on predictors of cancer mortality using data from four population-based cohorts in Northern Italy (n=6976 men and women aged 25-74 years at baseline and 17 years of median follow-up). Results: Type-I error for the standard Cox model was above the 5% nominal level in 15 out of 16 simulated settings; for n=10,000, the chances of a null association with hazard ratio=1.05 having a p-value<0.05 were 42.5%. Conversely, EPCV with 10 negative controls preserved the 5% nominal level in all the simulation settings, reducing bias in the point estimate by 80-90% when its main assumption was verified. In the real case, 15 out of 21 (71%) blood markers with no association with cancer mortality according to literature had a p-value<0.05 in age- and gender-adjusted Cox models. After calibration, only 1 (4.8%) remained statistically significant. Conclusions. In the analyses of large observational studies prone to selection bias, the use of empirical distribution to calibrate p-values can substantially reduce the number of trivial results needing further screening for relevance and external validity

Big data, observational research and p-value: a recipe for false positive findings? A study of simulated and real prospective cohorts

Giovanni Veronesi^Primo;Guido Grassi;Giordano Savelli;Piero Quatto;Antonella Zambon

2020-01-01

Abstract

Background: An increasing number of observational studies combine large sample sizes with low participation rates, which could lead to standard inference failing to control the false discovery rate. We investigated if the “empirical calibration of p-value” method (EPCV), reliant on negative controls, can preserve type-I error in the context of survival analysis. Methods. Simulated cohort studies with 50% participation rate and two different selection bias mechanisms, and a real-life application on predictors of cancer mortality using data from four population-based cohorts in Northern Italy (n=6976 men and women aged 25-74 years at baseline and 17 years of median follow-up). Results: Type-I error for the standard Cox model was above the 5% nominal level in 15 out of 16 simulated settings; for n=10,000, the chances of a null association with hazard ratio=1.05 having a p-value<0.05 were 42.5%. Conversely, EPCV with 10 negative controls preserved the 5% nominal level in all the simulation settings, reducing bias in the point estimate by 80-90% when its main assumption was verified. In the real case, 15 out of 21 (71%) blood markers with no association with cancer mortality according to literature had a p-value<0.05 in age- and gender-adjusted Cox models. After calibration, only 1 (4.8%) remained statistically significant. Conclusions. In the analyses of large observational studies prone to selection bias, the use of empirical distribution to calibrate p-values can substantially reduce the number of trivial results needing further screening for relevance and external validity

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno
	
				2020
			
	Rivista
	
				INTERNATIONAL JOURNAL OF EPIDEMIOLOGY
			
	DOI
	
				https://dx.doi.org/10.1093/ije/dyz206
			
	Codice PUBMED
	
				31620789
			
	Codice Web of Science
	
				WOS:000593364900021
			
	Codice Scopus
	
				2-s2.0-85089126835
			
	Parole chiave
	
				Big data; Calibration of P-value; Cohort studies; Observational studies; Selection bias; Survival analysis
			
	Tutti gli autori
	
						Veronesi, Giovanni; Grassi, Guido; Savelli, Giordano; Quatto, Piero; Zambon, Antonella
					
	Appare nelle tipologie:
	
				Articolo su Rivista

File in questo prodotto:

File	Dimensione	Formato
dyz206.pdf accesso aperto Tipologia: Versione Editoriale (PDF) Licenza: Creative commons Dimensione 543.67 kB Formato Adobe PDF Visualizza/Apri	543.67 kB	Adobe PDF	Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11383/2081252

Attenzione

L'Ateneo sottopone a validazione solo i file PDF allegati

Citazioni

3

7

6

social impact