Open data exploitation
Introduction
The last 3 years have seen an explosion in published manuscripts analyzing open-access health datasets, in many cases presenting misleading or biologically implausible findings. There is a growing evidence base to suggest that this is due in part to artificial intelligence-assisted and formulaic workflows, and publishers are responding by discouraging submissions employing open-access health datasets (Spick et al. 2026).
Key resources
Analysis of excess publications based on the NHANES database by Suchak et al. (2025).
Analysis of excess papers from 9 open datasets by Spick et al. (2026).
Analysis of publications duplicating the same analyses of the NHANES dataset by Maupin, Suchak, et al. (2025).
Potential remedies proposed by Maupin, Spick, et al. (2025).
