Introduction: While depositing omics data into public repositories is frequently required by grant funders and journals, the quality of deposited data is not uniform, limiting its reuse in large meta-analyses. Here we present a comprehensive evaluation of data quality across all publicly available datasets in Metabolomics Workbench.
Methods: We downloaded 7240 datasets from Metabolomics Workbench in July 2026. Needed harmonizations and repairs were made to downloaded analysis files.
For each dataset, we attempted to:
- convert to a summarized experiment R object containing metabolite metadata, abundances, and sample metadata;
- detect any quality-control samples;
- determine if the data were log-transformed;
- evaluate the number of metabolite features and usable samples in each group of subject-sample-factors (SSF);
- calculate sample-sample information-content-informed Kendall- (ICI-Kt) correlations;
- calculate the relative standard deviations (RSD) within each SSF group;
- perform principal component analysis (PCA) and calculate an analysis of variance (ANOVA) for sample PC scores based on SSF groups.
For each successful step, we assigned a color-based data quality grade: red, think twice before using; yellow, probably OK; and green, looks good for that step. In addition, we generated an HTML QC/QA report for each of the datasets.
Results: Of the 7240 datasets that could be parsed, 167 (2.3%) could not be coerced into a summarized experiment object for various reasons; 622 (8.6%) had <3 samples, and therefore could not undergo PCA and RSD calculations; another 5 (0.069%) started correlation but could not complete all steps, and 6446 (89%) were able to go through all steps to PCA. For successful PCA, 333 had either a single SSF, or every sample was its own SSF, making them inappropriate for differential analyses. Among the 6113 remaining, 2053 (34%) had all green statuses; 2077 (34%) had one or more yellow statuses; and 1983 (32%) had one or more red statuses across the QC/QA evaluations.