EventsThe 5th International Electronic Conference on Metabolomics
Published
This submission belongs to the session S3. Advanced Data Analysis and Integration in Metabolomics of the event The 5th International Electronic Conference on Metabolomics
Published date
09 Oct, 2026
Academic Editor
author-avatarReza Salek
Citation
Huaxu Yu, Jun Ding, Tong Shen, Min Liu, Yuanyue Li, Oliver Fiehn, MassCube: a Python framework for end-to-end metabolomics data processing from raw files to phenotype classifiers, in Proceedings of The 5th International Electronic Conference on Metabolomics, 14 October–16 October 2026, MDPI: Basel, Switzerland
Share
Email
Facebook
Twitter
LinkedIn

MassCube: a Python framework for end-to-end metabolomics data processing from raw files to phenotype classifiers

Jun Ding 2
Min Liu 1
Yuanyue Li 3
image
1. West Coast Metabolomics Center, University of California, Davis, California, 95616, USA
2. Wuhan Botanical Garden, Chinese Academy of Sciences, Wuhan, 430074, China
3. School of Medicine, Zhejiang University, Hangzhou, China
Abstract

MS-based nontargeted metabolomics now involves thousands of data files across multiple assays and laboratories. Meanwhile, new mass spectrometers like Orbitrap Astral MS have increased raw data sizes eightfold, demanding more efficient processing. Conventional software, including MS-DIAL, MZmine3, and XCMS, struggles with false-positive peak detection, inefficient in-source fragment handling, and slow computation. MassCube, an open-source Python framework, addresses these challenges with advanced peak detection, adduct/in-source fragment recognition, and high-performance parallel computing. Leveraging array-based programming, MassCube enables scalable, high-throughput analysis for large metabolomics datasets. Systematic benchmarking shows MassCube outperforms leading tools in speed, accuracy, and robustness. MassCube’s peak detection was benchmarked against MS-DIAL, MZmine3, and XCMS using both synthetic and experimental LC-MS datasets. In synthetic benchmarking, MassCube achieved 96.5% accuracy, significantly outperforming MS-DIAL (85.4%), MZmine3 (88.4%), and XCMS (87.4%). Experimental benchmarking involved eight LC-MS datasets including human plasma, serum, urine, fecal material, mouse plasma, fruit fly, and plant extracts. A total of 722 manually labeled ion traces were evaluated, with MassCube demonstrating superior accuracy for double-peak (93.5%) and single-peak (97.3%) detection, outperforming MS-DIAL (63.0% and 89.8%), MZmine3 (30.4% and 81.8%), and XCMS (28.3% and 70.3%). Additionally, MassCube exhibited unmatched processing speed and scalability. On a 105 GB Orbitrap Astral dataset, it processed 636 files in 64 minutes on a MacBook M3 Pro (36 GB memory, 12 cores), whereas other software required 8–24 times longer. The framework effectively handled large datasets while maintaining a low memory footprint, enabling high-throughput metabolomics analysis on standard computing hardware. Its modular, object-oriented design also facilitates rapid integration of new algorithms, making it adaptable for evolving metabolomics workflows.

Keywords
metabolomics
mass spectrometry
data processing
large-scale metabolomics
feature detection
Fluorine-Tagging Strategies for Amino Acid Detection by Low-Field ¹⁹F NMR Spectroscopy
Metabolomic Signatures of Polyphenol Intervention in Non-Alcoholic Fatty Liver Disease: Pathway-Level Insights into Lipid and Energy Metabolism