Recent developments in omics platforms have generated large datasets that offer unprecedented opportunities for data mining and discovery. Among omics, metabolomics yields an immediate, functional readout of tumor biology and cellular state. Despite that, its integration into oncological research continues to fall behind genomics and transcriptomics. This delay is largely driven by a scarcity of extensive patient cohorts, compounded by technical hurdles in measuring diverse metabolites across wide dynamic ranges. For example, distinct metabolites require unique extraction protocols, whereas metabolomics cannot be run on archived tissues, since flash freezing is required, and paraffin-embedding is the most common method of tissue storage. Consequently, expanding access to large metabolomic repositories remains a critical priority for the field.
To bridge this data gap, we established a machine learning framework incorporating statistical models, deep learning architectures, and generative AI. Our strategy relies on two main pillars: (1) synthesizing realistic metabolomic profiles and (2) predicting metabolic abundance directly from transcriptomic signatures. Deploying this pipeline allowed us to identify patterns of tumor metabolic reprogramming associated with patient survival, and pinpoint distinct metabolic vulnerabilities across specific cancer subtypes. Crucially, these computationally prioritized targets were experimentally verified in vitro using targeted pharmacological and genetic perturbation assays.
In summary, this work presents a suite of computational tools that resolve longstanding bottlenecks in metabolomics, offering actionable pathways toward target discovery in oncology.