Introduction:
Identifying clinically meaningful biomarker signatures from complex, multiplatform metabolomics data remains a major challenge in precision oncology. Although individual platforms provide complementary insights into breast cancer biology, an integrated framework for subtype-specific feature extraction and biologically interpretable biomarker discovery remains limited. Building on previous work in lipidomic reprogramming, pan-subtype metabolic fingerprinting, and HER2-focused multimodal integration, this study presents a computational framework for breast cancer biomarker discovery.
Methods:
Breast cancer tissue metabolomics data from the METAcancer cohort were harmonized across LC–MS, GC–MS, and NMR platforms. After matching samples across all three platforms, the dataset comprised 253 tissue samples with 183 LC–MS, 161 GC–MS, and 180 NMR metabolite features. Training data were subsequently oversampled to address class imbalance. A feed-forward artificial neural network with feature-level attention was used to prioritize subtype-discriminative metabolites. Selected features were analyzed using an ANN-based transfer-learning architecture with platform-specific autoencoders to derive compact cross-platform latent representations. Encoder-derived contributions were mapped to individual metabolites, and SHapley Additive exPlanations (SHAP) provided class-specific, directionally resolved interpretation. Contrastive representation learning was incorporated to refine subtype-aware latent representations.
Results:
The framework integrates complementary metabolic signals across platforms while retaining metabolite-level interpretability. The HER2-rich profile is characterized by elevated phosphatidylcholine species PC 30:1, PC 32:1, and PC 32:2, and phosphatidylethanolamine PE 32:1, together with higher contributions from N-acetyl-D-mannosamine, glutamine, taurine, and aspartic acid. Selected triacylglycerol (TAG) species show reduced contributions. SHAP-based attribution provides subtype-resolved interpretation of these HER2-associated metabolic signatures.
Conclusions:
This framework integrates attention-based feature prioritization, ANN transfer learning, contrastive representation learning, oversampling, and SHAP interpretation to translate heterogeneous multiplatform metabolomics into biologically interpretable HER2-associated candidate biomarkers and testable hypotheses for precision oncology.