Professor
Supervisor of Doctorate Candidates
E-Mail:
School/Department:Department of Electronic Science, Xiamen University
Business Address:Room B409, Wenxuan Building, Xiang’an Campus, Xiamen University
The Last Update Time: ..
2. Metabolomics
Metabolomics is an omics technology that systematically studies small-molecule metabolites in biological systems and their patterns of change. It focuses on the composition, abundance changes, and pathway activities of metabolites in cells, tissues, organs, body fluids, or entire organisms. Compared with genomics and proteomics, metabolomics is closer to biological phenotypes. Genes and proteins reflect upstream regulatory information in biological processes, whereas metabolites are often located at the terminal or near-terminal positions of biochemical reactions and can more directly reflect functional changes of the body under a specific time, environment, and disease state. Therefore, metabolomics is often regarded as an important bridge connecting molecular mechanisms and disease phenotypes [1-3].
From a historical perspective, the concept of metabolomics did not emerge suddenly. Early researchers had already noticed that small-molecule components in urine, blood, and tissue extracts could reflect the physiological and pathological states of individuals. After the 1970s, the development of gas chromatography, mass spectrometry, and nuclear magnetic resonance gradually moved small-molecule analysis in complex biological samples from qualitative observation toward systematic quantitative analysis. After the 1990s, with advances in high-resolution mass spectrometry, nuclear magnetic resonance spectroscopy, multivariate statistics, and bioinformatics, global metabolite analysis gradually developed into a relatively complete research system. In 1999, Nicholson and colleagues proposed the concept of metabonomics, emphasizing the use of multivariate statistical methods to understand the dynamic metabolic responses of biological systems to pathological stimuli or genetic modifications [1]. In 2002, Fiehn further systematically discussed the role of metabolomics in linking genotypes and phenotypes, making metabolomics an important research direction in the post-genomic era [2].
The basic research objects of metabolomics are metabolites. Common metabolites include amino acids, organic acids, sugars, lipids, nucleotides, bile acids, hormones, drugs, and their metabolic products. These molecules participate in energy metabolism, redox balance, signal transduction, cell membrane construction, inflammatory responses, immune regulation, and cell fate determination. Because metabolites are diverse in type, span a wide concentration range, and differ greatly in physicochemical properties, a single analytical platform can hardly cover all metabolites. Therefore, metabolomics usually requires the combined use of platforms such as nuclear magnetic resonance, gas chromatography-mass spectrometry, liquid chromatography-mass spectrometry, capillary electrophoresis-mass spectrometry, high-resolution mass spectrometry, and ion mobility mass spectrometry. Different platforms have their own characteristics. Nuclear magnetic resonance has good reproducibility, simple sample preparation, and stable quantification, but its sensitivity is relatively low. Mass spectrometry has high sensitivity and broad coverage and is suitable for detecting low-abundance metabolites, but it places higher demands on sample preparation, chromatographic separation, ionization efficiency, and data processing methods [3-5].
In terms of research workflow, metabolomics usually includes sample collection, sample preparation, instrumental analysis, data preprocessing, metabolite annotation, statistical modeling, pathway analysis, and biological interpretation. Sample types may include serum, plasma, urine, cerebrospinal fluid, tissue, cells, feces, or tissue sections. The data obtained after instrumental analysis often contain a large number of peak signals, which require peak extraction, retention time correction, peak matching, normalization, missing value processing, batch effect correction, and quality control. Researchers then usually use principal component analysis, partial least squares discriminant analysis, linear models, machine learning, metabolic pathway enrichment analysis, and network analysis to identify metabolic features associated with diseases, drugs, environmental exposure, or physiological states from high-dimensional metabolic data. The emergence of tools such as XCMS has promoted the standardization and automation of LC-MS-based untargeted metabolomics data processing [6]. The development of databases such as HMDB, METLIN, KEGG, and LipidMaps has also provided an important foundation for metabolite annotation and pathway interpretation [7].
Metabolomics can be divided into targeted metabolomics and untargeted metabolomics. Targeted metabolomics usually focuses on a predefined group of metabolites. It provides accurate quantification and good reproducibility and is suitable for validating candidate biomarkers and studying specific metabolic pathways. Untargeted metabolomics aims to detect metabolite signals in a sample as broadly as possible and is suitable for discovering unknown metabolic alterations and new disease clues. The two approaches do not replace each other. In actual studies, untargeted analysis is often used to discover candidate metabolites, whereas targeted analysis is used for subsequent validation and quantification. In recent years, with advances in high-resolution mass spectrometry, chemical derivatization, stable isotope labeling, and machine learning methods, the detection coverage, quantitative capability, and structural annotation ability of metabolomics have continued to improve.
An important feature of metabolomics is its high sensitivity to changes in physiological state. Diet, age, sex, medication, gut microbiota, circadian rhythm, exercise, environmental exposure, and disease progression can all affect the metabolic profile. For this reason, metabolomics has high biological information content, but it also faces strong confounding factors. Proper experimental design, strict sample collection protocols, stable quality control samples, reliable normalization methods, and appropriate statistical models are prerequisites for obtaining credible conclusions. If metabolomics data remain only at the level of differential metabolite lists, the results can easily become fragmented. A more valuable analysis is to place metabolic changes back into metabolic pathways, tissue function, cellular states, and disease progression, thereby explaining the mechanisms underlying metabolic abnormalities [8].
The significance of metabolomics lies in its ability to observe changes in biological systems at the functional level. In many diseases, metabolic networks have already changed before clinical manifestations appear. Abnormal glucose metabolism, lipid metabolism disorders, altered amino acid metabolism, increased oxidative stress, and energy metabolic reprogramming may all serve as important clues for early diagnosis, subtyping, therapeutic evaluation, and prognostic assessment. In cancer research, metabolomics can be used to analyze tumor metabolic reprogramming, identify serum or tissue biomarkers, evaluate treatment responses, and investigate mechanisms of drug resistance. In neurodegenerative disease research, metabolomics can be used to observe changes in energy metabolism, lipid metabolism, and neurotransmitter-related pathways in brain regions, blood, or cerebrospinal fluid. In diabetes, diabetic nephropathy, and liver disease research, metabolomics can help explain the continuous process from metabolic disorder to organ injury. In drug research and toxicology, metabolomics can be used to evaluate drug efficacy, toxic responses, drug metabolism, and individual differences [8-9].
The application areas of metabolomics are already very broad. In basic life sciences, it can be used to interpret gene function, cellular metabolic regulation, and interactions among tissues and organs. In clinical medicine, it can be used for biomarker discovery, risk prediction, disease subtyping, and individualized treatment. In public health and environmental science, it can be used to study the effects of environmental pollutants, dietary patterns, and lifestyle on human metabolic states. In food science and traditional Chinese medicine research, it can be used for quality evaluation, origin identification, active compound discovery, and mechanism analysis. Against the background of rapid development in spatial omics, metabolomics can also be integrated with mass spectrometry imaging, pathological imaging, spatial transcriptomics, and immune imaging, moving from global metabolic profiling toward spatial metabolic mechanism research.
Metabolomics still faces several key challenges. First, metabolite structures are complex, many peak signals are difficult to annotate accurately, and unknown metabolites still account for a large proportion of the data. Second, different instrument platforms, experimental batches, and sample preparation methods can introduce systematic differences and affect reproducibility. Third, metabolites are highly correlated with one another, and the same metabolite may participate in multiple pathways. Simple univariate differential analysis is therefore insufficient for explaining complex metabolic regulation. Finally, metabolomics research often needs to be integrated with transcriptomics, proteomics, microbiomics, radiomics, and clinical data, which places higher demands on data fusion, statistical modeling, and interpretability. Therefore, the development of metabolomics is moving from simple screening of differential metabolites toward mechanism interpretation, network modeling, multi-omics integration, and intelligent analysis.
Our team’s research interests in metabolomics mainly focus on the integration of computational methods, metabolic network modeling, and disease applications. Based on our existing research foundation, our team has long focused on machine learning, computational mass spectrometry, spatial multi-omics, molecular network modeling, and health big data. Metabolomics is an important intersection among these directions. Our team aims to use signal processing, image processing, high-dimensional data analysis, and machine learning methods to improve the capabilities of metabolomics data processing, feature extraction, pattern recognition, and mechanism interpretation, and to promote metabolomics from data-driven biomarker screening toward interpretable disease mechanism analysis.
In terms of data processing methods, our team focuses on missing values, batch effects, normalization, and quality control in metabolomics data. Metabolomics data are often affected by sample dilution, instrumental drift, ion suppression, and inter-batch differences. If these issues are not properly handled, true biological signals may be weakened, or even incorrect conclusions may be produced. To address these problems, our team has developed methods for missing value imputation, batch effect correction, and large-scale data normalization in mass spectrometry-based metabolomics. For example, an NMF-based method was used to handle missing values in mass spectrometry metabolomics data, CordBat was developed for batch effect correction in large-scale metabolomics, and Local Neighbor Normalization and Local Sample Cohesion Normalization further focused on preserving biological heterogeneity during normalization [10-13]. These studies point to a shared core issue: improving data consistency while preserving real biological information carried by disease status, individual differences, and metabolic heterogeneity as much as possible.
In metabolic pathway and metabolic network analysis, our team focuses on moving from differences in individual metabolites toward an interpretation of system-level metabolic changes. Traditional metabolic pathway enrichment analysis is often affected by insufficient metabolite coverage, pathway overlap, and complex correlation structures among metabolites. To address these issues, our team has developed methods such as overlapping group PLS, iMSEA, dci-MSEA, iMS2Net, and SynNet to analyze metabolite sets, metabolic pathways, metabolic coordination relationships, and multiscale metabolic interaction networks [14-18]. These methods do not simply count which metabolites are increased or decreased. Instead, they further analyze coordinated changes among metabolites, relationships among pathways, and metabolic network reconstruction under disease states. For problems such as cancer, Alzheimer’s disease, diabetic nephropathy, and multi-organ metabolic responses, these network-based methods can help reveal deeper metabolic regulatory features.
In disease applications, our team focuses on the use of metabolomics in the diagnosis, subtyping, prognostic evaluation, and mechanistic study of major diseases. Related research involves hepatocellular carcinoma, nasopharyngeal carcinoma, gastric cancer, Alzheimer’s disease, diabetic nephropathy, and other metabolism-related diseases. The team previously undertook a project on new methods for metabolomics data fusion and modeling and their application in diabetic nephropathy research, as well as research on new NMR-based metabolomics methods and their applications in the early diagnosis of diabetes, gestational diabetes, and diabetic nephropathy. In recent years, the team has further integrated metabolomics with machine learning, spatial metabolomics, and multimodal analysis for research on spatial metabolic mechanisms in Alzheimer’s disease, metabolic network analysis of hepatocellular carcinoma, and multimodal cross-scale intelligent metabolic subtyping of nasopharyngeal carcinoma. In these studies, metabolomics is used not only to identify candidate biomarkers, but also to explain disease progression, tissue heterogeneity, and abnormalities in metabolic regulation.
In spatial metabolomics and multi-omics integration, our team focuses on how to combine conventional metabolomics based on body fluids or tissue extracts with mass spectrometry imaging, histopathology, medical imaging, and clinical data. Conventional metabolomics can obtain relatively comprehensive metabolic profiles, but usually lacks spatial location information. Mass spectrometry imaging can observe the spatial distribution of metabolites in tissues, but also faces problems such as high noise, limited resolution, and difficult molecular annotation. Through spatial metabolomics and multimodal data fusion, metabolite changes can be linked with tissue structure, cellular states, and disease regions, providing more complete information for disease mechanism research. The team’s work on mass spectrometry imaging data segmentation, denoising, ion image representation learning, multimodal image fusion, and spatial heterogeneity analysis also provides a methodological foundation for the spatial, network-based, and intelligent development of metabolomics.
Overall, our team’s research goal in metabolomics is to develop computational metabolomics methods for complex disease research. Specifically, on the one hand, it is necessary to address missing values, noise, batch effects, normalization, and high-dimensional modeling problems in metabolomics data, thereby improving the stability and reproducibility of data processing. On the other hand, through metabolic pathway analysis, molecular network modeling, machine learning, and multi-omics integration, it is necessary to interpret metabolic heterogeneity and systemic metabolic reconstruction in disease. Through these studies, our team hopes to promote metabolomics from metabolite detection and biomarker screening toward an interpretable tool for disease mechanism analysis, providing new methodological support for early disease diagnosis, molecular subtyping, prognostic evaluation, and precision medicine research.
References
[1] Nicholson, J. K.; Lindon, J. C.; Holmes, E. “Metabonomics”: Understanding the Metabolic Responses of Living Systems to Pathophysiological Stimuli via Multivariate Statistical Analysis of Biological NMR Spectroscopic Data. Xenobiotica, 1999, 29(11): 1181–1189.
[2] Fiehn, O. Metabolomics: The Link between Genotypes and Phenotypes. Plant Molecular Biology, 2002, 48(1–2): 155–171.
[3] Patti, G. J.; Yanes, O.; Siuzdak, G. Metabolomics: The Apogee of the Omics Trilogy. Nature Reviews Molecular Cell Biology, 2012, 13(4): 263–269.
[4] Dettmer, K.; Aronov, P. A.; Hammock, B. D. Mass Spectrometry-Based Metabolomics. Mass Spectrometry Reviews, 2007, 26(1): 51–78.
[5] Beckonert, O.; Keun, H. C.; Ebbels, T. M. D.; Bundy, J.; Holmes, E.; Lindon, J. C.; Nicholson, J. K. Metabolic Profiling, Metabolomic and Metabonomic Procedures for NMR Spectroscopy of Urine, Plasma, Serum and Tissue Extracts. Nature Protocols, 2007, 2(11): 2692–2703.
[6] Smith, C. A.; Want, E. J.; O’Maille, G.; Abagyan, R.; Siuzdak, G. XCMS: Processing Mass Spectrometry Data for Metabolite Profiling Using Nonlinear Peak Alignment, Matching, and Identification. Analytical Chemistry, 2006, 78(3): 779–787.
[7] Wishart, D. S.; Tzur, D.; Knox, C.; et al. HMDB: The Human Metabolome Database. Nucleic Acids Research, 2007, 35(Database issue): D521–D526.
[8] Johnson, C. H.; Ivanisevic, J.; Siuzdak, G. Metabolomics: Beyond Biomarkers and towards Mechanisms. Nature Reviews Molecular Cell Biology, 2016, 17(7): 451–459.
[9] Wishart, D. S. Emerging Applications of Metabolomics in Drug Discovery and Precision Medicine. Nature Reviews Drug Discovery, 2016, 15(7): 473–484.
[10] Xu, J.; Wang, Y.; Xu, X.; Cheng, K. K.; Raftery, D.; Dong, J. NMF-Based Approach for Missing Values Imputation of Mass Spectrometry Metabolomics Data. Molecules, 2021, 26(19): 5787.
[11] Guo, F.; Lin, G.; Dong, L.; Cheng, K. K.; Deng, L.; Xu, X.; Raftery, D.; Dong, J. Concordance-Based Batch Effect Correction for Large-Scale Metabolomics. Analytical Chemistry, 2023, 95(18): 7220–7228.
[12] Lu, K.; Liu, Y.; Cheng, K. K.; Guo, F.; Deng, L.; Dong, J. Local Neighbor Normalization: Reconciling Accurate Normalization and Heterogeneity Recovery in Large-Scale Metabolomics. Analytica Chimica Acta, 2025, 1372: 344440.
[13] Guo, F.; Deng, L.; Cheng, K. K.; Lu, K.; Wang, Y.; Guo, L.; Raftery, D.; Dong, J. Local Sample Cohesion Normalization: Preserving Inherent Biological Heterogeneity in Metabolomics Data. Analytical Chemistry, 2026, 98(1): 343–353.
[14] Deng, L.; Ma, L.; Cheng, K. K.; Xu, X.; Raftery, D.; Dong, J. A Sparse PLS-Based Method for Overlapping Metabolite Set Enrichment Analysis. Journal of Proteome Research, 2021, 20(6): 3204–3213.
[15] Dong, J.; Peng, Q.; Deng, L.; Liu, J.; Huang, W.; Zhou, X.; Zhao, C.; Cai, Z. iMS2Net: A Multiscale Networking Methodology to Decipher Metabolic Synergy of Organism. iScience, 2022, 25(9): 104896.
[16] Wang, Y.; Liu, X.; Dong, L.; Cheng, K. K.; Lin, C.; Wang, X.; Dong, J.; Deng, L.; Raftery, D. iMSEA: A Novel Metabolite Set Enrichment Analysis Strategy to Decipher Drug Interactions. Analytical Chemistry, 2023, 95(15): 6203–6211.
[17] Lin, G.; Dong, L.; Cheng, K. K.; Xu, X.; Wang, Y.; Deng, L.; Raftery, D.; Dong, J. Differential Correlations Informed Metabolite Set Enrichment Analysis to Decipher Metabolic Heterogeneity of Disease. Analytical Chemistry, 2023, 95(33): 12505–12513.
[18] Wang, Y.; Zhu, Z.; Deng, L.; Cheng, K. K.; Guo, F.; Lin, G.; Raftery, D.; Dong, J. Multiscale Synergy Networks Offer Insights into Disease and Comorbidity Mechanisms. Analytical Chemistry, 2025, 97(6): 3633–3642.