Professor
Supervisor of Doctorate Candidates
E-Mail:
School/Department:Department of Electronic Science, Xiamen University
Business Address:Room B409, Wenxuan Building, Xiang’an Campus, Xiamen University
The Last Update Time: ..
3. Computatial Mass Spectrometry
Computational mass spectrometry is an interdisciplinary research field that integrates mass spectrometry, computer science, statistics, cheminformatics, and machine learning. Its central concern is not whether a mass spectrometer can acquire signals, but how complex mass spectrometric signals can be converted into reliable molecular information, structural information, and biomedical interpretation. A mass spectrometer directly records information such as mass-to-charge ratio, signal intensity, retention time, tandem mass spectra, and ion mobility. Researchers, however, are usually more interested in compound structures, molecular identities, quantitative results, and disease-associated molecular features. Computational mass spectrometry serves as a key link between raw instrumental signals and scientific conclusions.
With the development of liquid chromatography-mass spectrometry, high-resolution mass spectrometry, tandem mass spectrometry, and ion mobility mass spectrometry, mass spectrometry data have evolved from relatively simple peak lists into high-dimensional and multiscale data systems. Modern mass spectrometry experiments can generate thousands to millions of spectra in a single study. These data contain noise, drift, missing values, co-elution, isotope peaks, adduct ions, in-source fragments, and complex fragmentation relationships. Manual spectral interpretation is no longer sufficient for high-throughput research. Automated, standardized, and intelligent data analysis methods have therefore become an essential foundation for mass spectrometry research [1-6].
The development of computational mass spectrometry was initially driven by the needs of proteomics and small-molecule structural analysis. In proteomics, software tools such as MaxQuant and OpenMS have improved the automation of peptide identification, protein inference, and quantitative analysis [2,3]. In small-molecule analysis, tools such as XCMS, MZmine, and MS-DIAL have established relatively complete workflows for peak detection, peak matching, retention time correction, deconvolution, and quantitative analysis [4-6]. Subsequently, methods such as GNPS, SIRIUS, and MetFrag further promoted spectral library sharing, molecular formula prediction, in silico fragmentation, and structural candidate ranking [7-9]. These advances have expanded computational mass spectrometry from data preprocessing to molecular identification and structural interpretation.
The main tasks of computational mass spectrometry include data format conversion, quality control, peak detection, peak alignment, isotope peak recognition, adduct ion annotation, retention time correction, tandem spectrum interpretation, spectral library search, molecular formula prediction, structural candidate ranking, quantitative modeling, statistical analysis, and result visualization. Different types of mass spectrometry data raise different computational problems. LC-MS data emphasize peak shape recognition and cross-sample matching. MS/MS data emphasize fragmentation relationships and structural interpretation. Ion mobility mass spectrometry further emphasizes collision cross section and gas-phase conformational information. Although these application scenarios differ, their common goal is to extract reliable, stable, and interpretable molecular evidence from complex signals.
One important feature of computational mass spectrometry is that it requires an understanding of both signal processing and chemical rules. Mass spectrometry data are not ordinary high-dimensional matrices. They are complex signals shaped by instrumental response, ionization processes, separation processes, and molecular structures. Peak shapes, isotope distributions, adduct ion types, fragmentation patterns, retention time drift, ion suppression, and matrix effects can all affect subsequent interpretation. If general statistical models are used without considering the chemical origin of mass spectrometry signals, false matching and overinterpretation can easily occur. Computational mass spectrometry methods therefore need to integrate signal processing, statistical modeling, and chemical priors to balance automated analysis with result reliability.
Another important feature of computational mass spectrometry is its strong dependence on databases and knowledge bases. Proteomics usually relies on protein sequence databases, whereas small-molecule and metabolite analysis often relies on resources such as HMDB, METLIN, MassBank, GNPS, KEGG, LipidMaps, and PubChem. Databases can improve annotation efficiency, but they also introduce problems such as limited coverage, isomeric ambiguity, and excessive structural candidates. Many untargeted mass spectrometry signals cannot be fully matched in existing spectral libraries. For this reason, computational mass spectrometry is moving beyond simple spectral library matching toward molecular formula prediction, in silico fragmentation, structural similarity search, machine learning-based prediction, and generative model-assisted inference [7-9].
The significance of computational mass spectrometry lies in the fact that it directly determines whether mass spectrometry data can be transformed into reliable scientific knowledge. High-resolution mass spectrometers can generate large amounts of data, but raw data do not automatically produce conclusions. Whether a peak is real, whether peaks from different samples correspond to the same compound, whether a signal comes from an isotope peak or an adduct ion, whether a candidate structure is credible, and whether a set of differences has biological significance all require computational assessment. For large-scale cohort studies, clinical mass spectrometry, drug metabolism research, and multi-omics integration, computational mass spectrometry directly affects the reproducibility and interpretive depth of the results.
Computational mass spectrometry still faces several bottlenecks. Complex samples contain many low-abundance, co-eluting, and isomeric signals, which can interfere with feature extraction and quantification. Molecular annotation remains insufficient in untargeted studies. Many signals can only be assigned candidate molecular formulas or candidate structures, and they are still far from rigorous structural confirmation. Data differences across platforms, batches, and laboratories also affect cross-cohort analysis and model generalization. Deep learning and large models are now entering mass spectrometry research, but their predictions still need to be constrained by chemical rules, experimental validation, and uncertainty assessment.
Our team's research interests in computational mass spectrometry mainly focus on mass spectrometry data quality control, untargeted molecular identification, multidimensional mass spectrometry information modeling, and engineering-oriented mass spectrometry software development. With an interdisciplinary background in electronic information, physics, analytical chemistry, mass spectrometry, and biomedical data analysis, our team has long been interested in using signal processing, high-dimensional data analysis, and machine learning to solve key problems in real-world mass spectrometry data. In our research, computational mass spectrometry is not merely software tool development. Rather, it serves as a foundational methodology supporting metabolomics, multi-omics data analysis, and health big data research.
In data quality control, our team focuses on missing values, batch effects, normalization, and the preservation of biological heterogeneity in mass spectrometry data. Mass spectrometry data are easily affected by ion suppression, instrumental drift, sample dilution, and inter-batch variation. Simple correction methods may improve data consistency but can also weaken true disease-related and individual differences. To address this issue, our team has developed an NMF-based missing value imputation method, the CordBat batch effect correction method, Local Neighbor Normalization, and Local Sample Cohesion Normalization. These methods aim to improve data stability while preserving true biological heterogeneity as much as possible [10-13].
In untargeted small-molecule identification, our team focuses on combining chemical derivatization strategies, high-resolution mass spectrometric features, and intelligent software to improve the discovery of unknown molecules in complex samples. MS-IDF uses a narrow mass defect filtering strategy for the untargeted identification of endogenous metabolites after chemical isotope labeling [14]. MS-TDF combines a triple-dimensional combinatorial derivatization strategy with mass spectrometry data processing software to enhance the identification of compounds containing specific functional groups [15]. MS-DDF is designed for animal-derived medicines and complex natural products. It provides an intelligent high-resolution mass spectrometry data processing workflow for deciphering bioactive small molecules [16]. These studies reflect our team’s strategy of integrating chemical reaction design, mass spectrometric signal patterns, and computational workflows.
In multidimensional mass spectrometry information modeling, our team focuses on the relationship among ion mobility mass spectrometry, collision cross section, and molecular structural information. Ion mobility mass spectrometry provides an additional separation dimension beyond mass-to-charge ratio and retention time. Collision cross section reflects the gas-phase conformational properties of ions and is important for distinguishing isomers and assisting small-molecule annotation. Our team has developed a large language model-empowered method for compound collision cross section prediction, aiming to integrate chemical structure representation, deep learning, and mass spectrometry prior knowledge to improve CCS prediction and support small-molecule structural identification [17].
In engineering-oriented software development, our team emphasizes the translation of algorithmic methods into practical tools. Complex mass spectrometry data analysis requires not only theoretical models, but also stable, user-friendly, and reproducible software systems. Our team has continuously developed mass spectrometry data processing tools for real research scenarios, covering untargeted identification, batch effect correction, data normalization, molecular structure-assisted prediction, and batch processing. These tools emphasize algorithmic efficiency, batch processing capability, and result interpretability, enabling computational mass spectrometry methods to better serve metabolomics, drug analysis, clinical mass spectrometry, and molecular identification in complex samples.
Overall, our team's goal in computational mass spectrometry is to develop intelligent mass spectrometry data analysis methods for complex biomedical problems. On the one hand, we focus on intrinsic data problems such as noise, missing values, batch effects, peak recognition, spectral interpretation, and molecular annotation, aiming to establish more robust and more automated data processing workflows. On the other hand, by integrating chemical derivatization, high-resolution mass spectrometry, ion mobility mass spectrometry, deep learning, and engineering-oriented software development, we aim to advance mass spectrometry data from raw spectra toward reliable molecular identification and interpretable biomedical applications. Through these studies, computational mass spectrometry can serve as an important methodological foundation linking mass spectrometry instrumentation, molecular identification, and precision medicine, providing new technical support for disease mechanism research, biomarker discovery, drug evaluation, and multi-omics integration.
References
[1] Aebersold, R.; Mann, M. Mass-Spectrometric Exploration of Proteome Structure and Function. Nature, 2016, 537(7620): 347–355.
[2] Cox, J.; Mann, M. MaxQuant Enables High Peptide Identification Rates, Individualized p.p.b.-Range Mass Accuracies and Proteome-Wide Protein Quantification. Nature Biotechnology, 2008, 26(12): 1367–1372.
[3] Röst, H. L.; Sachsenberg, T.; Aiche, S.; Bielow, C.; Weisser, H.; Aicheler, F.; Andreotti, S.; Ehrlich, H.-C.; Gutenbrunner, P.; Kenar, E.; et al. OpenMS: A Flexible Open-Source Software Platform for Mass Spectrometry Data Analysis. Nature Methods, 2016, 13(9): 741–748.
[4] Smith, C. A.; Want, E. J.; O’Maille, G.; Abagyan, R.; Siuzdak, G. XCMS: Processing Mass Spectrometry Data for Metabolite Profiling Using Nonlinear Peak Alignment, Matching, and Identification. Analytical Chemistry, 2006, 78(3): 779–787.
[5] Pluskal, T.; Castillo, S.; Villar-Briones, A.; Orešič, M. MZmine 2: Modular Framework for Processing, Visualizing, and Analyzing Mass Spectrometry-Based Molecular Profile Data. BMC Bioinformatics, 2010, 11: 395.
[6] Tsugawa, H.; Cajka, T.; Kind, T.; Ma, Y.; Higgins, B.; Ikeda, K.; Kanazawa, M.; VanderGheynst, J.; Fiehn, O.; Arita, M. MS-DIAL: Data-Independent MS/MS Deconvolution for Comprehensive Metabolome Analysis. Nature Methods, 2015, 12(6): 523–526.
[7] Wang, M.; Carver, J. J.; Phelan, V. V.; Sanchez, L. M.; Garg, N.; Peng, Y.; Nguyen, D. D.; Watrous, J.; Kapono, C. A.; Luzzatto-Knaan, T.; et al. Sharing and Community Curation of Mass Spectrometry Data with Global Natural Products Social Molecular Networking. Nature Biotechnology, 2016, 34(8): 828–837.
[8] Dührkop, K.; Fleischauer, M.; Ludwig, M.; Aksenov, A. A.; Melnik, A. V.; Meusel, M.; Dorrestein, P. C.; Rousu, J.; Böcker, S. SIRIUS 4: A Rapid Tool for Turning Tandem Mass Spectra into Metabolite Structure Information. Nature Methods, 2019, 16(4): 299–302.
[9] Wolf, S.; Schmidt, S.; Müller-Hannemann, M.; Neumann, S. In Silico Fragmentation for Computer Assisted Identification of Metabolite Mass Spectra. BMC Bioinformatics, 2010, 11: 148.
[10] Xu, J.; Wang, Y.; Xu, X.; Cheng, K. K.; Raftery, D.; Dong, J. NMF-Based Approach for Missing Values Imputation of Mass Spectrometry Metabolomics Data. Molecules, 2021, 26(19): 5787.
[11] Guo, F.; Lin, G.; Dong, L.; Cheng, K. K.; Deng, L.; Xu, X.; Raftery, D.; Dong, J. Concordance-Based Batch Effect Correction for Large-Scale Metabolomics. Analytical Chemistry, 2023, 95(18): 7220–7228.
[12] Lu, K.; Liu, Y.; Cheng, K. K.; Guo, F.; Deng, L.; Dong, J. Local Neighbor Normalization: Reconciling Accurate Normalization and Heterogeneity Recovery in Large-Scale Metabolomics. Analytica Chimica Acta, 2025, 1372: 344440.
[13] Guo, F.; Deng, L.; Cheng, K. K.; Lu, K.; Wang, Y.; Guo, L.; Raftery, D.; Dong, J. Local Sample Cohesion Normalization: Preserving Inherent Biological Heterogeneity in Metabolomics Data. Analytical Chemistry, 2026, 98(1): 343–353.
[14] Wang, S.; Jiang, X.; Ding, R.; Chen, B.; Lyu, H.; Liu, J.; Zhu, C.; Shen, R.; Chen, J.; Hong, Y.; Wu, Y.-L.; Dong, J.; Wu, C. MS-IDF: A Software Tool for Nontargeted Identification of Endogenous Metabolites after Chemical Isotope Labeling Based on a Narrow Mass Defect Filter. Analytical Chemistry, 2022, 94(7): 3194–3202.
[15] Yuan, C.; Jin, Y.; Zhang, H.; Chen, S.; Yi, J.; Xie, Q.; Dong, J.; Wu, C. Strategy to Empower Nontargeted Metabolomics by Triple-Dimensional Combinatorial Derivatization with MS-TDF Software. Analytical Chemistry, 2024, 96(19): 7634–7642.
[16] Zhang, H.; Zhang, D.; Li, J.; Wei, Y.; He, Y.; Yuan, Y.; Ding, A.; Zhang, J.; Zhou, Y.; Dong, J.; Wu, C. MS-DDF: An Intelligent High-Resolution Mass Spectrometry Data-Processing Tool for Deciphering Bioactive Small Molecules in Animal-Derived Medicines. Acta Pharmaceutica Sinica B, 2025, DOI: 10.1016/j.apsb.2025.12.033.
[17] Zhu, Z.; Xie, C.; Lin, S.; Xu, X.; Niu, S.; Guo, L.; Dong, J. Large Language Model-Empowered Compound Collision Cross-Section Prediction. Analytical Chemistry, 2025, 97(36): 19791–19800.