Qr code
中文
Jiyang Dong

Professor

Supervisor of Doctorate Candidates


E-Mail:

School/Department:Department of Electronic Science, Xiamen University

Business Address:Room B409, Wenxuan Building, Xiang’an Campus, Xiamen University

Click:Times

The Last Update Time: ..

Current position: Home >>Research Focus

5. Artificial Intelligence +

"Artificial Intelligence Plus" refers to an interdisciplinary research direction that applies machine learning, deep learning, statistical modeling, and intelligent optimization methods to specific scientific problems and complex data analysis scenarios. In biomedical research, its focus is not on developing general-purpose AI algorithms in isolation, but on using AI methods to analyze high-dimensional, complex, noisy, and heterogeneous biomedical data, extract stable features, build interpretable models, and support disease subtyping, risk prediction, mechanistic analysis, and decision support.

A defining feature of this direction is that it is both problem-driven and data-driven. Biomedical data are often characterized by a large number of variables, limited sample size, pronounced batch variation, complex noise sources, and substantial individual differences. Traditional statistical methods can be limited when handling such data because of dimensionality, correlation structures, and nonlinear relationships. Machine learning methods can learn complex patterns from high-dimensional data, but they may also suffer from overfitting, poor generalizability, and limited interpretability. Therefore, the key issue of "Artificial Intelligence Plus" in biomedicine is not merely to improve prediction accuracy, but also to ensure that the selected features are stable, the models are interpretable, the results are verifiable, and the models remain reliable across different cohorts and application settings.

The significance of "Artificial Intelligence Plus" lies in its ability to provide a new analytical framework for complex biomedical data. Disease onset and progression often involve multiple variables, processes, and levels of biological information, and cannot be fully explained by a single indicator. Machine learning and AI methods can integrate multi-source data, identify hidden disease subtypes, discover feature combinations associated with diagnosis or prognosis, and help researchers generate new scientific hypotheses from complex data. In the context of health big data and precision medicine, these methods can help biomedical research move from empirical judgment and single-variable analysis toward a new stage that combines data-driven modeling, model-assisted analysis, and mechanistic interpretation.

Our team's research interests in "Artificial Intelligence Plus" mainly focus on machine learning algorithm design and its application to biomedical data analysis. The team has long worked on machine learning, high-dimensional data analysis and modeling, spatial multi-omics, computational mass spectrometry, molecular network modeling, and health big data. Our work does not simply apply AI tools in a generic manner. Instead, it focuses on real biomedical data problems involving noise, heterogeneity, nonlinear relationships, and interpretability, with the aim of developing more robust modeling methods that are better suited to specific data structures.

In machine learning algorithm design, our team focuses on feature selection, pattern recognition, data transformation, heterogeneity preservation, and robust modeling in high-dimensional biomedical data. These studies emphasize improving the ability of models to capture true biological differences under limited sample sizes and complex variable structures, while reducing misinterpretation caused by noise, batch effects, or overfitting. In tasks such as high-dimensional data analysis, metabolic phenotype modeling, and disease subtyping, the team attaches importance to balancing algorithmic performance, statistical stability, and biological interpretation.

In AI-assisted biomedical data analysis, our team focuses on the application of machine learning methods to disease diagnosis, subtyping, prognostic evaluation, and mechanistic studies. The research and publications listed on the personal profile show that the team has applied explainable machine learning to the robust evaluation of extracted ion chromatograms in LC-MS metabolomics, used machine learning to identify serum biomarkers for nasopharyngeal carcinoma, and developed a density-based resampling strategy to reveal metabolic heterogeneity and subtypes during liver disease progression. These studies reflect the team's research approach of applying AI methods to the analysis of complex disease data.

In deep learning and large model applications, our team focuses on the potential of AI methods in complex biomolecular and instrumental data. The MEMpre work listed on the personal profile uses AlphaFold-driven protein large language models for membrane protein type prediction, while the CCS prediction work applies large language models to compound collision cross-section prediction. These studies indicate that the team is interested not only in conventional machine learning models, but also in how large models, deep representation learning, and domain prior knowledge can be combined to improve the prediction and interpretation of complex biomedical data.

In health big data and multi-source data modeling, our team focuses on how AI methods can be used for integrated analysis of multidimensional medical data, omics data, and clinically relevant data. The projects listed on the personal profile, including multimodal cross-scale intelligent metabolic subtyping of nasopharyngeal carcinoma, early diagnosis and molecular mechanism studies of Alzheimer’s disease, and metabolomics modeling of diabetic nephropathy, reflect the team’s application-oriented direction in combining machine learning with disease research. The goal of these studies is not to build isolated classifiers, but to use models to help identify disease-associated features, interpret individual differences, and support personalized diagnostic and therapeutic decision-making.

Overall, our team's goal in "Artificial Intelligence Plus" is to develop machine learning modeling methods and intelligent analytical strategies for complex biomedical data. On the one hand, the team focuses on methodological issues such as feature selection, high-dimensional modeling, heterogeneity analysis, explainable machine learning, and large model applications. On the other hand, the team emphasizes that these methods must serve real biomedical problems and be able to address challenges arising from noise, batch variation, individual differences, and multi-source data integration. Through these studies, "Artificial Intelligence Plus" can serve as an important methodological foundation connecting electronic information, machine learning, biomedical data analysis, and precision medicine, providing technical support for disease subtyping, risk assessment, mechanistic research, and decision support.


References

[1] LeCun, Y.; Bengio, Y.; Hinton, G. Deep Learning. Nature, 2015, 521(7553): 436–444.

[2] Yu, K.-H.; Beam, A. L.; Kohane, I. S. Artificial Intelligence in Healthcare. Nature Biomedical Engineering, 2018, 2(10): 719–731.

[3] Topol, E. J. High-Performance Medicine: The Convergence of Human and Artificial Intelligence. Nature Medicine, 2019, 25(1): 44–56.

[4] Rajpurkar, P.; Chen, E.; Banerjee, O.; Topol, E. J. AI in Health and Medicine. Nature Medicine, 2022, 28(1): 31–38.

[5] Wang, H.; Fu, T.; Du, Y.; Gao, W.; Huang, K.; Liu, Z.; Chandak, P.; Liu, S.; Van Katwyk, P.; Deac, A.; Anandkumar, A.; Bergen, K.; Gomes, C. P.; Ho, S.; Kohli, P.; Lasenby, J.; Leskovec, J.; Liu, T.-Y.; Manrai, A.; Marks, D.; Ramsundar, B.; Song, L.; Sun, J.; Tang, J.; Veličković, P.; Welling, M.; Zhang, L.; Coley, C. W.; Bengio, Y.; Zitnik, M. Scientific Discovery in the Age of Artificial Intelligence. Nature, 2023, 620(7972): 47–60.

[6] Guo, L.; Zhu, Z.; Lam, T. K.-Y.; Zhong, Y.; Fang, J.; Wang, X.; Dong, J.; Chen, W. MEMpre: Enhanced Membrane Protein Type Prediction via AlphaFold-Driven Protein Large Language Models. Chemometrics and Intelligent Laboratory Systems, 2026, 269: 105738.

[7] Dai, J.; Dong, L.; Xu, J.; Deng, L.; Guo, L.; Dong, J. Explainable Machine Learning Enables Robust Evaluation of Extracted Ion Chromatograms in LC–MS Metabolomics. Chemometrics and Intelligent Laboratory Systems, 2026, 269: 105591.

[8] Wang, Y.; Liu, X.; Lin, S.; Cheng, K. K.; Deng, L.; Luo, Z.; Dong, J. Density-Based Resampling Strategy Unveils Metabolic Heterogeneity and Subtypes in Liver Diseases Progression. Journal of Analysis and Testing, 2026, DOI: 10.1007/s41664-026-00427-9.

[9] Yang, J.; Gao, S.; Yin, X.; Ding, J.; Qiu, X.; Fei, Z.; Dong, J. Metabolomics and Machine Learning Identify Novel Serum Biomarkers for Nasopharyngeal Carcinoma Diagnosis. International Journal of Radiation Oncology, Biology, Physics, 2025, 123(1, Suppl. 1): e334.