南方医科大学学报 ›› 2026, Vol. 46 ›› Issue (9): 2149-2158.doi: 10.12122/j.issn.1673-4254.2026.09.14

• • 上一篇    

基于血浆代谢组学与机器学习的缺血性卒中亚型判别及隐源性卒中病因预测

自胤余(), 刘子淼, 邓镇()   

  1. 南方医科大学南方医院神经内科,广东 广州 510515
  • 收稿日期:2025-12-20 出版日期:2026-09-20 发布日期:2026-09-30
  • 通讯作者: 邓镇 E-mail:yinyuzi_cn@163.com;dengzhencn83@ smu.edu.cn
  • 作者简介:自胤余,硕士,E-mail: yinyuzi_cn@163.com
  • 基金资助:
    国家自然科学基金(82271305);广东省自然科学基金(2021A1515011027);广东省自然科学基金(2022A1515012525)

Plasma metabolomics combined with machine learning for classifying ischemic stroke subtypes and predicting the etiology of cryptogenic stroke

Yinyu ZI(), Zimiao LIU, Zhen DENG()   

  1. Department of Neurology, Nanfang Hospital, Southern Medical University, Guangzhou 510515, China
  • Received:2025-12-20 Online:2026-09-20 Published:2026-09-30
  • Contact: Zhen DENG E-mail:yinyuzi_cn@163.com;dengzhencn83@ smu.edu.cn
  • Supported by:
    National Natural Science Foundation of China(82271305)

摘要:

目的 分析缺血性卒中不同亚型之间血浆代谢特征的差异,并通过10种机器学习算法构建识别缺血性脑卒中不同亚型的模型,为隐源性卒中的辨识提供客观化依据,以提高临床诊断率。 方法 本研究共招募了来自3家医院的277例急性缺血性脑卒中(AIS)患者作为研究对象,并根据TOAST分型对患者进行分组。所有患者采集静脉血样并进行1H-NMR代谢组检测,比较分析并鉴定了46种代谢物。采用代谢组学和孟德尔随机化分析来识别潜在的生物标志物, PyCaret 用于开发和优化机器学习(ML)模型。开发血浆代谢物机器学习 (PMML) 模型用以区分缺血性卒中的不同亚型,并将隐源性卒中(SUE)队列用以外部验证模型效能。 结果 在大动脉动脉粥样硬化 (LAA)、小动脉闭塞 (SVO) 和心源性栓塞 (CE)之间观察到代谢物谱的显著差异。孟德尔随机化分析表明,葡萄糖和脯氨酸可能作为 SVO 和 LAA的潜在生物标志物。基于3种诊断良好的亚型(LAA、SVO 和 CE)的代谢谱,本研究开发了血浆代谢物机器学习(PMML) 模型,该模型在区分上述缺血性卒中亚型方面表现出较高的预测准确性(准确率:0.8630;ROC曲线下面积:0.9586)。SUE患者被用作验证队列,其中,被预测为 LAA 的患者,23.9%出现易损斑块;而被预测为SVO或CE的患者有10.9%出现易损斑块。此外,5名被预测患有 CE 的患者通过长程心电图监测被诊断为阵发性心房颤动。 结论 不同的血浆代谢特征与不同的缺血性卒中亚型相关。本研究的 PMML 模型基于诊断良好的卒中亚型的血浆代谢特征,在识别这些亚型方面表现出较高的准确率,并且在预测隐源性卒中的病因学方面具有巨大潜力。

关键词: 急性缺血性卒中, 隐源性卒中, 血浆代谢物, 机器学习, 孟德尔随机化

Abstract:

Objective To analyze plasma metabolic profiles of different ischemic stroke subtypes and construct machine learning (ML)-based models to identify these subtypes for accurate diagnosis of cryptogenic stroke. Methods A total of 277 patients with acute ischemic stroke (AIS) from 3 hospitals were prospectively enrolled and stratified according to the TOAST classification criteria. Venous blood samples were collected for ¹H-NMR metabolomics detection, resulting in the identification of 46 metabolites. Metabolomics and Mendelian randomization analyses were employed to identify the potential biomarkers, and PyCaret was used to develop and optimize the ML models. A Plasma Metabolite Machine Learning (PMML) model was developed to differentiate the ischemic stroke subtypes, and the efficacy of the model was validated using a cryptogenic stroke (SUE) cohort. Results Significantly different metabolite profiles were observed among patients with large artery atherosclerosis (LAA), small vessel occlusion (SVO), and cardioembolism (CE). Mendelian randomization analysis indicated that glucose and proline may serve as potential biomarkers for SVO and LAA. The PMML model developed based on the metabolic profiles of LAA, SVO, and CE demonstrated a high predictive accuracy for distinguishing these ischemic stroke subtypes with an accuracy of 0.8630 and an area under the ROC curve of 0.9586. In the SUE patients (the validation cohort), 23.9% of the patients predicted to have LAA exhibited vulnerable plaques, as compared with a rate of 10.9% in those predicted to have SVO or CE. Five patients predicted to have CE were diagnosed with paroxysmal atrial fibrillation via long-term electrocardiographic monitoring. Conclusion Different ischemic stroke subtypes have distinct plasma metabolic profiles. The PMML model developed based on plasma metabolic features of the 3 well-diagnosed stroke subtypes demonstrates a high accuracy in identifying these subtypes with a great potential for predicting the etiology of cryptogenic stroke.

Key words: acute ischemic stroke, cryptogenic stroke, plasma metabolites, machine learning, Mendelian randomization