南方医科大学学报 ›› 2026, Vol. 46 ›› Issue (6): 1444-1454.doi: 10.12122/j.issn.1673-4254.2026.06.24

• • 上一篇    

心音异常高检出率模型MS-DTNet:融合多尺度卷积与双塔注意力的方法

温耀棋1(), 卢学麒1, 吴秋岑1, 陈华元2(), 陈超敏1()   

  1. 1.南方医科大学生物医学工程学院,广东 广州 510515
    2.广州南方医疗设备综合检测有限责任公司,广东 广州 510515
  • 收稿日期:2025-09-30 出版日期:2026-06-20 发布日期:2026-06-24
  • 通讯作者: 陈华元,陈超敏 E-mail:1365909788@qq.com;415081161@qq.com;571611621@qq.com
  • 作者简介:温耀棋,在读硕士研究生,E-mail: 1365909788@qq.com
  • 基金资助:
    国家重点研发计划(2023YFC2414500);国家重点研发计划(2023YFC2414502)

MS-DTNet: an efficient model combining multi-scale convolution and dual-tower attention for detecting abnormal heart sounds

Yaoqi WEN1(), Xueqi LU1, Qiucen WU1, Huayuan CHEN2(), Chaomin CHEN1()   

  1. 1.School of Biomedical Engineering, Southern Medical University, Guangzhou 510515, China
    2.Guangzhou Southern Medical Equipment Comprehensive testing Co. , Ltd. , Guangzhou 510515, China
  • Received:2025-09-30 Online:2026-06-20 Published:2026-06-24
  • Contact: Huayuan CHEN, Chaomin CHEN E-mail:1365909788@qq.com;415081161@qq.com;571611621@qq.com

摘要:

目的 提高心音异常的检出率,为自动化心音分析提供技术支持。 方法 本文提出了一种基于多尺度卷积与双塔注意力机制的模型(MS-DTNet),其通过深度可分离卷积以降低计算复杂度的同时结合多尺度卷积以并行获取不同感受野的特征,然后采用双塔结构兼顾短时与长时信息提取,并引入翻转注意力机制以进一步增强特征表征能力,从而提高对心音图分类的精度。 结果 使用公开数据集CinC2016对提出的模型进行五折交叉验证实验。在心音异常分类任务中,模型在测试集上的平均准确率、平均特异性、平均召回率、平均精确率、平均F1分数与平均ROC分别达到94.49%、93.44%、95.49%、93.81%、94.64%和98.65%,结果均优于其他几种模型。 结论 本研究提出的MS-DTNet模型在心音图分类表现出色,为后续的多源心音数据分析与临床辅助决策研究提供了有效的技术基础。

关键词: 心音图分类, 轻量化神经网络, 深度可分离卷积, 多尺度特征卷积, 注意力机制, 双塔机制

Abstract:

Objective To develop an efficient model for detecting abnormal heart sounds and providing technical support for automated phonocardiogram (PCG) analysis. Methods The proposed model, MS-DTNet, integrates multi-scale convolution and a dual-tower attention mechanism. The network employs depthwise separable convolution to achieve a lightweight architecture, applies multi-scale convolution to capture features across receptive fields in parallel, utilizes a dual-tower structure to extract both short- and long-term temporal information, and incorporates a flipped-attention mechanism to further enhance feature representation. Experiments were conducted on the public CinC2016 dataset using 5-fold cross-validation. Results In the abnormal heart sound classification task, the proposed model achieved an average accuracy of 94.49%, a specificity of 93.44%, a sensitivity of 95.49%, a precision of 93.81%, F1-score of 94.64%, and ROC-AUC of 98.65%, outperforming all the comparison models. Conclusion The proposed MS-DTNet model demonstrates excellent performance in PCG classification and provides technical assistance for multi-source heart sound analysis and clinical decision support.

Key words: phonocardiogram classification, lightweight neural network, depthwise separable convolution, multi-scale convolution, attention mechanism, dual-tower architecture