南方医科大学学报 ›› 2026, Vol. 46 ›› Issue (9): 2266-2275.doi: 10.12122/j.issn.1673-4254.2026.09.24
• • 上一篇
赵涛涛(
), 焦粤豪, 倪铭(
), 罗友璐, 夏顺兴, 徐颖颖, 何亚婷
收稿日期:2026-02-13
出版日期:2026-09-20
发布日期:2026-09-30
通讯作者:
倪铭
E-mail:1425067966@qq.com;nm@sicau.edu.cn
作者简介:赵涛涛,在读硕士研究生,E-mail: 1425067966@qq.com
基金资助:
Taotao ZHAO(
), Yuehao JIAO, Ming NI(
), Youlu LUO, Shunxing XIA, Yingying XU, Yating HE
Received:2026-02-13
Online:2026-09-20
Published:2026-09-30
Contact:
Ming NI
E-mail:1425067966@qq.com;nm@sicau.edu.cn
摘要:
目的 提出一种以前期高精度算法 YOLOv11-TDSP 为基准的全场景端侧量化部署与多模态协同诊疗框架。 方法 构建基于 KL 散度的数据感知型训练后量化(PTQ)机制,通过校准集锁定特征最佳分布区间,将FP32模型近乎无损压缩至 INT8 精度。 建立跨平台异构编译流水线,针对 Intel Hybrid Architecture(大小核架构)、Rockchip RK3588及Ascend 310P进行算子深度融合与指令集优化,实现软硬件极致解耦。构建面向资源受限边缘终端的“检诊一体化”协同推理架构,利用极致量化的检测模型释放关键算力,驱动本地部署的 Qwen2.5-3B 轻量化大模型,结合专家规则库实现从视觉检测到文本报告的端侧闭环。 结果 INT8 量化后的模型权重仅为 6.1 MB(压缩率 57.3%),mAP@0.5 保持在95.4%,精度损失仅为0.4%。模型在标准办公终端与资源受限的移动诊疗设备上均表现出卓越的泛化性能。特别是在算力与功耗双重受限的移动巡诊场景(Intel i5-12450H)中,推理速度提升至22.91 FPS,实现了3.64倍的加速比。系统在完全断网且仅有 8GB 内存的普通设备上,成功实现了毫秒级病灶检测与符合临床规范的报告生成。 结论 该框架有效解决了高精度算法在低配硬件上的部署瓶颈,在确保患者隐私安全的前提下,实现了口腔全景片分析从“云端依赖”向“边缘普惠”的范式转变,为基层医疗机构的智能化升级提供了一种低成本、高可靠的工程化解决方案。
赵涛涛, 焦粤豪, 倪铭, 罗友璐, 夏顺兴, 徐颖颖, 何亚婷. 端侧量化与多模态协同实现资源受限下的口腔全景片高效智能诊断[J]. 南方医科大学学报, 2026, 46(9): 2266-2275.
Taotao ZHAO, Yuehao JIAO, Ming NI, Youlu LUO, Shunxing XIA, Yingying XU, Yating HE. Efficient dental panoramic radiographic diagnosis under resource constraints via edge quantization and multi-modal synergy[J]. Journal of Southern Medical University, 2026, 46(9): 2266-2275.
| Species | Sample count |
|---|---|
| Decayed tooth | 30 |
| Impacted tooth | 41 |
| Ectopic Eruption | 32 |
| Root canal therapy | 25 |
| Dental crown | 63 |
| Dental implant root system | 39 |
| Dental implant | 26 |
| Dental filling | 11 |
表1 测试集牙齿异常类别分布
Tab.1 Distribution of dental anomaly categories in the test dataset
| Species | Sample count |
|---|---|
| Decayed tooth | 30 |
| Impacted tooth | 41 |
| Ectopic Eruption | 32 |
| Root canal therapy | 25 |
| Dental crown | 63 |
| Dental implant root system | 39 |
| Dental implant | 26 |
| Dental filling | 11 |
| Component | Software/Library | Version | Role in system |
|---|---|---|---|
| Hardware platform | Rockchip RK3588 | NPU 6.0 TOPS | Edge inference device |
| Operating system | Linaro linux | Alip (Debian) | Host OS |
| Deep learning framework | PyTorch | 2.2.0 | Model training & pruning |
| Model conversion | RKNN-Toolkit2 | 2.3.2 | Quantization & export |
| Inference engine | RKNN-Toolkit-Lite2 | 2.3.2 | On-device inference |
| Intermediate format | ONNX | 1.16.1 | Model exchange |
| Computer vision | OpenCV-Python | 4.11.0 | Image preprocessing |
| Scientific computing | NumPy | 1.26.4 | Tensor operations |
表2 RK3588 边缘端实验环境配置
Tab.2 Experimental environment and software configuration (Rockchip RK3588)
| Component | Software/Library | Version | Role in system |
|---|---|---|---|
| Hardware platform | Rockchip RK3588 | NPU 6.0 TOPS | Edge inference device |
| Operating system | Linaro linux | Alip (Debian) | Host OS |
| Deep learning framework | PyTorch | 2.2.0 | Model training & pruning |
| Model conversion | RKNN-Toolkit2 | 2.3.2 | Quantization & export |
| Inference engine | RKNN-Toolkit-Lite2 | 2.3.2 | On-device inference |
| Intermediate format | ONNX | 1.16.1 | Model exchange |
| Computer vision | OpenCV-Python | 4.11.0 | Image preprocessing |
| Scientific computing | NumPy | 1.26.4 | Tensor operations |
| Component | Software/Library | Version | Role in system |
|---|---|---|---|
| Hardware platform | Huawei Ascend server | Ascend 310P (NPU) | Ai inference accelerator |
| Operating system | openEuler | 22.03 (LTS-SP3) aarch64 | Host operating system |
| Deeplearning framework | PyTorch | 2.1.0/2.2.0 (compatible) | Model training & export |
| Development toolkit | Ascend CANN Toolkit | 8.0.0 | Npu driver & compiler stack |
| Model conversion | ATC (ascend tensor compiler) | 8.0.0 | ONNX to OM conversion |
| Inference engine | AscendCL (ACL) | 8.0.0 | On-device C++ inference |
| Intermediate format | ONNX | 1.14.0+ | Model exchange format |
| Video processing | FFmpeg & x264 | 4.1.3 / Latest | Video decoding & encoding |
| Computer vision | OpenCV | 4.x (C++) | Image preprocessing |
| Compiler | GCC/G++ | 7.3.0+ (system default) | Sample code compilation |
表3 昇腾边缘端实验环境配置
Tab.3 Experimental environment and software configuration (Huawei Ascend)
| Component | Software/Library | Version | Role in system |
|---|---|---|---|
| Hardware platform | Huawei Ascend server | Ascend 310P (NPU) | Ai inference accelerator |
| Operating system | openEuler | 22.03 (LTS-SP3) aarch64 | Host operating system |
| Deeplearning framework | PyTorch | 2.1.0/2.2.0 (compatible) | Model training & export |
| Development toolkit | Ascend CANN Toolkit | 8.0.0 | Npu driver & compiler stack |
| Model conversion | ATC (ascend tensor compiler) | 8.0.0 | ONNX to OM conversion |
| Inference engine | AscendCL (ACL) | 8.0.0 | On-device C++ inference |
| Intermediate format | ONNX | 1.14.0+ | Model exchange format |
| Video processing | FFmpeg & x264 | 4.1.3 / Latest | Video decoding & encoding |
| Computer vision | OpenCV | 4.x (C++) | Image preprocessing |
| Compiler | GCC/G++ | 7.3.0+ (system default) | Sample code compilation |
| Deployment scenario | Hardware specification | Precision | Latency (ms) ↓ | Throughput (FPS) ↑ | Speedup |
|---|---|---|---|---|---|
| Research center (lab server) | Intel i7-12700 | FP32 | 72.50 | 13.79 | 1.00 |
| INT8 (Ours) | 38.15 | 26.21 | 1.9 | ||
| Doctor's office | Intel i5-12400 | FP32 | 98.6 | 10.14 | 1.00 |
| INT8 (Ours) | 56.34 | 17.75 | 1.75 | ||
| Mobile clinic | Intel i5-12450H | FP32 | 158.71 | 6.30 | 1.00 |
| INT8 (Ours) | 43.66 | 22.91 | 3.64 | ||
| Doctor's office | Intel i5-12600KF | FP32 | 79.04 | 12.65 | 1.00 |
| INT8 (Ours) | 43.12 | 23.19 | 1.83 | ||
| Mobile clinic | Intel i7-10750H | FP32 | 139.12 | 7.19 | 1.00 |
| INT8 (Ours) | 92.02 | 10.87 | 1.51 | ||
| Doctor's office | Intel i7-7700 | FP32 | 125.00 | 8.00 | 1.00 |
| INT8 (Ours) | 79.36 | 12.60 | 1.58 | ||
| Rural clinic | Intel i5-6500 | FP32 | 172.41 | 5.80 | 1.00 |
| INT8 (Ours) | 106.38 | 9.40 | 1.62 | ||
| Rural clinic | Intel i3-8100T | FP32 | 205.60 | 4.86 | 1.00 |
| INT8 (Ours) | 135.80 | 7.36 | 1.51 |
表4 FP32与INT8量化模型在不同临床硬件平台上的性能对比)
Tab.4 Performance comparison of FP32 and quantized INT8 models across various clinical scenarios
| Deployment scenario | Hardware specification | Precision | Latency (ms) ↓ | Throughput (FPS) ↑ | Speedup |
|---|---|---|---|---|---|
| Research center (lab server) | Intel i7-12700 | FP32 | 72.50 | 13.79 | 1.00 |
| INT8 (Ours) | 38.15 | 26.21 | 1.9 | ||
| Doctor's office | Intel i5-12400 | FP32 | 98.6 | 10.14 | 1.00 |
| INT8 (Ours) | 56.34 | 17.75 | 1.75 | ||
| Mobile clinic | Intel i5-12450H | FP32 | 158.71 | 6.30 | 1.00 |
| INT8 (Ours) | 43.66 | 22.91 | 3.64 | ||
| Doctor's office | Intel i5-12600KF | FP32 | 79.04 | 12.65 | 1.00 |
| INT8 (Ours) | 43.12 | 23.19 | 1.83 | ||
| Mobile clinic | Intel i7-10750H | FP32 | 139.12 | 7.19 | 1.00 |
| INT8 (Ours) | 92.02 | 10.87 | 1.51 | ||
| Doctor's office | Intel i7-7700 | FP32 | 125.00 | 8.00 | 1.00 |
| INT8 (Ours) | 79.36 | 12.60 | 1.58 | ||
| Rural clinic | Intel i5-6500 | FP32 | 172.41 | 5.80 | 1.00 |
| INT8 (Ours) | 106.38 | 9.40 | 1.62 | ||
| Rural clinic | Intel i3-8100T | FP32 | 205.60 | 4.86 | 1.00 |
| INT8 (Ours) | 135.80 | 7.36 | 1.51 |
图6 YOLOv11-TDSP 算法量化前后 P-R 曲线及各类别 AP 值对比
Fig.6 Comparison of P-R curves and class-wise AP values for YOLOv11-TDSP before and after quantization. A: FP32. B: INT8.
| Evaluation metrics | Mean±SD | Score≥4 |
|---|---|---|
| Accuracy | 4.2±0.6 | 86.0% |
| Completeness | 4.8±0.5 | 93.0% |
| Standardization | 4.1±0.7 | 80.0% |
表5 边缘端轻量化大模型生成诊断报告的临床盲评量化结果 (5分制)
Tab.5 Quantitative results of clinical blind review for diagnostic reports generated by the edge-side lightweight LLM (5-point scale)
| Evaluation metrics | Mean±SD | Score≥4 |
|---|---|---|
| Accuracy | 4.2±0.6 | 86.0% |
| Completeness | 4.8±0.5 | 93.0% |
| Standardization | 4.1±0.7 | 80.0% |
| [1] | World Health Organization. Global oral health status report: towards universal health coverage for oral health by 2030[R]. Geneva: WHO, 2022. doi:10.1596/36724 |
| [2] | 陆海霞, 陶丹英, 卢展民, 等. 第四次全国口腔健康流行病学调查-背景和方法[C]//2018年中华口腔医学会第十八次口腔预防医学学术年会论文集. 西安, 2018: 23. |
| [3] | Ahmed N, Abbasi MS, Zuberi F, et al. Artificial intelligence techniques: analysis, application, and outcome in dentistry: a systematic review[J]. BioMed Res Int, 2021, 2021: 9751564. doi:10.1155/2021/9751564 |
| [4] | Tuzoff DV, Tuzova LN, Bornstein MM, et al. Tooth detection and numbering in panoramic radiographs using convolutional neural networks[J]. Dentomaxillofac Radiol, 2019, 48(4): 20180051. doi:10.1259/dmfr.20180051 |
| [5] | Ossowska A, Kusiak A, Świetlik D. Artificial intelligence in dentistry: narrative review[J]. Int J Environ Res Public Health, 2022, 19(6): 3449. doi:10.3390/ijerph19063449 |
| [6] | Zhao XT, Xu TK, Peng L, et al. Recognition and segmentation of teeth and mandibular nerve canals in panoramic dental X-rays by Mask RCNN[J]. Displays, 2023, 78: 102447. doi:10.1016/j.displa.2023.102447 |
| [7] | Esteva A, Chou K, Yeung S, et al. Deep learning-enabled medical computer vision[J]. NPJ Digit Med, 2021, 4(1): 5. doi:10.1038/s41746-020-00376-2 |
| [8] | Cui ZM, Li CJ, Wang WP. ToothNet: automatic tooth instance segmentation and identification from cone beam CT images[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 15-20, 2019, Long Beach, CA, USA. IEEE, 2020: 6361-70. doi:10.1109/cvpr.2019.00653 |
| [9] | Sheng C, Wang L, Huang ZH, et al. Transformer-based deep learning network for tooth segmentation on panoramic radiographs[J]. J Syst Sci Complex, 2023, 36(1): 257-72. doi:10.1007/s11424-022-2057-9 |
| [10] | Ragab MG, Abdulkadir SJ, Qaid N, et al. Periodontitis bone loss detection in panoramic radiographs using modified YOLOv7[J]. PeerJ Comput Sci, 2025, 11: e3102. doi:10.7717/peerj-cs.3102 |
| [11] | 梁 洪, 邱定乾, 丁世宇, 等. 基于改进YOLOv8的异常牙齿和修复体X射线影像检测[J]. 中国激光, 2025, 52(3): 0307106. doi:10.3788/CJL241265 |
| [12] | 赵涛涛, 倪 铭, 夏顺兴, 等. YOLOv11-TDSP: 口腔全景片的轻量化高精度异常牙检测模型[J]. 南方医科大学学报, 2025, 45(8): 1791-9. |
| [13] | Rancea A, Anghel I, Cioara T. Edge computing in healthcare: innovations, opportunities, and challenges[J]. Future Internet, 2024, 16(9): 329. doi:10.3390/fi16090329 |
| [14] | Chen JS, Ran XK. Deep learning with edge computing: a review[J]. Proc IEEE, 2019, 107(8): 1655-74. doi:10.1109/jproc.2019.2921977 |
| [15] | Shi WS, Cao J, Zhang Q, et al. Edge computing: vision and challenges[J]. IEEE Internet Things J, 2016, 3(5): 637-46. doi:10.1109/jiot.2016.2579198 |
| [16] | Wang RJ, Lai JS, Zhang ZY, et al. Privacy-preserving federated learning for Internet of medical things under edge computing[J]. IEEE J Biomed Health Inform, 2023, 27(2): 854-65. doi:10.1109/jbhi.2022.3157725 |
| [17] | Rieke N, Hancox J, Li WQ, et al. The future of digital health with federated learning[J]. npj Digit Med, 2020, 3: 119. doi:10.1038/s41746-020-00323-1 |
| [18] | Yang Q, Liu Y, Chen TJ, et al. Federated machine learning: concept and applications[J]. ACM Trans Intell Syst Technol, 2019, 10(2): 1-19. doi:10.1145/3298981 |
| [19] | Singhal K, Azizi S, Tu T, et al. Large language models encode clinical knowledge[J]. Nature, 2023, 620(7972): 172-80. doi:10.1038/s41586-023-06291-2 |
| [20] | Sha YY, Pan HX, Meng WY, et al. Contrastive knowledge-guided large language models for medical report generation[M]//Medical Image Computing and Computer Assisted Intervention – MICCAI 2025. Cham: Springer Nature Switzerland, 2025: 111-20. doi:10.1007/978-3-032-04978-0_11 |
| [21] | Moor M, Banerjee O, Abad ZSH, et al. Foundation models for generalist medical artificial intelligence[J]. Nature, 2023, 616(7956): 259-65. doi:10.1038/s41586-023-05881-4 |
| [22] | Thirunavukarasu AJ, Ting DSJ, Elangovan K, et al. Large language models in medicine[J]. Nat Med, 2023, 29(8): 1930-40. doi:10.1038/s41591-023-02448-8 |
| [23] | Tu T, Azizi S, Driess D, et al. Towards generalist biomedical AI[J]. Nejm Ai, 2024, 1(3):14334. doi:10.1056/aioa2300138 |
| [24] | Qu CY, Zhao R, Yu Y, et al. Post-training quantization for 3D medical image segmentation: a practical study on real inference engines[EB/OL]. 2025: arXiv: 2501.17343. . doi:10.1117/1.jmi.13.1.014006 |
| [25] | Jacob B, Kligys S, Chen B, et al. Quantization and training of neural networks for efficient integer-arithmetic-only inference[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition. June 18-23, 2018, Salt Lake City, UT, USA. IEEE, 2018: 2704-13. doi:10.1109/cvpr.2018.00286 |
| [26] | Cai H, Gan C, Wang TZ, et al. Once-for-all: train one network and specialize it for efficient deployment[EB/OL]. 2019: arXiv: 1908.09791. . |
| [27] | Yang A, Li AF, Yang BS, et al. Qwen3 technical report[EB/OL]. 2025: arXiv: 2505.09388. |
| [28] | 孙召飞, 俞经虎, 朱行飞, 等. 基于改进YOLOv5s的口腔全景片牙齿病症识别算法[J]. 中国激光, 2024, 51(15): 1507106. |
| [29] | Lin SY, Hao XJ, Liu Y, et al. Lightweight deep learning methods for panoramic dental X-ray image segmentation[J]. Neural Comput Appl, 2023, 35(11): 8295-306. doi:10.1007/s00521-022-08102-7 |
| [30] | 杨 春, 张睿尧, 黄 泷, 等. 深度神经网络模型量化方法综述[J]. 工程科学学报, 2023, 45(10): 1613-29. |
| [31] | Gholami A, Kim S, Dong Z, et al. A survey of quantization methods for efficient neural network inference[M]//Low-power computer vision. Chapman and Hall/CRC, 2022: 291-326. doi:10.1201/9781003162810-13 |
| [32] | 李 斌, 汪黎君, 曹厚德, 等. 基于上海市场十年调查数据看医疗设备“中国制造”的提升[J]. 中国医疗设备, 2018, 33(2): 23-6. doi:10.3969/j.issn.1674-1633.2018.02.007 |
| [33] | 林 岚, 武雨桐. 大型语言模型在医疗领域的应用现状与展望[J]. 医疗卫生装备, 2024, 45(8): 102-9. doi:10.19745/j.1003-8868.2024161 |
| [34] | Frantar E, Ashkboos S, Hoefler T, et al. GPTQ: Accurate post-training quantization for generative pre-trained transformers[C]// International Conference on Learning Representations (ICLR). 2023. |
| [35] | Lin J, Tang J, Tang H, et al. AWQ: Activation-aware weight quantization for LLM compression and acceleration[C]// Proceedings of Machine Learning and Systems (MLSys). 2024. doi:10.1145/3714983.3714987 |
| [36] | McMahan B, Moore E, Ramage D, et al. Communication-efficient learning of deep networks from decentralized data[C]// Artificial Intelligence and Statistics (AISTATS). PMLR, 2017: 1273-82. |
| [1] | 赵涛涛, 倪铭, 夏顺兴, 焦粤豪, 何亚婷. YOLOv11-TDSP:口腔全景片的轻量化高精度异常牙检测模型[J]. 南方医科大学学报, 2025, 45(8): 1791-1799. |
| 阅读次数 | ||||||
|
全文 |
|
|||||
|
摘要 |
|
|||||