南方医科大学学报 ›› 2026, Vol. 46 ›› Issue (7): 1703-1713.doi: 10.12122/j.issn.1673-4254.2026.07.23
• • 上一篇
收稿日期:2025-12-02
出版日期:2026-07-20
发布日期:2026-07-20
通讯作者:
曾栋
E-mail:1284897384@qq.com;zd1989@smu.edu.cn
作者简介:宋含笑,在读硕士研究生,E-mail: 1284897384@qq.com
基金资助:
Hanxiao SONG(
), Huaxian SHI, Ying LI, Zhaoying BIAN, Dong ZENG(
)
Received:2025-12-02
Online:2026-07-20
Published:2026-07-20
Contact:
Dong ZENG
E-mail:1284897384@qq.com;zd1989@smu.edu.cn
摘要:
目的 为解决在医学影像数据量呈指数级增长的背景下数据存储与传输受限问题,对隐式神经网络(INR)进行优化,提出频域自适应隐式神经网络(FAINC)的医学影像压缩方法。 方法 基于KiTS19和AVT数据集的356例患者CT数据进行回顾性分析,构建由多子网络协作的FAINC压缩方法。该方法通过光谱稀疏指数(SSI)评估数据块频域复杂度,接着采用频域门控机制将其动态分配至不同容量的子网络,最终结合参数量化与熵编码实现压缩输出。为验证性能,将所提方法与主流商业压缩标准(H.265/HEVC、JPEG2000)、隐式神经表示方法NeRV以及深度学习压缩方法DVC进行系统对比。 结果 在高压缩率时FAINC在KiTS19和AVT数据集均表现出最优重建性能。当比特率分别为BPV=0.32和BPV=0.34时,FAINC分别实现最高的PSNR(47.03和50.76)、最高的SSIM指标(0.9853和0.9930)以及最低的RMSE指标(0.0045和0.0029)。主观图像质量评分亦显著优于其他方法(P<0.05)。消融实验表明,频域门控与参数动态分配机制分别贡献约2.05dB和1.87dB的PSNR提升。 结论 实验结果证明,该方法在高压缩比下能够显著提高医学影像的重建质量,并在结构保真度和压缩效率方面优于现有主流方法。为医学影像的高效存储及低带宽远程传输提供技术支撑。
宋含笑, 石华仙, 李颖, 边兆英, 曾栋. 基于频域自适应隐式神经表征的医学影像压缩方法[J]. 南方医科大学学报, 2026, 46(7): 1703-1713.
Hanxiao SONG, Huaxian SHI, Ying LI, Zhaoying BIAN, Dong ZENG. A frequency-adaptive implicit neural representation method for medical image compression[J]. Journal of Southern Medical University, 2026, 46(7): 1703-1713.
图2 在KiTS19数据集上, 不同压缩方法的率失真曲线
Fig.2 Rate distortion curves of different compression methods on the KiTS19 dataset. A: BPV-PSNR curves. B: BPV-SSIM curves.
| Methods | PSNR | PSNR (Rel. ↑ %) | SSIM | SSIM (Rel. ↑ %) | RMSE | RMSE (Rel. ↓ %) |
|---|---|---|---|---|---|---|
| FAINC | 47.0346±0.7289 | - | 0.9853±0.0015 | - | 0.0045±0.0004 | - |
| JPEG2000 | 45.7261±0.4892 | 2.86 | 0.9800±0.0016 | 0.54 | 0.0052±0.0003 | 13.46 |
| H.265 | 43.8230±0.2628 | 7.34 | 0.9629±0.0039 | 2.32 | 0.0064±0.0002 | 29.69 |
| H.264 | 42.0655±0.3790 | 11.81 | 0.9539±0.0038 | 3.30 | 0.0079±0.0004 | 43.04 |
| JPEG | 42.0379±0.3820 | 11.87 | 0.9539±0.0039 | 3.30 | 0.0079±0.0004 | 43.04 |
| NeRV | 43.7035±1.2676 | 7.62 | 0.9716±0.0116 | 1.41 | 0.0124±0.0016 | 63.71 |
表1 KiTS19 数据集上, 各方法在20例三维CT影像上的重建指标均值及方差(BPV约为0.32-0.34)
Tab.1 Mean and variance of reconstruction metrics for different methods on 20 3D CT volumes from the KiTS19 dataset with BPV of 0.32-0.34
| Methods | PSNR | PSNR (Rel. ↑ %) | SSIM | SSIM (Rel. ↑ %) | RMSE | RMSE (Rel. ↓ %) |
|---|---|---|---|---|---|---|
| FAINC | 47.0346±0.7289 | - | 0.9853±0.0015 | - | 0.0045±0.0004 | - |
| JPEG2000 | 45.7261±0.4892 | 2.86 | 0.9800±0.0016 | 0.54 | 0.0052±0.0003 | 13.46 |
| H.265 | 43.8230±0.2628 | 7.34 | 0.9629±0.0039 | 2.32 | 0.0064±0.0002 | 29.69 |
| H.264 | 42.0655±0.3790 | 11.81 | 0.9539±0.0038 | 3.30 | 0.0079±0.0004 | 43.04 |
| JPEG | 42.0379±0.3820 | 11.87 | 0.9539±0.0039 | 3.30 | 0.0079±0.0004 | 43.04 |
| NeRV | 43.7035±1.2676 | 7.62 | 0.9716±0.0116 | 1.41 | 0.0124±0.0016 | 63.71 |
图3 在相近压缩比(BPV=0.34-0.36)下不同方法于KiTS19数据集上的影像重建结果视觉对比
Fig.3 Visual comparison of reconstructed images produced by different methods on the KiTS19 dataset under similar compression ratios (BPV=0.34-0.36). A1-A3: Ground Truth; B1-B3: FAINC; C1-C3: H.265; D1-D3: H.264; E1-E3: NeRV; F1-F3: JPEG. The display window is [-160, 240]HU.
| Method | Mean±SD | P |
|---|---|---|
| FAINC | 4.72±0.45 | |
| JPEG2000 | 4.62±0.58 | 0.019 |
| H.265 | 4.23±0.70 | <0.001 |
| H.264 | 3.97±0.64 | <0.001 |
| JPEG | 3.45±0.75 | <0.001 |
| NeRV | 2.35±0.61 | <0.001 |
表2 FAINC压缩KiTS19数据集后重建图像的整体质量评分统计结果
Tab.2 Overall image quality scores for reconstructed images after FAINC compression of the KiTS19 dataset
| Method | Mean±SD | P |
|---|---|---|
| FAINC | 4.72±0.45 | |
| JPEG2000 | 4.62±0.58 | 0.019 |
| H.265 | 4.23±0.70 | <0.001 |
| H.264 | 3.97±0.64 | <0.001 |
| JPEG | 3.45±0.75 | <0.001 |
| NeRV | 2.35±0.61 | <0.001 |
| Methods | PSNR | PSNR (Rel. ↑ %) | SSIM | SSIM (Rel. ↑ %) | RMSE | RMSE (Rel. ↓ %) |
|---|---|---|---|---|---|---|
| FAINC | 50.7630±0.6820 | - | 0.9930±0.0010 | - | 0.0029±0.0004 | - |
| JPEG2000 | 50.9773±1.2549 | -0.42 | 0.9918±0.0019 | 0.12 | 0.0031±0.0004 | 6.45 |
| H.265 | 49.1827±0.9431 | 3.21 | 0.9872±0.0085 | 0.59 | 0.0035±0.0004 | 17.14 |
| H.264 | 44.8626±1.0237 | 13.15 | 0.9564±0.0117 | 3.83 | 0.0058±0.0008 | 50.00 |
| JPEG | 44.8287±1.0157 | 13.24 | 0.9563±0.0117 | 3.84 | 0.0058±0.0008 | 50.00 |
| NeRV | 43.1059±4.5098 | 17.77 | 0.9701±0.0220 | 2.36 | 0.0047±0.0014 | 38.30 |
表3 AVT数据集上, 各方法在18例三维CT影像上的重建指标均值及方差(BPV约为 0.33-0.35)
Tab.3 Mean and variance of reconstruction metrics for different methods on 18 3D CT volumes from the AVT dataset (BPV=0.33-0.35)
| Methods | PSNR | PSNR (Rel. ↑ %) | SSIM | SSIM (Rel. ↑ %) | RMSE | RMSE (Rel. ↓ %) |
|---|---|---|---|---|---|---|
| FAINC | 50.7630±0.6820 | - | 0.9930±0.0010 | - | 0.0029±0.0004 | - |
| JPEG2000 | 50.9773±1.2549 | -0.42 | 0.9918±0.0019 | 0.12 | 0.0031±0.0004 | 6.45 |
| H.265 | 49.1827±0.9431 | 3.21 | 0.9872±0.0085 | 0.59 | 0.0035±0.0004 | 17.14 |
| H.264 | 44.8626±1.0237 | 13.15 | 0.9564±0.0117 | 3.83 | 0.0058±0.0008 | 50.00 |
| JPEG | 44.8287±1.0157 | 13.24 | 0.9563±0.0117 | 3.84 | 0.0058±0.0008 | 50.00 |
| NeRV | 43.1059±4.5098 | 17.77 | 0.9701±0.0220 | 2.36 | 0.0047±0.0014 | 38.30 |
图5 在相近压缩比(BPV=0.34~0.36)下不同方法于AVT数据集上的影像重建结果视觉对比
Fig.5 Visual comparison of reconstructed images produced by different methods on the AVT dataset under similar compression ratios (BPV=0.34-0.36). A1-A3: Ground Truth; B1-B3: FAINC; C1-C3: H.265; D1-D3: H.264; E1-E3: NeRV; F1-F3: JPEG. The display window is [-160, 240]HU.
| Method | Mean ± SD | P |
|---|---|---|
| FAINC | 4.77±0.43 | |
| JPEG2000 | 4.55±0.50 | 0.005 |
| H.265 | 4.20±0.58 | 0.002 |
| H.264 | 3.92±0.46 | <0.001 |
| JPEG | 3.55±0.65 | <0.001 |
| NeRV | 2.72±0.69 | <0.001 |
表4 FAINC压缩AVT数据集后重建图像的整体质量评分统计结果
Tab.4 Overall image quality scores for reconstructed images after FAINC compression of the AVT dataset
| Method | Mean ± SD | P |
|---|---|---|
| FAINC | 4.77±0.43 | |
| JPEG2000 | 4.55±0.50 | 0.005 |
| H.265 | 4.20±0.58 | 0.002 |
| H.264 | 3.92±0.46 | <0.001 |
| JPEG | 3.55±0.65 | <0.001 |
| NeRV | 2.72±0.69 | <0.001 |
图6 FAINC与两种消融实验模型的医学影像重建视觉效果对比
Fig.6 Visual comparison of reconstructed medical images by FAINC and two ablation models. A1-A2: Ground Truth; B1-B2: FAINC; C1-C2: W/o Frequency-Guided Partition; D1-D2: W/o Adaptive Allocation. The display window is [-160, 240]HU.
图8 FAINC与两种消融模型在不同子网络上的PSNR与参数规模对比
Fig.8 PSNR and parameter scale comparison across multiple subnetworks for FAINC and two ablation models. Top: Reconstruction quality (PSNR); bottom: Sub-network memory usage (MB).
| [1] | Minopoulou I, Kleyer A, Yalcin-Mutlu M, et al. Imaging in inflammatory arthritis: progress towards precision medicine[J]. Nat Rev Rheumatol, 2023, 19(10): 650-65. doi:10.1038/s41584-023-01016-1 |
| [2] | Moore J, Allan C, Besson S, et al. OME-NGFF: a next-generation file format for expanding bioimaging data-access strategies[J]. Nat Methods, 2021, 18(12): 1496-8. doi:10.1038/s41592-021-01326-w |
| [3] | Wiseman S. A FAIR platform for data-sharing[J]. Nat Neurosci, 2021, 24(12): 1640. doi:10.1038/s41593-021-00976-5 |
| [4] | Boergens KM, Berning M, Bocklisch T, et al. webKnossos: efficient online 3D data annotation for connectomics[J]. Nat Methods, 2017, 14(7): 691-4. doi:10.1038/nmeth.4331 |
| [5] | Wallace GK. The JPEG still picture compression standard[J]. Commun ACM, 1991, 34(4): 30-44. doi:10.1145/103085.103089 |
| [6] | Wiegand T, Sullivan GJ, Bjontegaard G, et al. Overview of the H.264/AVC video coding standard[J]. IEEE Trans Circuits Syst Video Technol, 2003, 13(7): 560-76. doi:10.1109/tcsvt.2003.815165 |
| [7] | Wien M. High efficiency video coding[J]. Coding Tools and specification, 2015, 24: 1. doi:10.1007/978-3-662-44276-0 |
| [8] | Dong C, Deng YB, Loy CC, et al. Compression artifacts reduction by a deep convolutional network[C]//2015 IEEE International Conference on Computer Vision (ICCV). December 7-13, 2015. Santiago, Chile. IEEE, 2015: 576-84. doi:10.1109/iccv.2015.73 |
| [9] | Rufai AM, Anbarjafari G, Demirel H. Lossy medical image compression using Huffman coding and singular value deco-mposition[C]//2013 21st Signal Processing and Communications Applications Conference (SIU). April 24-26, 2013, Haspolat, Turkey. IEEE, 2013: 1-4. doi:10.1109/siu.2013.6531592 |
| [10] | Song CY, Hui C, Lin Q, et al. LVPNet: a latent-variable-based prediction-driven end-to-end framework for lossless compression of medical images[C]//2025 Medical Image Computing and Computer Assisted Intervention(MICCAI). Sept 23-27, 2025, Daejeon, Republic of Korea. Springer Nature Switzerland, 2025: 288-98. doi:10.1007/978-3-032-04984-1_28 |
| [11] | Duan XT, Liu JJ, Zhang E. Efficient image encryption and compression based on a VAE generative model[J]. J Real Time Image Process, 2019, 16(3): 765-73. doi:10.1007/s11554-018-0826-4 |
| [12] | Sitzmann V, Martel JNP, Bergman AW, et al. Implicit neural representations with periodic activation functions[C]//2020 Neural Information Processing Systems Conference (NeurIPS). Dec 6-12, 2020, San Diego, USA. Curran Associates, Inc, 2020: 7462-7473. doi:10.48550/arXiv.2006.09661 |
| [13] | Dai GL, Zhang RY, Wuwu Q, et al. Implicit neural image field for biological microscopy image compression[J]. Nat Comput Sci, 2025, 5(11): 1041-50. doi:10.1038/s43588-025-00889-4 |
| [14] | Mildenhall B, Srinivasan PP, Tancik M, et al. NeRF: representing scenes as neural radiance fields for view synthesis[J]. Commun ACM, 2022, 65(1): 99-106. doi:10.1145/3503250 |
| [15] | Dupont E, Goliński A, Alizadeh M, et al. Coin: Compression with implicit neural representations[EB/OL]. 2021: arXiv:2103.03123. . doi:10.48550/arXiv.2103.03123 |
| [16] | Chen H, He B, Wang H, et al. Nerv: Neural representations for videos[C]//2021 Neural Information Processing Systems Conference (NeurIPS). Dec 6-14, 2021, San Diego, USA. Curran Associates, Inc, 2021: 21557-68. doi:10.48550/arXiv.2110.13903 |
| [17] | Lu Y, Jiang K, Levine JA, et al. Compressive neural representations of volumetric scalar fields[J]. Comput Graph Forum, 2021, 40(3): 135-46. doi:10.1111/cgf.14295 |
| [18] | Bird T, Ballé J, Singh S, et al. 3D scene compression through entropy penalized neural representation functions[C]//2021 Picture Coding Symposium (PCS). June 29-July 2, 2021, Bristol, United Kingdom. IEEE, 2021: 1-5. doi:10.1109/PCS50896.2021.9477505 |
| [19] | Yang RZ, Xiao TX, Cheng YX, et al. Sharing massive biomedical data at magnitudes lower bandwidth using implicit neural function[J]. Proc Natl Acad Sci USA, 2024, 121(28): e2320870121. doi:10.1073/pnas.2320870121 |
| [20] | Ma Y, Yi C, Zhou Y, et al. Semantic redundancy-aware implicit neural compression for multidimensional biomedical image data[J]. Communications Biology, 2024, 7(1): 1081. doi:10.1038/s42003-024-06788-0 |
| [21] | Yang RZ, Xiao TX, Cheng YX, et al. SCI: a spectrum concentrated implicit neural compression for biomedical data[J]. Proc AAAI Conf Artif Intell, 2023, 37(4): 4774-82. doi:10.1609/aaai.v37i4.25602 |
| [22] | Dupont E, Loya H, Alizadeh M, et al. Coin++: Neural compression across modalities[EB/OL]. 2022: arXiv: 2201.12904. . doi:10.48550/arXiv.2201.12904 |
| [23] | Sheibanifard A, Yu HC. A novel implicit neural representation for volume data[J]. Appl Sci, 2023, 13(5): 3242. doi:10.3390/app13053242 |
| [24] | Liang R, Sun H, Vijaykumar N. Coordx: Accelerating implicit neural representation with a split mlp architecture[EB/OL]. 2022: arXiv:2201.12425. . doi:10.48550/arXiv.2201.12425 |
| [25] | Xu JS, Moyer D, Gagoski B, et al. NeSVoR: implicit neural representation for slice-to-volume reconstruction in MRI[J]. IEEE Trans Med Imag, 2023, 42(6): 1707-19. doi:10.1109/tmi.2023.3236216 |
| [26] | Jacobs RA, Jordan MI, Nowlan SJ, et al. Adaptive mixtures of local experts[J]. Neural Comput, 1991, 3(1): 79-87. doi:10.1162/neco.1991.3.1.79 |
| [27] | Heller N, Sathianathen N, Kalapara A, et al. The KiTS19 challenge data: 300 kidney tumor cases with clinical context, CT semantic segmentations, and surgical outcomes[EB/OL]. 2019: arXiv: 1904.00445. . doi:10.48550/arXiv.1904.00445 |
| [28] | Radl L, Jin Y, Pepe A, et al. AVT: Multicenter aortic vessel tree CTA dataset collection with ground truth segmentation masks[J]. Data Brief, 2022, 40: 107801. doi:10.1016/j.dib.2022.107801 |
| [29] | Gong YC, Liu L, Yang M, et al. Compressing deep convolutional networks using vector quantization[EB/OL]. 2014: arXiv: 1412.6115. . |
| [30] | Moffat A. Huffman coding[J]. ACM Comput Surv, 2020, 52(4): 1-35. doi:10.1145/3342555 |
| [31] | Skodras A, Christopoulos C, Ebrahimi T. The JPEG 2000 still image compression standard[J]. IEEE Signal Process Mag, 2001, 18(5): 36-58. doi:10.1109/79.952804 |
| [32] | SJLjTomar. Converting video formats with FFmpeg[J]. Linux journal, 2006, 2006(146):10. |
| [33] | Lu G, Ouyang WL, Xu D, et al. DVC: an end-to-end deep video compression framework[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 15-20, 2019. Long Beach, CA, USA. IEEE, 2019: 10998-1007. doi:10.1109/cvpr.2019.01126 |
| [34] | McDole K, Guignard L, Amat F, et al. In toto imaging and reconstruction of post-implantation mouse development at the single-cell level[J]. Cell, 2018, 175(3): 859-76. e33. doi:10.1016/j.cell.2018.09.031 |
| No related articles found! |
| 阅读次数 | ||||||
|
全文 |
|
|||||
|
摘要 |
|
|||||