Multi-modal diffusion network with controllable variability for medical image segmentation

Abstract

In diffusion-based medical segmentation models, stochastic sampling is commonly used to generate multiple masks. However, the inherent variability in diffusion models can lead to significant biases in some masks, resulting in the fused mask deviating from the true mask. In this study, we propose a novel multi-modal diffusion segmentation network (MMDSN) with controllable variability, specifically designed to address the issue of variability in diffusion models. MMDSN achieves multi-modal conditional control through medical text annotations, thereby enhancing consistency of visual semantic representation and establishing a correspondence between vision and language for diffusion models. Additionally, MMDSN constrains the uncertainty distributions of multiple timesteps within the latent Gaussian space, controlling the variability at each denoising timestep. Extensive experiments on the Qata-Covid19 and …

Publication
In BIBM

本文提出一种具有可控变异性的多模态扩散分割网络(MMDSN),用于医学图像分割。该方法通过医学文本注释实现多模态条件控制,增强视觉语义表示的一致性,并建立视觉与语言之间的对应关系。同时,MMDSN 在潜在高斯空间中对多个时间步的不确定性分布进行约束,从而控制每个去噪时间步的变异性,减少扩散模型随机采样带来的分割偏差。在 Qata-Covid19 等数据集上的实验验证了其有效性。