<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>多模态学习 | ViLab</title>
    <link>https://vilab.team/tag/%E5%A4%9A%E6%A8%A1%E6%80%81%E5%AD%A6%E4%B9%A0/</link>
      <atom:link href="https://vilab.team/tag/%E5%A4%9A%E6%A8%A1%E6%80%81%E5%AD%A6%E4%B9%A0/index.xml" rel="self" type="application/rss+xml" />
    <description>多模态学习</description>
    <generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Sun, 03 May 2026 00:00:00 +0000</lastBuildDate>
    <image>
      <url>https://vilab.team/media/icon_hu2896232876136423579.png</url>
      <title>多模态学习</title>
      <link>https://vilab.team/tag/%E5%A4%9A%E6%A8%A1%E6%80%81%E5%AD%A6%E4%B9%A0/</link>
    </image>
    
    <item>
      <title>Salient Diagnostic Value Perception For Preoperative Posterior Fossa Tumor Diagnosis</title>
      <link>https://vilab.team/publication/salient-diagnostic-value-perception-for-preoperative-posteri/</link>
      <pubDate>Sun, 03 May 2026 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/salient-diagnostic-value-perception-for-preoperative-posteri/</guid>
      <description>&lt;p&gt;本文提出显著诊断价值感知方法（SDVP），用于后颅窝肿瘤的术前准确诊断。该方法整合MRI影像与放射学报告，从三个互补视角学习关键诊断线索：通过对抗性样本内对比学习增强跨中心与设备差异的鲁棒性；借助知识增强的样本内对比学习提取专家引导的样本特异性特征；并在干净MRI样本上进行监督式类间对比学习以强化类别特征。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>祝贺实验室 2 项科研成果发表在ICASSP 2026！</title>
      <link>https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4-2-%E9%A1%B9%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8/</link>
      <pubDate>Sun, 03 May 2026 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4-2-%E9%A1%B9%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8/</guid>
      <description>&lt;p&gt;热烈祝贺陈政同学、吴小满同学！近期，实验室共有 2 项科研成果正式发表。&lt;/p&gt;
&lt;h2 id=&#34;salient-diagnostic-value-perception-for-preoperative-posterior-fossa-tumor-diagnosis&#34;&gt;Salient Diagnostic Value Perception For Preoperative Posterior Fossa Tumor Diagnosis&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;祝贺陈政同学！该论文已发表在 &lt;em&gt;ICASSP 2026&lt;/em&gt;。&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;作者：&lt;/strong&gt; Zheng Chen、Peng Li、Zheyu Zhang、Siying Wu、Jing Zhang、Cuiling Hu、Yunwei Ou、Xiaoyan Sun&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表载体：&lt;/strong&gt; &lt;em&gt;ICASSP&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表时间：&lt;/strong&gt; 2026年5月3日&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;相关链接：&lt;/strong&gt; &lt;a href=&#34;https://ieeexplore.ieee.org/abstract/document/11460465/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;论文链接&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;论文介绍&#34;&gt;论文介绍&lt;/h3&gt;
&lt;p&gt;本文提出显著诊断价值感知方法（SDVP），用于后颅窝肿瘤的术前准确诊断。该方法整合MRI影像与放射学报告，从三个互补视角学习关键诊断线索：通过对抗性样本内对比学习增强跨中心与设备差异的鲁棒性；借助知识增强的样本内对比学习提取专家引导的样本特异性特征；并在干净MRI样本上进行监督式类间对比学习以强化类别特征。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;sscm-a-spatial-semantic-consistent-model-for-multi-contrast-mri-super-resolution&#34;&gt;SSCM: A Spatial-Semantic Consistent Model for Multi-Contrast MRI Super-Resolution&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;祝贺吴小满同学！该论文已发表在 &lt;em&gt;ICASSP 2026&lt;/em&gt;。&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img alt=&#34;SSCM: A Spatial-Semantic Consistent Model for Multi-Contrast MRI Super-Resolution&#34; srcset=&#34;
               /event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4-2-%E9%A1%B9%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8/images/paper-02_hu16643979777475813095.webp 400w,
               /event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4-2-%E9%A1%B9%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8/images/paper-02_hu13958872581769715241.webp 760w,
               /event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4-2-%E9%A1%B9%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8/images/paper-02_hu13117256722096246496.webp 1200w&#34;
               src=&#34;https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4-2-%E9%A1%B9%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8/images/paper-02_hu16643979777475813095.webp&#34;
               width=&#34;760&#34;
               height=&#34;438&#34;
               loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;作者：&lt;/strong&gt; Xiaoman Wu、Lubin Gan、Siying Wu、Jing Zhang、Yunwei Ou、Xiaoyan Sun&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表载体：&lt;/strong&gt; &lt;em&gt;ICASSP&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表时间：&lt;/strong&gt; 2026年5月3日&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;相关链接：&lt;/strong&gt; &lt;a href=&#34;https://ieeexplore.ieee.org/abstract/document/11464902/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;论文链接&lt;/a&gt; · &lt;a href=&#34;https://arxiv.org/pdf/2509.18593&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;PDF&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;论文介绍-1&#34;&gt;论文介绍&lt;/h3&gt;
&lt;p&gt;本文提出空间语义一致模型（SSCM），用于多对比度磁共振成像超分辨率。该方法通过动态空间扭曲模块实现对比度间空间对齐，利用语义感知令牌聚合块建模长程依赖，并结合空间-频率融合块恢复高频细节，从而在结构差异和运动干扰下保持解剖结构的空间语义一致性。实验表明SSCM在效率和性能上优于现有方法。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Enhancing zero-shot brain tumor subtype classification via fine-grained patch-text alignment</title>
      <link>https://vilab.team/publication/enhancing-zero-shot-brain-tumor-subtype-classification-via-f/</link>
      <pubDate>Mon, 27 Oct 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/enhancing-zero-shot-brain-tumor-subtype-classification-via-f/</guid>
      <description>&lt;p&gt;本文提出细粒度补丁对齐网络（FG-PAN），用于脑肿瘤亚型的零样本分类。该方法包含局部特征细化模块，通过建模代表性补丁间的空间关系增强视觉特征；以及细粒度文本描述生成模块，利用大语言模型生成病理感知的类别语义原型。通过对齐细粒度视觉与语义特征，并引入坐标感知聚合机制，FG-PAN在整张病理切片级别实现了更准确的亚型判别，缓解了标注数据稀缺和形态差异细微带来的挑战。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>MeDKCoOp: Dual Knowledge-guided Graph Prompt Learning for Biomedical Vision-Language Models</title>
      <link>https://vilab.team/publication/medkcoop-dual-knowledge-guided-graph-prompt-learning-for-bio/</link>
      <pubDate>Mon, 27 Oct 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/medkcoop-dual-knowledge-guided-graph-prompt-learning-for-bio/</guid>
      <description>&lt;p&gt;本文提出MeDKCoOp，一种面向生物医学视觉语言模型的双知识引导图提示学习方法。该方法系统整合医学领域知识，从文本与视觉分支提取专门知识并构建图结构表示，通过知识引导的关系转移实现跨模态融合，并动态优化可学习提示，以增强CLIP等模型在医学下游任务中的适应能力。实验表明其在多个生物医学基准上取得优异性能。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>祝贺实验室科研成果发表于 Expert Systems with Applications 130161！</title>
      <link>https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-expert-systems-with-applications-130161/</link>
      <pubDate>Mon, 27 Oct 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-expert-systems-with-applications-130161/</guid>
      <description>&lt;p&gt;热烈祝贺甘鲁斌同学！论文《Enhancing zero-shot brain tumor subtype classification via fine-grained patch-text alignment》已发表在 &lt;em&gt;Expert Systems with Applications&lt;/em&gt; 130161。&lt;/p&gt;
&lt;h2 id=&#34;enhancing-zero-shot-brain-tumor-subtype-classification-via-fine-grained-patch-text-alignment&#34;&gt;Enhancing zero-shot brain tumor subtype classification via fine-grained patch-text alignment&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;祝贺甘鲁斌同学！该论文已发表在 &lt;em&gt;Expert Systems with Applications&lt;/em&gt; 130161。&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img alt=&#34;Enhancing zero-shot brain tumor subtype classification via fine-grained patch-text alignment&#34; srcset=&#34;
               /event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-expert-systems-with-applications-130161/images/paper-01_hu479104807647607719.webp 400w,
               /event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-expert-systems-with-applications-130161/images/paper-01_hu7525044382441938419.webp 760w,
               /event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-expert-systems-with-applications-130161/images/paper-01_hu10830315984298057718.webp 1200w&#34;
               src=&#34;https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-expert-systems-with-applications-130161/images/paper-01_hu479104807647607719.webp&#34;
               width=&#34;760&#34;
               height=&#34;445&#34;
               loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;作者：&lt;/strong&gt; Lubin Gan、Jing Zhang、Linhao Qu、Yijun Wang、Siying Wu、Xiaoyan Sun&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表载体：&lt;/strong&gt; &lt;em&gt;Expert Systems with Applications&lt;/em&gt; 130161&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表时间：&lt;/strong&gt; 2025年10月27日&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;相关链接：&lt;/strong&gt; &lt;a href=&#34;https://www.sciencedirect.com/science/article/pii/S0957417425037765&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;论文链接&lt;/a&gt; · &lt;a href=&#34;https://arxiv.org/pdf/2508.01602&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;PDF&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;论文介绍&#34;&gt;论文介绍&lt;/h3&gt;
&lt;p&gt;本文提出细粒度补丁对齐网络（FG-PAN），用于脑肿瘤亚型的零样本分类。该方法包含局部特征细化模块，通过建模代表性补丁间的空间关系增强视觉特征；以及细粒度文本描述生成模块，利用大语言模型生成病理感知的类别语义原型。通过对齐细粒度视觉与语义特征，并引入坐标感知聚合机制，FG-PAN在整张病理切片级别实现了更准确的亚型判别，缓解了标注数据稀缺和形态差异细微带来的挑战。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Enhancing Visual Question Answering Via Clustered In-Context Sequence Configuration</title>
      <link>https://vilab.team/publication/enhancing-visual-question-answering-via-clustered-in-context/</link>
      <pubDate>Sun, 14 Sep 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/enhancing-visual-question-answering-via-clustered-in-context/</guid>
      <description>&lt;p&gt;本文针对多模态大语言模型在多模态上下文学习中的演示序列配置问题，提出一种基于聚类的上下文配置方法。该方法自适应地对候选数据进行分组，并从每个簇中选取演示样本，以增强序列内多样性并保持语义一致性，从而减少高相似演示带来的归纳偏置，使模型更关注演示的主要意图。在OK-VQA、VQAv2、VizWiz和TextVQA四个视觉问答基准上的实验验证了其有效性。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>祝贺实验室科研成果发表于 ICIP 2025！</title>
      <link>https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-icip/</link>
      <pubDate>Sun, 14 Sep 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-icip/</guid>
      <description>&lt;p&gt;热烈祝贺贺子龙同学！论文《Enhancing Visual Question Answering Via Clustered In-Context Sequence Configuration》已发表在 &lt;em&gt;ICIP&lt;/em&gt;。&lt;/p&gt;
&lt;h2 id=&#34;enhancing-visual-question-answering-via-clustered-in-context-sequence-configuration&#34;&gt;Enhancing Visual Question Answering Via Clustered In-Context Sequence Configuration&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;祝贺贺子龙同学！该论文已发表在 &lt;em&gt;ICIP&lt;/em&gt;。&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;作者：&lt;/strong&gt; Zilong He、Yijun Pan、Hebei Li、Feipeng Ma、Yansong Peng、Siying Wu、Xiaoyan Sun&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表载体：&lt;/strong&gt; &lt;em&gt;ICIP&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表时间：&lt;/strong&gt; 2025年9月14日&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;相关链接：&lt;/strong&gt; &lt;a href=&#34;https://ieeexplore.ieee.org/abstract/document/11084650/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;论文链接&lt;/a&gt; · &lt;a href=&#34;https://drive.google.com/file/d/1rivgSv2AEK1_1Q40pHnp-VK9QCfzskTK/view&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;PDF&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;论文介绍&#34;&gt;论文介绍&lt;/h3&gt;
&lt;p&gt;本文针对多模态大语言模型在多模态上下文学习中的演示序列配置问题，提出一种基于聚类的上下文配置方法。该方法自适应地对候选数据进行分组，并从每个簇中选取演示样本，以增强序列内多样性并保持语义一致性，从而减少高相似演示带来的归纳偏置，使模型更关注演示的主要意图。在OK-VQA、VQAv2、VizWiz和TextVQA四个视觉问答基准上的实验验证了其有效性。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>MMSupcon: An image fusion-based multi-modal supervised contrastive method for brain tumor diagnosis</title>
      <link>https://vilab.team/publication/mmsupcon-an-image-fusion-based-multi-modal-supervised-contra/</link>
      <pubDate>Thu, 28 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/mmsupcon-an-image-fusion-based-multi-modal-supervised-contra/</guid>
      <description>&lt;p&gt;本文针对脑肿瘤多模态MRI诊断中融合策略受限于样本稀缺的问题，提出多模态监督对比学习方法MMSupcon。该方法通过多模态医学图像融合生成信息丰富的样本，并设计多模态监督对比损失，引导模型有效整合互补模态信息，提升诊断准确性。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>祝贺实验室科研成果发表于 Artificial Intelligence in Medicine！</title>
      <link>https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-artificial-intelligence-in-medicine/</link>
      <pubDate>Thu, 28 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-artificial-intelligence-in-medicine/</guid>
      <description>&lt;p&gt;热烈祝贺王浩宇同学！论文《MMSupcon: An image fusion-based multi-modal supervised contrastive method for brain tumor diagnosis》已发表在 &lt;em&gt;Artificial Intelligence in Medicine&lt;/em&gt;。&lt;/p&gt;
&lt;h2 id=&#34;mmsupcon-an-image-fusion-based-multi-modal-supervised-contrastive-method-for-brain-tumor-diagnosis&#34;&gt;MMSupcon: An image fusion-based multi-modal supervised contrastive method for brain tumor diagnosis&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;祝贺王浩宇同学！该论文已发表在 &lt;em&gt;Artificial Intelligence in Medicine&lt;/em&gt;。&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;作者：&lt;/strong&gt; Haoyu Wang、Jing Zhang、Siying Wu、Haoran Wei、Xun Chen、Yunwei Ou、Xiaoyan Sun&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表载体：&lt;/strong&gt; &lt;em&gt;Artificial Intelligence in Medicine&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表时间：&lt;/strong&gt; 2025年8月28日&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;相关链接：&lt;/strong&gt; &lt;a href=&#34;https://www.sciencedirect.com/science/article/pii/S0933365725001885&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;论文链接&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;论文介绍&#34;&gt;论文介绍&lt;/h3&gt;
&lt;p&gt;本文针对脑肿瘤多模态MRI诊断中融合策略受限于样本稀缺的问题，提出多模态监督对比学习方法MMSupcon。该方法通过多模态医学图像融合生成信息丰富的样本，并设计多模态监督对比损失，引导模型有效整合互补模态信息，提升诊断准确性。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Incomplete multi-modal brain tumor segmentation via learnable sorting state space model</title>
      <link>https://vilab.team/publication/incomplete-multi-modal-brain-tumor-segmentation-via-learnabl/</link>
      <pubDate>Tue, 10 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/incomplete-multi-modal-brain-tumor-segmentation-via-learnabl/</guid>
      <description>&lt;p&gt;本文提出一种可学习排序状态空间模型（LS3M），用于不完整多模态脑肿瘤分割。该方法基于Mamba架构高效建模长距离依赖，并引入可微置换矩阵，根据模态特定特征对输入序列进行动态重排序，从而保留3D脑MRI中关键的空间归纳偏置与长程语义相关性。LS3M能够充分利用可用模态信息，提升分割性能。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>祝贺实验室科研成果发表于 CVPR 2025！</title>
      <link>https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-cvpr/</link>
      <pubDate>Tue, 10 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-cvpr/</guid>
      <description>&lt;p&gt;热烈祝贺张哲宇同学！论文《Incomplete multi-modal brain tumor segmentation via learnable sorting state space model》已发表在 &lt;em&gt;CVPR&lt;/em&gt;。&lt;/p&gt;
&lt;h2 id=&#34;incomplete-multi-modal-brain-tumor-segmentation-via-learnable-sorting-state-space-model&#34;&gt;Incomplete multi-modal brain tumor segmentation via learnable sorting state space model&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;祝贺张哲宇同学！该论文已发表在 &lt;em&gt;CVPR&lt;/em&gt;。&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img alt=&#34;Incomplete multi-modal brain tumor segmentation via learnable sorting state space model&#34; srcset=&#34;
               /event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-cvpr/images/paper-01_hu12446196012194697519.webp 400w,
               /event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-cvpr/images/paper-01_hu11624591397241064252.webp 760w,
               /event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-cvpr/images/paper-01_hu8784721236978794445.webp 1200w&#34;
               src=&#34;https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-cvpr/images/paper-01_hu12446196012194697519.webp&#34;
               width=&#34;760&#34;
               height=&#34;210&#34;
               loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;作者：&lt;/strong&gt; Zheyu Zhang、Yayuan Lu、Feipeng Ma、Yueyi Zhang、Huanjing Yue、Xiaoyan Sun&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表载体：&lt;/strong&gt; &lt;em&gt;CVPR&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表时间：&lt;/strong&gt; 2025年6月10日&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;相关链接：&lt;/strong&gt; &lt;a href=&#34;https://ieeexplore.ieee.org/abstract/document/11094296/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;论文链接&lt;/a&gt; · &lt;a href=&#34;https://openaccess.thecvf.com/content/CVPR2025/papers/Zhang_Incomplete_Multi-modal_Brain_Tumor_Segmentation_via_Learnable_Sorting_State_Space_CVPR_2025_paper.pdf&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;PDF&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;论文介绍&#34;&gt;论文介绍&lt;/h3&gt;
&lt;p&gt;本文提出一种可学习排序状态空间模型（LS3M），用于不完整多模态脑肿瘤分割。该方法基于Mamba架构高效建模长距离依赖，并引入可微置换矩阵，根据模态特定特征对输入序列进行动态重排序，从而保留3D脑MRI中关键的空间归纳偏置与长程语义相关性。LS3M能够充分利用可用模态信息，提升分割性能。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Multi-modal diffusion network with controllable variability for medical image segmentation</title>
      <link>https://vilab.team/publication/multi-modal-diffusion-network-with-controllable-variability-/</link>
      <pubDate>Tue, 03 Dec 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/multi-modal-diffusion-network-with-controllable-variability-/</guid>
      <description>&lt;p&gt;本文提出一种具有可控变异性的多模态扩散分割网络（MMDSN），用于医学图像分割。该方法通过医学文本注释实现多模态条件控制，增强视觉语义表示的一致性，并建立视觉与语言之间的对应关系。同时，MMDSN 在潜在高斯空间中对多个时间步的不确定性分布进行约束，从而控制每个去噪时间步的变异性，减少扩散模型随机采样带来的分割偏差。在 Qata-Covid19 等数据集上的实验验证了其有效性。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Asymmetric event-guided video super-resolution</title>
      <link>https://vilab.team/publication/asymmetric-event-guided-video-super-resolution/</link>
      <pubDate>Mon, 28 Oct 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/asymmetric-event-guided-video-super-resolution/</guid>
      <description>&lt;p&gt;本文首次提出非对称事件引导的视频超分辨率任务，针对事件相机与RGB相机难以严格标定的实际场景，构建了非对称事件引导视频超分辨率网络（AsEVSRN）。该网络通过专门设计的事件特征利用与跨模态融合机制，充分发挥事件相机高时间分辨率优势，有效提升视频超分辨率性能，拓展了事件相机在双摄手机、无人机等新兴高分辨率设备上的应用潜力。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Event-assisted low-light video object segmentation</title>
      <link>https://vilab.team/publication/event-assisted-low-light-video-object-segmentation/</link>
      <pubDate>Sun, 16 Jun 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/event-assisted-low-light-video-object-segmentation/</guid>
      <description>&lt;p&gt;本文针对低光照条件下视频目标分割（VOS）性能严重下降的问题，提出一种利用事件相机数据辅助分割的新框架。该方法包含两个关键模块：自适应跨模态融合（ACMF）模块，用于提取并融合图像与事件模态特征以抑制噪声干扰；事件引导记忆匹配（EGMM）模块，用于修正低光下查询帧与记忆帧之间的相似度计算误差。实验表明，该方法在合成和真实低光数据集上均能显著提升分割精度，生成更准确的目标掩码。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Multi-modal generative embedding model</title>
      <link>https://vilab.team/publication/multi-modal-generative-embedding-model/</link>
      <pubDate>Wed, 29 May 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/multi-modal-generative-embedding-model/</guid>
      <description>&lt;p&gt;本文提出多模态生成嵌入模型MM-GEM，将生成与嵌入两种目标统一于单个大语言模型中，实现每个模态仅需一个模型。通过引入PoolAggregator提升效率并支持细粒度嵌入与生成。实验表明，生成与嵌入目标并不显著冲突，模型在跨模态检索、零样本分类和图像描述等任务上表现优异，同时具备区域级描述生成与检索能力，并在长文本图像检索中取得显著提升。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Tmformer: Token merging transformer for brain tumor segmentation with missing modalities</title>
      <link>https://vilab.team/publication/tmformer-token-merging-transformer-for-brain-tumor-segmentat/</link>
      <pubDate>Sun, 24 Mar 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/tmformer-token-merging-transformer-for-brain-tumor-segmentat/</guid>
      <description>&lt;p&gt;本文提出 TMFormer，一种用于缺失模态脑肿瘤分割的 Token 合并 Transformer。该方法通过提取并合并可用模态为更紧凑的 token 序列，解决现有方法以零图填充缺失模态带来的特征偏差与冗余计算问题。其核心包括单模态 Token 合并块（UMB）和多模态 Token 合并块（MMB），分别增强单模态表示并缓解多模态融合偏差。在 BraTS 2018 和 2020 数据集上的实验表明，TMFormer 在缺失模态场景下优于现有方法。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Anatomical consistency distillation and inconsistency synthesis for brain tumor segmentation with missing modalities</title>
      <link>https://vilab.team/publication/anatomical-consistency-distillation-and-inconsistency-synthe/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/anatomical-consistency-distillation-and-inconsistency-synthe/</guid>
      <description>&lt;p&gt;本文提出ACDIS框架，用于解决脑肿瘤分割中MRI模态缺失的问题。通过解剖一致性蒸馏将多模态图像中的共享解剖结构迁移至单模态表示，并利用模态特征合成块生成模态特定特征，从而增强单模态图像在特定区域的组织表现。该方法有效缓解了模态缺失带来的性能下降，提升了分割精度。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Text-Only Image Captioning with Multi-Context Data Generation.</title>
      <link>https://vilab.team/publication/text-only-image-captioning-with-multi-context-data-generatio/</link>
      <pubDate>Sun, 01 Jan 2023 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/text-only-image-captioning-with-multi-context-data-generatio/</guid>
      <description>&lt;p&gt;本文针对仅使用文本数据训练图像描述模型的任务，提出一种多上下文数据生成方法。通过构造多样化的文本上下文，生成合成图像-描述训练对，使模型学习跨模态对齐与语义描述能力，减少对真实图像标注的依赖。实验表明该方法在多个图像描述基准上取得有效性能，为数据稀缺场景下的多模态学习提供了新思路。&lt;/p&gt;
</description>
    </item>
    
  </channel>
</rss>
