<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>跨模态融合 | ViLab</title>
    <link>https://vilab.team/tag/%E8%B7%A8%E6%A8%A1%E6%80%81%E8%9E%8D%E5%90%88/</link>
      <atom:link href="https://vilab.team/tag/%E8%B7%A8%E6%A8%A1%E6%80%81%E8%9E%8D%E5%90%88/index.xml" rel="self" type="application/rss+xml" />
    <description>跨模态融合</description>
    <generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Sat, 14 Mar 2026 00:00:00 +0000</lastBuildDate>
    <image>
      <url>https://vilab.team/media/icon_hu2896232876136423579.png</url>
      <title>跨模态融合</title>
      <link>https://vilab.team/tag/%E8%B7%A8%E6%A8%A1%E6%80%81%E8%9E%8D%E5%90%88/</link>
    </image>
    
    <item>
      <title>Seeing the unseen: Zooming in the dark with event cameras</title>
      <link>https://vilab.team/publication/seeing-the-unseen-zooming-in-the-dark-with-event-cameras/</link>
      <pubDate>Sat, 14 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/seeing-the-unseen-zooming-in-the-dark-with-event-cameras/</guid>
      <description>&lt;p&gt;本文提出RetinexEVSR，首个事件驱动的低光视频超分辨率框架。该框架利用高对比度事件信号与Retinex先验，通过双向跨模态融合策略，有效整合噪声事件数据与退化RGB帧中的有用信息。其中，照明引导事件增强模块利用Retinex模型导出的光照图逐步细化事件特征，抑制低光伪影并保留高对比度细节；事件引导反射率增强模块则通过多尺度融合机制动态恢复反射率细节。实验表明，该方法在三个数据集上达到最优性能，在SDSD基准上相比先前事件方法提升2.95 dB，并减少65%运行时间。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>祝贺实验室科研成果发表于 AAAI 2026！</title>
      <link>https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-aaai/</link>
      <pubDate>Sat, 14 Mar 2026 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-aaai/</guid>
      <description>&lt;p&gt;热烈祝贺开大纯同学！论文《Seeing the unseen: Zooming in the dark with event cameras》已发表在 &lt;em&gt;AAAI 2026&lt;/em&gt;。&lt;/p&gt;
&lt;h2 id=&#34;seeing-the-unseen-zooming-in-the-dark-with-event-cameras&#34;&gt;Seeing the unseen: Zooming in the dark with event cameras&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;祝贺开大纯同学！该论文已发表在 &lt;em&gt;AAAI&lt;/em&gt;。&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img alt=&#34;Seeing the unseen: Zooming in the dark with event cameras&#34; srcset=&#34;
               /event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-aaai/images/paper-01_hu13636376305985334282.webp 400w,
               /event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-aaai/images/paper-01_hu16933069341322044714.webp 760w,
               /event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-aaai/images/paper-01_hu642809347018091764.webp 1200w&#34;
               src=&#34;https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-aaai/images/paper-01_hu13636376305985334282.webp&#34;
               width=&#34;760&#34;
               height=&#34;391&#34;
               loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;作者：&lt;/strong&gt; Dachun Kai、Zeyu Xiao、Huyue Zhu、Jiaxiao Wang、Yueyi Zhang、Xiaoyan Sun&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表载体：&lt;/strong&gt; &lt;em&gt;AAAI&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表时间：&lt;/strong&gt; 2026年3月14日&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;相关链接：&lt;/strong&gt; &lt;a href=&#34;https://ojs.aaai.org/index.php/AAAI/article/view/37478&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;论文链接&lt;/a&gt; · &lt;a href=&#34;https://ojs.aaai.org/index.php/AAAI/article/download/37478/41440&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;PDF&lt;/a&gt; · &lt;a href=&#34;https://github.com/DachunKai/RetinexEVSR&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;代码&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;论文介绍&#34;&gt;论文介绍&lt;/h3&gt;
&lt;p&gt;本文提出RetinexEVSR，首个事件驱动的低光视频超分辨率框架。该框架利用高对比度事件信号与Retinex先验，通过双向跨模态融合策略，有效整合噪声事件数据与退化RGB帧中的有用信息。其中，照明引导事件增强模块利用Retinex模型导出的光照图逐步细化事件特征，抑制低光伪影并保留高对比度细节；事件引导反射率增强模块则通过多尺度融合机制动态恢复反射率细节。实验表明，该方法在三个数据集上达到最优性能，在SDSD基准上相比先前事件方法提升2.95 dB，并减少65%运行时间。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>EvTexture&#43;&#43;: Event-Driven Texture Enhancement for Video Super-Resolution</title>
      <link>https://vilab.team/publication/evtexture-event-driven-texture-enhancement-for-video-super-r/</link>
      <pubDate>Mon, 02 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/evtexture-event-driven-texture-enhancement-for-video-super-r/</guid>
      <description>&lt;p&gt;本文提出EvTexture++，一种事件驱动的视频超分辨率纹理增强框架。与以往将事件用于运动估计不同，该方法利用事件的高频时空细节显式恢复纹理，通过定制纹理增强分支和迭代纹理增强模块，逐步挖掘高时间分辨率事件信息，实现纹理区域的渐进细化，从而生成更精确、细节更丰富的高分辨率视频。该框架还可作为即插即用模块提升现有VSR模型性能。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>祝贺实验室科研成果发表于 IEEE TPAMI！</title>
      <link>https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-ieee-tpami/</link>
      <pubDate>Mon, 02 Feb 2026 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-ieee-tpami/</guid>
      <description>&lt;p&gt;热烈祝贺开大纯同学！论文《EvTexture++: Event-Driven Texture Enhancement for Video Super-Resolution》已发表在 &lt;em&gt;IEEE TPAMI&lt;/em&gt;。&lt;/p&gt;
&lt;h2 id=&#34;evtexture-event-driven-texture-enhancement-for-video-super-resolution&#34;&gt;EvTexture++: Event-Driven Texture Enhancement for Video Super-Resolution&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;祝贺开大纯同学！该论文已发表在 &lt;em&gt;IEEE TPAMI&lt;/em&gt;。&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img alt=&#34;EvTexture&amp;#43;&amp;#43;: Event-Driven Texture Enhancement for Video Super-Resolution&#34; srcset=&#34;
               /event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-ieee-tpami/images/paper-01_hu133827975853785128.webp 400w,
               /event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-ieee-tpami/images/paper-01_hu10222371703703860375.webp 760w,
               /event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-ieee-tpami/images/paper-01_hu15191880168499112270.webp 1200w&#34;
               src=&#34;https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-ieee-tpami/images/paper-01_hu133827975853785128.webp&#34;
               width=&#34;760&#34;
               height=&#34;348&#34;
               loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;作者：&lt;/strong&gt; Dachun Kai、Jiayao Lu、Yueyi Zhang、Xiaoyan Sun&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表载体：&lt;/strong&gt; &lt;em&gt;IEEE TPAMI&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表时间：&lt;/strong&gt; 2026年2月2日&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;相关链接：&lt;/strong&gt; &lt;a href=&#34;https://ieeexplore.ieee.org/abstract/document/11369964/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;论文链接&lt;/a&gt; · &lt;a href=&#34;https://arxiv.org/pdf/2606.13580&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;PDF&lt;/a&gt; · &lt;a href=&#34;https://github.com/DachunKai/EvTexture&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;代码&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;论文介绍&#34;&gt;论文介绍&lt;/h3&gt;
&lt;p&gt;本文提出EvTexture++，一种事件驱动的视频超分辨率纹理增强框架。与以往将事件用于运动估计不同，该方法利用事件的高频时空细节显式恢复纹理，通过定制纹理增强分支和迭代纹理增强模块，逐步挖掘高时间分辨率事件信息，实现纹理区域的渐进细化，从而生成更精确、细节更丰富的高分辨率视频。该框架还可作为即插即用模块提升现有VSR模型性能。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>MeDKCoOp: Dual Knowledge-guided Graph Prompt Learning for Biomedical Vision-Language Models</title>
      <link>https://vilab.team/publication/medkcoop-dual-knowledge-guided-graph-prompt-learning-for-bio/</link>
      <pubDate>Mon, 27 Oct 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/medkcoop-dual-knowledge-guided-graph-prompt-learning-for-bio/</guid>
      <description>&lt;p&gt;本文提出MeDKCoOp，一种面向生物医学视觉语言模型的双知识引导图提示学习方法。该方法系统整合医学领域知识，从文本与视觉分支提取专门知识并构建图结构表示，通过知识引导的关系转移实现跨模态融合，并动态优化可学习提示，以增强CLIP等模型在医学下游任务中的适应能力。实验表明其在多个生物医学基准上取得优异性能。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>MMSupcon: An image fusion-based multi-modal supervised contrastive method for brain tumor diagnosis</title>
      <link>https://vilab.team/publication/mmsupcon-an-image-fusion-based-multi-modal-supervised-contra/</link>
      <pubDate>Thu, 28 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/mmsupcon-an-image-fusion-based-multi-modal-supervised-contra/</guid>
      <description>&lt;p&gt;本文针对脑肿瘤多模态MRI诊断中融合策略受限于样本稀缺的问题，提出多模态监督对比学习方法MMSupcon。该方法通过多模态医学图像融合生成信息丰富的样本，并设计多模态监督对比损失，引导模型有效整合互补模态信息，提升诊断准确性。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>祝贺实验室科研成果发表于 Artificial Intelligence in Medicine！</title>
      <link>https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-artificial-intelligence-in-medicine/</link>
      <pubDate>Thu, 28 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/event/%E7%A5%9D%E8%B4%BA%E5%AE%9E%E9%AA%8C%E5%AE%A4%E7%A7%91%E7%A0%94%E6%88%90%E6%9E%9C%E5%8F%91%E8%A1%A8%E4%BA%8E-artificial-intelligence-in-medicine/</guid>
      <description>&lt;p&gt;热烈祝贺王浩宇同学！论文《MMSupcon: An image fusion-based multi-modal supervised contrastive method for brain tumor diagnosis》已发表在 &lt;em&gt;Artificial Intelligence in Medicine&lt;/em&gt;。&lt;/p&gt;
&lt;h2 id=&#34;mmsupcon-an-image-fusion-based-multi-modal-supervised-contrastive-method-for-brain-tumor-diagnosis&#34;&gt;MMSupcon: An image fusion-based multi-modal supervised contrastive method for brain tumor diagnosis&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;祝贺王浩宇同学！该论文已发表在 &lt;em&gt;Artificial Intelligence in Medicine&lt;/em&gt;。&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;作者：&lt;/strong&gt; Haoyu Wang、Jing Zhang、Siying Wu、Haoran Wei、Xun Chen、Yunwei Ou、Xiaoyan Sun&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表载体：&lt;/strong&gt; &lt;em&gt;Artificial Intelligence in Medicine&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表时间：&lt;/strong&gt; 2025年8月28日&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;相关链接：&lt;/strong&gt; &lt;a href=&#34;https://www.sciencedirect.com/science/article/pii/S0933365725001885&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;论文链接&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;论文介绍&#34;&gt;论文介绍&lt;/h3&gt;
&lt;p&gt;本文针对脑肿瘤多模态MRI诊断中融合策略受限于样本稀缺的问题，提出多模态监督对比学习方法MMSupcon。该方法通过多模态医学图像融合生成信息丰富的样本，并设计多模态监督对比损失，引导模型有效整合互补模态信息，提升诊断准确性。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Efficient event-based semantic segmentation via exploiting frame-event fusion: A hybrid neural network approach</title>
      <link>https://vilab.team/publication/efficient-event-based-semantic-segmentation-via-exploiting-f/</link>
      <pubDate>Fri, 11 Apr 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/efficient-event-based-semantic-segmentation-via-exploiting-f/</guid>
      <description>&lt;p&gt;本文提出一种高效的混合神经网络框架，用于事件相机语义分割。该框架包含处理事件流的脉冲神经网络（SNN）分支和处理帧图像的人工神经网络（ANN）分支，并设计了自适应时间加权（ATW）注入器、事件驱动稀疏（EDS）注入器和通道选择融合（CSF）模块，以充分融合帧与事件的互补时空信息。在DDD17-Seg、DSEC-Semantic和M3ED-Semantic数据集上取得了最先进精度，并在DSEC-Semantic上降低63%能耗。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>祝贺实验室 3 项科研成果发表在 AAAI 2025！</title>
      <link>https://vilab.team/event/publication-news-021641c63fc12e96/</link>
      <pubDate>Fri, 11 Apr 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/event/publication-news-021641c63fc12e96/</guid>
      <description>&lt;p&gt;热烈祝贺开大纯同学、李和倍同学、吴沛熹同学！近期，实验室共有 3 项科研成果正式发表。&lt;/p&gt;
&lt;h2 id=&#34;event-enhanced-blurry-video-super-resolution&#34;&gt;Event-enhanced blurry video super-resolution&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;祝贺开大纯同学！该论文已发表在 Proceedings of the AAAI Conference on Artificial Intelligence 39 (4), 4175-4183。&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;作者：&lt;/strong&gt; Dachun Kai、Yueyi Zhang、Jin Wang、Zeyu Xiao、Zhiwei Xiong、Xiaoyan Sun&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表载体：&lt;/strong&gt; Proceedings of the AAAI Conference on Artificial Intelligence 39 (4), 4175-4183&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表时间：&lt;/strong&gt; 2025年4月11日&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;相关链接：&lt;/strong&gt; &lt;a href=&#34;https://ojs.aaai.org/index.php/AAAI/article/view/32438&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;论文链接&lt;/a&gt; · &lt;a href=&#34;https://ojs.aaai.org/index.php/AAAI/article/download/32438/34593&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;PDF&lt;/a&gt; · &lt;a href=&#34;https://github.com/DachunKai/Ev-DeblurVSR&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;代码&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;论文介绍&#34;&gt;论文介绍&lt;/h3&gt;
&lt;p&gt;In this paper, we tackle the task of blurry video super-resolution (BVSR), aiming to generate high-resolution (HR) videos from low-resolution (LR) and blurry inputs. Current BVSR methods often fail to restore sharp details at high resolutions, resulting in noticeable artifacts and jitter due to insufficient motion information for deconvolution and the lack of high-frequency details in LR frames. To address these challenges, we introduce event signals into BVSR and propose a novel event-enhanced network, Ev-DeblurVSR. To effectively fuse information from frames and events for feature deblurring, we introduce a reciprocal feature deblurring module that leverages motion information from intra-frame events to deblur frame features while reciprocally using global scene context from the frames to enhance event features. Furthermore, to enhance temporal consistency, we propose a hybrid deformable alignment module that fully exploits the complementary motion information from inter-frame events and optical flow to improve motion estimation in the deformable alignment process. Extensive evaluations demonstrate that Ev-DeblurVSR establishes a new state-of-the-art performance on both synthetic and real-world datasets. Notably, on real data, our method is 2.59 dB more accurate and 7.28× faster than the recent best BVSR baseline FMA-Net.&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;efficient-event-based-semantic-segmentation-via-exploiting-frame-event-fusion-a-hybrid-neural-network-approach&#34;&gt;Efficient event-based semantic segmentation via exploiting frame-event fusion: A hybrid neural network approach&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;祝贺李和倍同学！该论文已发表在 &lt;em&gt;AAAI&lt;/em&gt; 39(17)。&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img alt=&#34;Efficient event-based semantic segmentation via exploiting frame-event fusion: A hybrid neural network approach&#34; srcset=&#34;
               /event/publication-news-021641c63fc12e96/images/paper-02_hu7998433290983327509.webp 400w,
               /event/publication-news-021641c63fc12e96/images/paper-02_hu5742706483688984605.webp 760w,
               /event/publication-news-021641c63fc12e96/images/paper-02_hu5093957591575336652.webp 1200w&#34;
               src=&#34;https://vilab.team/event/publication-news-021641c63fc12e96/images/paper-02_hu7998433290983327509.webp&#34;
               width=&#34;760&#34;
               height=&#34;237&#34;
               loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;作者：&lt;/strong&gt; Hebei Li、Yansong Peng、Jiahui Yuan、Peixi Wu、Jin Wang、Yueyi Zhang、Xiaoyan Sun&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表载体：&lt;/strong&gt; &lt;em&gt;AAAI&lt;/em&gt; 39(17)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表时间：&lt;/strong&gt; 2025年4月11日&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;相关链接：&lt;/strong&gt; &lt;a href=&#34;https://ojs.aaai.org/index.php/AAAI/article/view/34013&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;论文链接&lt;/a&gt; · &lt;a href=&#34;https://ojs.aaai.org/index.php/AAAI/article/download/34013/36168&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;PDF&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;论文介绍-1&#34;&gt;论文介绍&lt;/h3&gt;
&lt;p&gt;本文提出一种高效的混合神经网络框架，用于事件相机语义分割。该框架包含处理事件流的脉冲神经网络（SNN）分支和处理帧图像的人工神经网络（ANN）分支，并设计了自适应时间加权（ATW）注入器、事件驱动稀疏（EDS）注入器和通道选择融合（CSF）模块，以充分融合帧与事件的互补时空信息。在DDD17-Seg、DSEC-Semantic和M3ED-Semantic数据集上取得了最先进精度，并在DSEC-Semantic上降低63%能耗。&lt;/p&gt;
&lt;hr&gt;
&lt;h2 id=&#34;spiking-point-transformer-for-point-cloud-classification&#34;&gt;Spiking point transformer for point cloud classification&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;祝贺吴沛熹同学！该论文已发表在 &lt;em&gt;AAAI&lt;/em&gt; 39(20)。&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;















&lt;figure  &gt;
  &lt;div class=&#34;d-flex justify-content-center&#34;&gt;
    &lt;div class=&#34;w-100&#34; &gt;&lt;img alt=&#34;Spiking point transformer for point cloud classification&#34; srcset=&#34;
               /event/publication-news-021641c63fc12e96/images/paper-03_hu7240971347687367392.webp 400w,
               /event/publication-news-021641c63fc12e96/images/paper-03_hu15235344109779677276.webp 760w,
               /event/publication-news-021641c63fc12e96/images/paper-03_hu7343380604242636640.webp 1200w&#34;
               src=&#34;https://vilab.team/event/publication-news-021641c63fc12e96/images/paper-03_hu7240971347687367392.webp&#34;
               width=&#34;760&#34;
               height=&#34;418&#34;
               loading=&#34;lazy&#34; data-zoomable /&gt;&lt;/div&gt;
  &lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;作者：&lt;/strong&gt; Peixi Wu、Bosong Chai、Hebei Li、Menghua Zheng、Yansong Peng、Zeyu Wang、Xuan Nie、Yueyi Zhang、Xiaoyan Sun&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表载体：&lt;/strong&gt; &lt;em&gt;AAAI&lt;/em&gt; 39(20)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;发表时间：&lt;/strong&gt; 2025年4月11日&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;相关链接：&lt;/strong&gt; &lt;a href=&#34;https://ojs.aaai.org/index.php/AAAI/article/view/35459&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;论文链接&lt;/a&gt; · &lt;a href=&#34;https://ojs.aaai.org/index.php/AAAI/article/view/35459/37614&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;PDF&lt;/a&gt; · &lt;a href=&#34;https://github.com/PeppaWu/SPT&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;代码&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;论文介绍-2&#34;&gt;论文介绍&lt;/h3&gt;
&lt;p&gt;本文提出Spiking Point Transformer（SPT），首个基于Transformer的脉冲神经网络框架，用于三维点云分类。SPT设计队列驱动采样直接编码，在降低计算成本的同时保留关键支撑点；并引入混合动力学积分发放神经元（HD-IF），模拟选择性神经元激活，减少对特定人工神经元的过度依赖。在多个真实与合成点云基准上取得领先结果，理论能耗较ANN对应模型降低至少6.4倍。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Multi-modal diffusion network with controllable variability for medical image segmentation</title>
      <link>https://vilab.team/publication/multi-modal-diffusion-network-with-controllable-variability-/</link>
      <pubDate>Tue, 03 Dec 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/multi-modal-diffusion-network-with-controllable-variability-/</guid>
      <description>&lt;p&gt;本文提出一种具有可控变异性的多模态扩散分割网络（MMDSN），用于医学图像分割。该方法通过医学文本注释实现多模态条件控制，增强视觉语义表示的一致性，并建立视觉与语言之间的对应关系。同时，MMDSN 在潜在高斯空间中对多个时间步的不确定性分布进行约束，从而控制每个去噪时间步的变异性，减少扩散模型随机采样带来的分割偏差。在 Qata-Covid19 等数据集上的实验验证了其有效性。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Asymmetric event-guided video super-resolution</title>
      <link>https://vilab.team/publication/asymmetric-event-guided-video-super-resolution/</link>
      <pubDate>Mon, 28 Oct 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/asymmetric-event-guided-video-super-resolution/</guid>
      <description>&lt;p&gt;本文首次提出非对称事件引导的视频超分辨率任务，针对事件相机与RGB相机难以严格标定的实际场景，构建了非对称事件引导视频超分辨率网络（AsEVSRN）。该网络通过专门设计的事件特征利用与跨模态融合机制，充分发挥事件相机高时间分辨率优势，有效提升视频超分辨率性能，拓展了事件相机在双摄手机、无人机等新兴高分辨率设备上的应用潜力。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Event-assisted low-light video object segmentation</title>
      <link>https://vilab.team/publication/event-assisted-low-light-video-object-segmentation/</link>
      <pubDate>Sun, 16 Jun 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/event-assisted-low-light-video-object-segmentation/</guid>
      <description>&lt;p&gt;本文针对低光照条件下视频目标分割（VOS）性能严重下降的问题，提出一种利用事件相机数据辅助分割的新框架。该方法包含两个关键模块：自适应跨模态融合（ACMF）模块，用于提取并融合图像与事件模态特征以抑制噪声干扰；事件引导记忆匹配（EGMM）模块，用于修正低光下查询帧与记忆帧之间的相似度计算误差。实验表明，该方法在合成和真实低光数据集上均能显著提升分割精度，生成更准确的目标掩码。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Tmformer: Token merging transformer for brain tumor segmentation with missing modalities</title>
      <link>https://vilab.team/publication/tmformer-token-merging-transformer-for-brain-tumor-segmentat/</link>
      <pubDate>Sun, 24 Mar 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/tmformer-token-merging-transformer-for-brain-tumor-segmentat/</guid>
      <description>&lt;p&gt;本文提出 TMFormer，一种用于缺失模态脑肿瘤分割的 Token 合并 Transformer。该方法通过提取并合并可用模态为更紧凑的 token 序列，解决现有方法以零图填充缺失模态带来的特征偏差与冗余计算问题。其核心包括单模态 Token 合并块（UMB）和多模态 Token 合并块（MMB），分别增强单模态表示并缓解多模态融合偏差。在 BraTS 2018 和 2020 数据集上的实验表明，TMFormer 在缺失模态场景下优于现有方法。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Video super-resolution via event-driven temporal alignment</title>
      <link>https://vilab.team/publication/video-super-resolution-via-event-driven-temporal-alignment/</link>
      <pubDate>Sun, 08 Oct 2023 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/video-super-resolution-via-event-driven-temporal-alignment/</guid>
      <description>&lt;p&gt;本文提出一种事件驱动的双向视频超分辨率框架（EBVSR），利用事件相机的高时间分辨率特性捕捉非线性运动，并设计事件辅助的时间对齐模块，以补充光流法在快速光照变化下的不足。同时构建基于事件的帧合成模块，通过双向跨模态融合增强网络对光照变化的鲁棒性。在合成和真实数据上的实验验证了该方法在视频超分辨率任务中的有效性。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Multimodal sentiment analysis with preferential fusion and distance-aware contrastive learning</title>
      <link>https://vilab.team/publication/multimodal-sentiment-analysis-with-preferential-fusion-and-d/</link>
      <pubDate>Mon, 10 Jul 2023 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/multimodal-sentiment-analysis-with-preferential-fusion-and-d/</guid>
      <description>&lt;p&gt;本文针对多模态情感分析中文本模态与情感标签之间存在的虚假关联问题，提出了一种名为PriSA的新框架。该框架首先通过优先跨模态融合方法，利用文本模态引导计算跨模态相关性；随后引入距离感知对比学习，利用情感标签之间的距离信息进一步计算混合模态相关性；最终基于混合模态相关性和判别性类内特征识别情感信息。实验表明该方法能有效缓解文本虚假关联带来的影响，提升多模态情感分析性能。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Enriching optical flow with appearance information for action recognition</title>
      <link>https://vilab.team/publication/enriching-optical-flow-with-appearance-information-for-actio/</link>
      <pubDate>Wed, 01 Jan 2020 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/enriching-optical-flow-with-appearance-information-for-actio/</guid>
      <description>&lt;p&gt;本文提出一种利用外观信息丰富光流表示的动作识别方法。通过将RGB外观特征与光流特征进行融合，增强运动表征的判别能力，从而提升视频动作识别的准确率。该方法在多个基准数据集上验证了有效性，表明外观信息能够有效补充光流在动作识别中的不足。&lt;/p&gt;
</description>
    </item>
    
  </channel>
</rss>
