<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Zheyu Zhang | ViLab</title>
    <link>https://vilab.team/author/zheyu-zhang/</link>
      <atom:link href="https://vilab.team/author/zheyu-zhang/index.xml" rel="self" type="application/rss+xml" />
    <description>Zheyu Zhang</description>
    <generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Sun, 03 May 2026 00:00:00 +0000</lastBuildDate>
    <image>
      <url>https://vilab.team/media/icon_hu2896232876136423579.png</url>
      <title>Zheyu Zhang</title>
      <link>https://vilab.team/author/zheyu-zhang/</link>
    </image>
    
    <item>
      <title>Salient Diagnostic Value Perception For Preoperative Posterior Fossa Tumor Diagnosis</title>
      <link>https://vilab.team/publication/salient-diagnostic-value-perception-for-preoperative-posteri/</link>
      <pubDate>Sun, 03 May 2026 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/salient-diagnostic-value-perception-for-preoperative-posteri/</guid>
      <description>&lt;p&gt;本文提出显著诊断价值感知方法（SDVP），用于后颅窝肿瘤的术前准确诊断。该方法整合MRI影像与放射学报告，从三个互补视角学习关键诊断线索：通过对抗性样本内对比学习增强跨中心与设备差异的鲁棒性；借助知识增强的样本内对比学习提取专家引导的样本特异性特征；并在干净MRI样本上进行监督式类间对比学习以强化类别特征。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>MeDKCoOp: Dual Knowledge-guided Graph Prompt Learning for Biomedical Vision-Language Models</title>
      <link>https://vilab.team/publication/medkcoop-dual-knowledge-guided-graph-prompt-learning-for-bio/</link>
      <pubDate>Mon, 27 Oct 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/medkcoop-dual-knowledge-guided-graph-prompt-learning-for-bio/</guid>
      <description>&lt;p&gt;本文提出MeDKCoOp，一种面向生物医学视觉语言模型的双知识引导图提示学习方法。该方法系统整合医学领域知识，从文本与视觉分支提取专门知识并构建图结构表示，通过知识引导的关系转移实现跨模态融合，并动态优化可学习提示，以增强CLIP等模型在医学下游任务中的适应能力。实验表明其在多个生物医学基准上取得优异性能。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Efficient spiking point mamba for point cloud analysis</title>
      <link>https://vilab.team/publication/efficient-spiking-point-mamba-for-point-cloud-analysis/</link>
      <pubDate>Sun, 19 Oct 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/efficient-spiking-point-mamba-for-point-cloud-analysis/</guid>
      <description>&lt;p&gt;本文提出 Spiking Point Mamba (SPM)，这是首个将 Mamba 引入三维点云分析的脉冲神经网络。针对直接适配 Mamba 时存在的时序动态不匹配和脉冲引起的信息损失问题，作者设计了层次动态编码 (HDE) 以增强动态时序建模，并提出 Spiking Mamba Block (SMB) 来学习跨时间步特征并减少脉冲信息丢失。此外，采用非对称 SNN 训练策略进一步提升性能。SPM 可作为高效骨干网络，适用于点云分类、部件分割与重建等任务。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Incomplete multi-modal brain tumor segmentation via learnable sorting state space model</title>
      <link>https://vilab.team/publication/incomplete-multi-modal-brain-tumor-segmentation-via-learnabl/</link>
      <pubDate>Tue, 10 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/incomplete-multi-modal-brain-tumor-segmentation-via-learnabl/</guid>
      <description>&lt;p&gt;本文提出一种可学习排序状态空间模型（LS3M），用于不完整多模态脑肿瘤分割。该方法基于Mamba架构高效建模长距离依赖，并引入可微置换矩阵，根据模态特定特征对输入序列进行动态重排序，从而保留3D脑MRI中关键的空间归纳偏置与长程语义相关性。LS3M能够充分利用可用模态信息，提升分割性能。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Multi-modal diffusion network with controllable variability for medical image segmentation</title>
      <link>https://vilab.team/publication/multi-modal-diffusion-network-with-controllable-variability-/</link>
      <pubDate>Tue, 03 Dec 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/multi-modal-diffusion-network-with-controllable-variability-/</guid>
      <description>&lt;p&gt;本文提出一种具有可控变异性的多模态扩散分割网络（MMDSN），用于医学图像分割。该方法通过医学文本注释实现多模态条件控制，增强视觉语义表示的一致性，并建立视觉与语言之间的对应关系。同时，MMDSN 在潜在高斯空间中对多个时间步的不确定性分布进行约束，从而控制每个去噪时间步的变异性，减少扩散模型随机采样带来的分割偏差。在 Qata-Covid19 等数据集上的实验验证了其有效性。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Ee-mllm: A data-efficient and compute-efficient multimodal large language model</title>
      <link>https://vilab.team/publication/ee-mllm-a-data-efficient-and-compute-efficient-multimodal-la/</link>
      <pubDate>Wed, 21 Aug 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/ee-mllm-a-data-efficient-and-compute-efficient-multimodal-la/</guid>
      <description>&lt;p&gt;Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated satisfactory performance across various vision-language tasks. Current approaches for vision and language interaction fall into two categories: self-attention-based and cross-attention-based methods. However, both approaches present inherent limitations, forcing a trade-off between data and computational efficiency. To address this issue, we introduce the Data-$\textbf{E}$fficient and Compute-$\textbf{E}$fficient $\textbf{MLLM}$ ($\textbf{EE-MLLM}$). Specifically, we modify the original self-attention mechanism in MLLM to a composite attention mechanism. This mechanism has two key characteristics: 1) eliminating the computational overhead of self-attention among visual tokens to achieve $\textbf{compute efficiency}$, and 2) reusing the weights from each layer of LLM to facilitate effective vision-language modality alignment for $\textbf{data efficiency}$. As a result, EE-MLLM significantly outperforms Flamingo with limited training data, and reduces the prefilling time to 79 ms on an H800 GPU, compared to LLaVA&amp;rsquo;s 277 ms. To further investigate the efficiency of EE-MLLM, we present a training-free variant named EE-MLLM-F, which reduces the computation cost of self-attention-based method without additional training. Experimental results demonstrate the effectiveness of EE-MLLM across a range of benchmarks, including general-purpose datasets like MMBench and SeedBench, as well as fine-grained tasks such as TextVQA and DocVQA.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>A micro-expression recognition system with event cameras</title>
      <link>https://vilab.team/publication/a-micro-expression-recognition-system-with-event-cameras/</link>
      <pubDate>Mon, 15 Jul 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/a-micro-expression-recognition-system-with-event-cameras/</guid>
      <description>&lt;p&gt;本文提出了一种基于事件相机的微表情识别系统。针对微表情持续时间短、幅度微弱、难以用传统相机捕捉的问题，系统利用事件相机的高时间分辨率特性，设计了事件增强运动提取器（EEME）以放大细微运动，并引入事件引导注意力（EGA）聚焦关键面部区域，从而提升微表情识别的准确性与鲁棒性。该系统为情感计算领域提供了有效工具。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Estme: Event-driven spatio-temporal motion enhancement for micro-expression recognition</title>
      <link>https://vilab.team/publication/estme-event-driven-spatio-temporal-motion-enhancement-for-mi/</link>
      <pubDate>Mon, 15 Jul 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/estme-event-driven-spatio-temporal-motion-enhancement-for-mi/</guid>
      <description>&lt;p&gt;本文针对微表情识别中动作幅度小、持续时间短、难以捕捉的问题，提出了一种事件驱动的时空运动增强网络。该方法引入事件相机捕获的高时间分辨率事件信号，设计事件增强运动提取模块以增强细微运动细节，并利用事件引导注意力模块聚焦特定区域的微小变化，从而获取更精确的空间特征。在合成和真实数据集上的实验结果表明，该方法在微表情识别任务上具有优越性能。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Task navigator: Decomposing complex tasks for multimodal large language models</title>
      <link>https://vilab.team/publication/task-navigator-decomposing-complex-tasks-for-multimodal-larg/</link>
      <pubDate>Mon, 17 Jun 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/task-navigator-decomposing-complex-tasks-for-multimodal-larg/</guid>
      <description>&lt;p&gt;本文提出一种名为 Task Navigator 的框架，利用大语言模型作为导航器，将复杂多模态任务逐步分解为更易处理的子问题，并引导多模态大语言模型按步骤求解。该方法无需重新训练模型，而是系统化地调用 MLLM 已有的多种能力，如 OCR、识别、推理等，从而提升复杂任务的处理效果。作者还构建了包含数学推理、嵌入式文本问答和视觉规划等任务的基准，验证了框架的有效性。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Semi-supervised medical image segmentation via dynamic pseudo-label refinement</title>
      <link>https://vilab.team/publication/semi-supervised-medical-image-segmentation-via-dynamic-pseud/</link>
      <pubDate>Mon, 27 May 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/semi-supervised-medical-image-segmentation-via-dynamic-pseud/</guid>
      <description>&lt;p&gt;本文提出一种基于动态伪标签优化的半监督医学图像分割框架。针对双视角方法易丢失重要数据且伪标签不准确的问题，设计分层伪标签生成（HPLG）与动态伪标签校正（DPLC）两个互补模块，按可靠性生成分层像素级伪标签，并利用双视角的一致性与差异进行动态修正，从而提升分割性能与标签质量。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Tmformer: Token merging transformer for brain tumor segmentation with missing modalities</title>
      <link>https://vilab.team/publication/tmformer-token-merging-transformer-for-brain-tumor-segmentat/</link>
      <pubDate>Sun, 24 Mar 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/tmformer-token-merging-transformer-for-brain-tumor-segmentat/</guid>
      <description>&lt;p&gt;本文提出 TMFormer，一种用于缺失模态脑肿瘤分割的 Token 合并 Transformer。该方法通过提取并合并可用模态为更紧凑的 token 序列，解决现有方法以零图填充缺失模态带来的特征偏差与冗余计算问题。其核心包括单模态 Token 合并块（UMB）和多模态 Token 合并块（MMB），分别增强单模态表示并缓解多模态融合偏差。在 BraTS 2018 和 2020 数据集上的实验表明，TMFormer 在缺失模态场景下优于现有方法。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Anatomical consistency distillation and inconsistency synthesis for brain tumor segmentation with missing modalities</title>
      <link>https://vilab.team/publication/anatomical-consistency-distillation-and-inconsistency-synthe/</link>
      <pubDate>Mon, 01 Jan 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/anatomical-consistency-distillation-and-inconsistency-synthe/</guid>
      <description>&lt;p&gt;本文提出ACDIS框架，用于解决脑肿瘤分割中MRI模态缺失的问题。通过解剖一致性蒸馏将多模态图像中的共享解剖结构迁移至单模态表示，并利用模态特征合成块生成模态特定特征，从而增强单模态图像在特定区域的组织表现。该方法有效缓解了模态缺失带来的性能下降，提升了分割精度。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Eoformer: Edge-oriented transformer for brain tumor segmentation</title>
      <link>https://vilab.team/publication/eoformer-edge-oriented-transformer-for-brain-tumor-segmentat/</link>
      <pubDate>Sun, 01 Oct 2023 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/eoformer-edge-oriented-transformer-for-brain-tumor-segmentat/</guid>
      <description>&lt;p&gt;本文提出边缘导向Transformer（EoFormer），用于脑肿瘤MRI图像分割。该方法采用CNN-Transformer混合编码器，CNN提取局部低级特征，Transformer建模长距离依赖以生成全局高级特征；解码器集成边缘导向Sobel与Laplacian锐化模块，增强边缘信息。同时引入高效注意力与重参数化技术，提升特征表示能力与分割精度。&lt;/p&gt;
</description>
    </item>
    
  </channel>
</rss>
