<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Siying Wu | ViLab</title>
    <link>https://vilab.team/author/siying-wu/</link>
      <atom:link href="https://vilab.team/author/siying-wu/index.xml" rel="self" type="application/rss+xml" />
    <description>Siying Wu</description>
    <generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Sun, 03 May 2026 00:00:00 +0000</lastBuildDate>
    <image>
      <url>https://vilab.team/media/icon_hu2896232876136423579.png</url>
      <title>Siying Wu</title>
      <link>https://vilab.team/author/siying-wu/</link>
    </image>
    
    <item>
      <title>Salient Diagnostic Value Perception For Preoperative Posterior Fossa Tumor Diagnosis</title>
      <link>https://vilab.team/publication/salient-diagnostic-value-perception-for-preoperative-posteri/</link>
      <pubDate>Sun, 03 May 2026 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/salient-diagnostic-value-perception-for-preoperative-posteri/</guid>
      <description>&lt;p&gt;本文提出显著诊断价值感知方法（SDVP），用于后颅窝肿瘤的术前准确诊断。该方法整合MRI影像与放射学报告，从三个互补视角学习关键诊断线索：通过对抗性样本内对比学习增强跨中心与设备差异的鲁棒性；借助知识增强的样本内对比学习提取专家引导的样本特异性特征；并在干净MRI样本上进行监督式类间对比学习以强化类别特征。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>SSCM: A Spatial-Semantic Consistent Model for Multi-Contrast MRI Super-Resolution</title>
      <link>https://vilab.team/publication/sscm-a-spatial-semantic-consistent-model-for-multi-contrast-/</link>
      <pubDate>Sun, 03 May 2026 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/sscm-a-spatial-semantic-consistent-model-for-multi-contrast-/</guid>
      <description>&lt;p&gt;本文提出空间语义一致模型（SSCM），用于多对比度磁共振成像超分辨率。该方法通过动态空间扭曲模块实现对比度间空间对齐，利用语义感知令牌聚合块建模长程依赖，并结合空间-频率融合块恢复高频细节，从而在结构差异和运动干扰下保持解剖结构的空间语义一致性。实验表明SSCM在效率和性能上优于现有方法。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Enhancing zero-shot brain tumor subtype classification via fine-grained patch-text alignment</title>
      <link>https://vilab.team/publication/enhancing-zero-shot-brain-tumor-subtype-classification-via-f/</link>
      <pubDate>Mon, 27 Oct 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/enhancing-zero-shot-brain-tumor-subtype-classification-via-f/</guid>
      <description>&lt;p&gt;本文提出细粒度补丁对齐网络（FG-PAN），用于脑肿瘤亚型的零样本分类。该方法包含局部特征细化模块，通过建模代表性补丁间的空间关系增强视觉特征；以及细粒度文本描述生成模块，利用大语言模型生成病理感知的类别语义原型。通过对齐细粒度视觉与语义特征，并引入坐标感知聚合机制，FG-PAN在整张病理切片级别实现了更准确的亚型判别，缓解了标注数据稀缺和形态差异细微带来的挑战。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>MeDKCoOp: Dual Knowledge-guided Graph Prompt Learning for Biomedical Vision-Language Models</title>
      <link>https://vilab.team/publication/medkcoop-dual-knowledge-guided-graph-prompt-learning-for-bio/</link>
      <pubDate>Mon, 27 Oct 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/medkcoop-dual-knowledge-guided-graph-prompt-learning-for-bio/</guid>
      <description>&lt;p&gt;本文提出MeDKCoOp，一种面向生物医学视觉语言模型的双知识引导图提示学习方法。该方法系统整合医学领域知识，从文本与视觉分支提取专门知识并构建图结构表示，通过知识引导的关系转移实现跨模态融合，并动态优化可学习提示，以增强CLIP等模型在医学下游任务中的适应能力。实验表明其在多个生物医学基准上取得优异性能。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Enhancing Visual Question Answering Via Clustered In-Context Sequence Configuration</title>
      <link>https://vilab.team/publication/enhancing-visual-question-answering-via-clustered-in-context/</link>
      <pubDate>Sun, 14 Sep 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/enhancing-visual-question-answering-via-clustered-in-context/</guid>
      <description>&lt;p&gt;本文针对多模态大语言模型在多模态上下文学习中的演示序列配置问题，提出一种基于聚类的上下文配置方法。该方法自适应地对候选数据进行分组，并从每个簇中选取演示样本，以增强序列内多样性并保持语义一致性，从而减少高相似演示带来的归纳偏置，使模型更关注演示的主要意图。在OK-VQA、VQAv2、VizWiz和TextVQA四个视觉问答基准上的实验验证了其有效性。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>MMSupcon: An image fusion-based multi-modal supervised contrastive method for brain tumor diagnosis</title>
      <link>https://vilab.team/publication/mmsupcon-an-image-fusion-based-multi-modal-supervised-contra/</link>
      <pubDate>Thu, 28 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/mmsupcon-an-image-fusion-based-multi-modal-supervised-contra/</guid>
      <description>&lt;p&gt;本文针对脑肿瘤多模态MRI诊断中融合策略受限于样本稀缺的问题，提出多模态监督对比学习方法MMSupcon。该方法通过多模态医学图像融合生成信息丰富的样本，并设计多模态监督对比损失，引导模型有效整合互补模态信息，提升诊断准确性。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Semamil: Semantic reordering with retrieval-guided state space modeling for whole slide image classification</title>
      <link>https://vilab.team/publication/semamil-semantic-reordering-with-retrieval-guided-state-spac/</link>
      <pubDate>Fri, 01 Aug 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/semamil-semantic-reordering-with-retrieval-guided-state-spac/</guid>
      <description>&lt;p&gt;本文针对全切片图像分类中多实例学习忽略上下文、Transformer计算复杂、状态空间模型打乱语义顺序的问题，提出SemaMIL方法。该方法包含语义重排模块，通过可逆置换将语义相似的图像块聚类排列；以及语义引导检索状态空间模块，选择代表性查询子集调整状态空间参数，实现高效全局建模。在四个WSI数据集上验证了方法的有效性。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Create anything anywhere: Layout-controllable personalized diffusion model for multiple subjects</title>
      <link>https://vilab.team/publication/create-anything-anywhere-layout-controllable-personalized-di/</link>
      <pubDate>Mon, 30 Jun 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/create-anything-anywhere-layout-controllable-personalized-di/</guid>
      <description>&lt;p&gt;Diffusion models have significantly advanced text-to-image generation, laying the foundation for the development of personalized generative frameworks. However, existing methods lack precise layout controllability and overlook the potential of dynamic features of reference subjects in improving fidelity. In this work, we propose Layout-Controllable Personalized Diffusion (LCP-Diffusion) model, a novel framework that integrates subject identity preservation with flexible layout guidance in a tuning-free approach. Our model employs a Dynamic-Static Complementary Visual Refining module to comprehensively capture the intricate details of reference subjects, and introduces a Dual Layout Control mechanism to enforce robust spatial control across both training and inference stages. Extensive experiments validate that LCP-Diffusion excels in both identity preservation and layout controllability. To the best of our …&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Hierarchical Task-aware Temporal Modeling and Matching for few-shot action recognition</title>
      <link>https://vilab.team/publication/hierarchical-task-aware-temporal-modeling-and-matching-for-f/</link>
      <pubDate>Tue, 01 Apr 2025 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/hierarchical-task-aware-temporal-modeling-and-matching-for-f/</guid>
      <description>&lt;p&gt;本文针对少样本动作识别中训练样本稀缺且视频结构复杂的问题，提出分层任务感知时间建模与匹配方法（HTTMM）。该方法通过分层结构充分建模时空特征，并利用任务感知机制增强对关键运动模式的感知，从而提升查询样本与支持样本之间的匹配效果。在多个基准数据集上的实验验证了其有效性，尤其适用于需要局部运动感知的细粒度动作分类任务。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Visual perception by large language model’s weights</title>
      <link>https://vilab.team/publication/visual-perception-by-large-language-models-weights/</link>
      <pubDate>Mon, 16 Dec 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/visual-perception-by-large-language-models-weights/</guid>
      <description>&lt;p&gt;Existing Multimodal Large Language Models (MLLMs) follow the paradigm that perceives visual information by aligning visual features with the input space of Large Language Models (LLMs) and concatenating visual tokens with text tokens to form a unified sequence input for LLMs. These methods demonstrate promising results on various vision-language tasks but are limited by the high computational effort due to the extended input sequence resulting from the involvement of visual tokens. In this paper, instead of input space alignment, we propose a novel parameter space alignment paradigm that represents visual information as model weights. For each input image, we use a vision encoder to extract visual features, convert features into perceptual weights, and merge the perceptual weights with LLM&amp;rsquo;s weights. In this way, the input of LLM does not require visual tokens, which reduces the length of the input sequence and greatly improves efficiency. Following this paradigm, we propose VLoRA with the perceptual weights generator. The perceptual weights generator is designed to convert visual features to perceptual weights with low-rank property, exhibiting a form similar to LoRA. The experimental results show that our VLoRA achieves comparable performance on various benchmarks for MLLMs, while significantly reducing the computational costs for both training and inference. Code and models are released at\url {https://github. com/FeipengMa6/VLoRA}.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Semantic-enhanced point-box joint prompting for video object segmentation</title>
      <link>https://vilab.team/publication/semantic-enhanced-point-box-joint-prompting-for-video-object/</link>
      <pubDate>Sun, 27 Oct 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/semantic-enhanced-point-box-joint-prompting-for-video-object/</guid>
      <description>&lt;p&gt;本文提出基于SAM的语义增强点框联合提示框架SAM-SPB，用于视频对象分割。该框架通过点跟踪分支维持对象局部结构信息，并利用语义感知的基于记忆的框跟踪分支跨帧传播对象语义一致性，从而结合局部与全局线索实现鲁棒分割。在主流VOS基准上取得了领先性能，验证了点框联合提示相比仅用点提示的优势。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Ee-mllm: A data-efficient and compute-efficient multimodal large language model</title>
      <link>https://vilab.team/publication/ee-mllm-a-data-efficient-and-compute-efficient-multimodal-la/</link>
      <pubDate>Wed, 21 Aug 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/ee-mllm-a-data-efficient-and-compute-efficient-multimodal-la/</guid>
      <description>&lt;p&gt;Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated satisfactory performance across various vision-language tasks. Current approaches for vision and language interaction fall into two categories: self-attention-based and cross-attention-based methods. However, both approaches present inherent limitations, forcing a trade-off between data and computational efficiency. To address this issue, we introduce the Data-$\textbf{E}$fficient and Compute-$\textbf{E}$fficient $\textbf{MLLM}$ ($\textbf{EE-MLLM}$). Specifically, we modify the original self-attention mechanism in MLLM to a composite attention mechanism. This mechanism has two key characteristics: 1) eliminating the computational overhead of self-attention among visual tokens to achieve $\textbf{compute efficiency}$, and 2) reusing the weights from each layer of LLM to facilitate effective vision-language modality alignment for $\textbf{data efficiency}$. As a result, EE-MLLM significantly outperforms Flamingo with limited training data, and reduces the prefilling time to 79 ms on an H800 GPU, compared to LLaVA&amp;rsquo;s 277 ms. To further investigate the efficiency of EE-MLLM, we present a training-free variant named EE-MLLM-F, which reduces the computation cost of self-attention-based method without additional training. Experimental results demonstrate the effectiveness of EE-MLLM across a range of benchmarks, including general-purpose datasets like MMBench and SeedBench, as well as fine-grained tasks such as TextVQA and DocVQA.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Task navigator: Decomposing complex tasks for multimodal large language models</title>
      <link>https://vilab.team/publication/task-navigator-decomposing-complex-tasks-for-multimodal-larg/</link>
      <pubDate>Mon, 17 Jun 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/task-navigator-decomposing-complex-tasks-for-multimodal-larg/</guid>
      <description>&lt;p&gt;本文提出一种名为 Task Navigator 的框架，利用大语言模型作为导航器，将复杂多模态任务逐步分解为更易处理的子问题，并引导多模态大语言模型按步骤求解。该方法无需重新训练模型，而是系统化地调用 MLLM 已有的多种能力，如 OCR、识别、推理等，从而提升复杂任务的处理效果。作者还构建了包含数学推理、嵌入式文本问答和视觉规划等任务的基准，验证了框架的有效性。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Multi-modal generative embedding model</title>
      <link>https://vilab.team/publication/multi-modal-generative-embedding-model/</link>
      <pubDate>Wed, 29 May 2024 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/multi-modal-generative-embedding-model/</guid>
      <description>&lt;p&gt;本文提出多模态生成嵌入模型MM-GEM，将生成与嵌入两种目标统一于单个大语言模型中，实现每个模态仅需一个模型。通过引入PoolAggregator提升效率并支持细粒度嵌入与生成。实验表明，生成与嵌入目标并不显著冲突，模型在跨模态检索、零样本分类和图像描述等任务上表现优异，同时具备区域级描述生成与检索能力，并在长文本图像检索中取得显著提升。&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>Uncertainty-aware label rectification for domain adaptive mitochondria segmentation</title>
      <link>https://vilab.team/publication/uncertainty-aware-label-rectification-for-domain-adaptive-mi/</link>
      <pubDate>Fri, 01 Jan 2021 00:00:00 +0000</pubDate>
      <guid>https://vilab.team/publication/uncertainty-aware-label-rectification-for-domain-adaptive-mi/</guid>
      <description>&lt;p&gt;本文提出一种不确定性感知的标签修正方法，用于域自适应线粒体分割。针对跨域场景下伪标签噪声导致的监督偏差问题，通过不确定性估计识别并修正不可靠标签，从而提升模型在目标域上的分割性能。在电子显微镜线粒体分割任务上验证了该方法的有效性。&lt;/p&gt;
</description>
    </item>
    
  </channel>
</rss>
