通过双分支强化学习,提升病理图像推理能力并降低计算开销。
Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology Reasoning
- 双分支强化学习:一个学诊断逻辑,一个动态分配计算资源。
- 平均性能提升41.7%,推理成本降低70.3%。
- 适合需要高效精准病理分析的临床研究与系统部署。
多模态病理图像理解因融合视觉与文本信息以提升诊断准确性和实现个性化治疗而备受关注。然而,现有方法推理能力有限,难以应对复杂诊断场景;同时,病理图像体量巨大,带来严重计算负担,限制实际应用。为此,我们提出一种新型双边强化学习框架,包含两个协同分支:一分支通过标签直接学习任务相关的决策过程(即病理推理依据),无需显式推理监督;另一分支根据图像视觉内容和任务上下文动态分配适配的令牌数量,优化计算效率。该方法应用于视觉问答、癌症亚型分类和病灶检测等多种病理任务。大量实验表明,相比基线模型,平均性能提升41.7个百分点,推理成本降低70.3%,在推理准确率与计算效率上均取得显著进步。
原文摘要 · Abstract (English)
Multimodal pathological image understanding has garnered widespread interest due to its potential to improve diagnostic accuracy and enable personalized treatment through integrated visual and textual data. However, existing methods exhibit limited reasoning capabilities, which hamper their ability to handle complex diagnostic scenarios. Additionally, the enormous size of pathological images leads to severe computational burdens, further restricting their practical deployment. To address these limitations, we introduce a novel bilateral reinforcement learning framework comprising two synergistic branches. One reinforcement branch enhances the reasoning capability by enabling the model to learn task-specific decision processes, i.e., pathology rationales, directly from labels without explicit reasoning supervision. While the other branch dynamically allocates a tailored number of tokens to different images based on both their visual content and task context, thereby optimizing computational efficiency. We apply our method to various pathological tasks such as visual question answering, cancer subtyping, and lesion detection. Extensive experiments show an average +41.7 absolute performance improvement with 70.3% lower inference costs over the base models, achieving both reasoning accuracy and computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。