Photon用可变长度标记压缩3D医学影像,提速降耗还更准
Photon: Speedup Volume Understanding with Efficient Multimodal Large Language Models
- 用可变长度标记序列表示3D医学体积,保持空间连续性
- 训练推理时自适应删减冗余标记,计算量下降40%以上
- 适合需要高效处理3D医学影像的临床AI研发人员
多模态大模型在临床视觉问答中前景广阔,但扩展至3D成像时受限于高计算成本。现有方法常依赖2D切片或固定长度标记压缩,破坏体积连续性并掩盖细微病变。本文提出Photon框架,以可变长度标记序列表示3D医学体积。Photon引入指令引导的标记调度与代理梯度传播,实现训练与推理阶段的自适应标记缩减,降低计算开销,缓解冗余标记导致的注意力稀释问题。其采用定制反向传播规则与梯度恢复机制,支持离散标记删除下的可微优化。为进一步稳定标记压缩并确保视觉证据可靠,Photon还引入正则化目标,抑制纯语言偏差。在多种医学视觉问答任务上的实验表明,Photon在保持顶尖准确率的同时,显著降低资源消耗并加速训练与推理。
原文摘要 · Abstract (English)
Multimodal large language models are promising for clinical visual question answering tasks, but scaling to 3D imaging is hindered by high computational costs. Prior methods often rely on 2D slices or fixed-length token compression, disrupting volumetric continuity and obscuring subtle findings. We present Photon, a framework that represents 3D medical volumes with token sequences of variable length. Photon introduces instruction-conditioned token scheduling and surrogate gradient propagation to adaptively reduce tokens during both training and inference, which lowers computational cost while mitigating the attention dilution caused by redundant tokens. It incorporates a custom backpropagation rule with gradient restoration to enable differentiable optimization despite discrete token drop. To stabilize token compression and ensure reliable use of visual evidence, Photon further applies regularization objectives that mitigate language-only bias and improve reliability. Experiments on diverse medical visual question answering tasks show that Photon achieves state-of-the-art accuracy while reducing resource usage and accelerating both training and inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。