简单直接的答题微调,比复杂模型更鲁棒。
Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA
- 直接对答案进行监督微调,紧扣任务目标
- 在固定数据集上准确率显著优于冻结基线
- 适合追求稳定性能而非复杂架构的研究者
多帧医学视觉问答任务中,现有方法趋向于采用复杂适应机制:控制器推理、定位感知重排序、静态难例混合及分阶段延续等。我们测试了一个更简单的假设:在控制评估条件(固定划分、匹配预算、重复种子、校准)下,与评测最终答案目标高度对齐的方法应是最具鲁棒性的适配家族。对比了基于控制器的方法、支架演化、静态混合监督、延续主导变体以及仅答案监督的微调(SFT)。结果显示,基于MedGemma-1.5-4B的直接解码器答案SFT表现最佳。实证表明,该方法在保留样本报告准确率上显著超越冻结基线,且在重复种子和对照实验中保持高度稳定,确保结论反映家族级鲁棒性而非单一超参峰值。此外,后处理校准可有效修复置信度估计而不损失准确率,核心方法还能一致迁移至Qwen2.5-VL-3B等其他骨干模型。因此,关键发现并非复杂辅助机制胜出,而是目标对齐的直接答案SFT是当前最强的鲁棒适配家族。通过建立这一简约但强大的基线,我们希望引导社区关注根本性鲁棒优化,而非架构复杂度。
原文摘要 · Abstract (English)
Multi-frame medical VQA appears to reward increasingly complex adaptation: controller-style inference, localization-aware reranking, static hard-negative mixing, and staged continuation all appear plausible from first principles. We test a simpler competing hypothesis on MedFrameQA: methods that remain tightly aligned with the benchmark's final answer objective should be the strongest \emph{robust} adaptation family once evaluation is controlled across fixed splits, matched budgets, repeated seeds, and calibration. We compare controller-based methods, scaffold evolution, static mixed supervision, continuation-heavy variants, and direct answer-only supervised fine-tuning (SFT). The strongest robust family is direct decoder-only answer SFT on MedGemma-1.5-4B. Empirically, this family yields substantial improvements in held-out report accuracy over frozen baselines while remaining remarkably stable across repeated seeds and matched controls, ensuring our claims reflect true family-level robustness rather than an isolated hyperparameter peak. Furthermore, post-hoc calibration effectively repairs confidence estimation without compromising accuracy, and the core approach transfers consistently to secondary backbones like Qwen2.5-VL-3B. The main result is therefore not that a complex auxiliary mechanism wins, but that objective-aligned direct answer SFT is the strongest robust adaptation family we found for MedFrameQA. By establishing this strong, minimalist baseline, we hope to redirect community focus toward fundamentally robust optimization rather than architectural complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。