arXiv:2607.27564cs.CV2026-07

简单投票规则比复杂搜索更有效,提升医学图像问答准确率

Inference-Time Agentic Decision Rules Beat Longer Evolving Search for Multi-Image Medical Reasoning

论文配图:Inference-Time Agentic Decision Rules Beat Longer Evolving Search for Multi-Image Medical Reasoning
图 1 · 摘自论文原文
  • 用投票机制整合多图证据,抗顺序干扰
  • 57.89%准确率显著高于基线和复杂策略
  • 适合追求稳定性能的医疗AI研发者

多图像医学视觉问答不仅是提示长度问题,更是代理决策的根本挑战。医学视觉语言代理需在有序图像间聚合证据,对答案顺序扰动保持鲁棒性,并避免过度依赖搜索时反馈。我们通过控制对比五种推理时代理策略,在相同高预算ShinkaEvolve配置下优化,并在可复现的内部冻结数据集(1,331次进化,665个保留集,855个最终测试集)上评估。五次独立重复实验中,最有效的策略是结构最简单的鲁棒聚合器: extbf{顺序投票}策略取得57.89±0.65%的最终测试准确率,显著优于固定基线(52.73±0.42%)和更复杂但脆弱的顺序重排变体(55.79±0.43%)。配对自举分析证实了这些显著提升。将进化搜索预算从50代增至100代并未带来泛化收益:保留集性能略有提升,但最终测试准确率从57.89%降至56.02%。研究结果表明,对于多图像医学推理,定义正确的代理决策规则远比扩大优化搜索预算更重要。

原文摘要 · Abstract (English)

Multi-image medical VQA is not merely a prompt-length problem; it is a fundamental challenge of agentic decision-making. Medical vision-language agents must aggregate evidence across ordered images, remain robust to answer-order perturbations, and avoid overfitting to noisy search-time feedback. We study MedFrameQA through a controlled comparison of five inference-time agentic strategies, optimized using the same high-budget ShinkaEvolve configuration and evaluated on a reproducible internal frozen split (1,331 evolution, 665 holdout, 855 final test). Across five independent repeated runs, the strongest method emerges as the simplest robust aggregator: the \textbf{order-vote} policy achieves $57.89 \pm 0.65\%$ final-test accuracy, significantly outperforming the fixed baseline ($52.73 \pm 0.42\%$) and the more complex, albeit brittle, order-rerank variant ($55.79 \pm 0.43\%$). Paired bootstrap analysis confirms these significant gains. Extending the evolutionary search budget from 50 to 100 generations yields no generalization benefit: while holdout performance marginally increases, final-test accuracy drops from $57.89\%$ to $56.02\%$. Our findings suggest that for multi-image medical reasoning, defining the correct agentic decision rule is substantially more impactful than expanding the optimization search budget.

医学视觉问答代理决策多图像推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。