让大模型的红外图像回答既准确又有热力依据。
Right Answer, Wrong Heat: Explanation-Aware Evaluation and Thermal-Grounded Feedback for MLLMs on Infrared Images

- 用双大模型共识判断解释是否基于红外热力证据
- 模型答对但解释可能依赖可见光,关掉红外图影响不大
- 无需训练的反馈机制可改进解释可信度
通用多模态大语言模型(MLLMs)在红外图像上的应用日益广泛,但通常仅以答案正确率评分。然而,正确答案并不保证其解释基于红外热力证据。我们提出一种解释感知评估框架,区分答案正确性、输出层面解释的合理性与热力证据的关联性。通过双大模型共识裁判并结合初步人工锚定校验,发现正确答案仍可能依赖弱热力或可见光信息;若仅提供类可见光渲染图而隐藏原始红外图,热力关联性显著下降但准确率变化甚微;这一现象在更强大的模型中尤为明显,但当红外图可用时则消失。我们进一步提出热力接地反馈(TGF),一种无需训练的反馈循环,能诊断解释失败并修正解释,同时保留原答案。在配对输入验证中,TGF提升解释侧的热力关联性而不改变答案。结果表明,未来可靠的红外场景理解模型应重视热力接地解释,而非仅追求答案准确。
原文摘要 · Abstract (English)
General-purpose multimodal large language models (MLLMs) are increasingly applied to infrared images, where they are commonly scored by answer accuracy alone. However, a correct answer does not ensure that the model's explanation is grounded in infrared thermal evidence. We introduce an explanation-aware evaluation framework that separates answer correctness, output-level explanation groundedness, and thermal grounding for infrared visual questions. Using a Dual-LLM Consensus Judge with a preliminary human-anchor calibration check, we find that correct answers can still rely on weak or visible-light evidence; withholding the original infrared image and showing only a visible-like rendering erodes thermal grounding with little accuracy change; and this erosion is observed most strongly for more capable models but disappears when infrared remains available. We further propose Thermal-Grounded Feedback (TGF), a training-free feedback loop that diagnoses explanation-side failures and revises the explanation while preserving the selected answer. On local paired-input validation, TGF improves explanation-side grounding without changing answers. These findings suggest that future trustworthy MLLMs for infrared scene understanding should be evaluated and developed to produce thermally grounded explanations rather than merely accurate answers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。