现有方法降低幻觉但更保守,反而损害模型能力
Does Playing it Safe Count as Faithfulness? Reassessing LVLM Hallucination Mitigation Methods

- 用多模型多基准测试发现,降幻觉常伴随信息量减少
- 幻觉指标提升不等于实际能力增强,细粒度任务表现下降
- 应综合评估真实性、信息量和综合能力,而非只看幻觉分数
近期针对大视觉语言模型(LVLMs)的推理时幻觉缓解方法在幻觉评测上报告了显著提升。然而,低幻觉分数是否反映更好的多模态对齐,还是仅体现更保守生成仍不明确。我们评估了六种缓解方法在三种LVLMs和四个基准上的表现,包括专注幻觉的评估与多能力基准MMStar。分析揭示两个一致模式:第一,幻觉降低常伴随信息量减少——方法虽降低幻觉率,但也导致物体召回率、视觉覆盖范围或回答详尽度下降;第二,幻觉基准上的改进无法可靠迁移至更广泛的多模态能力,方法在细粒度感知与推理任务中表现不一致或退化。结果表明,当前评估协议可能因奖励保守生成而夸大进展。我们主张,幻觉缓解应作为真实性-信息量-能力的权衡来评估,而非仅依赖幻觉得分。
原文摘要 · Abstract (English)
Recent inference-time hallucination mitigation methods for large vision-language models (LVLMs) report strong gains on hallucination benchmarks. However, it remains unclear whether lower hallucination scores reflect improved multimodal grounding or more conservative generation. We evaluate six mitigation methods across three LVLMs and four benchmarks, including hallucination-focused evaluation and the diverse capability benchmark MMStar. Our analysis reveals two consistent patterns. First, hallucination reduction is often coupled with reduced informativeness: methods that lower hallucination rates also reduce object recall, visual coverage, or response detailedness. Second, improvements on hallucination benchmarks do not reliably transfer to broader multimodal capabilities, with methods showing inconsistent or degraded performance on fine-grained perception and reasoning tasks. Our findings suggest that current evaluation protocols may overestimate progress by rewarding conservative generation. We argue that hallucination mitigation should be evaluated as a faithfulness--informativeness--capability trade-off rather than through hallucination scores alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。