arXiv:2601.06460cs.CVcs.AI2026-01被引 3

不同语气提示会显著影响视觉模型幻觉,强压提示反而降低幻觉率。

Tone Matters: The Impact of Linguistic Tone on Hallucination in VLMs

  • 用五级语气强度框架测试提示压力对幻觉的影响。
  • 高压力提示下幻觉率下降,但模型响应存在差异。
  • 适合关注模型安全对齐与提示工程的研究者。

视觉语言模型(VLMs)在需要可靠视觉定位的安全关键应用中日益普及,但常因满足用户提示而生成图像中不存在的细节,产生幻觉。尽管已有数据集和基准用于评估系统性幻觉,许多幻觉行为仍缺乏充分描述。以往研究多聚焦物体存在与否,未揭示提示措辞与结构约束如何系统诱导幻觉。本文探究不同提示压力形式对幻觉行为的影响,提出基于程序生成的合成场景数据集Ghost-100,其中关键视觉细节被刻意移除,以实现对缺失型幻觉的可控分析。采用五级提示强度框架,将提示从中性查询逐步升级为有毒要求及严格格式约束。评估MiniCPM-V 2.6-8B、Qwen2-VL-7B和Qwen3-VL-8B三款开源模型。结果显示,所有模型的幻觉率并未随提示强度单调上升;在不同阈值下,各模型均出现幻觉率下降,但并非所有模型在最高强制压力下维持稳定降低。这表明当前安全对齐更擅长识别语义敌意,而在应对结构胁迫时表现不一,揭示了模型在合规压力下的特定局限性。数据集已公开:https://github.com/bli1/tone-matters

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) are increasingly used in safety-critical applications that require reliable visual grounding. However, these models often hallucinate details that are not present in the image to satisfy user prompts. While recent datasets and benchmarks have been introduced to evaluate systematic hallucinations in VLMs, many hallucination behaviors remain insufficiently characterized. In particular, prior work primarily focuses on object presence or absence, leaving it unclear how prompt phrasing and structural constraints can systematically induce hallucinations. In this paper, we investigate how different forms of prompt pressure influence hallucination behavior. We introduce Ghost-100, a procedurally generated dataset of synthetic scenes in which key visual details are deliberately removed, enabling controlled analysis of absence-based hallucinations. Using a structured 5-Level Prompt Intensity Framework, we vary prompts from neutral queries to toxic demands and rigid formatting constraints. We evaluate three representative open-weight VLMs: MiniCPM-V 2.6-8B, Qwen2-VL-7B, and Qwen3-VL-8B. Across all three models, hallucination rates do not increase monotonically with prompt intensity. All models exhibit reductions at higher intensity levels at different thresholds, though not all show sustained reduction under maximum coercion. These results suggest that current safety alignment is more effective at detecting semantic hostility than structural coercion, revealing model-specific limitations in handling compliance pressure. Our dataset is available at: https://github.com/bli1/tone-matters

视觉模型幻觉检测提示工程安全对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。