测试AI安全解释的语义对齐,发现看似有依据实则误导风险判断。
Grounded but Misleading: Evaluating Semantic Alignment in AI-Generated Security Explanations
- 构建可控测试框架VEXA,独立控制证据引用与语义表达
- 人类评估显示误导性解释仍获高可信度评分(3.66分)
- 警示:仅看是否引用证据不足以判断解释可靠性
在线诈骗日益采用流畅且情境感知的社会工程策略,推动对AI生成风险解释的需求。然而,基于检测器证据的解释可能在语义上削弱或偏离原始风险判断。我们提出VEXA:Verifying Semantic Explanation Alignment,一个用于研究AI生成欺诈风险解释中词汇锚定与语义风险对齐差距的受控测试平台。VEXA通过独立控制证据锚定和语义框架,生成无锚定、风险对齐及风险稀释型解释。通过大语言模型为裁判和人工评估,我们发现即使语义解释弱化了检测器的原意,解释仍可能显得相对有据。在人工评估中,风险稀释型解释虽帮助性(3.00)与推理支持(3.14)较低,但感知证据锚定得分仍较高(3.66)。这些结果提供了关于AI生成安全解释中锚定幻觉效应的受控证据,表明可信赖的解释评估必须不仅验证证据引用,还需审查其解释方式。
原文摘要 · Abstract (English)
Online scams increasingly leverage fluent and context-aware social engineering strategies, creating growing demand for AI systems that explain why a message may be risky. However, explanations that cite detector-derived evidence may still semantically weaken or redirect the intended risk interpretation. We introduce VEXA: Verifying Semantic Explanation Alignment, a controlled testbed for studying the gap between lexical grounding and semantic risk alignment in AI-generated scam-risk explanations. VEXA generates ungrounded, risk-aligned, and risk-diluting explanations by independently controlling evidence grounding and semantic framing. Through LLM-as-a-judge and human evaluations, we show that explanations may continue to appear comparatively grounded even when their semantic interpretation weakens the detector's intended risk assessment. In human evaluation, risk-diluting XAI-grounded explanations retained comparatively elevated Perceived Evidence Grounding scores (3.66) despite lower Helpfulness (3.00) and Reasoning Support (3.14) scores. These findings provide controlled evidence of grounding illusion effects in AI-generated security explanations and suggest that trustworthy explanation evaluation must verify not only whether evidence is cited, but also how that evidence is interpreted.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。