arXiv:2603.09989cs.CLcs.AI2026-03

用用户视角评估大模型幻觉行为,轻量易用。

The System Hallucination Scale (SHS): A Minimal yet Effective Human-Centered Instrument for Evaluating Hallucination-Related Behavior in Large Language Models

  • 基于用户交互设计,衡量事实错误、逻辑混乱等表现
  • 210人测试显示信度高(Cronbach's alpha=0.87)
  • 适合迭代开发与部署监控,非自动检测工具

我们提出系统幻觉量表(SHS),一种轻量且以用户为中心的评估工具,用于衡量大型语言模型(LLMs)中的幻觉相关行为。受系统可用性量表(SUS)和系统因果性量表(SCS)等成熟心理测量工具启发,SHS 可在真实交互场景下,快速、可解释地评估事实不可靠性、逻辑不连贯、误导性呈现及对用户引导的响应能力。SHS 并非自动幻觉检测器或基准指标,而是从用户视角捕捉幻觉表现。210名参与者的真实世界评估表明,该量表具有高清晰度、一致的响应行为以及结构效度,统计分析支持内部一致性(Cronbach's alpha = 0.87)和维度间显著相关性(p < 0.001)。与 SUS、SCS 的对比分析揭示了互补的测量特性,证实 SHS 在比较分析、迭代系统开发和部署监控中的实用性。

原文摘要 · Abstract (English)

We introduce the System Hallucination Scale (SHS), a lightweight and human-centered measurement instrument for assessing hallucination-related behavior in large language models (LLMs). Inspired by established psychometric tools such as the System Usability Scale (SUS) and the System Causability Scale (SCS), SHS enables rapid, interpretable, and domain-agnostic evaluation of factual unreliability, incoherence, misleading presentation, and responsiveness to user guidance in model-generated text. SHS is explicitly not an automatic hallucination detector or benchmark metric; instead, it captures how hallucination phenomena manifest from a user perspective under realistic interaction conditions. A real-world evaluation with 210 participants demonstrates high clarity, coherent response behavior, and construct validity, supported by statistical analysis including internal consistency (Cronbach's alpha = 0.87$) and significant inter-dimension correlations (p < 0.001$). Comparative analysis with SUS and SCS reveals complementary measurement properties, supporting SHS as a practical tool for comparative analysis, iterative system development, and deployment monitoring.

幻觉评估用户中心心理测量LLM评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。