研究人类如何评价AI的纠错效果,发现语言习惯和经验影响判断。
Exploring Human Perceptions of AI Responses: Insights from a Mixed-Methods Study on Risk Mitigation in Generative Models
- 通过混合方法实验,对比有无安全防护的AI回应
- 57人参与,发现语言错误被严格惩罚,语义保留受奖励
- 提出新评估指标,适合做人类感知研究
随着生成式AI的快速普及,研究人类对生成内容的感知变得至关重要。主要挑战在于模型易产生幻觉和有害内容。尽管已部署多种防护机制,但人们对这些策略的实际感知仍不明确。本研究采用混合方法实验,从忠实性、公平性、有害内容消除能力与相关性四个维度评估一种防护策略。在被试内设计中,57名参与者分别评估了含有害内容的回复及其防护后的版本,以及仅防护后的回复。结果表明,参与者的母语、AI使用经验及标注熟悉度显著影响评价。参与者对语言和上下文特征极为敏感,轻微语法错误即遭惩罚,而语义上下文的保留则获奖励。这与当前大模型量化评估中对语言的处理方式形成对比。研究还提出了用于训练与评估防护策略的新指标,并为人类-智能体评估研究提供了洞见。
原文摘要 · Abstract (English)
With the rapid uptake of generative AI, investigating human perceptions of generated responses has become crucial. A major challenge is their `aptitude' for hallucinating and generating harmful contents. Despite major efforts for implementing guardrails, human perceptions of these mitigation strategies are largely unknown. We conducted a mixed-method experiment for evaluating the responses of a mitigation strategy across multiple-dimensions: faithfulness, fairness, harm-removal capacity, and relevance. In a within-subject study design, 57 participants assessed the responses under two conditions: harmful response plus its mitigation and solely mitigated response. Results revealed that participants' native language, AI work experience, and annotation familiarity significantly influenced evaluations. Participants showed high sensitivity to linguistic and contextual attributes, penalizing minor grammar errors while rewarding preserved semantic contexts. This contrasts with how language is often treated in the quantitative evaluation of LLMs. We also introduced new metrics for training and evaluating mitigation strategies and insights for human-AI evaluation studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。