arXiv:2502.14975cs.CLcs.AI2025-02被引 2

评测大模型情感边界处理能力,发现英文响应拒绝率远高于非英文。

Beyond No: Quantifying AI Over-Refusal and Emotional Attachment Boundaries

  • 构建多语言评测框架,量化模型七类情感回应模式。
  • 克劳德-3.5表现最优(8.69/10),英文拒绝率高达43.20%。
  • 适合关注AI情感智能与跨文化交互的研究者。

我们提出一个开源基准和评估框架,用于衡量大语言模型(LLMs)在情感边界处理方面的能力。基于涵盖六种语言的1156个提示数据集,对GPT-4o、Claude-3.5 Sonnet和Mistral-large三款领先模型进行了评估,通过模式匹配分析其响应中的情感边界处理方式。框架量化了七类关键响应模式:直接拒绝、道歉、解释、转移、承认、边界设定和情绪觉察。结果表明,各模型在边界处理上存在显著差异,其中Claude-3.5得分最高(8.69/10),平均响应长度达86.51词。英语交互平均得分为25.62,非英语交互则低于0.22,且英语拒绝率高达43.20%,而非英语不足1%。模式分析显示,Mistral更倾向使用转移策略(4.2%),而所有模型的情感共鸣评分均低于0.06。局限性包括模式匹配可能过度简化、缺乏上下文理解及对复杂情绪反应的二分类问题。未来工作应探索更精细的评分方法,扩展语言覆盖范围,并研究文化差异对情感边界期望的影响。本研究为系统评估大模型情感智能与边界设定能力提供了基础。

原文摘要 · Abstract (English)

We present an open-source benchmark and evaluation framework for assessing emotional boundary handling in Large Language Models (LLMs). Using a dataset of 1156 prompts across six languages, we evaluated three leading LLMs (GPT-4o, Claude-3.5 Sonnet, and Mistral-large) on their ability to maintain appropriate emotional boundaries through pattern-matched response analysis. Our framework quantifies responses across seven key patterns: direct refusal, apology, explanation, deflection, acknowledgment, boundary setting, and emotional awareness. Results demonstrate significant variation in boundary-handling approaches, with Claude-3.5 achieving the highest overall score (8.69/10) and producing longer, more nuanced responses (86.51 words on average). We identified a substantial performance gap between English (average score 25.62) and non-English interactions (< 0.22), with English responses showing markedly higher refusal rates (43.20% vs. < 1% for non-English). Pattern analysis revealed model-specific strategies, such as Mistral's preference for deflection (4.2%) and consistently low empathy scores across all models (< 0.06). Limitations include potential oversimplification through pattern matching, lack of contextual understanding in response analysis, and binary classification of complex emotional responses. Future work should explore more nuanced scoring methods, expand language coverage, and investigate cultural variations in emotional boundary expectations. Our benchmark and methodology provide a foundation for systematic evaluation of LLM emotional intelligence and boundary-setting capabilities.

情感智能大模型评测跨语言边界处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。