测试大模型当调解员效果,表现优于人类。
Robots in the Middle: Evaluating LLMs in Dispute Resolution
- 用真实纠纷案例评估大模型调解能力
- 62%情况下干预策略优于或等同于人类
- 84%的回应质量达到或超过人类水平
调解是一种由中立第三方介入帮助双方解决争端的方法。本文研究大语言模型(LLMs)作为调解员的可行性,考察其分析纠纷对话、选择合适干预方式及生成恰当回应的能力。基于一个手工构建的50个纠纷场景数据集,我们对大模型与人类标注者进行了盲评对比。结果显示,大模型整体表现优异,甚至在多个维度上超越人类。具体而言,在62%的案例中,大模型选择的干预类型被评价为优于或等同于人类;在84%的案例中,其生成的干预语句质量不低于人类。大模型在公正性、理解力和情境化方面也表现良好。结果表明,将AI整合至在线纠纷解决(ODR)平台具有巨大潜力。
原文摘要 · Abstract (English)
Mediation is a dispute resolution method featuring a neutral third-party (mediator) who intervenes to help the individuals resolve their dispute. In this paper, we investigate to which extent large language models (LLMs) are able to act as mediators. We investigate whether LLMs are able to analyze dispute conversations, select suitable intervention types, and generate appropriate intervention messages. Using a novel, manually created dataset of 50 dispute scenarios, we conduct a blind evaluation comparing LLMs with human annotators across several key metrics. Overall, the LLMs showed strong performance, even outperforming our human annotators across dimensions. Specifically, in 62% of the cases, the LLMs chose intervention types that were rated as better than or equivalent to those chosen by humans. Moreover, in 84% of the cases, the intervention messages generated by the LLMs were rated as better than or equal to the intervention messages written by humans. LLMs likewise performed favourably on metrics such as impartiality, understanding and contextualization. Our results demonstrate the potential of integrating AI in online dispute resolution (ODR) platforms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。