arXiv:2512.03005cs.AI2025-12KDD被引 3

用大模型化解网络争吵,不只识别有害内容,还能引导和解。

From Moderation to Mediation: Can LLMs Serve as Mediators in Online Flame Wars?

  • 将调解拆解为判断情绪与公平性、生成共情回应两步。
  • 基于Reddit数据集测试,商用API模型表现优于开源模型。
  • 适合研究社交对话、负责任AI或人机协作的学者与开发者。

大语言模型(LLMs)的快速发展为人工智能向善应用开辟了新路径。随着LLMs越来越多地介入在线交流,其促进共情与建设性对话的潜力成为负责任AI研究的重要方向。本文探讨了LLMs能否不仅作为检测有害内容的监管者,更作为理解并缓和网络冲突的调解者。我们提出一个框架,将调解分为两个子任务:判断(评估对话的公平性与情感动态)与引导(生成共情、降激化的回应以推动解决)。为评估调解质量,我们构建了一个基于Reddit的大规模数据集,并设计了多阶段评估流程,结合原则评分、用户模拟和人工对比。实验表明,在执行调解时,基于API的模型在推理与干预对齐方面均优于开源模型。研究结果揭示了当前LLMs作为在线社会调解代理的潜力与局限。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has opened new possibilities for AI for good applications. As LLMs increasingly mediate online communication, their potential to foster empathy and constructive dialogue becomes an important frontier for responsible AI research. This work explores whether LLMs can serve not only as moderators that detect harmful content, but as mediators capable of understanding and de-escalating online conflicts. Our framework decomposes mediation into two subtasks: judgment, where an LLM evaluates the fairness and emotional dynamics of a conversation, and steering, where it generates empathetic, de-escalatory messages to guide participants toward resolution. To assess mediation quality, we construct a large Reddit-based dataset and propose a multi-stage evaluation pipeline combining principle-based scoring, user simulation, and human comparison. Experiments show that API-based models outperform open-source counterparts in both reasoning and intervention alignment when doing mediation. Our findings highlight both the promise and limitations of current LLMs as emerging agents for online social mediation.

大模型社交调解共情对话

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。