arXiv:2503.16505cs.HCcs.CL2025-03被引 1

用小模型生成虚拟讨论,省钱又高效,还能提前发现大模型干预失效问题。

Designing Synthetic Discussion Generation Systems: A Case Study for Online Facilitation

  • 用7B-8B量化小模型模拟讨论,成本低于大厂模型44倍以上。
  • 实验证明小模型能有效模拟真实讨论,且可揭示大模型干预时机判断缺陷。
  • 适合做线上引导实验的学者,尤其关注低成本、可复现的社科研究者。

社会科学研究中,涉及人类参与的实验成本高昂。本文提出合成讨论生成(SDG)这一新型NLP方向,旨在构建可替代真实讨论的模拟系统,实现低成本预实验。我们指出,尽管当前研究普遍使用OpenAI GPT系列等专有模型,但其在成本与能力上往往并不合理。实验表明,7B-8B量化的较小模型在生成有效模拟讨论方面表现优异,成本低于专有模型44倍以上。以在线引导为应用场景,我们将该框架用于下游任务,发现合成模拟能提前揭示潜在问题,如大模型难以判断干预时机,导致频繁介入并引发类似人类互动中的脱轨现象。此外,不同引导策略对对话动态有一定影响。除理论框架外,本文还提供成本对比方法、可用模型与算法调研、开源Python工具包,以及涵盖多个模型的大规模公开讨论数据集。

原文摘要 · Abstract (English)

A critical challenge in social science research is the high cost associated with experiments involving human participants. We identify Synthetic Discussion Generation (SDG), a novel Natural Language Processing (NLP) direction aimed at creating simulated discussions that enable cost-effective pilot experiments and develop a theoretical, task-agnostic framework for designing, evaluating, and implementing these simulations. We argue that the use of proprietary models such as the OpenAI GPT family for such experiments is often unjustified in terms of both cost and capability, despite its prevalence in current research. Our experiments demonstrate that smaller quantized models (7B-8B) can produce effective simulations at a cost more than 44 times lower compared to their proprietary counterparts. We use our framework in the context of online facilitation, where humans actively engage in discussions to improve them, unlike more conventional content moderation. By treating this problem as a downstream task for our framework, we show that synthetic simulations can yield generalizable results at least by revealing limitations before engaging human discussants. In LLM facilitators, a critical limitation is that they are unable to determine when to intervene in a discussion, leading to undesirable frequent interventions and, consequently, derailment patterns similar to those observed in human interactions. Additionally, we find that different facilitation strategies influence conversational dynamics to some extent. Beyond our theoretical SDG framework, we also present a cost-comparison methodology for experimental design, an exploration of available models and algorithms, an open-source Python framework, and a large, publicly available dataset of LLM-generated discussions across multiple models.

合成讨论低成本实验大模型评估在线引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。