用智能代理生成逼真在线讨论数据,解决真实数据难获取问题。
CHORUS: An Agentic Framework for Generating Realistic Deliberation Data

- 用带记忆的智能体模拟有性格的发言者,按真实用户节奏参与对话。
- 30位专家评估显示内容真实、逻辑连贯、分析价值高。
- 支持调用外部工具,可对接网页平台,适合研究网络对话的学者。
理解在线话语的复杂动态依赖大规模的协商数据,但受访问限制、伦理问题和质量不一影响,交互式网络平台上的此类数据仍十分稀缺。本文提出Chorus,一种智能体框架,通过具备行为一致人格的LLM驱动代理生成真实感强的协商讨论。每个代理由自主智能体管理,拥有对讨论演进的记忆;参与时间由基于泊松过程的时序模型控制,近似真实用户的异构参与模式。框架还支持结构化工具调用,使代理可访问外部资源,并便于与互动网页平台集成。该框架部署于 extsc{Deliberate}平台,经30位专家在内容真实性、讨论连贯性和分析实用性三个维度评估,证实Chorus是生成高质量协商数据的实用工具,适用于在线话语分析。
原文摘要 · Abstract (English)
Understanding the intricate dynamics of online discourse depends on large-scale deliberation data, a resource that remains scarce across interactive web platforms due to restrictive accessibility policies, ethical concerns and inconsistent data quality. In this paper, we propose Chorus, an agentic framework, which orchestrates LLM-powered actors with behaviorally consistent personas to generate realistic deliberation discussions. Each actor is governed by an autonomous agent equipped with memory of the evolving discussion, while participation timing is governed by a principled Poisson process-based temporal model, which approximates the heterogeneous engagement patterns of real users. The framework is further supported by structured tool usage, enabling actors to access external resources and facilitating integration with interactive web platforms. The framework was deployed on the \textsc{Deliberate} platform and evaluated by 30 expert participants across three dimensions: content realism, discussion coherence and analytical utility, confirming Chorus as a practical tool for generating high-quality deliberation data suitable for online discourse analysis
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。