arXiv:2608.01724cs.CLcs.AI2026-08中稿 · COLM

TIDES记录12支团队一学期的中英双语对话,用于建模长期社交动态。

TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics

论文配图:TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics
图 1 · 摘自论文原文
  • 构建纵向双语数据集,追踪12支团队一学期的面对面会议
  • 在预测下一位发言者上提升13.8个百分点,接近顶尖模型表现
  • 适合研究团队演化、多角色交互与自然对话生成的学者

群体对话是人类协作的基础,但现有大语言模型仍难以应对多方互动的复杂性。主要原因在于现有数据集多为短期实验室环境下的虚构任务,无法反映真实团队的长期社会动态。为此,我们提出TIDES,一个高分辨率纵向数据集,记录了12支大学项目团队在完整学期内的交流。数据包含75,971条中英文面对面会议语句,并配有社会结构标注,涵盖互动类型、涌现角色与发展阶段,支持对团队演化的建模。实验表明,在TIDES上微调可使下一说话人预测准确率相比基线(64.53%)提升13.8个百分点,性能接近强大多数私有零样本模型。同时,在AMI会议语料库上表现接近现有最佳成果,仅需约42%的训练数据。然而,人工评估显示,尽管预测更准,微调模型生成的话语却不如原始模型自然流畅,提示需进一步探索结构建模如何促进自然对话生成。

原文摘要 · Abstract (English)

Group conversations are fundamental to human collaboration, yet standard large language models (LLMs) still struggle with the complexities of multi-party interaction. This challenge persists in part because existing group conversation datasets are often limited to short-term lab settings with contrived tasks, failing to capture the long-term social dynamics of real-world teams. To bridge this gap, we introduce TIDES, a high-resolution longitudinal dataset tracking 12 university project teams over a full semester. Comprising 75,971 utterances in both English and Korean from in-person meetings, TIDES provides a naturalistic record of teams working on self-managed projects. Our socio-structural annotations-covering interaction types, emergent roles, and development stages-allow for modeling of team evolution over months. Experiments show that fine-tuning on TIDES improves next-speaker prediction by 13.8 percentage points over a bigram baseline (64.53%) and yields performance comparable to strong proprietary zero-shot models. The model also comes within 2.1 percentage points of the published state of the art on the AMI Meeting Corpus while using approximately 42% less training data. However, human evaluations suggest that better next-speaker prediction does not necessarily yield more natural or coherent utterances, as fine-tuned models were generally less preferred than vanilla models. This potential mismatch motivates further study of how structural modeling can support natural multi-party generation.

多说话人社会动态纵向数据双语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。