arXiv:2606.07893cs.CL2026-06

让合成对话更贴近真实人群行为分布,提升数据真实性。

Beyond Individual Personas: Aligning Synthetic Dialogue to Population-Level Behavior Distributions

论文配图:Beyond Individual Personas: Aligning Synthetic Dialogue to Population-Level Behavior Distributions
图 1 · 摘自论文原文
  • 将人群行为统计转为生成控制,分離核心行为与附带影响。
  • 合成对话与真实数据的分布差异降低24.4%,达最佳效果。
  • 适合需要真实人群行为数据的对话系统训练与评估。

合成对话语料库日益被用作目标对话数据的替代品,但基于个人角色的生成器仅优化单个对话,导致整体语料库的人群行为分布失真。我们提出GroupPersona框架,使合成语料库对齐参考语料库的人群行为分布。该方法将群体统计数据转化为生成控制:分离每个对话的核心行为特征与可预测的附带效应,并利用这些行为组来引导用户代理,使其遵循定义参考人群的交互模式。我们在四个语料库上评估了GroupPersona,涵盖两种对话来源(助手型和Reddit衍生)及两种构建变体(结构保持型与变异增强型)。相较于最强基线,GroupPersona将合成数据与参考数据在12项行为属性上的Jensen-Shannon散度从0.234降至0.177,降幅达24.4%,且在所有四个语料库中表现最佳或并列最佳;同时保持结构对齐。其生成对话的质量评分与参考对话的偏差最小,平均绝对偏差降至0.63,优于次优基线的0.91。

原文摘要 · Abstract (English)

Synthetic dialogue corpora are increasingly used as proxies for target dialogue data, yet persona-grounded generators optimize individual conversations rather than corpus composition, yielding locally plausible dialogues with distorted population-level behavior mixes. We introduce GroupPersona, a framework that aligns synthetic dialogue corpora to the behavior distribution of a reference corpus. GroupPersona turns population statistics into generation controls: it separates each dialogue's core behavioral signature from predictable side effects, and uses the resulting behavioral groups to condition user agents on the interaction patterns that define the reference population. We evaluate GroupPersona on four corpora crossing two dialogue sources, assistant-style and Reddit-derived, with two construction variants: structure-preserving and variation-enhanced. GroupPersona lowers Jensen-Shannon divergence between synthetic and reference distributions over 12 behavior attributes from 0.234 to 0.177 relative to the strongest average baseline, a 24.4% reduction, and is best or tied-best on all four corpora while preserving structural alignment. It also achieves the closest calibration to reference-conversation quality scores, reducing mean absolute deviation from the reference-conversation profile to 0.63 versus 0.91 for the next-best baseline.

对话生成行为对齐数据合成分布匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。