arXiv:2509.15808cs.SDeess.AS2025-09被引 3

让多人对话更真实:考虑说话人差异与互动规律的时序模拟

From Independence to Interaction: Speaker-Aware Simulation of Multi-Speaker Conversational Timing

  • 基于说话人特异性偏差分布,保证个体发言节奏一致
  • 用马尔可夫链建模换人规则,统一处理停顿与重叠
  • 在Switchboard数据上显著提升对话真实性,适合语音合成研究

本文提出一种说话人感知的多说话人对话时序模拟方法,能捕捉时间一致性与真实的换言动态。以往方法通常在说话人和发言回合间假设独立性,而本方法通过说话人特异性偏差分布确保发言者内部时间一致性,使用马尔可夫链建模换言行为,并固定房间冲激响应以保持空间真实性。同时将停顿与重叠统一为单一间隙分布,采用核密度估计实现平滑连续。在Switchboard数据集上的评估显示,该方法在全局间隙统计、连续间隙相关性、高阶依赖关系(基于柯普拉)、换言熵及间隙存活函数等内在指标上均优于基线,更贴近真实对话模式,捕获细粒度时间依赖与自然说话人交替,但仍面临长程对话结构建模的挑战。

原文摘要 · Abstract (English)

We present a speaker-aware approach for simulating multi-speaker conversations that captures temporal consistency and realistic turn-taking dynamics. Prior work typically models aggregate conversational statistics under an independence assumption across speakers and turns. In contrast, our method uses speaker-specific deviation distributions enforcing intra-speaker temporal consistency, while a Markov chain governs turn-taking and a fixed room impulse response preserves spatial realism. We also unify pauses and overlaps into a single gap distribution, modeled with kernel density estimation for smooth continuity. Evaluation on Switchboard using intrinsic metrics - global gap statistics, correlations between consecutive gaps, copula-based higher-order dependencies, turn-taking entropy, and gap survival functions - shows that speaker-aware simulation better aligns with real conversational patterns than the baseline method, capturing fine-grained temporal dependencies and realistic speaker alternation, while revealing open challenges in modeling long-range conversational structure.

对话建模时序模拟多说话人语音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。