arXiv:2607.18076cs.CL2026-07

对比真人与AI生成对话中的沉默时长,发现性别和制作方式影响交流节奏。

Modeling turn-taking with distant viewing: investigating silence thresholds in human and AI-generated discourse

  • 用音高阈值区分性别,分析三十部美剧与五十一段AI播客的沉默间隔
  • 男性发言者沉默时长平均比女性多0.8秒,且真人剧比AI生成更不规律
  • 适合关注人机对话差异、语音交互设计的研究者与开发者

本研究分析了三十部美国情景喜剧和五十一段由Google NotebookLM生成的合成播客中的沉默间隔。通过在Praat中基于基频阈值判定发言者性别,并比较不同制作场景下的沉默特征。结果显示,男性发言者的沉默时间平均比女性长0.8秒,真人录制节目中的沉默分布较AI生成内容更为不规则。研究揭示了人类与人工智能生成话语在交流节奏上的系统性差异。

原文摘要 · Abstract (English)

This study investigates silence gaps in two kinds of audiovisual material. We analysed thirty US situational comedies and fifty-one synthetic podcasts generated with Google NotebookLM. Gaps were compared across speaker gender, assigned from a fundamental-frequency threshold estimated in Praat, and across production settings.

对话建模沉默分析人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。