对比人类与模型在混语对话中的模仿行为,发现模型偏好与人不同。
A Cross-lingual Comparison of Human and Classification Model Entrainment Behavior in Code-switched Speech Settings

- 分析三种混语对话中人类的模仿模式
- 模型虽能识别模仿但更依赖非人类关注特征
- 为多语言语音模型评估提供人性化新框架
会话中的语言模仿在单语和书面语境中已有深入研究,但在口语混语(CSW)场景下仍缺乏探索。本文首次对汉语-英语、印地语-英语、西班牙语-英语三组对话中的跨语言模仿行为进行分析,发现词汇层面的模仿具有跨语言通用性,而声学-韵律及混语风格层面的模仿则呈现情境特异性。基于此,我们进一步考察分类模型是否捕捉人类模仿行为。通过特征重要性和消融分析发现,传统与Transformer类分类器虽能合理识别模仿现象,但始终更倾向于使用对人类行为不具显著性的特征。本研究提出一种以人类行为为基准的多语言风格化对话模型评估框架,揭示了未来构建自然混语对话系统所面临的挑战。
原文摘要 · Abstract (English)
Conversational entrainment is well-studied in monolingual and written contexts, but remains underexplored in spoken code-switching (CSW). We present a novel cross-lingual analysis of entrainment in Mandarin-English, Hindi-English, and Spanish-English dialogue and show that, while lexical entrainment generalizes across language pairs, entrainment over acoustic-prosodic and CSW style aspects exhibits context-specific variation. We build on these findings by asking whether classification models capture these human behavioral patterns. Applying feature importance and ablation analyses, we find that classical and Transformer-based classifiers detect entrainment reasonably well but consistently prioritize features other than those most salient to human entraining behavior. Our approach introduces a human-grounded framework for evaluating model decision-making in multilingual stylistic contexts, and suggests future challenges for developing conversational agents capable of producing naturalistic code-switched speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。