arXiv:2601.08342cs.CL2026-01中稿 · IWSDS 2026被引 1

首个研究语音中精神操控检测的工作,发现语音识别更难捕捉隐性操纵。

Detecting Mental Manipulation in Speech via Synthetic Multi-Speaker Dialogue

  • 构建合成多说话人语音数据集,模拟真实对话中的精神操控
  • 模型在语音上召回率显著低于文本,因缺少音调等关键声学线索
  • 适合关注语音安全、多模态对话系统设计的研究者

精神操控指通过语言策略隐蔽影响或利用他人,是计算社交推理领域的新任务。以往研究仅聚焦文本对话,忽略了语音中操纵策略的表现形式。本文首次开展语音对话中的精神操控检测研究,提出合成多说话人基准数据集SPEECHMENTALMANIP,通过高质量、语音一致的文生语音技术扩充原有文本数据集。基于少样本大音频-语言模型与人工标注,评估模态对检测准确率和感知的影响。结果表明,模型在语音上的特异性高但召回率明显低于文本,说明训练中缺失声学或语调线索导致敏感性下降。人类评估者在语音情境下也表现出相似不确定性,凸显操纵性语音本身的模糊性。上述发现强调了多模态对话系统需进行模态感知的评估与安全对齐。

原文摘要 · Abstract (English)

Mental manipulation, the strategic use of language to covertly influence or exploit others, is a newly emerging task in computational social reasoning. Prior work has focused exclusively on textual conversations, overlooking how manipulative tactics manifest in speech. We present the first study of mental manipulation detection in spoken dialogues, introducing a synthetic multi-speaker benchmark SPEECHMENTALMANIP that augments a text-based dataset with high-quality, voice-consistent Text-to-Speech rendered audio. Using few-shot large audio-language models and human annotation, we evaluate how modality affects detection accuracy and perception. Our results reveal that models exhibit high specificity but markedly lower recall on speech compared to text, suggesting sensitivity to missing acoustic or prosodic cues in training. Human raters show similar uncertainty in the audio setting, underscoring the inherent ambiguity of manipulative speech. Together, these findings highlight the need for modality-aware evaluation and safety alignment in multimodal dialogue systems.

语音分析精神操控多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。