arXiv:2501.04138cs.CL2025-01

大模型能从语音文本中迁移隐性欺骗识别能力

"Yeah Right!" -- Do LLMs Exhibit Multimodal Feature Transfer?

  • 对比语音+文本与纯文本模型,测试多模态迁移能力
  • 语音文本模型在识别隐性欺骗上表现更优
  • 专注人类对话训练的模型更具优势,适合安全检测场景

人类交流是多方面且多模态的技能。理解话语不仅需要表面文字内容,还需把握其隐含意图。人类学习沟通意图始于口语交流,之后将这些能力迁移到书面交流中。本文评估了语音+文本模型以及专注于人类对话训练的纯文本模型,在执行隐蔽欺骗检测任务时的多模态技能迁移能力。结果表明,无需特殊提示,语音+文本大模型在该任务上优于单模态模型;同样,经过人类对话训练的模型也表现出更强的识别能力。

原文摘要 · Abstract (English)

Human communication is a multifaceted and multimodal skill. Communication requires an understanding of both the surface-level textual content and the connotative intent of a piece of communication. In humans, learning to go beyond the surface level starts by learning communicative intent in speech. Once humans acquire these skills in spoken communication, they transfer those skills to written communication. In this paper, we assess the ability of speech+text models and text models trained with special emphasis on human-to-human conversations to make this multimodal transfer of skill. We specifically test these models on their ability to detect covert deceptive communication. We find that with no special prompting speech+text LLMs have an advantage over unimodal LLMs in performing this task. Likewise, we find that human-to-human conversation-trained LLMs are also advantaged in this skill.

多模态大模型欺骗检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。