用语音对话中的文本、声音和行为线索预测人格特质
PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs
- 整合语音、文本与行为信号,构建对话级人格标注数据集
- 基于大模型的预测系统与人工评估高度一致
- 为个性化对话代理提供可落地的人格感知能力
尽管神经语音对话系统取得显著进展,但因语音数据集缺乏人格标注,具备人格感知能力的对话代理仍研究不足。本文提出一套流程,将原始音频转化为带时间戳、回应类型及情绪/情感标签的对话数据集。通过自动语音识别(ASR)提取转录文本与时间信息,生成对话层级的标注。利用这些标注,设计基于大语言模型的人格预测系统。邀请人工评估者识别对话特征并分配人格标签。分析表明,该系统在与人类判断的一致性上优于现有方法。
原文摘要 · Abstract (English)
Despite significant progress in neural spoken dialog systems, personality-aware conversation agents -- capable of adapting behavior based on personalities -- remain underexplored due to the absence of personality annotations in speech datasets. We propose a pipeline that preprocesses raw audio recordings to create a dialogue dataset annotated with timestamps, response types, and emotion/sentiment labels. We employ an automatic speech recognition (ASR) system to extract transcripts and timestamps, then generate conversation-level annotations. Leveraging these annotations, we design a system that employs large language models to predict conversational personality. Human evaluators were engaged to identify conversational characteristics and assign personality labels. Our analysis demonstrates that the proposed system achieves stronger alignment with human judgments compared to existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。