构建多模态对话转折点数据集,助力识别情感与决策关键变化。
MTP: A Dataset for Multi-Modal Turning Points in Casual Conversations
- 提出基于视觉语言模型的叙事构建与转折点检测框架
- 分类任务F1达0.88,检测任务达0.61,结果可解释
- 适合研究情感分析、人机交互与行为预测的学者
识别对话中的关键转折时刻(如情绪爆发或决策改变)对理解人类行为转变及其后果至关重要。本文提出一个新问题设定,聚焦于这些时刻作为转折点(TPs),并构建了一个精心标注、高共识度的人类标注多模态数据集。数据集包含精确的时间戳、描述及可视化-文本证据,突出显示情绪、行为、观点和决策的变化。我们提出TPMaven框架,利用先进的视觉-语言模型从视频中构建叙事,再通过大语言模型对转折点进行分类与检测。评估结果显示,该框架在分类任务中取得0.88的F1分数,在检测任务中为0.61,且额外生成的解释符合人类预期。
原文摘要 · Abstract (English)
Detecting critical moments, such as emotional outbursts or changes in decisions during conversations, is crucial for understanding shifts in human behavior and their consequences. Our work introduces a novel problem setting focusing on these moments as turning points (TPs), accompanied by a meticulously curated, high-consensus, human-annotated multi-modal dataset. We provide precise timestamps, descriptions, and visual-textual evidence high-lighting changes in emotions, behaviors, perspectives, and decisions at these turning points. We also propose a framework, TPMaven, utilizing state-of-the-art vision-language models to construct a narrative from the videos and large language models to classify and detect turning points in our multi-modal dataset. Evaluation results show that TPMaven achieves an F1-score of 0.88 in classification and 0.61 in detection, with additional explanations aligning with human expectations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。