用迭代对话框架提升AI理解人类互动的能力。
Leveraging LLMs with Iterative Loop Structure for Enhanced Social Intelligence in Video Question Answering
- 设计循环辩论机制,融合视觉与语言模型理解视频中的人际互动
- 在Social-IQ 2.0上零微调即达顶尖性能
- 适合需要高社交理解力的陪护、教育类AI系统
社交智能,即解读情绪、意图和行为的能力,对有效沟通和自适应回应至关重要。随着机器人和AI在护理、医疗和教育领域的普及,对能自然与人类交互的AI需求日益增长。然而,实现视觉与语音等多模态信息的无缝融合仍是挑战。现有基于视频的社交智能方法多依赖通用视频识别或情感识别技术,常忽略人际互动中的独特要素。为此,我们提出循环视频辩论(Looped Video Debating, LVD)框架,将大语言模型(LLMs)与面部表情、肢体动作等视觉信息结合,提升涉及人类互动视频的问答任务的透明度与可靠性。在Social-IQ 2.0基准测试中,LVD实现无需微调的最先进性能。此外,对现有数据集的补充人工标注揭示了模型准确率,为未来AI社交智能改进提供指导。
原文摘要 · Abstract (English)
Social intelligence, the ability to interpret emotions, intentions, and behaviors, is essential for effective communication and adaptive responses. As robots and AI systems become more prevalent in caregiving, healthcare, and education, the demand for AI that can interact naturally with humans grows. However, creating AI that seamlessly integrates multiple modalities, such as vision and speech, remains a challenge. Current video-based methods for social intelligence rely on general video recognition or emotion recognition techniques, often overlook the unique elements inherent in human interactions. To address this, we propose the Looped Video Debating (LVD) framework, which integrates Large Language Models (LLMs) with visual information, such as facial expressions and body movements, to enhance the transparency and reliability of question-answering tasks involving human interaction videos. Our results on the Social-IQ 2.0 benchmark show that LVD achieves state-of-the-art performance without fine-tuning. Furthermore, supplementary human annotations on existing datasets provide insights into the model's accuracy, guiding future improvements in AI-driven social intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。