构建首个多语言混合对话视频语料库,推动视频对话系统多样化发展
KwaiChat: A Large-Scale Video-Driven Multilingual Mixed-Type Dialogue Corpus
- 构建人类生成的跨语言多类型视频对话数据集
- 涵盖4类对话、30个领域、13个主题,共9.3万段视频24.6万条对话
- 揭示当前大模型在复杂视频对话中仍表现不足,适合多模态对话研究者
基于视频的对话系统在教育助手等领域具有广阔应用前景,但现有系统受限于单一对话类型,难以适应问答、情感交流等多种场景。本文提出生成视频驱动的多语言混合类型对话的新任务,并构建了名为KwaiChat的人类生成视频对话语料库,包含93,209段视频和246,080条对话,覆盖4种对话类型、30个领域、4种语言和13个主题。同时在该数据集上建立了基线模型。对7种不同LLM的广泛分析表明,尽管GPT-4o表现最优,但在上下文学习和微调支持下仍表现不佳,说明该任务具有挑战性,亟需进一步研究。
原文摘要 · Abstract (English)
Video-based dialogue systems, such as education assistants, have compelling application value, thereby garnering growing interest. However, the current video-based dialogue systems are limited by their reliance on a single dialogue type, which hinders their versatility in practical applications across a range of scenarios, including question-answering, emotional dialog, etc. In this paper, we identify this challenge as how to generate video-driven multilingual mixed-type dialogues. To mitigate this challenge, we propose a novel task and create a human-to-human video-driven multilingual mixed-type dialogue corpus, termed KwaiChat, containing a total of 93,209 videos and 246,080 dialogues, across 4 dialogue types, 30 domains, 4 languages, and 13 topics. Additionally, we establish baseline models on KwaiChat. An extensive analysis of 7 distinct LLMs on KwaiChat reveals that GPT-4o achieves the best performance but still cannot perform well in this situation even with the help of in-context learning and fine-tuning, which indicates that the task is not trivial and needs further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。