构建37小时人机对话数据集,用于检测和研究具身AI Agent
The DeepSpeak-Agentic Dataset

- 构建可扩展的视频采集系统,自动匹配人类与AI Agent进行对话
- 包含超37小时半结构化对话,支持音视频文本多模态检测
- 适合研究人机交互、AI生成内容识别及大模型应用的团队
我们提出DeepSpeak-Agentic数据集,包含超过37小时的人类与具身AI Agent之间的半结构化对话。该数据集用于评估音频、视频或文本层面的AI Agent自动取证能力,研究人机交互特性,并为推动大型语言模型及驱动具身AI Agent的语音与人脸生成技术提供基准。我们还贡献了一个可扩展的数据采集系统,能够自动生成代理,自动配对人类众包工作者,在指定场景下录制音视频对话,并从混合流中识别与分离人类与代理的声音和画面。
原文摘要 · Abstract (English)
We present DeepSpeak-Agentic, a dataset of videos comprising over 37 hours of semi-structured conversations between a human and an embodied AI agent. We use this dataset to evaluate the automatic forensic identification (audio, video, or text) of AI agents, study the nature of human-agent interactions, and provide a benchmark for future advances in the large-language models and AI-generated voices and faces that power embodied AI agents. We also contribute a scalable data-capture system that creates agents, automatically pairs them with human crowd workers, records audiovisual conversations across specified scenarios, and identifies and separates the human and agent in the combined stream.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。