VC-Agent让用户用最少输入快速收集定制化视频数据。
VC-Agent: An Interactive Agent for Customized Video Dataset Collection
- 通过自然语言理解用户需求,自动检索匹配视频片段。
- 交互过程中动态优化过滤策略,提升筛选准确率。
- 适合需要高效构建个性化视频数据集的研究者使用。
面对数据规模增长的规律,互联网视频数据变得愈发重要。然而,收集满足特定需求的大量视频极其耗时费力。本文研究如何加速该过程,提出首个交互式代理系统 VC-Agent,能够理解用户查询与反馈,并在最小用户输入下检索/扩展相关视频片段。针对用户界面,设计多种基于文本描述和确认的友好操作方式;针对代理功能,利用现有的多模态大模型将用户需求与视频内容关联。更重要的是,提出两种可随用户交互持续更新的新型过滤策略。最后,构建了个性化视频数据集收集的新基准,并通过细致的用户研究验证了该代理在多种真实场景中的有效性。大量实验表明,该代理在定制化视频数据收集中兼具高效性与有效性。项目页面:https://allenyidan.github.io/vcagent_page/
原文摘要 · Abstract (English)
Facing scaling laws, video data from the internet becomes increasingly important. However, collecting extensive videos that meet specific needs is extremely labor-intensive and time-consuming. In this work, we study the way to expedite this collection process and propose VC-Agent, the first interactive agent that is able to understand users' queries and feedback, and accordingly retrieve/scale up relevant video clips with minimal user input. Specifically, considering the user interface, our agent defines various user-friendly ways for the user to specify requirements based on textual descriptions and confirmations. As for agent functions, we leverage existing multi-modal large language models to connect the user's requirements with the video content. More importantly, we propose two novel filtering policies that can be updated when user interaction is continually performed. Finally, we provide a new benchmark for personalized video dataset collection, and carefully conduct the user study to verify our agent's usage in various real scenarios. Extensive experiments demonstrate the effectiveness and efficiency of our agent for customized video dataset collection. Project page: https://allenyidan.github.io/vcagent_page/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。