7B参数模型实现短视频多粒度结构化理解,提升推荐与搜索效果。
ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts
- 端到端处理视频、音频、文本,支持时间戳级描述与推理
- 在1分钟视频上推理仅需10秒,零样本或少量微调即适用
- 适用于短视频推荐、搜索等真实场景,已显著提升用户满意度
真实世界用户生成的短视频(如微信视频号、TikTok)主导移动互联网。当前大模型缺乏对这类视频的时序结构化、细节化与深度理解能力,而这是有效视频搜索与推荐的基础。理解此类短视频极具挑战:视觉元素复杂,视听信息密度高,节奏快且重情感表达与观点传递,需高级跨模态推理整合视觉、音频与文本。本文提出ARC-Hunyuan-Video-7B,一个可端到端处理原始视频输入中视觉、音频和文本信号的多模态模型,支持多粒度时间戳视频描述、摘要、开放问答、时间定位与视频推理。基于自动化标注数据集,该7B参数小模型通过预训练、指令微调、冷启动、强化学习后训练及最终指令微调的完整流程训练。在自建基准ShortVid-Bench上的定量评估与定性对比显示其在真实视频理解中表现优异,支持零样本或少量样本微调适配多种下游任务。模型已投入真实生产部署,显著提升用户参与度与满意度,具备极高效率:在H20 GPU上,1分钟视频推理仅需10秒。
原文摘要 · Abstract (English)
Real-world user-generated short videos, especially those distributed on platforms such as WeChat Channel and TikTok, dominate the mobile internet. However, current large multimodal models lack essential temporally-structured, detailed, and in-depth video comprehension capabilities, which are the cornerstone of effective video search and recommendation, as well as emerging video applications. Understanding real-world shorts is actually challenging due to their complex visual elements, high information density in both visuals and audio, and fast pacing that focuses on emotional expression and viewpoint delivery. This requires advanced reasoning to effectively integrate multimodal information, including visual, audio, and text. In this work, we introduce ARC-Hunyuan-Video, a multimodal model that processes visual, audio, and textual signals from raw video inputs end-to-end for structured comprehension. The model is capable of multi-granularity timestamped video captioning and summarization, open-ended video question answering, temporal video grounding, and video reasoning. Leveraging high-quality data from an automated annotation pipeline, our compact 7B-parameter model is trained through a comprehensive regimen: pre-training, instruction fine-tuning, cold start, reinforcement learning (RL) post-training, and final instruction fine-tuning. Quantitative evaluations on our introduced benchmark ShortVid-Bench and qualitative comparisons demonstrate its strong performance in real-world video comprehension, and it supports zero-shot or fine-tuning with a few samples for diverse downstream applications. The real-world production deployment of our model has yielded tangible and measurable improvements in user engagement and satisfaction, a success supported by its remarkable efficiency, with stress tests indicating an inference time of just 10 seconds for a one-minute video on H20 GPU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。