arXiv:2603.20354cs.MMcs.AI2026-03

提出六维结构化视频表示框架,让AI理解短视频的节奏与编排逻辑。

Leum-VL Technical Report

  • 用六个可观察维度分解短视频结构,如镜头语言、剪辑节奏、传播策略等
  • 模型在多个评测中表现优异,最高达70.8分,尤其擅长时序关键单元识别
  • 适合需要精准控制视频生成、推荐或编辑的场景,如带字幕和图文叠加的内容

一段短视频的成功不仅取决于内容本身,更在于注意力调度的组织方式——但现有多模态模型缺乏解析或生成这种结构的能力。当前模型能描述画面、回答事件相关问题、读取屏幕文字,却难以可靠识别时间线上的关键单元,如开头钩子、剪辑理由、镜头引发的张力、平台适配提示等。我们提出SV6D(六维结构化视频),受影视分镜实践启发,将互联网原生视频分解为六个互补的结构维度:主体、美学、镜头语言、剪辑、叙事、传播,每个标签均对应时间轴上可观测的证据。我们建立统一优化目标,融合匈牙利匹配的时间对齐、维度间语义距离及质量正则化。基于此框架,我们构建了Leum-VL-8B,一个80亿参数的视频-语言模型,通过专家驱动的后训练流程与可验证的强化学习进一步优化,在无字幕条件下取得VideoMME 70.8分、MVBench 70.0分、MotionBench 61.6分的成绩,同时保持在MMBench-EN等通用多模态任务上的竞争力。我们还构建了FeedBench,用于评估结构敏感的短视频理解能力。结果表明,视频AI缺失的并非像素生成,而是结构化表征:需基于时间线、关联可见证据,并可直接服务于剪辑、检索、推荐与生成控制等下游流程,尤其适用于含叠加字幕和图文布局的文本密集型视频格式。

原文摘要 · Abstract (English)

A short video succeeds not simply because of what it shows, but because of how it schedules attention -- yet current multimodal models lack the structural grammar to parse or produce this organization. Existing models can describe scenes, answer event-centric questions, and read on-screen text, but they are far less reliable at identifying timeline-grounded units such as hooks, cut rationales, shot-induced tension, and platform-facing packaging cues. We propose SV6D (Structured Video in Six Dimensions), inspired by professional storyboard practice in film and television production, a representation framework that decomposes internet-native video into six complementary structural dimensions -- subject, aesthetics, camera language, editing, narrative, and dissemination -- with each label tied to physically observable evidence on the timeline. We formalize a unified optimization objective over SV6D that combines Hungarian-matched temporal alignment, dimension-wise semantic label distance, and quality regularization. Building on this framework, we present Leum-VL-8B, an 8B video-language model that realizes the SV6D objective through an expert-driven post-training pipeline, further refined through verifiable reinforcement learning on perception-oriented tasks. Leum-VL-8B achieves 70.8 on VideoMME (w/o subtitles), 70.0 on MVBench, and 61.6 on MotionBench, while remaining competitive on general multimodal evaluations such as MMBench-EN. We also construct FeedBench, a benchmark for structure-sensitive short-video understanding. Our results indicate that the missing layer in video AI is not pixel generation but structural representation: grounded on the timeline, linked to visible evidence, and directly consumable by downstream workflows such as editing, retrieval, recommendation, and generation control, including text-heavy internet video formats with overlays and image-text layouts.

视频理解结构化表示多模态生成控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。