用语音引导压缩技术,让电商视频自动生成精准时间对齐的分章故事。
HiVid-Narrator: Hierarchical Video Narrative Generation with Scene-Primed ASR-anchored Compression
- 分阶段生成:先收集语音与画面证据,再基于时间链构建章节。
- 输入令牌减少40%,在E-HVC数据集上叙事质量超越现有方法。
- 适合需要自动剪辑电商视频的创作者或平台使用。
生成真实电商视频的结构化叙述需同时感知细粒度视觉细节并组织成连贯高层故事,现有方法难以统一实现。我们构建了双粒度、时间对齐的电商视频分层字幕数据集E-HVC,包含事件级观察的时序思维链(Temporal Chain-of-Thought)和凝练的故事中心摘要(Chapter Summary)。不直接生成章节,而是分步进行:先通过筛选后的语音识别(ASR)与帧级描述收集可靠语义与视觉证据,再基于时序思维链精炼粗略标注,生成精确章节边界与标题,确保事实依据与时间对齐。我们还发现电商视频节奏快、信息密集,视觉令牌主导输入序列。为此提出场景引导的语音锚定压缩器(SPA-Compressor),利用语音语义线索将多模态令牌压缩为层级化的场景与事件表示。基于此设计,HiVid-Narrator框架在更少输入令牌下实现更优叙事质量。
原文摘要 · Abstract (English)
Generating structured narrations for real-world e-commerce videos requires models to perceive fine-grained visual details and organize them into coherent, high-level stories--capabilities that existing approaches struggle to unify. We introduce the E-commerce Hierarchical Video Captioning (E-HVC) dataset with dual-granularity, temporally grounded annotations: a Temporal Chain-of-Thought that anchors event-level observations and Chapter Summary that compose them into concise, story-centric summaries. Rather than directly prompting chapters, we adopt a staged construction that first gathers reliable linguistic and visual evidence via curated ASR and frame-level descriptions, then refines coarse annotations into precise chapter boundaries and titles conditioned on the Temporal Chain-of-Thought, yielding fact-grounded, time-aligned narratives. We also observe that e-commerce videos are fast-paced and information-dense, with visual tokens dominating the input sequence. To enable efficient training while reducing input tokens, we propose the Scene-Primed ASR-anchored Compressor (SPA-Compressor), which compresses multimodal tokens into hierarchical scene and event representations guided by ASR semantic cues. Built upon these designs, our HiVid-Narrator framework achieves superior narrative quality with fewer input tokens compared to existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。