arXiv:2509.20961cs.CVcs.AI2025-09被引 2

用多模态技术自动总结金融短视频,让长内容变易懂。

Unlocking Financial Insights: An advanced Multimodal Summarization with Multimodal Output Framework for Financial Advisory Videos

  • 融合视觉、语音、文字特征,自动生成精准摘要
  • 在470段视频上表现优于大模型,事实准确率高
  • 适合金融研究者和智能投顾开发者使用

社交媒体的快速发展拓宽了金融咨询类播客视频的传播范围,但从30-40分钟的多模态内容中提取有效信息仍具挑战。本文提出FASTER(Financial Advisory Summariser with Textual Embedded Relevant images)框架,解决三大问题:(1)提取跨模态特征,(2)生成简洁优化摘要,(3)对齐关键帧与文本内容。该框架采用BLIP生成视觉语义描述,OCR识别文字模式,基于Whisper的语音转录与说话人分离作为基础输入特征。通过改进的直接偏好优化(DPO)损失函数并加入针对基础输出的验证机制,确保摘要在事实性、相关性和一致性上符合人工标准。此外,引入基于排序的检索机制,实现关键帧与摘要内容的精准对齐,提升可解释性与跨模态一致性。为缓解数据稀缺问题,构建了包含470个公开可获取的金融咨询类演讲视频的数据集Fin-APT。跨领域实验表明,FASTER在性能、鲁棒性与泛化能力上均优于主流大语言模型(LLMs)与视觉-语言模型(VLMs)。该工作为多模态摘要树立新标准,使金融内容更易获取与应用。代码与数据集已开源。

原文摘要 · Abstract (English)

The dynamic propagation of social media has broadened the reach of financial advisory content through podcast videos, yet extracting insights from lengthy, multimodal segments (30-40 minutes) remains challenging. We introduce FASTER (Financial Advisory Summariser with Textual Embedded Relevant images), a modular framework that tackles three key challenges: (1) extracting modality-specific features, (2) producing optimized, concise summaries, and (3) aligning visual keyframes with associated textual points. FASTER employs BLIP for semantic visual descriptions, OCR for textual patterns, and Whisper-based transcription with Speaker diarization as BOS features. A modified Direct Preference Optimization (DPO)-based loss function, equipped with BOS-specific fact-checking, ensures precision, relevance, and factual consistency against the human-aligned summary. A ranker-based retrieval mechanism further aligns keyframes with summarized content, enhancing interpretability and cross-modal coherence. To acknowledge data resource scarcity, we introduce Fin-APT, a dataset comprising 470 publicly accessible financial advisory pep-talk videos for robust multimodal research. Comprehensive cross-domain experiments confirm FASTER's strong performance, robustness, and generalizability when compared to Large Language Models (LLMs) and Vision-Language Models (VLMs). By establishing a new standard for multimodal summarization, FASTER makes financial advisory content more accessible and actionable, thereby opening new avenues for research. The dataset and code are available at: https://github.com/sarmistha-D/FASTER

多模态金融分析摘要生成视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。