用视觉语言模型自动生成个性化药学视频摘要,提升效率与合规性。
From Understanding to Engagement: Personalized pharmacy Video Clips via Vision Language Models (VLMs)
- 融合音视频语言模型,自动剪辑长视频生成连贯片段。
- 速度提升3-4倍,成本降低4倍,关键指标优于现有模型。
- 支持营销、培训等场景个性化输出,适合生命科学领域使用。
视觉语言模型(VLMs)有望推动制药行业的数字化转型,实现多模态内容的智能、可扩展和自动化处理。传统人工标注异构数据(文本、图像、视频、音频、网页链接)易出现不一致、质量下降和利用效率低的问题,尤其面对长视频与音频数据(如临床试验访谈、教育讲座)时挑战更大。本文提出一种针对药学领域的视频到片段生成框架,整合音频语言模型(ALMs)与视觉语言模型(VLMs),生成高质量亮点片段。贡献包括:(i) 可复现的剪切合并算法,支持淡入淡出与时间戳归一化,确保画面与音频同步;(ii) 基于角色定义与提示注入的个性化机制,可生成面向市场、培训、监管等不同用途的内容;(iii) 成本高效的端到端流程,平衡ALM/VLM增强处理能力。在Video MME基准(900条)及自有16,159条跨14个疾病领域的药学视频数据集上评估,实现3-4倍提速、4倍成本降低,且片段质量具有竞争力。此外,生成片段的连贯性得分(0.348)与信息量得分(0.721)均优于主流VLM基线(如Gemini 2.5 Pro),展现了透明、定制化、合规支持的视频摘要在生命科学中的潜力。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) are poised to revolutionize the digital transformation of pharmacyceutical industry by enabling intelligent, scalable, and automated multi-modality content processing. Traditional manual annotation of heterogeneous data modalities (text, images, video, audio, and web links), is prone to inconsistencies, quality degradation, and inefficiencies in content utilization. The sheer volume of long video and audio data further exacerbates these challenges, (e.g. long clinical trial interviews and educational seminars). Here, we introduce a domain adapted Video to Video Clip Generation framework that integrates Audio Language Models (ALMs) and Vision Language Models (VLMs) to produce highlight clips. Our contributions are threefold: (i) a reproducible Cut & Merge algorithm with fade in/out and timestamp normalization, ensuring smooth transitions and audio/visual alignment; (ii) a personalization mechanism based on role definition and prompt injection for tailored outputs (marketing, training, regulatory); (iii) a cost efficient e2e pipeline strategy balancing ALM/VLM enhanced processing. Evaluations on Video MME benchmark (900) and our proprietary dataset of 16,159 pharmacy videos across 14 disease areas demonstrate 3 to 4 times speedup, 4 times cost reduction, and competitive clip quality. Beyond efficiency gains, we also report our methods improved clip coherence scores (0.348) and informativeness scores (0.721) over state of the art VLM baselines (e.g., Gemini 2.5 Pro), highlighting the potential of transparent, custom extractive, and compliance supporting video summarization for life sciences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。