arXiv:2601.04891cs.CVcs.LG2026-01被引 1

在工业级约束下,实测多模态模型处理长视频的效率与瓶颈。

Scaling Vision Language Models for Pharmaceutical Long Form Video Reasoning on Industrial GenAI Platform

论文配图:Scaling Vision Language Models for Pharmaceutical Long Form Video Reasoning on Industrial GenAI Platform
图 1 · 摘自论文原文
  • 构建工业级框架,支持超20万份PDF、2.5万段多格式视频处理
  • 发现商品级GPU上使用SDPA注意力可提升3-8倍效率,多模态提升8/12任务表现
  • 揭示时序对齐与关键帧检测是当前模型主要瓶颈,适合医药领域开发者参考

视觉语言模型(VLMs)在多模态推理任务中表现强劲,但现有评估多集中于短视频且不考虑计算资源限制。在制药内容理解等工业场景中,需在严格GPU、延迟和成本约束下处理长视频,许多现有方法难以扩展。本文提出一个工业级GenAI框架,处理超过20万份PDF、25,326段跨8种格式(如MP4、M4V等)的视频,以及888个跨20多种语言的多语言音频文件。研究贡献包括:(i) 面向制药领域的工业级大规模多模态推理架构;(ii) 在Video-MME和MMBench两个主流基准及自建包含25,326段视频、覆盖14种疾病领域的数据集上,对40余种VLM进行实证分析;(iii) 得出四个关于长视频推理的关键发现:多模态作用、注意力机制权衡、时序推理极限,以及视频分片在GPU约束下的挑战。结果显示,商品级GPU上使用SDPA注意力可实现3-8倍效率提升,多模态在8/12任务域中表现更优(尤其依赖长度的任务),且开源自研模型均存在时序对齐与关键帧检测瓶颈。本文未提出新模型,而是刻画当前VLM在真实部署条件下的实际局限、权衡与失败模式,为研究人员和从业者提供可落地的长视频多模态系统设计指导。

原文摘要 · Abstract (English)

Vision Language Models (VLMs) have shown strong performance on multimodal reasoning tasks, yet most evaluations focus on short videos and assume unconstrained computational resources. In industrial settings such as pharmaceutical content understanding, practitioners must process long-form videos under strict GPU, latency, and cost constraints, where many existing approaches fail to scale. In this work, we present an industrial GenAI framework that processes over 200,000 PDFs, 25,326 videos across eight formats (e.g., MP4, M4V, etc.), and 888 multilingual audio files in more than 20 languages. Our study makes three contributions: (i) an industrial large-scale architecture for multimodal reasoning in pharmaceutical domains; (ii) empirical analysis of over 40 VLMs on two leading benchmarks (Video-MME and MMBench) and proprietary dataset of 25,326 videos across 14 disease areas; and (iii) four findings relevant to long-form video reasoning: the role of multimodality, attention mechanism trade-offs, temporal reasoning limits, and challenges of video splitting under GPU constraints. Results show 3-8 times efficiency gains with SDPA attention on commodity GPUs, multimodality improving up to 8/12 task domains (especially length-dependent tasks), and clear bottlenecks in temporal alignment and keyframe detection across open- and closed-source VLMs. Rather than proposing a new "A+B" model, this paper characterizes practical limits, trade-offs, and failure patterns of current VLMs under realistic deployment constraints, and provide actionable guidance for both researchers and practitioners designing scalable multimodal systems for long-form video understanding in industrial domains.

多模态长视频工业应用推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。