arXiv:2507.18311cs.CV2025-07中稿 · Machine Intelligen…被引 1

让大模型读懂流场数据,提升科学计算理解能力。

Improving Large Vision-Language Models' Understanding for Flow Field Data

  • 用物理特征提取生成结构化文本,指导模型理解流场。
  • 压缩数据输入,保留关键信息,提升模型学习效率。
  • 适合科研人员在流体力学等领域的智能分析应用。

大型视觉语言模型(LVLM)在图像描述和视觉问答等多模态任务中表现优异,得益于其在大规模图文配对数据上的训练。然而,其在自然科学研究中的复杂场数据(如流场)理解仍不充分。本文提出FieldLVLM框架,包含两个核心组件:一是基于专用机器学习管道提取流场的关键物理特征(如流动分类、雷诺数、涡旋模式),转化为结构化文本数据;二是通过数据压缩策略降低场数据复杂度,仅保留最具信息量的值,以适配模型的语言解码器并引导有效学习。在新构建的基准数据集上的实验表明,FieldLVLM在科学场数据任务中显著优于现有方法。结果表明,该方法为将大模型应用于科学发现提供了新路径,弥合了通用模型与领域知识之间的鸿沟。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have shown impressive capabilities across a range of tasks that integrate visual and textual understanding, such as image captioning and visual question answering. These models are trained on large-scale image and video datasets paired with text, enabling them to bridge visual perception and natural language processing. However, their application to scientific domains, especially in interpreting complex field data commonly used in the natural sciences, remains underexplored. In this work, we introduce FieldLVLM, a novel framework designed to improve large vision-language models' understanding of field data. FieldLVLM consists of two main components: a field-aware language generation strategy and a data-compressed multimodal model tuning. The field-aware language generation strategy leverages a special-purpose machine learning pipeline to extract key physical features from field data, such as flow classification, Reynolds number, and vortex patterns. This information is then converted into structured textual descriptions that serve as a dataset. The data-compressed multimodal model tuning focuses on LVLMs with these generated datasets, using a data compression strategy to reduce the complexity of field inputs and retain only the most informative values. This ensures compatibility with the models language decoder and guides its learning more effectively. Experimental results on newly proposed benchmark datasets demonstrate that FieldLVLM significantly outperforms existing methods in tasks involving scientific field data. Our findings suggest that this approach opens up new possibilities for applying large vision-language models to scientific research, helping bridge the gap between large models and domain-specific discovery.

视觉语言模型流场分析科学计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。