arXiv:2511.05577cs.LGcond-mat.mtrl-sci2025-11

用多模态数据微调视觉语言模型,预测高分子性质

Fine-Tuning Vision-Language Models for Multimodal Polymer Property Prediction

  • 用指令对微调视觉语言模型,融合图像与文本信息
  • 微调后模型在性质预测上超越单模态方法
  • 一模型可预测多种性质,降低部署维护成本

视觉语言模型(VLMs)在视觉问答和多模态文本生成等任务中表现优异,但在材料科学等科学领域应用仍有限。尽管已有部分机器学习方法解决该领域的特定挑战,但缺乏针对高分子性质预测等广泛任务设计的通用基础模型。本文构建了一个多模态高分子数据集,通过指令微调对VLM进行训练,并评估多模态信息对预测性能的影响。使用LoRA微调的模型优于单模态及基线方法,证明了多模态学习的优势。此外,该方法无需为不同性质单独训练模型,显著降低了部署与维护成本。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have shown strong performance in tasks like visual question answering and multimodal text generation, but their effectiveness in scientific domains such as materials science remains limited. While some machine learning methods have addressed specific challenges in this field, there is still a lack of foundation models designed for broad tasks like polymer property prediction using multimodal data. In this work, we present a multimodal polymer dataset to fine-tune VLMs through instruction-tuning pairs and assess the impact of multimodality on prediction performance. Our fine-tuned models, using LoRA, outperform unimodal and baseline approaches, demonstrating the benefits of multimodal learning. Additionally, this approach reduces the need to train separate models for different properties, lowering deployment and maintenance costs.

多模态学习高分子预测视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。