用视觉信息辅助文本判断,提升多语言多模态观点提取准确率
Towards Efficient Multimodal and Multilingual Opinion Extraction for STI: A QLoRA-Based Fine-Tuning Approach

- 用视觉证据锚定文本,结合QLoRA微调模型提升结构化输出
- 在2194样本上实现51.14% F1值,西班牙语和俄语提升超40个百分点
- 适合需要多语言多模态观点分析的科技情报筛选场景
大型语言模型推动了语义分析发展。科学与技术情报(STI)中的观点提取需从海量信息中提炼核心观点。现有模型在零样本多语言、多模态场景下噪声过滤能力差,结构化输出可靠性低。本文提出一种融合视觉证据的多模态核心观点提取框架,以视频语言模型VideoLLaMA2.1(VL2.1)为基础,采用量化低秩适配(QLoRA)在2,194个多语言多模态样本上进行微调。在图像增强设置下,微调后的VL2.1生成结构化JSON输出,精度64.98%,召回率42.15%,F1值51.14%,样本级准确率74.00%。相比零样本的VL2.1,西班牙语和俄语的F1值从4.83%、0.45%分别提升至46.05%、51.93%。框架还引入基于模糊累积前景理论的后处理模块,实现案例级价值评估,为下游科技情报筛选提供价值信号。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have reshaped semantic analysis. Opinion Extraction (OE) for Science and Technology Intelligence (STI) requires concise core opinions from large information streams. Off-the-shelf models struggle to filter noise from these streams and show limited structured-output reliability in zero-shot multilingual and multi-modal settings. To address information overload and extraction defocus, this study proposes a multimodal core-opinion extraction framework in which visual evidence serves as a contextual anchor for textual judgment. Using VideoLLaMA2 (VL2) and VideoLLaMA2.1 (VL2.1) as the base models, we apply Quantized Low-Rank Adaptation (QLoRA) fine-tuning on a curated dataset of 2,194 multilingual and multimodal samples. Under the selected Image-Augmented setting, fine-tuned VL2.1 generates structured JSON core-opinion outputs, achieving 64.98% Precision, 42.15% Recall, 51.14% F1-score, and 74.00% sample-level accuracy. Relative to the zero-shot VL2.1 setting, it raises the F1-scores of Spanish and Russian from 4.83% and 0.45% to 46.05% and 51.93%, respectively. The framework further incorporates a Fuzzy Cumulative Prospect Theory-based post-extraction triage module for case-level value assessment, providing a case-level value signal for downstream STI screening.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。