轻量级日语视觉语言模型,适配边缘设备的工厂与基建场景应用。
PLaMo 2.1-VL Technical Report

- 基于合成数据生成与日语资源构建,支持本地部署的轻量视觉语言模型。
- 在日语VQA与视觉定位任务上超越同类开源模型,最高达85.2%准确率。
- 适用于工厂工具识别与电力设施异常检测,微调后检测性能显著提升。
我们提出PLaMo 2.1-VL,一款面向自主设备的轻量级视觉语言模型(VLM),提供2B和8B两种版本,支持本地与边缘部署,并具备日语操作能力。聚焦视觉问答(VQA)与视觉定位(Visual Grounding)两大核心功能,我们在两个真实场景中评估模型:通过工具识别进行工厂任务分析,以及基础设施异常检测。同时构建了大规模合成数据生成流水线及全面的日语训练与评测资源。PLaMo 2.1-VL在日语与英文基准上优于可比开源模型,在JA-VG-VQA-500上取得61.5 ROUGE-L,在日语Ref-L4上达到85.2%准确率。在工厂任务分析中实现53.9%零样本准确率;对电厂数据微调后,异常检测的bbox+label F1分数从39.7提升至64.9。
原文摘要 · Abstract (English)
We introduce PLaMo 2.1-VL, a lightweight Vision Language Model (VLM) for autonomous devices, available in 8B and 2B variants and designed for local and edge deployment with Japanese-language operation. Focusing on Visual Question Answering (VQA) and Visual Grounding as its core capabilities, we develop and evaluate the models for two real-world application scenarios: factory task analysis via tool recognition, and infrastructure anomaly detection. We also develop a large-scale synthetic data generation pipeline and comprehensive Japanese training and evaluation resources. PLaMo 2.1-VL outperforms comparable open models on Japanese and English benchmarks, achieving 61.5 ROUGE-L on JA-VG-VQA-500 and 85.2% accuracy on Japanese Ref-L4. For the two application scenarios, it achieves 53.9% zero-shot accuracy on factory task analysis, and fine-tuning on power plant data improves anomaly detection bbox + label F1-score from 39.7 to 64.9.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。