arXiv:2605.11634cs.CVcs.AI2026-05被引 1

构建首个UML类图视觉问答基准,提升模型理解软件设计图能力

Unlocking UML Class Diagram Understanding in Vision Language Models

  • 设计UML类图专用的视觉问答任务,填补领域空白
  • 建立1.6万条图文问答数据集,训练后性能超越270亿参数大模型
  • 适合关注程序理解、AI辅助软件工程的研究者

尽管视觉语言模型在各类应用中取得显著进展,但在回答图表相关问题方面仍远落后于图像。虽然柱状图、折线图等已有研究,但针对计算机科学领域其他类型图表(如UML类图)的研究仍寥寥无几。本文提出一个基于UML类图的视觉问答基准,兼具挑战性与可操作性。我们构建了一个包含1.6万张图像-问题-答案三元组的大规模训练数据集,并证明采用LoRA微调的方法能轻松超越最新且表现优异的Qwen 3.5 27B模型,该模型在多个基准测试中均表现良好。

原文摘要 · Abstract (English)

Although Vision Language Models (VLMs) have seen tremendous progress across all kinds of use cases, they still fall behind in answering questions regard-ing diagrams compared to photos. Although progress has been made in the area of bar charts, line charts and other diagrams like that there is still few research concerned with other types of diagrams, e.g. in the computer science domain. Our work presents a benchmark for visual question answering based on UML class diagrams which is both challenging and manageable. We further construct a large-scale training dataset with 16.000 image-question-answer triples and show that a LoRA-based finetune easily outperforms Qwen 3.5 27B, which is a recent and well-performing VLM in many other benchmarks.

UML视觉问答软件工程模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。