arXiv:2509.23715cs.CLcs.LG2025-09中稿 · @ CONSILR 2025 Buc…

评测大模型对罗马尼亚交通法规的理解能力,探索多模态与微调效果。

Do LLMs Understand Romanian Driving Laws? A Study on Multimodal and Fine-Tuned Question Answering

  • 构建1208题数据集,含387个图文题,对比文本与多模态模型表现。
  • 微调后的Llama 3.1-8B在法规问答中表现接近顶尖模型,文本描述优于直接图像输入。
  • 发现大模型自评解释质量存在偏好偏差,适用于小语种可解释问答研究。

确保新手与老司机掌握现行交通规则对道路安全至关重要。本文评估大语言模型(LLMs)在罗马尼亚交通法规问答任务中的表现,并生成解释。我们发布了包含1,208道题的语料库(其中387题为多模态),比较了纯文本与多模态的SOTA系统,并测试了针对Llama 3.1-8B-Instruct和RoLlama 3.1-8B-Instruct进行领域微调的效果。结果显示,现有模型表现良好,微调后的8B模型具备竞争力;文本描述图像的表现优于直接输入视觉信息。最后,通过一个以LLM为裁判的评估框架,发现模型在解释质量自评中存在自我偏好偏差。本研究为低资源语言的可解释问答提供重要参考。

原文摘要 · Abstract (English)

Ensuring that both new and experienced drivers master current traffic rules is critical to road safety. This paper evaluates Large Language Models (LLMs) on Romanian driving-law QA with explanation generation. We release a 1{,}208-question dataset (387 multimodal) and compare text-only and multimodal SOTA systems, then measure the impact of domain-specific fine-tuning for Llama 3.1-8B-Instruct and RoLlama 3.1-8B-Instruct. SOTA models perform well, but fine-tuned 8B models are competitive. Textual descriptions of images outperform direct visual input. Finally, an LLM-as-a-Judge assesses explanation quality, revealing self-preference bias. The study informs explainable QA for less-resourced languages.

大模型法规问答多模态小语种

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。