arXiv:2507.19666cs.CL2025-07Conference of the …

构建罗马尼亚驾驶执照考试多模态数据集,评测大模型法律理解能力

RoD-TAL: A Benchmark for Answering Questions in Romanian Driving License Exams

  • 构建含图文的罗马尼亚驾考数据集,附专家标注法律依据
  • 领域微调显著提升检索效果,思维链提示使问答准确率超及格线
  • 适合研究法律教育AI、多模态推理或低资源语言应用者

人工智能与法律系统的结合日益迫切,尤其在罗马尼亚等资源有限的语言中更显重要。本文旨在评估大语言模型(LLMs)和视觉语言模型(VLMs)对罗马尼亚驾驶法规的理解与推理能力,涵盖文本与图像问答任务。为此,我们提出RoD-TAL——一个全新的多模态数据集,包含罗马尼亚驾驶考试题目、文本与图像形式的问题,以及由人类专家标注的法律条文引用与解释。我们实现并评估了检索增强生成(RAG)流程、密集检索器及优化推理的模型,涵盖信息检索(IR)、问答(QA)、视觉信息检索与视觉问答任务。实验表明,领域特定微调显著提升检索性能;思维链提示与专用推理模型可提高问答准确率,超过驾驶考试及格标准。本文揭示了大模型在法律教育中的潜力与局限,并开源代码与资源。

原文摘要 · Abstract (English)

The intersection of AI and legal systems presents a growing need for tools that support legal education, particularly in under-resourced languages such as Romanian. In this work, we aim to evaluate the capabilities of Large Language Models (LLMs) and Vision-Language Models (VLMs) in understanding and reasoning about the Romanian driving law through textual and visual question-answering tasks. To facilitate this, we introduce RoD-TAL, a novel multimodal dataset comprising Romanian driving test questions, text-based and image-based, along with annotated legal references and explanations written by human experts. We implement and assess retrieval-augmented generation (RAG) pipelines, dense retrievers, and reasoning-optimized models across tasks, including Information Retrieval (IR), Question Answering (QA), Visual IR, and Visual QA. Our experiments demonstrate that domain-specific fine-tuning significantly enhances retrieval performance. At the same time, chain-of-thought prompting and specialized reasoning models improve QA accuracy, surpassing the minimum passing grades required for driving exams. We highlight the potential and limitations of applying LLMs and VLMs to legal education. We release the code and resources through the GitHub repository.

法律AI多模态低资源语言评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。