首个面向拉丁语-英语双语问答的基准数据集,用于评估模型在古典语言中的理解能力。
RespondeoQA: a Benchmark for Bilingual Latin-English Question Answering

- 构建包含7800组问答对的双语拉丁语-英语数据集,涵盖多种题型。
- 大模型在技能类问题上表现较差,推理模型对诗律和修辞任务略有提升。
- 适合研究古典语言处理、多语言AI评估及教育AI的学者与开发者。
我们引入了一个用于双语拉丁语-英语问答与翻译的基准数据集,包含约7800个问答对。题目源自拉丁语教学材料,包括考试、智力问答和教材,时间跨度从19世纪至今。经过自动化提取、清洗与人工审核,数据集涵盖知识性、技能性、多跳推理、受限翻译及混合语言对等多样化题型。据我们所知,这是首个聚焦拉丁语的问答基准。作为案例研究,我们评估了LLaMa 3、Qwen QwQ和OpenAI o3-mini三个大模型,发现所有模型在技能导向问题上表现不佳。尽管推理模型在诗律和文学修辞任务上表现更好,但整体提升有限。Qwen QwQ在拉丁语提问时略优,而LLaMa3和o3-mini更具任务依赖性。该数据集为评估模型在特定语言与文化领域的能力提供了新资源,其构建流程可推广至其他语言。数据集已开源:https://github.com/slanglab/RespondeoQA。
原文摘要 · Abstract (English)
We introduce a benchmark dataset for question answering and translation in bilingual Latin and English settings, containing about 7,800 question-answer pairs. The questions are drawn from Latin pedagogical sources, including exams, quizbowl-style trivia, and textbooks ranging from the 1800s to the present. After automated extraction, cleaning, and manual review, the dataset covers a diverse range of question types: knowledge- and skill-based, multihop reasoning, constrained translation, and mixed language pairs. To our knowledge, this is the first QA benchmark centered on Latin. As a case study, we evaluate three large language models -- LLaMa 3, Qwen QwQ, and OpenAI's o3-mini -- finding that all perform worse on skill-oriented questions. Although the reasoning models perform better on scansion and literary-device tasks, they offer limited improvement overall. QwQ performs slightly better on questions asked in Latin, but LLaMa3 and o3-mini are more task dependent. This dataset provides a new resource for assessing model capabilities in a specialized linguistic and cultural domain, and the creation process can be easily adapted for other languages. The dataset is available at: https://github.com/slanglab/RespondeoQA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。