构建首个多模态人物知识评估框架,提升大模型在人物事实推理中的准确性。
ADAM: A Diverse Archive of Mankind for Evaluating and Enhancing LLMs in Biographical Reasoning
- 构建覆盖400万+人物的跨语言多模态数据集,支持多维度人物知识评估。
- 引入AdamRAG检索增强生成系统,显著降低模型对冷门人物的幻觉错误。
- 适用于需要准确人物知识推理的研究者,尤其关注多语言与多模态场景。
我们提出ADAM(A Diverse Archive of Mankind),一个用于评估和提升多模态大语言模型(MLLMs)在人物事实推理能力的框架。据我们所知,这是首个系统性研究人物知识这一关键但被忽视的事实知识维度的工作。核心部分包括:AdamDB,一个涵盖400万+个人、跨地理、时间与职业的多语言多模态数据集;AdamBench,基于布卢姆分类学设计的认知结构化评测体系,覆盖六种推理层级,支持英文及本土语言。为缓解幻觉问题,特别是针对知名度较低的人物,我们提出专用于人物场景的AdamRAG检索增强生成系统。实验表明,AdamRAG显著提升开源模型性能,对闭源模型有适度改善,且在低阶推理任务中收益最大。模型准确性受人物流行度显著影响,而面部图像等多模态输入带来的提升较小且不一致。ADAM建立了首个基于认知、文化与多模态融合的人物知识评估基准,推动多语言、高准确率、抗幻觉的MLLM发展。
原文摘要 · Abstract (English)
We introduce ADAM (A Diverse Archive of Mankind), a framework for evaluating and improving multimodal large language models (MLLMs) in biographical reasoning. To the best of our knowledge, this is the first work to systematically examine LLM capabilities in biography, a critical yet underexplored dimension of factual knowledge. At its core, AdamDB is a multilingual and multimodal dataset covering over 4 million individuals across geography, time, and profession, while AdamBench provides cognitively structured evaluations based on Bloom's taxonomy, spanning six reasoning levels in both English and native languages. To address hallucinations, particularly for lesser-known individuals, we propose AdamRAG, a retrieval-augmented generation system tailored to biographical contexts. Experiments show that AdamRAG substantially improves open-source models and modestly benefits closed-source ones, with the largest gains on lower-order reasoning. Popularity strongly mediates accuracy, and multimodal input via face images offers smaller, less consistent improvements than retrieval. ADAM establishes the first benchmark and framework for cognitively, culturally, and multimodally grounded biographical evaluation, advancing the development of multilingual, accurate, and hallucination-resistant MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。