对比人与AI答题能力,发现各有所长。
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
- 用项目反应理论量化评估人机问答能力差异
- 人类在概念推理上胜过AI,AI在查信息上更强
- 适合关注人机协作或认知模型的研究者
大型语言模型(LLMs)的进展引发关于其在自然语言处理任务中超越人类的讨论。本文提出CAIMIRA框架,基于项目反应理论(IRT),对问答(QA)代理——人类与AI系统——的问题解决能力进行量化评估。通过分析超过30万条来自约70个AI系统和155名人类在数千道测验题上的回答,发现两者在知识领域与推理技能上表现出不同优势:人类在基于知识的溯因推理和概念推理上表现更优;而如GPT-4、LLaMA等先进大模型在目标明确的信息检索与事实推理任务中更具优势,尤其当信息缺口可通过模式匹配或数据检索填补时。研究提示未来问答任务应聚焦于挑战高阶推理、科学思维以及微妙语言理解与跨情境知识应用的问题,以推动更贴近或互补人类认知能力的AI发展。
原文摘要 · Abstract (English)
Recent advancements of large language models (LLMs) have led to claims of AI surpassing humans in natural language processing (NLP) tasks such as textual understanding and reasoning. This work investigates these assertions by introducing CAIMIRA, a novel framework rooted in item response theory (IRT) that enables quantitative assessment and comparison of problem-solving abilities of question-answering (QA) agents: humans and AI systems. Through analysis of over 300,000 responses from ~70 AI systems and 155 humans across thousands of quiz questions, CAIMIRA uncovers distinct proficiency patterns in knowledge domains and reasoning skills. Humans outperform AI systems in knowledge-grounded abductive and conceptual reasoning, while state-of-the-art LLMs like GPT-4 and LLaMA show superior performance on targeted information retrieval and fact-based reasoning, particularly when information gaps are well-defined and addressable through pattern matching or data retrieval. These findings highlight the need for future QA tasks to focus on questions that challenge not only higher-order reasoning and scientific thinking, but also demand nuanced linguistic interpretation and cross-contextual knowledge application, helping advance AI developments that better emulate or complement human cognitive abilities in real-world problem-solving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。