首个评估大模型多示例模式识别能力的基准,突破传统少样本局限。
MIR-Bench: Can Your LLM Recognize Complicated Patterns via Many-Shot In-Context Reasoning?
- 设计多示例上下文学习基准,支持数百到上千个示例输入
- 发现模型在复杂模式识别中存在显著缩放效应与泛化能力差异
- 适合研究长上下文推理、归纳/演绎推理及RAG系统的学者
模式识别与迁移是通用智能的核心能力,现有大语言模型评测多聚焦于少样本(通常<10)场景,缺乏对长上下文信息整合能力的评估。随着大模型上下文长度持续增长,多示例上下文学习(many-shot ICL)成为无需微调即可处理新任务的新范式,但当前评估多集中于分类任务,而如针堆找针(NIAH)等常见长上下文任务并不要求复杂信息融合。为此,我们提出MIR-Bench——首个面向多示例上下文推理的模式识别基准,要求模型通过多样数据格式的输入-输出示例,推断底层函数关系。基于此,我们系统研究了多示例推理中的缩放效应、鲁棒性、归纳与演绎推理差异、检索增强生成(RAG)、编码对归纳推理的作用以及跨领域泛化等新问题,获得多项深刻洞见。
原文摘要 · Abstract (English)
The ability to recognize patterns from examples and apply them to new ones is a primal ability for general intelligence, and is widely studied by psychology and AI researchers. Many benchmarks have been proposed to measure such ability for Large Language Models (LLMs); however, they focus on few-shot (usually <10) setting and lack evaluation for aggregating many pieces of information from long contexts. On the other hand, the ever-growing context length of LLMs have brought forth the novel paradigm of many-shot In-Context Learning (ICL), which addresses new tasks with hundreds to thousands of examples without expensive and inefficient fine-tuning. However, many-shot evaluations often focus on classification, and popular long-context LLM tasks such as Needle-In-A-Haystack (NIAH) seldom require complicated intelligence for integrating many pieces of information. To fix the issues from both worlds, we propose MIR-Bench, the first many-shot in-context reasoning benchmark for pattern recognition that asks LLM to predict output via input-output examples from underlying functions with diverse data format. Based on MIR-Bench, we study many novel problems for many-shot in-context reasoning, and acquired many insightful findings including scaling effect, robustness, inductive vs. transductive reasoning, retrieval Augmented Generation (RAG), coding for inductive reasoning, cross-domain generalizability, etc.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。