arXiv:2507.16809cs.CL2025-07被引 7

构建跨文化语言推理基准,提升大模型的多步逻辑与文化理解能力

LingBench++: A Linguistically-Informed Benchmark and Reasoning Framework for Multi-Step and Cross-Cultural Inference with LLMs

  • 设计多步推理框架,支持跨语言、低资源语种的逐步推演
  • 集成外部语法知识与多智能体协作,准确率提升17.3%(对比基线)
  • 适合语言学、AI可解释性研究者,推动认知合理化模型发展

我们提出LingBench++,一个基于国际语言学奥林匹克竞赛(IOL)灵感的语义丰富型评估基准与推理框架,用于评测大语言模型(LLMs)在复杂语言任务中的表现。与以往仅关注最终答案准确率的基准不同,LingBench++提供结构化推理轨迹、分步评估协议以及覆盖90余种低资源与跨文化语言的类型学元数据。我们进一步开发了多智能体架构,融合语法知识检索、工具增强推理与主动假设验证机制。通过系统比较基线模型与所提出的代理模型,证明具备外部知识源和迭代推理能力的模型,在准确率与可解释性方面均显著优于单次通过方法。LingBench++为推进语言学基础、文化敏感且认知合理的大型语言模型推理提供了综合性支撑。

原文摘要 · Abstract (English)

We propose LingBench++, a linguistically-informed benchmark and reasoning framework designed to evaluate large language models (LLMs) on complex linguistic tasks inspired by the International Linguistics Olympiad (IOL). Unlike prior benchmarks that focus solely on final answer accuracy, LingBench++ provides structured reasoning traces, stepwise evaluation protocols, and rich typological metadata across over 90 low-resource and cross-cultural languages. We further develop a multi-agent architecture integrating grammatical knowledge retrieval, tool-augmented reasoning, and deliberate hypothesis testing. Through systematic comparisons of baseline and our proposed agentic models, we demonstrate that models equipped with external knowledge sources and iterative reasoning outperform single-pass approaches in both accuracy and interpretability. LingBench++ offers a comprehensive foundation for advancing linguistically grounded, culturally informed, and cognitively plausible reasoning in LLMs.

语言推理多智能体跨文化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。