提升多语言大模型在等价与继承推理中的一致性
Towards Robust Knowledge Representations in Multilingual LLMs for Equivalence and Inheritance based Consistent Reasoning
- 用组合表示法让跨语言词汇保持等价关系
- 多语言测试中错误率最高达57.5%,继承约束违反率达37.2%
- 适合关注多语言推理一致性的研究者
推理与语言能力是人类智能的核心,支撑问题解决与决策。近期大型语言模型(LLMs)展现出卓越的语言能力与涌现的推理行为,推动其在多个领域的广泛应用。然而,当前模型在复杂推理任务上仍表现不佳,暴露出系统性局限。本文聚焦评估模型是否具备基于“等价”与“继承”两类基础关系进行推理所需的表征能力。我们设计了涵盖六种语言的新任务与基准测试,发现当前顶尖模型在不同语言间对同一问题的回答存在冲突的情况占比17.3%-57.5%,且最多有37.2%的案例违反继承约束。为提升跨语言一致性,我们提出“组合表示”方法,将词元表示为跨语言等价词元的组合,实验显示该方法可使冲突率降低最高达4.7%,验证了共享表征的有效性。
原文摘要 · Abstract (English)
Reasoning and linguistic skills form the cornerstone of human intelligence, facilitating problem-solving and decision-making. Recent advances in Large Language Models (LLMs) have led to impressive linguistic capabilities and emergent reasoning behaviors, fueling widespread adoption across application domains. However, LLMs still struggle with complex reasoning tasks, highlighting their systemic limitations. In this work, we focus on evaluating whether LLMs have the requisite representations to reason using two foundational relationships: "equivalence" and "inheritance". We introduce novel tasks and benchmarks spanning six languages and observe that current SOTA LLMs often produce conflicting answers to the same questions across languages in 17.3-57.5% of cases and violate inheritance constraints in up to 37.2% cases. To enhance consistency across languages, we propose novel "Compositional Representations" where tokens are represented as composition of equivalent tokens across languages, with resulting conflict reduction (up to -4.7%) indicating benefits of shared LLM representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。