用句子收敛性构建文本,能显著提升大模型答推理题准确率
Context Convergence Improves Answering Inferential Questions
- 以句子间消除错误答案的能力(收敛性)选句构造问答文本
- 高收敛性句子组合使模型答对率明显高于相似度选句方法
- 按收敛性降序排列句子可微调性能,适合研究模型推理机制
尽管大语言模型广泛用于开放域问答,其处理需推断的问答题(答案需推理得出而非直接检索)的能力仍不充分。本文研究段落结构与质量对模型在推理类问题上表现的影响,聚焦‘收敛性’——即句子帮助排除错误答案的有效程度——作为段落构建标准。基于TriviaHG数据集子集,我们使用不同收敛性水平的句子组合成段落,并评估六种不同规模与架构的大模型。结果表明,由高收敛性句子构成的段落显著提升答案准确率,优于基于余弦相似度选择的段落;且按收敛性降序排列句子可小幅提升性能,表明模型更依赖早期信息丰富的线索。这些发现凸显收敛性作为指导段落构建与分析大模型推理行为的有效信号。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) are widely used in open-domain Question Answering (QA), their ability to handle inferential questions-where answers must be derived rather than directly retrieved-remains still underexplored. This study investigates how the structure and quality of passages influence LLM performance on such questions. We focus on convergence, a measure of how effectively sentences (hints) eliminate incorrect answers, as a criterion for constructing passages. Using subsets of the TriviaHG dataset, we form passages by combining sentences with varying convergence levels and evaluate six LLMs of different sizes and architectures. Our results show that passages built from higher convergence sentences lead to substantially better answer accuracy than those selected by cosine similarity, indicating that convergence captures meaningful relevance for inferential reasoning. Additionally, ordering sentences by descending convergence slightly improves performance, suggesting that LLMs tend to prioritize earlier, information-rich cues. These findings highlight convergence as a practical signal for guiding passage construction and analyzing inferential reasoning behavior in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。