arXiv:2508.18444cs.CLcs.AI2025-08被引 1

探究不同训练方法对大模型重排推理可靠性的影响

How Reliable are LLMs for Reasoning on the Re-ranking task?

  • 用小规模地球科学数据集测试大模型重排能力
  • 发现部分训练方法虽提升效果但缺乏真实语义理解
  • 适合关注模型可解释性与小样本场景的研究者

随着大语言模型(LLMs)语义理解能力的提升,其在重排任务中展现出更强的人类价值观对齐性,但透明度下降。尽管实验结果令人鼓舞,但要理解模型内部运作机制以解释重排决策仍至关重要。尤其在用户参与少、排名数据不足的新系统中,准确重排仍是挑战。我们分析发现,不同训练方法对语义理解影响显著,部分方法仅优化评估指标而未获得真实知识,引发对模型可靠性的质疑。为此,本研究在小规模地球科学领域数据集上评估多种训练方法对重排任务的影响,考察模型是否能生成可信的文本推理,以应对透明性缺失和数据稀缺问题。

原文摘要 · Abstract (English)

With the improving semantic understanding capability of Large Language Models (LLMs), they exhibit a greater awareness and alignment with human values, but this comes at the cost of transparency. Although promising results are achieved via experimental analysis, an in-depth understanding of the LLM's internal workings is unavoidable to comprehend the reasoning behind the re-ranking, which provides end users with an explanation that enables them to make an informed decision. Moreover, in newly developed systems with limited user engagement and insufficient ranking data, accurately re-ranking content remains a significant challenge. While various training methods affect the training of LLMs and generate inference, our analysis has found that some training methods exhibit better explainability than others, implying that an accurate semantic understanding has not been learned through all training methods; instead, abstract knowledge has been gained to optimize evaluation, which raises questions about the true reliability of LLMs. Therefore, in this work, we analyze how different training methods affect the semantic understanding of the re-ranking task in LLMs and investigate whether these models can generate more informed textual reasoning to overcome the challenges of transparency or LLMs and limited training data. To analyze the LLMs for re-ranking tasks, we utilize a relatively small ranking dataset from the environment and the Earth science domain to re-rank retrieved content. Furthermore, we also analyze the explainable information to see if the re-ranking can be reasoned using explainability.

大模型推理重排任务可解释性小样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。