arXiv:2511.10234cs.LGcs.AI2025-11被引 3

LLM图推理对节点重编号等变化不鲁棒,大模型更抗干扰

Lost in Serialization: Invariance and Generalization of LLM Graph Reasoners

  • 分解图序列化为节点标签、边编码和语法三部分分析鲁棒性
  • 未微调的大模型对节点重编号更鲁棒,微调反而增加格式敏感度
  • 提出谱任务评估模型泛化能力,适合关注图推理稳定性的研究者

基于大语言模型(LLMs)的图推理方法缺乏对图表示对称性的内在不变性。由于在序列化图上运行,当节点重新索引、边顺序改变或格式调整时,LLM可能产生不同输出,引发鲁棒性问题。我们系统分析了这些影响,研究微调如何改变编码敏感性以及在未见任务上的泛化能力。提出将图序列化分解为节点标签、边编码和语法三部分,并在全面基准测试套件上评估模型对各因素变化的鲁棒性。同时引入一组新的谱任务以进一步检验微调推理器的泛化能力。结果表明,更大的非微调模型更具鲁棒性;微调可降低对节点重命名的敏感性,但可能增加对结构与格式变化的敏感性,且并未一致提升未见任务表现。

原文摘要 · Abstract (English)

While promising, graph reasoners based on Large Language Models (LLMs) lack built-in invariance to symmetries in graph representations. Operating on sequential graph serializations, LLMs can produce different outputs under node reindexing, edge reordering, or formatting changes, raising robustness concerns. We systematically analyze these effects, studying how fine-tuning impacts encoding sensitivity as well generalization on unseen tasks. We propose a principled decomposition of graph serializations into node labeling, edge encoding, and syntax, and evaluate LLM robustness to variations of each of these factors on a comprehensive benchmarking suite. We also contribute a novel set of spectral tasks to further assess generalization abilities of fine-tuned reasoners. Results show that larger (non-fine-tuned) models are more robust. Fine-tuning reduces sensitivity to node relabeling but may increase it to variations in structure and format, while it does not consistently improve performance on unseen tasks.

图神经网络大模型鲁棒性推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。