小模型微调后在图结构推理上表现稳定,能泛化到更大更陌生的图。
Generalization Boundaries of Fine-Tuned Small Language Models for Graph Structural Inference
- 用指令微调的小模型评估图结构推理能力,测试超训练范围的图规模和类型
- 模型对不同图家族保持强排序一致性,即使输入远大于训练数据仍有效
- 不同模型架构有各自退化模式,揭示了可靠泛化的边界
微调后的小语言模型在图属性估计任务中表现出色,但在训练条件外的泛化能力尚不明确。本文系统研究微调小模型在图规模和图家族分布两个维度上的泛化边界,并在真实世界图基准上评估其领域学习能力。采用三款3-4B参数量的指令微调模型与两种图序列化格式,在控制实验环境下评估模型对显著超出训练范围的图以及未见随机图家族的表现。结果表明,微调模型在结构差异大的图家族间保持强序关系,即使输入图远大于训练数据仍能正确排序图的结构属性,且各模型呈现特定的架构退化特征。这些发现明确了微调小语言模型在图推理任务中的可靠泛化范围,为实际应用提供实证依据。
原文摘要 · Abstract (English)
Small language models fine-tuned for graph property estimation have demonstrated strong in-distribution performance, yet their generalization capabilities beyond training conditions remain poorly understood. In this work, we systematically investigate the boundaries of structural inference in fine-tuned small language models along two generalization axes - graph size and graph family distribution - and assess domain-learning capability on real-world graph benchmarks. Using a controlled experimental setup with three instruction-tuned models in the 3-4B parameter class and two graph serialization formats, we evaluate performance on graphs substantially larger than the training range and across held-out random graph families. Our results show that fine-tuned models maintain strong ordinal consistency across structurally distinct graph families and continue to rank graphs by structural properties on inputs substantially larger than those seen during training, with distinct architecture-specific degradation profiles. These findings delineate where fine-tuned small language models generalize reliably, providing empirical grounding for their use in graph-based reasoning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。