构建首个统一评估文本图学习鲁棒性的基准,覆盖9种真实数据退化场景。
OpenRTAG: A Comprehensive Benchmark for Robust Text-Attributed Graph Learning under Data Quality Degradation

- 提出3×3分类体系,系统归类文本图的结构、文本、标签三方面退化
- 在9个数据集上验证模型对噪声、稀疏、不平衡等退化的敏感性
- 适合研究图神经网络鲁棒性或实际部署中数据质量差的场景
文本属性图(TAGs)结合了关系结构与丰富节点文本,但真实场景中常存在文本、结构、标签三方面的质量问题,表现为稀疏、噪声和不平衡。这些因素共形成九种典型退化场景,严重影响图学习效果。现有研究多针对特定退化类型,缺乏跨任务、数据集与模型的统一评估。为此,本文提出OpenRTAG,一个面向文本属性图学习的鲁棒性基准。该基准采用统一的3×3退化分类体系,支持在九个主流数据集和三个下游任务上进行标准化评测,系统评估各类退化场景的有效性与模型敏感度,对比传统GNN、LLM-GNN与代表性图生成模型(GFM),分析不同场景下基线方法的效率、有效性与鲁棒性,并探究复合退化下的模型行为。OpenRTAG为理解真实低质量环境下文本图学习的鲁棒性提供了标准化测试平台。
原文摘要 · Abstract (English)
Text-attributed graphs (TAGs) are an important graph data form that combine relational structure with rich node text. However, real-world TAGs are often imperfect, with quality issues arising from text, structure, and labels, and typically manifesting as sparsity, noise, and imbalance. These dimensions define nine representative degradation scenarios that can substantially affect TAG learning. Although prior studies have explored specific mitigation strategies, existing evidence remains fragmented across degradation types, datasets, tasks, and model families, leaving TAG robustness insufficiently understood. To address this gap, we present OpenRTAG, a robustness benchmark for text-attributed graph learning. OpenRTAG organizes TAG quality issues into a unified 3 * 3 taxonomy and supports standardized evaluation across nine TAG datasets and three downstream tasks. It systematically evaluates scenario validity and model sensitivity, compares traditional GNNs, LLM-GNNs, and a representative GFM, investigates the effectiveness, efficiency, and robustness of scenario-matched baselines, and further examines model behavior under composite degradation scenarios. OpenRTAG provides a standardized testbed for understanding robustness in TAG learning under realistic low-quality settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。