首个融合文本与超图的基准,助力语言模型与超图学习结合研究
TAHB: A Comprehensive Benchmark for Text-Attributed Hypergraph Learning

- 构建首个公开文本属性超图基准TAHB,含10个真实世界数据集
- 实验证明语言模型增强文本语义可显著提升超图学习性能
- 适合研究超图表示学习与大模型融合的学者使用
超图能有效建模超越成对关系的高阶群体关系,而预训练语言模型(PLMs)和大语言模型(LLMs)则提供丰富的文本语义理解。然而,由于缺乏公开的文本属性超图基准,将语言模型与超图学习结合的研究仍受限。为此,我们提出TAHB(Text-Attributed Hypergraph Benchmark),首个整合超图结构与原始文本属性的公开基准。TAHB包含来自电商、学术、影视和政治网络四个领域的10个真实世界数据集,支持文本感知超图表示学习的系统评估。实验表明,TAHB保留了真实超图的关键结构特性,并一致复现了现有基准中的性能趋势。在LLM作为增强器和预测器两种设置下,实验均显示语言模型增强的文本语义能提升超图学习表现,且结构与文本信息联合使用时达到最佳预测效果。本基准为超图学习与语言模型交叉领域研究提供了基础。
原文摘要 · Abstract (English)
Hypergraphs effectively model higher-order groupwise relationships beyond pairwise interactions, while pretrained language models (PLMs) and large language models (LLMs) provide rich semantic understanding from textual attributes. However, research on combining language models with hypergraph learning remains limited due to the lack of public text-attributed hypergraph benchmarks. To address this limitation, we present TAHB (Text-Attributed Hypergraph Benchmark), the first public benchmark integrating hypergraph structures and raw textual attributes. TAHB contains 10 real-world datasets from four domains - e-commerce, academia, movies, and politics networks - enabling systematic evaluation of text-aware hypergraph representation learning. Experimental results show that TAHB preserves key structural properties of real-world hypergraphs and consistently reproduces performance tendencies observed in existing benchmarks. Furthermore, experiments under both LLM-as-Enhancer and LLM-as-Predictor settings demonstrate that LLM-enhanced textual semantics improve hypergraph learning performance, while structural and textual information jointly provide the best setting for LLM-based prediction. Our benchmark provides a foundation for future research at the intersection of hypergraph learning and language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。