用数学结构让任意文本集可自由跳转浏览,方法简单但效果不输大模型。
Make Any Collection Navigable: Methods for Constructing and Evaluating Hypergraph of Text
- 用文本相似度构建超图结构,实现跨文档灵活跳转。
- 提出新评估指标‘努力比’,发现简单方法能媲美大模型。
- 适合想提升文本检索与导航体验的研究者和开发者。
网页之所以比纯文档集合更易用,是因为超链接构建的结构支持灵活跳转。然而超链接通常需人工创建,难以捕捉语料的隐含语义结构。是否存在一种通用方法,使任意文本集合都可导航?近期研究将此问题形式化为构建文本超图(Hypergraph of Text, HoT),提供支持跳转与浏览的形式化数学结构。但如何构造和评估HoT仍是挑战。本文提出并研究多种构建HoT的方法,还提出一种新的量化评估指标——努力比(effort ratio),用于衡量构建的HoT结构质量。实验结果表明,即使简单的TF-IDF基线方法,在该指标上也能与基于大语言模型的方法相媲美。
原文摘要 · Abstract (English)
One reason the Web is more useful than a simple collection of documents is that the structure created by hyperlinks enables flexible navigation from one web page to another. However, hyperlinks are typically created manually and cannot fully capture a corpus' implicit semantic structures. Is there a general way to make an arbitrary collection navigable? Recent work has formalized this problem generally as constructing a Hypergraph of Text (HoT), which provides a formal mathematical structure for supporting navigation and browsing. However, how to construct and evaluate a Hypergraph of Text remains a challenge. In this paper, we propose and study several methods for constructing a HoT. We also propose a novel quantitative metric, effort ratio, for evaluating the structural quality of a constructed HoT. Experimental results show that even simple TF-IDF baselines can match LLM-based methods on our proposed effort ratio metric.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。