arXiv:2509.08304cs.CL2025-09被引 2

用问答集定义文本语义关系,实现可解释的文档对比

Explainable Semantic Text Relations: A Question-Answering Framework for Comparing Document Content

  • 以可回答问题集构建文本间语义关系的集合论模型
  • 通过问题集差异定位信息重叠与缺失的具体内容
  • 适合需要可解释性文档分析的研究者使用

理解两篇文本间的语义关系对信息管理至关重要,需判断内容是否完全重合、一方是否被另一方包含,或仅部分重叠且各具独特信息。除确定关系外,还需提供可解释输出,明确指出哪些信息存在、缺失或新增。本文通过文本可回答问题集(AQS)的集合论关系,正式定义文本间语义关系(STR),如等价、包含和互重叠。不同文本AQS的集合差值可作为解释工具,识别信息差异。研究构建了一个合成基准数据集,通过受控改写与刻意删减信息生成细粒度关系样本,并基于此评估多种判别与生成模型在区分文本关系类别上的表现,检验其超越表层相似性的语义理解能力。相关数据集与生成代码已公开。

原文摘要 · Abstract (English)

Understanding semantic relations between two texts is crucial for many information and document management tasks, in which one must determine whether the content fully overlaps, is completely superseded by another document, or overlaps only partially, with unique information in each. Beyond establishing this relation, it is equally important to provide explainable outputs that specify which pieces of information are present, missing, or newly added between the text pair. In this study, we formally define semantic relations between two texts through the set-theoretic relation between their respective Answerable Question Sets (AQS), the sets of questions each text can answer. Under this formulation, Semantic Text Relation (STR), such as equivalence, inclusion, and mutual overlap, becomes a well-defined set relation between the corresponding texts' AQSs. The set differences between the AQSs also serve as an explanation or diagnostic tool for identifying how the information in the texts diverges. Using this definition, we construct a synthetic benchmark that captures fine-grained informational relations through controlled paraphrasing and deliberate information removal supported by AQS manipulations. We then use this dataset to evaluate several discriminative and generative models for classifying text pairs into STR categories, assessing how well different model architectures capture semantic relations beyond surface-level similarity. We publicly release both the dataset and the data generation code to support further research.

可解释性文本对比问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。