arXiv:2608.29529cs.CLcs.AI2026-08

用论证结构提升规范文本对齐准确率,让机器更懂标准背后的逻辑。

Argument-Aware Semantic Alignment of Normative Texts: A Toulmin-Based Neuro-Symbolic Approach

论文配图:Argument-Aware Semantic Alignment of Normative Texts: A Toulmin-Based Neuro-Symbolic Approach
图 1 · 摘自论文原文
  • 结合神经网络与图灵论证模型,提取主张、论据、理由等逻辑要素
  • 在NERC-CIP到NIST-CSF映射任务中,比纯语义方法提升对齐效果
  • 只需主张-论据-理由三要素,即可保持良好性能,适合安全标准对齐

当规范性文本中的等效要求使用不同术语、语法和抽象层次时,语义对齐极具挑战。词法重叠、分布表示和语义相似性虽能捕捉主题相关性,但常忽略规范性主张的论证结构——即主张如何被支持、限定与证明。本文探究显式论证结构是否可为跨标准对齐提供互补信息。将跨标准控制映射视为论证感知的语义对齐问题,构建了一个融合神经文本表征与图灵特征的神经符号管道。通过大语言模型显式化识别主张、论据、理由、限定词与支撑,并重构省略推理(enthymemes)。这些信息以论证感知相似性和结构特征输入对齐模型。在NERC-CIP至NIST-CSF映射基准测试中,基于论证的特征优于神经符号语义基线。特征选择显示理由类特征信号最强,表明主张与其支持理由之间的关联无法由传统相似度捕获。一个仅包含主张—论据—理由的精简子集仍可媲美完整图灵特征集。结果初步表明,论证结构是规范性文本对齐的有用中间表示。以网络安全标准为受控测试平台,不证明其跨领域泛化能力。大语言模型生成的论证图也可用于后续检索、推理与解释任务。

原文摘要 · Abstract (English)

Semantic alignment between specialized normative texts is challenging when equivalent requirements use different terms, syntax, and levels of abstraction. Lexical overlap, distributional embeddings, and semantic similarity capture topical relatedness but often miss the argumentative structure by which normative claims are supported, qualified, and justified. This paper asks whether explicit argument structure adds information complementary to neural semantics for aligning requirements. We treat cross-standard control mapping as argument-aware semantic alignment and build a neuro-symbolic pipeline that combines neural text representations with Toulmin features. An LLM explicitation step identifies claims, grounds, warrants, qualifiers, and backing and reconstructs enthymemes. These feed an alignment model via argument-aware similarity and structural features. On a NERC-CIP to NIST-CSF mapping benchmark, argument-derived features improve alignment over a neuro-symbolic semantic baseline. Feature selection shows especially strong signal from warrant-related features, indicating that the link between a claim and its supporting reasoning is not captured by conventional similarity alone. A compact claim--grounds--warrant subset remains competitive with the full Toulmin feature set. The results give preliminary evidence that argument structure is a useful intermediate representation for aligning specialized normative texts. Cybersecurity standards are used as a controlled testbed, not as proof of domain-independent generalization. The argument graphs produced by LLM explicitation may also support later work on retrieval, reasoning, and explanation over normative text.

规范文本对齐论证结构神经符号安全标准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。