arXiv:2602.23656cs.CLcs.AI2026-02

用专利中的技术矛盾挖掘,提升专利分析的准确性。

TRIZ-RAGNER: A Retrieval-Augmented Large Language Model for TRIZ-Aware Named Entity Recognition in Patent-Based Contradiction Mining

  • 引入TRIZ知识库增强大模型,解决专利术语歧义问题。
  • 在PaTRIZ数据集上F1达84.2%,比最强基线高7.3个百分点。
  • 适合从事专利分析、技术创新研究的研究者使用。

基于TRIZ的技术矛盾挖掘是专利分析与系统性创新的基础任务,旨在识别推动发明问题求解的改进与恶化技术参数。现有方法多依赖规则系统或传统机器学习模型,在处理复杂专利语言时面临语义模糊、领域依赖和泛化能力弱的问题。近年来,大语言模型(LLM)展现出强大的语义理解能力,但其直接用于TRIZ参数提取仍受限于幻觉和缺乏结构化TRIZ知识的支撑。为此,本文提出TRIZ-RAGNER框架,一种面向专利中技术矛盾挖掘的检索增强型大语言模型。该框架将矛盾挖掘重构为语义级命名实体识别任务,整合稠密检索、交叉编码器重排序及结构化提示机制,从专利句子中提取改进与恶化参数。通过将领域特定的TRIZ知识注入大模型推理过程,有效降低语义噪声并提升提取一致性。在PaTRIZ数据集上的实验表明,TRIZ-RAGNER显著优于传统序列标注模型和基于LLM的基线。该框架在技术矛盾对识别中达到85.6%精度、82.9%召回率和84.2% F1值,相较最强提示增强型GPT基线提升7.3个百分点,验证了检索增强式TRIZ知识注入在鲁棒且精准的专利矛盾挖掘中的有效性。

原文摘要 · Abstract (English)

TRIZ-based contradiction mining is a fundamental task in patent analysis and systematic innovation, as it enables the identification of improving and worsening technical parameters that drive inventive problem solving. However, existing approaches largely rely on rule-based systems or traditional machine learning models, which struggle with semantic ambiguity, domain dependency, and limited generalization when processing complex patent language. Recently, large language models (LLMs) have shown strong semantic understanding capabilities, yet their direct application to TRIZ parameter extraction remains challenging due to hallucination and insufficient grounding in structured TRIZ knowledge. To address these limitations, this paper proposes TRIZ-RAGNER, a retrieval-augmented large language model framework for TRIZ-aware named entity recognition in patent-based contradiction mining. TRIZ-RAGNER reformulates contradiction mining as a semantic-level NER task and integrates dense retrieval over a TRIZ knowledge base, cross-encoder reranking for context refinement, and structured LLM prompting to extract improving and worsening parameters from patent sentences. By injecting domain-specific TRIZ knowledge into the LLM reasoning process, the proposed framework effectively reduces semantic noise and improves extraction consistency. Experiments on the PaTRIZ dataset demonstrate that TRIZ-RAGNER consistently outperforms traditional sequence labeling models and LLM-based baselines. The proposed framework achieves a precision of 85.6%, a recall of 82.9%, and an F1-score of 84.2% in TRIZ contradiction pair identification. Compared with the strongest baseline using prompt-enhanced GPT, TRIZ-RAGNER yields an absolute F1-score improvement of 7.3 percentage points, confirming the effectiveness of retrieval-augmented TRIZ knowledge grounding for robust and accurate patent-based contradiction mining.

专利分析TRIZ大模型信息抽取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。