arXiv:2502.13076cs.CL2025-02

用关键词画像提升专利分析效果,让技术洞察更清晰

KAPPA: A Generic Patent Analysis Framework with Keyphrase-Based Portraits

  • 基于提示的分层解码生成专利关键词,融合语义结构特征
  • 在基准数据集上超越现有方法,显著提升关键词生成效果
  • 适合专利分析、技术情报和知识产权研究者使用

专利分析高度依赖简洁且可解释的文档表示,即专利画像。关键词(包括存在与缺失的)因其简明、代表性与清晰性,是构建专利画像的理想选择。本文提出KAPPA框架,通过两阶段实现关键词驱动的专利画像构建与分析:第一阶段采用语义校准的关键词生成范式,结合预训练语言模型与基于提示的分层解码策略,利用专利的多层级结构特征;第二阶段构建基于画像的分析框架,实现高效精准的专利分析。在基准关键词生成数据集上的实验表明,该模型显著优于现有先进方法。真实专利申请案例的进一步实验验证了关键词画像能有效捕捉领域知识,并增强专利分析的语义表达能力。

原文摘要 · Abstract (English)

Patent analysis highly relies on concise and interpretable document representations, referred to as patent portraits. Keyphrases, both present and absent, are ideal candidates for patent portraits due to their brevity, representativeness, and clarity. In this paper, we introduce KAPPA, an integrated framework designed to construct keyphrase-based patent portraits and enhance patent analysis. KAPPA operates in two phases: patent portrait construction and portrait-based analysis. To ensure effective portrait construction, we propose a semantic-calibrated keyphrase generation paradigm that integrates pre-trained language models with a prompt-based hierarchical decoding strategy to leverage the multi-level structural characteristics of patents. For portrait-based analysis, we develop a comprehensive framework that employs keyphrase-based patent portraits to enable efficient and accurate patent analysis. Extensive experiments on benchmark datasets of keyphrase generation, the proposed model achieves significant improvements compared to state-of-the-art baselines. Further experiments conducted on real-world patent applications demonstrate that our keyphrase-based portraits effectively capture domain-specific knowledge and enrich semantic representation for patent analysis tasks.

专利分析关键词生成语义表示AI辅助创新

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。