用多维主题模型发现专利空白区,精准定位未开发但有价值的创新方向。
BLANC: Discovering Patent White Space via Changes in Normalized Pointwise Mutual Information Between Multi-View Clusters

- 通过三个维度的神经主题建模构建专利多视图表征
- 引入ΔNPMI指标,在关键词过滤后检测组合关联性骤降的空白区
- 在真实案例中成功识别专家确认的潜在创新方向,适合研发战略规划
识别专利空白区——即尚未探索但具有潜力的技术领域——对战略性研发规划至关重要。现有方法依赖人工映射或单视图聚类,缺乏量化缺口检测能力。本文提出BLANC(基于NPMI条件分析的空白区发现),分三阶段:(1) 从应用/用途、新颖性、创造性步骤三个语义维度进行多视图神经主题建模;(2) 使用归一化点互信息(NPMI)量化跨维度聚类关联;(3) 通过条件检测,标记在用户指定关键词过滤后NPMI下降的组合。该下降由新指标ΔNPMI捕捉,用于识别‘全局已建立、局部未探索’的组合。由于空白区无真值标签,我们在两个公开美国专利库上评估:机器学习/AI(5,417件,CPC G06N)和玻璃组成(1,982件,CPC C03C)。通过人为剔除目标组合的75%文档,BLANC分别恢复了34.1%(ML/AI)和27.3%(玻璃)的组合,而随机剔除或不同组合的剔除实验均未恢复,191次干扰测试无一成功。将三视图合并为单一视图则无法恢复任何组合,而传统共现度量在随机剔除下仍会误报,缺乏特异性。在私有案例中(302件浮法玻璃/玻璃陶瓷专利),关键词‘氟’揭示了‘氟表面处理×翘曲抑制’这一候选组合(ΔNPMI最高达0.48),与专家独立发现一致。
原文摘要 · Abstract (English)
Identifying white space --- the unexplored but potentially valuable regions of a patent landscape --- is essential for strategic R&D planning, yet existing methods rely on manual patent mapping or apply single-view clustering without quantitative gap detection. We propose BLANC (Blank Landscape Analysis through NPMI Conditioning), a three-phase pipeline combining (1) multi-view neural topic modeling along three semantic dimensions (application/use, novelty, inventive step); (2) Normalized Pointwise Mutual Information (NPMI) to quantify cross-dimensional cluster association; and (3) conditional detection that flags combinations whose NPMI drops when the corpus is filtered by a user-specified keyword. The drop is captured by a new metric, $Δ$NPMI, which identifies combinations "established globally, unexplored locally." Because white space has no ground truth, we evaluate BLANC on two public USPTO corpora --- machine learning/AI (5,417 patents, CPC G06N) and glass compositions (1,982 patents, CPC C03C) --- by artificially depleting known technology combinations and testing recovery. When three-quarters of a target pair's documents are removed, BLANC recovers 34.1% (ML/AI) and 27.3% (glass) of the depleted combinations, whereas size-matched removals not aimed at them (random documents, or those of a different established combination) essentially never do: the target is never recovered in 191 decoy trials. Collapsing the three semantic views into one recovers nothing, while prior co-occurrence measures also flag the target under random removal, offering no specificity. In a proprietary case (302 float glass / glass-ceramics patents), the keyword "fluorine" reveals a fluorine surface treatment $\times$ warpage suppression candidate ($Δ$NPMI up to 0.48) that experts had independently identified.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。