arXiv:2501.15120cs.IRcs.DB2025-01

用大模型精准识别企业技术栈,助力商业决策。

Technology Mapping with Large Language Models

  • 结合大模型与语义检索,从非结构化文本中提取技术信息
  • 在多个数据集上提升技术识别准确率,显著优于传统方法
  • 适合战略分析、市场洞察和跨行业技术追踪的从业者

在快速演进的商业环境中,掌握企业使用的技术栈对于建立合作、发现市场机会和制定战略至关重要。然而,传统技术映射依赖关键词搜索,在面对海量多样数据时往往难以捕捉新兴技术。为此,我们提出STARS(语义技术与检索系统),利用大语言模型(LLMs)和Sentence-BERT,从非结构化内容中识别相关技术,构建企业完整技术画像,并根据运营重要性对技术进行排序。通过融合实体抽取与思维链提示,以及语义排名机制,STARS实现了对企业技术组合的高精度映射。实验表明,该方法显著提升了检索准确率,为跨行业技术映射提供了高效通用的解决方案。

原文摘要 · Abstract (English)

In today's fast-evolving business landscape, having insight into the technology stacks that organizations use is crucial for forging partnerships, uncovering market openings, and informing strategic choices. However, conventional technology mapping, which typically hinges on keyword searches, struggles with the sheer scale and variety of data available, often failing to capture nascent technologies. To overcome these hurdles, we present STARS (Semantic Technology and Retrieval System), a novel framework that harnesses Large Language Models (LLMs) and Sentence-BERT to pinpoint relevant technologies within unstructured content, build comprehensive company profiles, and rank each firm's technologies according to their operational importance. By integrating entity extraction with Chain-of-Thought prompting and employing semantic ranking, STARS provides a precise method for mapping corporate technology portfolios. Experimental results show that STARS markedly boosts retrieval accuracy, offering a versatile and high-performance solution for cross-industry technology mapping.

技术映射大模型应用企业分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。