arXiv:2508.02430cs.CL2025-08

用大模型自动评估创新,替代人工专家判断。

AI-Based Measurement of Innovation: Mapping Expert Insight into Large Language Model Applications

  • 设计框架从非结构化文本中模拟专家对创新的判断。
  • 在软件更新和用户反馈中验证,准确率高于现有方法。
  • 适合企业研发、学术研究及论文评审人员使用。

创新测量常依赖特定场景的代理指标和专家评估,导致实证研究受限于数据可得性。本文探讨如何利用大语言模型(LLM)克服人工评估的局限,辅助创新测量。我们设计了一个LLM框架,能从非结构化文本中可靠地近似领域专家对创新的评价。通过两个不同情境的研究验证:(1) 软件应用更新的创新性;(2) 产品评论中用户生成反馈与改进建议的原创性。对比了该框架与以往研究中常用方法及先进机器学习/深度学习模型的性能(F1分数)与可靠性(一致性率)。结果表明,该框架在F1得分上优于其他方法,且结果高度一致(跨运行无变化)。本文为企业的研发人员、研究人员、审稿人和编辑提供了有效使用LLM测量创新的知识与工具,并讨论了模型选择、提示工程、训练数据规模与分布、参数设置等关键设计决策对性能与可靠性的影。鉴于人工评估与现有文本度量方法的固有挑战,本框架为将大模型作为可靠、日益可及且广泛适用的创新测量工具提供了重要启示。

原文摘要 · Abstract (English)

Measuring innovation often relies on context-specific proxies and on expert evaluation. Hence, empirical innovation research is often limited to settings where such data is available. We investigate how large language models (LLMs) can be leveraged to overcome the constraints of manual expert evaluations and assist researchers in measuring innovation. We design an LLM framework that reliably approximates domain experts' assessment of innovation from unstructured text data. We demonstrate the performance and broad applicability of this framework through two studies in different contexts: (1) the innovativeness of software application updates and (2) the originality of user-generated feedback and improvement ideas in product reviews. We compared the performance (F1-score) and reliability (consistency rate) of our LLM framework against alternative measures used in prior innovation studies, and to state-of-the-art machine learning- and deep learning-based models. The LLM framework achieved higher F1-scores than the other approaches, and its results are highly consistent (i.e., results do not change across runs). This article equips R&D personnel in firms, as well as researchers, reviewers, and editors, with the knowledge and tools to effectively use LLMs for measuring innovation and evaluating the performance of LLM-based innovation measures. In doing so, we discuss, the impact of important design decisions-including model selection, prompt engineering, training data size, training data distribution, and parameter settings-on performance and reliability. Given the challenges inherent in using human expert evaluation and existing text-based measures, our framework has important implications for harnessing LLMs as reliable, increasingly accessible, and broadly applicable research tools for measuring innovation.

创新测量大模型应用文本分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。