用AI自动评估创意新颖性,效率远超人工。
MuseScorer: Idea Originality Scoring At Scale
- 通过大模型与检索结合,自动识别相似创意并分组。
- 在1.6万条创意上,评分与人工一致度达0.89。
- 适合大规模创意研究,结果可信且符合人类判断。
衡量创意新颖性的客观方法是统计其在群体中的稀有程度,该方法在创造力研究中已有长期应用。然而,计算频率需手动对创意重述进行归类,过程主观、耗时、易错且难以规模化。我们提出MuseScorer,一个完全自动化、经过心理测量验证的基于频率的新颖性评分系统。该系统结合大语言模型(LLM)与外部检索:给定新创意,先检索语义相似的过往创意分组,再零样本提示LLM判断该创意是否属于现有分组或形成新组。这些分组支持无需人工标注的频率式新颖性评分。在五个数据集(共1143名参与者,16,294条创意)上,MuseScorer在创意聚类结构(AMI=0.59)和个体评分(r=0.89)方面与人工标注高度一致,并展现出良好的收敛效度与外部效度。该系统为创造力研究提供了可扩展、意图敏感且与人类对齐的新颖性评估能力。
原文摘要 · Abstract (English)
An objective, face-valid method for scoring idea originality is to measure each idea's statistical infrequency within a population -- an approach long used in creativity research. Yet, computing these frequencies requires manually bucketing idea rephrasings, a process that is subjective, labor-intensive, error-prone, and brittle at scale. We introduce MuseScorer, a fully automated, psychometrically validated system for frequency-based originality scoring. MuseScorer integrates a Large Language Model (LLM) with externally orchestrated retrieval: given a new idea, it retrieves semantically similar prior idea-buckets and zero-shot prompts the LLM to judge whether the idea fits an existing bucket or forms a new one. These buckets enable frequency-based originality scoring without human annotation. Across five datasets N_{participants}=1143, n_{ideas}=16,294), MuseScorer matches human annotators in idea clustering structure (AMI = 0.59) and participant-level scoring (r = 0.89), while demonstrating strong convergent and external validity. The system enables scalable, intent-sensitive, and human-aligned originality assessment for creativity research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。