个性化评估模型比平均共识更准,适合商业创意评价
Aggregate vs. Personalized Judges in Business Idea Evaluation: Evidence from Expert Disagreement

- 用专家评分历史定制评估模型,而非取平均意见
- 个性化模型与专家打分相关性高出30%以上
- 适合需要精准匹配专家偏好的商业创意评审场景
评估大模型生成的商业创意比生成本身更难规模化。不同于标准NLP基准,商业创意评估涉及可行性、新颖性、差异化、用户需求和市场规模等多维标准,且专家意见常不一致。本文构建了PBIG-DATA数据集,包含300个基于专利的创意,由领域专家在6个商业维度(具体性、技术合理性、创新性、竞争优势、需求有效性、市场规模)上给出约3,000个个体评分。分析显示,细粒度序数评分存在显著专家分歧,粗粒度筛选时一致性更高,表明分歧具有结构特征而非随机噪声。比较三种评估配置:仅依赖规则的零样本模型、基于混合评估者历史的聚合模型、以及基于目标评估者历史的个性化模型。在不同维度和模型规模下,个性化模型与对应评估者的打分一致性更高,且评估者间同意率仅在个性化条件下与模型推理相似性相关。结果表明,在多元评价环境中,聚合标签可能脆弱,应采用评估者条件化的设计。
原文摘要 · Abstract (English)
Evaluating LLM-generated business ideas is often harder to scale than generating them. Unlike standard NLP benchmarks, business idea evaluation relies on multi-dimensional criteria such as feasibility, novelty, differentiation, user need, and market size, and expert judgments often disagree. This paper studies a methodological question raised by such disagreement: should an automatic judge approximate an aggregate consensus, or model evaluators individually? We introduce PBIG-DATA, a dataset of approximately 3,000 individual scores across 300 patent-grounded product ideas, provided by domain experts on six business-oriented dimensions: specificity, technical validity, innovativeness, competitive advantage, need validity, and market size. Analyses show substantial expert disagreement on fine-grained ordinal scores, while agreement is higher under coarse selection, suggesting structured heterogeneity rather than random noise. We then compare three judge configurations: a rubric-only zero-shot judge, an aggregate judge conditioned on mixed evaluator histories, and a personalized judge conditioned on the target evaluator's scoring history. Across dimensions and model sizes, personalized judges align more closely with the corresponding evaluator than aggregate judges, and evaluator agreement correlates with similarity of judge-generated reasoning only under personalized conditioning. These results indicate that pooled labels can be a fragile target in pluralistic evaluation settings and motivate evaluator-conditioned judge designs for business idea assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。