用AI从球迷文字描述预测体验评分,准确率达67%。
LLM Predictive Scoring and Validation: Inferring Experience Ratings from Unstructured Text
- 仅凭一段开放文本,用GPT-4.1预测粉丝体验评分。
- 67%预测值与真实评分相差不超过1分,36%完全一致。
- 预测值反映关键情绪时刻,真实值包含整体评价,各有信息价值。
我们让GPT-4.1阅读棒球粉丝对观赛体验的开放式描述,并预测其在0-10分量表上的总体评分。模型仅接收单条文本回复,预测结果与来自五支大联盟球队约10,000条粉丝反馈的真实评分进行对比。结果显示,三分之二的预测评分与实际评分相差不超过1分(67%在±1内,36%完全匹配),且三次独立评分运行中预测一致性极高(87%完全一致,99.9%在±1内)。预测评分与整体体验评分相关性最强(r = 0.82),而非特定环节如停车、餐饮或工作人员表现。但预测值普遍比自评低约1分,且该偏差并非由单一因素导致。分析表明,自评体现的是对整个体验的整体判断,而预测值则聚焦于令人难忘、情感强烈、异常或可行动的关键片段。两者各有侧重,差异反映构念差异,不应视为误差消除,而是具有保留价值的测量差异。
原文摘要 · Abstract (English)
We tasked GPT-4.1 to read what baseball fans wrote about their game-day experience and predict the overall experience rating each fan gave on a 0-10 survey scale. The model received only the text of a single open-ended response. These AI predictions were compared with the actual experience ratings captured by the survey instrument across approximately 10,000 fan responses from five Major League Baseball teams. In total two-thirds of predicted ratings fell within one point of self-reported fan ratings (67% within +/-1, 36% exact match), and the predicted measurement was near-deterministic across three independent scoring runs (87% exact agreement, 99.9% within +/-1). Predicted ratings aligned most strongly with the overall experience rating (r = 0.82) rather than with any specific aspect of the game-day experience such as parking, concessions, staff, etc. However, predictions were systematically lower than self-reported ratings by approximately one point, and this gap was not driven by any single aspect. Rather, our analysis shows that self-reported ratings capture the fan's verdict, an overall evaluative judgment that integrates the entire experience. While predicted ratings quantify the impact of salient moments characterized as memorable, emotionally intense, unusual, or actionable. Each measure contains information the other misses. These baseline results establish that a simple, unoptimized prompt can directionally predict how fans rate their experience from the text a fan wrote and that a gap between the two numbers can be interpreted as a construct difference worth preserving rather than an error to eliminate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。