arXiv:2510.06242cs.CLcs.AI2025-10EMNLP综述

用AI自动评估用户问卷回复质量,更准更省力。

Transparent Reference-free Automated Evaluation of Open-Ended User Survey Responses

  • 分两阶段:先筛乱码,再从努力度、相关性、完整性三方面评估。
  • 在英韩数据集上与专家评分高度相关,优于现有方法。
  • 适合市场调研、问卷分析等需批量处理开放题的场景。

开放式问卷回复在市场营销研究中提供宝贵洞察,但低质量回复不仅增加研究人员手动筛选负担,还可能引发误导性结论,亟需高效评估手段。现有自动化评估方法主要针对大模型生成文本,难以有效衡量人类写作回复的独特特征。为此,我们提出一种专为人类问卷回复设计的两阶段评估框架:首先通过伪随机过滤剔除无意义回复;随后基于真实问卷数据的实证分析,利用大模型能力在努力度、相关性和完整性三个维度进行评估。在英语和韩语数据集上的验证表明,该框架不仅优于现有指标,且在回复质量预测与拒绝决策等实际应用中表现出高可用性,与专家评估具有强相关性。

原文摘要 · Abstract (English)

Open-ended survey responses provide valuable insights in marketing research, but low-quality responses not only burden researchers with manual filtering but also risk leading to misleading conclusions, underscoring the need for effective evaluation. Existing automatic evaluation methods target LLM-generated text and inadequately assess human-written responses with their distinct characteristics. To address such characteristics, we propose a two-stage evaluation framework specifically designed for human survey responses. First, gibberish filtering removes nonsensical responses. Then, three dimensions-effort, relevance, and completeness-are evaluated using LLM capabilities, grounded in empirical analysis of real-world survey data. Validation on English and Korean datasets shows that our framework not only outperforms existing metrics but also demonstrates high practical applicability for real-world applications such as response quality prediction and response rejection, showing strong correlations with expert assessment.

问卷评估大模型应用质量检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。