arXiv:2604.05163cs.CLcs.AI2026-04

分析14个真实研究中的343份访谈,找出真正有价值的回答特征。

What Makes a Good Response? An Empirical Analysis of Quality in Qualitative Interviews

  • 构建包含16940条回答的访谈语料库,验证10种质量评估指标
  • 直接回应核心研究问题的回答最能贡献研究发现
  • 清晰度和意外性无法预测回答质量,适合研究者参考

定性访谈在获取人类经验洞察时,依赖高质量的回答。尽管定性研究与自然语言处理领域提出多种访谈质量评估方法,但这些方法缺乏实证支持:高分回答是否真能促进研究目标达成。本文识别、实现并评估了10种已有响应质量度量方法,以确定哪些能有效预测回答对研究发现的贡献。为此,我们构建了定性访谈语料库(Qualitative Interview Corpus),包含来自14项真实研究的343份访谈转录本及16,940条参与者回答。结果表明,回答与核心研究问题的直接相关性是质量最强的预测因子。此外,常用于评估NLP访谈系统的清晰度与基于意外性的信息量指标,并不能预测实际质量。本研究为定性研究设计与自动化访谈系统评估提供了可落地的量化指标与实证依据。

原文摘要 · Abstract (English)

Qualitative interviews provide essential insights into human experiences when they elicit high-quality responses. While qualitative and NLP researchers have proposed various measures of interview quality, these measures lack validation that high-scoring responses actually contribute to the study's goals. In this work, we identify, implement, and evaluate 10 proposed measures of interview response quality to determine which are actually predictive of a response's contribution to the study findings. To conduct our analysis, we introduce the Qualitative Interview Corpus, a newly constructed dataset of 343 interview transcripts with 16,940 participant responses from 14 real research projects. We find that direct relevance to a key research question is the strongest predictor of response quality. We additionally find that two measures commonly used to evaluate NLP interview systems, clarity and surprisal-based informativeness, are not predictive of response quality. Our work provides analytic insights and grounded, scalable metrics to inform the design of qualitative studies and the evaluation of automated interview systems.

定性研究访谈分析评估指标语料库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。