测试LLM理解问卷结构的能力,发现格式和提示能提升8.8%准确率。
Questionnaire meets LLM: A Benchmark and Empirical Study of Structural Skills for Understanding Questions and Responses
- 设计六种问卷格式与提示策略,系统评估LLM理解能力
- 最优组合比差方案高8.8%准确率,轻量提示可再提3-4%
- 适合研究问卷自动化分析的学者与数据从业者
每天有数百万人参与调查,涵盖市场调研、学术研究、医疗问卷和客户反馈。这些数据集蕴含宝贵洞察,但其规模与结构对大型语言模型(LLMs)构成独特挑战,尽管它们在开放文本少样本推理上表现优异。现有工具(如Qualtrics、SPSS、REDCap)通常为人设计,限制了与LLM及AI自动化集成。本研究提出QASU(Questionnaire Analysis and Structural Understanding)基准,评估六种结构能力,包括答案查找、受访者数量统计和多跳推理,覆盖六种序列化格式与多种提示策略。实验显示,选择合适格式与提示组合可使准确率提升最高达8.8个百分点;针对特定任务,通过自增强提示加入轻量结构提示,平均还能提升3-4个百分点。通过分离格式与提示影响,该开源基准为推进基于LLM的问卷分析研究与实践提供可靠基础。
原文摘要 · Abstract (English)
Millions of people take surveys every day, from market polls and academic studies to medical questionnaires and customer feedback forms. These datasets capture valuable insights, but their scale and structure present a unique challenge for large language models (LLMs), which otherwise excel at few-shot reasoning over open-ended text. Yet, their ability to process questionnaire data or lists of questions crossed with hundreds of respondent rows remains underexplored. Current retrieval and survey analysis tools (e.g., Qualtrics, SPSS, REDCap) are typically designed for humans in the workflow, limiting such data integration with LLM and AI-empowered automation. This gap leaves scientists, surveyors, and everyday users without evidence-based guidance on how to best represent questionnaires for LLM consumption. We address this by introducing QASU (Questionnaire Analysis and Structural Understanding), a benchmark that probes six structural skills, including answer lookup, respondent count, and multi-hop inference, across six serialization formats and multiple prompt strategies. Experiments on contemporary LLMs show that choosing an effective format and prompt combination can improve accuracy by up to 8.8% points compared to suboptimal formats. For specific tasks, carefully adding a lightweight structural hint through self-augmented prompting can yield further improvements of 3-4% points on average. By systematically isolating format and prompting effects, our open source benchmark offers a simple yet versatile foundation for advancing both research and real-world practice in LLM-based questionnaire analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。