用大模型生成问卷回答,高效评估不同提问方式的效果
QSTN: A Modular Framework for Robust Questionnaire Inference with Large Language Models
- 模块化框架支持自动构建问卷并生成回答
- 4000万+次测试显示提问结构影响答案一致性
- 无需编程也能设置实验,适合非技术研究者
我们提出QSTN,一个开源Python框架,用于系统生成问卷类提示的响应,支持基于大语言模型(LLMs)的虚拟调查与标注任务。QSTN可稳健评估问卷呈现方式、提示扰动及响应生成方法的影响。超过4000万次调查响应的广泛评估表明,问题结构和生成方法显著影响生成答案与人类答案的一致性。通过改变呈现方式,可在极低计算成本下获得有效答案。此外,框架提供无代码用户界面,使研究人员无需编程即可开展可靠实验。我们希望QSTN能提升未来大模型研究的可复现性与可靠性。
原文摘要 · Abstract (English)
We introduce QSTN, an open-source Python framework for systematically generating responses from questionnaire-style prompts to support in-silico surveys and annotation tasks with large language models (LLMs). QSTN enables robust evaluation of questionnaire presentation, prompt perturbations, and response generation methods. Our extensive evaluation (>40 million survey responses) shows that question structure and response generation methods have a significant impact on the alignment of generated survey responses with human answers. We also find that answers can be obtained for a fraction of the compute cost, by changing the presentation method. In addition, we offer a no-code user interface that allows researchers to set up robust experiments with LLMs \emph{without coding knowledge}. We hope that QSTN will support the reproducibility and reliability of LLM-based research in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。