用大模型让测评更智能,少问20%问题却更准
TestAgent: An Adaptive and Intelligent Expert for Human Assessment
- 用大模型动态对话选题,实时捕捉回答异常
- 心理教育测评少20%题目,准确率更高
- 适合需要个性化评估的场景,如教育医疗
精准评估人类内在状态对理解偏好、提供个性化服务及识别现实应用中的挑战至关重要。源自心理测量学的自适应测试已成为主流的人类测评方法,广泛应用于教育、医疗、体育和社科领域。其通过精选最少题目实现定制化测评。然而,现有方法面临诸多挑战:多数算法机械化,导致猜测行为,难以处理开放题;主观评估存在响应数据噪声和粗粒度输出,进一步限制效果。为逼近理想自适应测试流程,我们提出TestAgent——一个基于大语言模型(LLM)的智能代理,首次将大模型应用于自适应测试。TestAgent支持个性化题目选择,捕捉答题者反应与异常,通过动态对话交互实现精确结果。在心理、教育和生活方式测评上的实验表明,该方法在少20%题目下仍达到更高准确率,且测评者在速度、流畅性等维度更青睐此方案。
原文摘要 · Abstract (English)
Accurately assessing internal human states is key to understanding preferences, offering personalized services, and identifying challenges in real-world applications. Originating from psychometrics, adaptive testing has become the mainstream method for human measurement and has now been widely applied in education, healthcare, sports, and sociology. It customizes assessments by selecting the fewest test questions . However, current adaptive testing methods face several challenges. The mechanized nature of most algorithms leads to guessing behavior and difficulties with open-ended questions. Additionally, subjective assessments suffer from noisy response data and coarse-grained test outputs, further limiting their effectiveness. To move closer to an ideal adaptive testing process, we propose TestAgent, a large language model (LLM)-powered agent designed to enhance adaptive testing through interactive engagement. This is the first application of LLMs in adaptive testing. TestAgent supports personalized question selection, captures test-takers' responses and anomalies, and provides precise outcomes through dynamic, conversational interactions. Experiments on psychological, educational, and lifestyle assessments show our approach achieves more accurate results with 20% fewer questions than state-of-the-art baselines, and testers preferred it in speed, smoothness, and other dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。