用强模型生成高质量对话数据,提升大模型处理数据任务时的交互能力。
Synthetic Clarification and Correction Dialogues about Data-Centric Tasks -- A Teacher-Student Approach
- 用教师大模型生成可控的多轮问答对话,模拟用户与AI协作解题
- 在TAT-QA和WikiTableQuestions上生成数据集,验证大模型仍难有效提问澄清
- 聚焦真实场景中的澄清与修正,适合研究人机交互与数据增强的学者
真实场景中用户与AI协作完成数据驱动任务时,因信息不全导致对话路径动态且不可预测,构建此类互动数据集既困难又耗时。本文提出一种新框架,可从已有完全标注的表格问答数据集中,合成可控的多轮对话,用于表格式问答任务。每轮对话通过协作解决一个表格推理问题,模拟两种现实场景:(1)由AI发起澄清,(2)由用户发起纠正。关键在于使用强大教师大模型验证合成对话的正确性,保障质量。我们在TAT-QA和WikiTableQuestions数据集上生成了合成数据集,用作前沿大模型的基准测试。结果发现,即使更大模型也难以有效提出澄清问题或准确整合用户反馈进行修正。
原文摘要 · Abstract (English)
Real dialogues with AI assistants for solving data-centric tasks often follow dynamic, unpredictable paths due to imperfect information provided by the user or in the data, which must be caught and handled. Developing datasets which capture such user-AI interactions is difficult and time-consuming. In this work, we develop a novel framework for synthetically generating controlled, multi-turn conversations between a user and AI assistant for the task of table-based question answering, which can be generated from an existing dataset with fully specified table QA examples for any target domain. Each conversation aims to solve a table-based reasoning question through collaborative effort, modeling one of two real-world scenarios: (1) an AI-initiated clarification, or (2) a user-initiated correction. Critically, we employ a strong teacher LLM to verify the correctness of our synthetic conversations, ensuring high quality. We demonstrate synthetic datasets generated from TAT-QA and WikiTableQuestions as benchmarks of frontier LLMs. We find that even larger models struggle to effectively issuing clarification questions and accurately integrate user feedback for corrections.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。