构建首个聚焦时间模糊性的问题检测数据集,提升开放域问答准确性。
Detecting Temporal Ambiguity in Questions
- 基于问题改写生成消歧版本,结合多策略搜索识别时间模糊性。
- 在8,162个问题上验证,新方法显著优于零样本与少样本基线。
- 适合研究问答系统鲁棒性、信息检索与自然语言理解的开发者。
开放域问答中,模糊问题的检测与回答始终是挑战,其答案依赖于语义解读,形式多样。时间模糊性是其中最常见的类型。本文提出TEMPAMBIQA,一个手工标注的时序模糊问答数据集,包含从现有数据集中提取的8,162个开放域问题。标注重点在于捕捉时间模糊性,以支持该任务的研究。我们提出一种新方法,通过使用基于问题消歧版本的多样化搜索策略实现检测。同时引入并测试了非搜索类基线,采用零样本和少样本方法评估时间模糊性检测效果。
原文摘要 · Abstract (English)
Detecting and answering ambiguous questions has been a challenging task in open-domain question answering. Ambiguous questions have different answers depending on their interpretation and can take diverse forms. Temporally ambiguous questions are one of the most common types of such questions. In this paper, we introduce TEMPAMBIQA, a manually annotated temporally ambiguous QA dataset consisting of 8,162 open-domain questions derived from existing datasets. Our annotations focus on capturing temporal ambiguity to study the task of detecting temporally ambiguous questions. We propose a novel approach by using diverse search strategies based on disambiguated versions of the questions. We also introduce and test non-search, competitive baselines for detecting temporal ambiguity using zero-shot and few-shot approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。