提出用户参与的检测框架,发现多语言问答中的事实与文化差异。
Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering
- 引入用户反馈机制,识别跨语言问答中的事实与文化分歧。
- 在母婴健康领域双语问答中成功检测出87%的不一致答案。
- 适用于医疗、教育等需文化敏感性的多语言系统开发。
多语言问答系统需确保事实一致性,尤其对客观问题(如“黄疸是什么?”)和主观回答中的文化差异保持敏感。本文提出MIND——一种用户参与的事实核查流程,用于检测多语言问答知识库中的事实与文化偏差。MIND能突出显示受地域与语境影响的答案差异(如“分娩时由谁协助?”)。我们在母婴健康领域的双语问答系统上评估MIND,并发布一个标注了事实与文化不一致的双语问题数据集。进一步在其他领域数据集上测试表明,MIND在各类场景中均能可靠识别不一致,支持构建更具文化意识与事实一致性的问答系统。
原文摘要 · Abstract (English)
Multilingual question answering (QA) systems must ensure factual consistency across languages, especially for objective queries such as What is jaundice?, while also accounting for cultural variation in subjective responses. We propose MIND, a user-in-the-loop fact-checking pipeline to detect factual and cultural discrepancies in multilingual QA knowledge bases. MIND highlights divergent answers to culturally sensitive questions (e.g., Who assists in childbirth?) that vary by region and context. We evaluate MIND on a bilingual QA system in the maternal and infant health domain and release a dataset of bilingual questions annotated for factual and cultural inconsistencies. We further test MIND on datasets from other domains to assess generalization. In all cases, MIND reliably identifies inconsistencies, supporting the development of more culturally aware and factually consistent QA systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。