通过追踪阅读行为,揭示标注者如何做出偏好判断。
Tracing How Annotators Think: Augmenting Preference Judgments with Reading Processes
- 用鼠标轨迹记录标注者阅读过程,捕捉注意力焦点与重读行为。
- 一半标注任务中会重读选项,且多重读最终选择项,重读提升标注一致性。
- 适合研究主观性NLP任务中的认知机制与标注可靠性。
我们提出一种标注方法,不仅记录标签,还捕获标注者决策背后的阅读过程,例如文本关注点、重读或略读行为。基于该框架,我们在偏好标注任务上开展案例研究,构建了包含细粒度阅读行为的数据集PreferRead,该数据来自鼠标轨迹。PreferRead使我们能够详细分析标注者在提示与两个候选回复之间导航并做出选择的过程。结果显示,在约一半的试验中,标注者会重读回复,且多数情况下重读的是最终选择的选项,很少返回提示。阅读行为与标注结果显著相关:重读与更高标注者间一致性相关,而长阅读路径和时间则与更低的一致性相关。这些结果表明,阅读过程为理解复杂主观性NLP任务中的标注可靠性、决策机制与分歧提供了互补的认知维度。代码与数据已公开。
原文摘要 · Abstract (English)
We propose an annotation approach that captures not only labels but also the reading process underlying annotators' decisions, e.g., what parts of the text they focus on, re-read or skim. Using this framework, we conduct a case study on the preference annotation task, creating a dataset PreferRead that contains fine-grained annotator reading behaviors obtained from mouse tracking. PreferRead enables detailed analysis of how annotators navigate between a prompt and two candidate responses before selecting their preference. We find that annotators re-read a response in roughly half of all trials, most often revisiting the option they ultimately choose, and rarely revisit the prompt. Reading behaviors are also significantly related to annotation outcomes: re-reading is associated with higher inter-annotator agreement, whereas long reading paths and times are associated with lower agreement. These results demonstrate that reading processes provide a complementary cognitive dimension for understanding annotator reliability, decision-making and disagreement in complex, subjective NLP tasks. Our code and data are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。