arXiv:2502.07404cs.HCcs.AI2025-02被引 4

研究人机协作标注中模型可靠性如何影响标注准确率。

Human-in-the-Loop Annotation for Image-Based Engagement Estimation: Assessing the Impact of Model Reliability on Annotation Accuracy

  • 将高精度图像情绪模型嵌入人机协作框架,测试不同可靠性下的标注表现。
  • 模型可靠时标注一致性高,不可靠时导致焦虑和判断波动,但增加批判性评估。
  • 负面提示会引发认知偏差,让人误以为不可靠模型更可信,适合心理机制研究者。

人机协同(HITL)框架在情绪估计系统中日益受到重视,通过融合机器预测与人类专业判断提升标注准确性。本研究将高性能图像情绪模型集成至HITL框架,评估人机协作潜力,并识别影响协作成功的关键心理与实践因素。重点探究模型可靠性与认知框架对人类信任、认知负荷及标注行为的影响。基于29名参与者在三种实验场景下的数据:基础可靠性(S1)、人为制造错误(S2)、负面框架引入认知偏差(S3),分析行为与质性数据。S1中可靠预测带来高信任度与标注一致性;S2中不可靠输出引发更多批判性评估,但伴随更高挫败感与响应变异性;S3中负面框架使参与者误认为模型更可靠且更易共情,即使其实际性能被误导。结果表明,机器输出的可靠性与心理因素共同塑造高效的人机协作。本研究构建可扩展的HITL情绪标注框架,为自适应学习与人机交互提供基础。

原文摘要 · Abstract (English)

Human-in-the-loop (HITL) frameworks are increasingly recognized for their potential to improve annotation accuracy in emotion estimation systems by combining machine predictions with human expertise. This study focuses on integrating a high-performing image-based emotion model into a HITL annotation framework to evaluate the collaborative potential of human-machine interaction and identify the psychological and practical factors critical to successful collaboration. Specifically, we investigate how varying model reliability and cognitive framing influence human trust, cognitive load, and annotation behavior in HITL systems. We demonstrate that model reliability and psychological framing significantly impact annotators' trust, engagement, and consistency, offering insights into optimizing HITL frameworks. Through three experimental scenarios with 29 participants--baseline model reliability (S1), fabricated errors (S2), and cognitive bias introduced by negative framing (S3)--we analyzed behavioral and qualitative data. Reliable predictions in S1 yielded high trust and annotation consistency, while unreliable outputs in S2 led to increased critical evaluations but also heightened frustration and response variability. Negative framing in S3 revealed how cognitive bias influenced participants to perceive the model as more relatable and accurate, despite misinformation regarding its reliability. These findings highlight the importance of both reliable machine outputs and psychological factors in shaping effective human-machine collaboration. By leveraging the strengths of both human oversight and automated systems, this study establishes a scalable HITL framework for emotion annotation and lays the foundation for broader applications in adaptive learning and human-computer interaction.

人机协作情绪识别标注优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。