构建平衡数据集B-RIGHT,让人类-物体交互评测更公平可靠
B-RIGHT: Benchmark Re-evaluation for Integrity in Generalized Human-Object Interaction Testing
- 用算法生成并筛选数据,实现每类交互样本数相等
- 新测试集使模型性能排名变化显著,评分方差大幅降低
- 适合评估通用交互模型,尤其关注评测公平性的研究者
人类-物体交互(HOI)是人工智能理解视觉世界的关键任务,但现有基准如HICO-DET存在严重类别不平衡及训练/测试集规模不一致的问题,可能导致模型性能评估失真。本文提出B-RIGHT:一个针对广义人类-物体交互测试的完整性再评估基准。通过平衡算法与自动化生成-过滤流程,B-RIGHT确保每个HOI类别具有相等的样本数量。此外,设计了均衡的零样本测试集,系统评估模型在未见场景下的表现。在该基准上重评现有模型发现,评分方差显著下降,性能排名发生明显变化。实验表明,在平衡条件下进行评估能带来更可靠、更公平的模型比较。
原文摘要 · Abstract (English)
Human-object interaction (HOI) is an essential problem in artificial intelligence (AI) which aims to understand the visual world that involves complex relationships between humans and objects. However, current benchmarks such as HICO-DET face the following limitations: (1) severe class imbalance and (2) varying number of train and test sets for certain classes. These issues can potentially lead to either inflation or deflation of model performance during evaluation, ultimately undermining the reliability of evaluation scores. In this paper, we propose a systematic approach to develop a new class-balanced dataset, Benchmark Re-evaluation for Integrity in Generalized Human-object Interaction Testing (B-RIGHT), that addresses these imbalanced problems. B-RIGHT achieves class balance by leveraging balancing algorithm and automated generation-and-filtering processes, ensuring an equal number of instances for each HOI class. Furthermore, we design a balanced zero-shot test set to systematically evaluate models on unseen scenario. Re-evaluating existing models using B-RIGHT reveals substantial the reduction of score variance and changes in performance rankings compared to conventional HICO-DET. Our experiments demonstrate that evaluation under balanced conditions ensure more reliable and fair model comparisons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。