arXiv:2608.18317cs.CVcs.RO2026-08中稿 · Workshop on Human-…

提出标准化文档框架,提升多模态可操作性预测的可复现性。

Reproducible Multimodal Affordance Prediction

论文配图:Reproducible Multimodal Affordance Prediction
图 1 · 摘自论文原文
  • 设计Affordance Sheet统一描述任务、数据与实验条件
  • 支持真实场景下模型泛化与安全评估的可靠验证
  • 适合研究者与开发者用于公平比较和部署模型

可操作性预测旨在从多模态输入中识别代理对目标物体可能执行的动作。由于问题定义不统一、数据标注不一致、实验协议报告不完整及部署条件信息缺失,现有方法难以评估与比较。为提升透明度,我们提出Affordance Sheet,一种包含任务定义、输入模态、模型架构、训练信息、数据集和实验协议的详细文档。该框架支持可复现的基准测试与真实场景下的可靠评估,涵盖新条件泛化与人机安全性能。

原文摘要 · Abstract (English)

Affordance prediction is the identification of potential actions an agent can perform on a target object from multimodal inputs. Affordance prediction methods are difficult to evaluate and compare due to heterogeneous problem formulations, inconsistent dataset annotations, incomplete reporting of experimental protocols, and limited information about deployment conditions. These limitations challenge fair benchmarking and performance comparison. To promote transparency, we propose the Affordance Sheet, a documentation detailing task formulation with its input modalities, model architectures and training information, datasets, and experimental protocols. Affordance Sheets enable reproducible benchmarking and reliable evaluation of affordance models for real-world scenarios, including generalisation to novel conditions and human safety.

多模态可复现性动作预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。