arXiv:2505.11108cs.ROcs.AI2025-05中稿 · IEEE ROMAN 2025被引 6

用场景上下文学习用户习惯,让机器人自动整理物品。

Personalized Robotic Object Rearrangement from Scene Context

  • 基于大模型分析场景上下文,支持灵活偏好下的物品摆放。
  • 在110K数据集上表现优于单一上下文模型,三类环境均排名前二。
  • 适合研究个性化家居机器人与智能交互系统的人参考。

物体重排是家庭机器人的重要任务,需在无明确指令下实现个性化、有意义的物品摆放,并泛化至未见物体与新环境。为此,我们提出PARSEC基准,基于72名用户的众包数据构建了包含110万例重排样本、93类物品和15个场景的数据集。为贴近真实组织习惯,提出ContextSortLM——一种基于LLM的个性化重排模型,通过显式处理多有效放置位置来适应复杂场景。在PARSEC上评估该模型及现有方法,并由108位在线评价者对预测结果按用户偏好匹配度打分。结果表明,融合多个场景上下文源的模型性能优于单源模型;ContextSortLM在复制目标用户布局方面表现最佳,在三类环境中均位列前二。评估还揭示了跨环境语义建模的挑战,并为未来工作提供建议。

原文摘要 · Abstract (English)

Object rearrangement is a key task for household robots requiring personalization without explicit instructions, meaningful object placement in environments occupied with objects, and generalization to unseen objects and new environments. To facilitate research addressing these challenges, we introduce PARSEC, an object rearrangement benchmark for learning user organizational preferences from observed scene context to place objects in a partially arranged environment. PARSEC is built upon a novel dataset of 110K rearrangement examples crowdsourced from 72 users, featuring 93 object categories and 15 environments. To better align with real-world organizational habits, we propose ContextSortLM, an LLM-based personalized rearrangement model that handles flexible user preferences by explicitly accounting for objects with multiple valid placement locations when placing items in partially arranged environments. We evaluate ContextSortLM and existing personalized rearrangement approaches on the PARSEC benchmark and complement these findings with a crowdsourced evaluation of 108 online raters ranking model predictions based on alignment with user preferences. Our results indicate that personalized rearrangement models leveraging multiple scene context sources perform better than models relying on a single context source. Moreover, ContextSortLM outperforms other models in placing objects to replicate the target user's arrangement and ranks among the top two in all three environment categories, as rated by online evaluators. Importantly, our evaluation highlights challenges associated with modeling environment semantics across different environment categories and provides recommendations for future work.

机器人个性化场景理解重排

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。