CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory

Nur Muhammad Mahi Shafiullah,Chris Paxton,Lerrel Pinto,Soumith Chintala,Arthur Szlam

arxiv（2022）

引用 24|浏览146

暂无评分

摘要

We propose CLIP-Fields, an implicit scene model that can be trained with no direct human supervision. This model learns a mapping from spatial locations to semantic embedding vectors. The mapping can then be used for a variety of tasks, such as segmentation, instance identification, semantic search over space, and view localization. Most importantly, the mapping can be trained with supervision coming only from web-image and web-text trained models such as CLIP, Detic, and Sentence-BERT. When compared to baselines like Mask-RCNN, our method outperforms on few-shot instance identification or semantic segmentation on the HM3D dataset with only a fraction of the examples. Finally, we show that using CLIP-Fields as a scene memory, robots can perform semantic navigation in real-world environments. Our code and demonstrations are available here: https://mahis.life/clip-fields

查看译文

关键词

weakly supervised semantic clip-fields,robotic memory

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要