TY - JOUR
T1 - ImageNet3D
T2 - 38th Conference on Neural Information Processing Systems, NeurIPS 2024
AU - Ma, Wufei
AU - Zhang, Guofeng
AU - Liu, Qihao
AU - Zeng, Guanning
AU - Kortylewski, Adam
AU - Liu, Yaoyao
AU - Yuille, Alan
N1 - Alan Yuille acknowledges support from the Army Research Laboratory W911NF2320008 and National Eye Institute (NEI) with Award ID: R01EY037193. Adam Kortylewski acknowledges support for his Emmy Noether Research Group funded by the German Science Foundation (DFG) under Grant No. 468670075. We would like to thank David Chauca, Olga Kuzmich, Yi Luo, Soh Kay Leen, Chenhao Lin, Hardik Shah, Jishuo Yang, Yuqi Li, Manqin Cai, Nanru Dai, Shen Wang, Wanyi Dai, Yifan Shuai, Zhangbo Cheng for their valuable help with the ImageNet3D data collection. We are also grateful to the anonymous reviewers for the valuable discussions and constructive feedback.
PY - 2024
Y1 - 2024
N2 - A vision model with general-purpose object-level 3D understanding should be capable of inferring both 2D (e.g., class name and bounding box) and 3D information (e.g., 3D location and 3D viewpoint) for arbitrary rigid objects in natural images. This is a challenging task, as it involves inferring 3D information from 2D signals and most importantly, generalizing to rigid objects from unseen categories. However, existing datasets with object-level 3D annotations are often limited by the number of categories or the quality of annotations. Models developed on these datasets become specialists for certain categories or domains, and fail to generalize. In this work, we present ImageNet3D, a large dataset for general-purpose object-level 3D understanding. ImageNet3D augments 200 categories from the ImageNet dataset with 2D bounding box, 3D pose, 3D location annotations, and image captions interleaved with 3D information. With the new annotations available in ImageNet3D, we could (i) analyze the object-level 3D awareness of visual foundation models, and (ii) study and develop general-purpose models that infer both 2D and 3D information for arbitrary rigid objects in natural images, and (iii) integrate unified 3D models with large language models for 3D-related reasoning. We consider two new tasks, probing of object-level 3D awareness and open vocabulary pose estimation, besides standard classification and pose estimation. Experimental results on ImageNet3D demonstrate the potential of our dataset in building vision models with stronger general-purpose object-level 3D understanding. Our dataset and project page are available here: https://imagenet3d.github.io.
AB - A vision model with general-purpose object-level 3D understanding should be capable of inferring both 2D (e.g., class name and bounding box) and 3D information (e.g., 3D location and 3D viewpoint) for arbitrary rigid objects in natural images. This is a challenging task, as it involves inferring 3D information from 2D signals and most importantly, generalizing to rigid objects from unseen categories. However, existing datasets with object-level 3D annotations are often limited by the number of categories or the quality of annotations. Models developed on these datasets become specialists for certain categories or domains, and fail to generalize. In this work, we present ImageNet3D, a large dataset for general-purpose object-level 3D understanding. ImageNet3D augments 200 categories from the ImageNet dataset with 2D bounding box, 3D pose, 3D location annotations, and image captions interleaved with 3D information. With the new annotations available in ImageNet3D, we could (i) analyze the object-level 3D awareness of visual foundation models, and (ii) study and develop general-purpose models that infer both 2D and 3D information for arbitrary rigid objects in natural images, and (iii) integrate unified 3D models with large language models for 3D-related reasoning. We consider two new tasks, probing of object-level 3D awareness and open vocabulary pose estimation, besides standard classification and pose estimation. Experimental results on ImageNet3D demonstrate the potential of our dataset in building vision models with stronger general-purpose object-level 3D understanding. Our dataset and project page are available here: https://imagenet3d.github.io.
UR - https://www.scopus.com/pages/publications/105000547604
UR - https://www.scopus.com/pages/publications/105000547604#tab=citedBy
U2 - 10.52202/079017-3046
DO - 10.52202/079017-3046
M3 - Conference article
SN - 1049-5258
VL - 37
SP - 96127
EP - 96149
JO - Advances in Neural Information Processing Systems
JF - Advances in Neural Information Processing Systems
Y2 - 9 December 2024 through 15 December 2024
ER -