TY - GEN
T1 - Dynamical Scene Representation and Control with Keypoint-Conditioned Neural Radiance Field
AU - Wang, Weiyao
AU - Morgan, Andrew S.
AU - Dollar, Aaron M.
AU - Hager, Gregory D.
N1 - Funding Information:
This work was supported by the U.S. National Science Foundation grants IIS-1900952 to Johns Hopkins University and grants IIS-1752134, IIS-1900681 to Yale University 1W. Wang and G. D. Hager are with the Department of Computer Science, Johns Hopkins University, Baltimore, MD 21218, USA {wwang121,hager} @ cs.jhu.edu 2A. Morgan and A. Dollar are with the Department of Mechanical Engineering and Materials Science, Yale University, New Haven, CT 06520, USA {andrew.morgan, aaron.dollar} @ yale.edu
Publisher Copyright:
© 2022 IEEE.
PY - 2022
Y1 - 2022
N2 - In this work, we present a method that can learn to model dynamic and arbitrary 3D scenes, purely from 2D visual observations. Our approach uses a keypoint-conditioned Neural Radiance Field (KP-NeRF) to capture and model these scenes with the overarching goal of supporting image-based robot manipulation. Differentiating this from previous methods, which typically condition the model on generic embedding vectors for representation, our implicit neural radiance function is conditioned on a set of keypoints that are inferred from a learned encoder given imagery observations. This implicitly separates the visual modeling components into object appearances and object pose configurations. Such inductive bias built into the architecture encourages discovered keypoints to capture state transitions in the robot's environment across time and space. We then learn a forward prediction model of the encoded keypoints, constructed over the keypoint representation space, and perform MPC control for challenging manipulation tasks including block pushing and door closing. We evaluate the performance of our method through various tasks: novel scene view synthesis, action-conditioned forward prediction, and robot manipulation tasks.
AB - In this work, we present a method that can learn to model dynamic and arbitrary 3D scenes, purely from 2D visual observations. Our approach uses a keypoint-conditioned Neural Radiance Field (KP-NeRF) to capture and model these scenes with the overarching goal of supporting image-based robot manipulation. Differentiating this from previous methods, which typically condition the model on generic embedding vectors for representation, our implicit neural radiance function is conditioned on a set of keypoints that are inferred from a learned encoder given imagery observations. This implicitly separates the visual modeling components into object appearances and object pose configurations. Such inductive bias built into the architecture encourages discovered keypoints to capture state transitions in the robot's environment across time and space. We then learn a forward prediction model of the encoded keypoints, constructed over the keypoint representation space, and perform MPC control for challenging manipulation tasks including block pushing and door closing. We evaluate the performance of our method through various tasks: novel scene view synthesis, action-conditioned forward prediction, and robot manipulation tasks.
UR - https://www.scopus.com/pages/publications/85141702068
UR - https://www.scopus.com/pages/publications/85141702068#tab=citedBy
U2 - 10.1109/CASE49997.2022.9926555
DO - 10.1109/CASE49997.2022.9926555
M3 - Conference contribution
AN - SCOPUS:85141702068
T3 - IEEE International Conference on Automation Science and Engineering
SP - 1138
EP - 1143
BT - 2022 IEEE 18th International Conference on Automation Science and Engineering, CASE 2022
PB - IEEE Computer Society
T2 - 18th IEEE International Conference on Automation Science and Engineering, CASE 2022
Y2 - 20 August 2022 through 24 August 2022
ER -