Great work. I have a few questions about the camera extrinsics.
During inference, the encoder uses only the intrinsics and does not use the extrinsics, while the decoder does use the extrinsics. This makes me curious: in which coordinate system are the extrinsics defined during rendering?
Does this design imply that, on an unseen dataset, if one directly uses the provided checkpoint without further training, the extrinsics might be inaccurate?
Great work. I have a few questions about the camera extrinsics.
During inference, the encoder uses only the intrinsics and does not use the extrinsics, while the decoder does use the extrinsics. This makes me curious: in which coordinate system are the extrinsics defined during rendering?
Does this design imply that, on an unseen dataset, if one directly uses the provided checkpoint without further training, the extrinsics might be inaccurate?