Why learn coordinates before VLA?
There is a longer path from located in the bottom right'' of the camera image to in which direction the robot moves its hand'' than it appears. Although VLA can learn behavior directly from images, instructions, and states, the coordinate system cannot be ignored for debugging, collision avoidance, porting to another camera, and connection with low-level control. Even if the image looks correct, if the camera moves a few centimeters, the target for the hand changes.
- 1World frame W
- 2robot base frame B
- 3end-effector frame E
- 4gripper-tip frame G
- 1World frame W
- 2camera frame C
- 3image pixels (u, v)
- 1Calibration parameters + synchronized timestamps
- 2transforms
- 3reachability and collision checks
- 4controller
Internal/external parameters
Camera intrinsic parameters represent the characteristics of a three-dimensional point projected onto an image, such as focal length, principal point, and lens distortion. OpenCV Calibration Explanation explains the procedure for estimating this from known patterns such as a chess board. The extrinsic parameters are where the camera is at a given time with respect to the world or the robot base, and which direction it is facing, i.e., the rotation R and the translation t.
Homogeneous transformation expands the point to [x,y,z,1] and moves the coordinate system using a 4x4 matrix. Since it has translation as well as rotation, the order of multiplication and the notation "A to B" are clearly specified. For example, if you define T_BC as a matrix that converts points in camera coordinates to base coordinates, do not reverse the definition. Even if the matrix is correct, if you mix up whether it is right-handed or left-handed, whether the angle is rad or degree, or whether the y-axis of the image is pointing downward, the hand will move unexpectedly.
eye-to-hand and eye-in-hand
Eye-to-hand, which places a fixed camera outside the desk, determines the relative position of the base and camera. In eye-in-hand, where a camera is attached to the wrist, the relative position of the hand and camera is determined and combined with the hand posture obtained from the joint angles. The latter allows you to get closer to the target, but the image moves and there are more cables and shields. In either case, if you connect the image object'' and the joint status'' when time synchronization is broken, they will appear to match in the still image, but will deviate only when in motion.
The average reprojection error alone is not sufficient for evaluating calibration. Even if the hand position error is small in a plane with a pattern, it may be large at the edge of the work area, at different depths, or in different postures. Place a known safe target at multiple locations, compare the estimated location and independent measurements, and save the maximum error and distribution. On the actual machine, this check is performed in a low-speed, no-load, isolated environment to ensure that the system stops in the event of an abnormality.
Division of labor with VLA
It is possible for VLA to interpret high-level conditions such as ``blue block to left box'' and suggest sequential actions. When reading research such as OpenVLA, check the definition of output action, observation system, and body. Separately, the safety layer monitors speed, acceleration, joint limits, collision area, and communication loss. Observations that the calibration value has been updated, the camera has been blocked, or the reliability is low can be set as conditions for stopping the system or returning it to humans, rather than leaving it up to countermeasures.
Exercises without robot
On the coordinate plane of the paper, place camera C at (2,1), robot reference B at (0,0), and object at (3,2). If C is rotated 90 degrees, calculate the camera coordinates of the object, reverse the sign of the rotation, and compare the results. Next, write down the camera is off by 2 cm'' and the timestamp is 100 ms late'' as failure hypotheses, and think about whether you can discover it through images, joints, or logs. This is a calculation and design exercise, not a claim to have produced calibrated hardware.
Numerical example: Tracing rotation and translation by hand
Let the world coordinate object be p_W=(3,2) and the camera center be c_W=(2,1). Assuming that the camera is oriented 90 degrees counterclockwise to the world, we define the rotation from the world to the camera as R_CW=[[0,1],[-1,0]]. Then, the difference p_W-c_W=(1,1) becomes p_C=R_CW(1,1)=(1,-1). This code depends on the definition of the coordinate system adopted, so do not memorize the value. Be sure to write tests such as Is the matrix W→C or C→W?'' and Is it a column vector or a row vector?''
Furthermore, assuming that the in-camera parameters are focal length f_x=f_y=500px, principal point (320,240), and camera coordinates (X,Y,Z)=(0.1,-0.05,1.0)m, the projection without distortion is u=500*0.1/1+320=370, v=500*(-0.05)/1+240=215 pixel. If Z is halved, the distance on the image will double even with the same X. This is why absolute distance cannot be determined uniquely from pixel movement in a monocular image. It requires depth, multiple viewpoints, a known plane, or robot state.
Calibration experiment design and failure examples
A typical verification plan when using a real machine is to place a known pattern not only in the center of the working area, but also at a different depth from the edges, and check for reprojections in positions not used for training. Leave image_id, board_pose, estimated_pose, reprojection_error_px, robot_pose, timestamp in the log. Using only the average value will eliminate errors that are large only at the edges or that are biased in one direction. We often make mistakes such as ignoring lens distortion, moving the camera after fixing it, not synchronizing the image and joint angle times, and hiding the pattern with our hands.
If you don't have a robot, try implementing the above formula in a spreadsheet and moving a random point back and forth from W→C→W to see if it returns to the original state. Deliberately forget to transpose the rotation matrix and see the test fail. This is a standalone verification of the conversion code, and is not proof that the camera or robot has been calibrated.
Put point_W, T_CW_version, point_C_expected, point_C_observed, round_trip_error, pass in the test table. Use three or more known points and do not trust the conversion if only one point matches. Unless points at different depths are included, it is difficult to notice misunderstandings in the projection. Separate quizzes will be given for sections that use matrix inversion, unit conversion, and time interpolation. If the numbers do not match, first check the coordinate system definition table and the direction of transformation, rather than object recognition.
Calibration values are also version-managed in the same way as datasets. Set camera_serial, lens_setting, image_resolution, mounting_pose, calibration_date, method, residual as a set and do not reuse old values after changing resolution or remounting. When we replace only the input image of VLA and it does not work, we first check the possibility that this assumption is broken rather than the generalization of the model.
In a peer review that does not use actual equipment, two people independently create a conversion table for the same point cloud, and then check out loud which coordinate system the points were moved to. Even if the expressions look the same, if one is a millimeter and the other is a meter, the translation will be off by a factor of 1,000. By writing the unit, origin, positive axis direction, and rotation order in the data string and test name instead of in the README, you can reduce tacit knowledge at the implementation stage.
Calibration evaluation is a residual distribution
Do not accept calibration from one visually plausible overlay. Hold out target positions across the reachable workspace, command no motion, and compare projected versus observed fiducials. Report median and worst residual, the camera distance and angle range, and whether the error grows near image edges. Then inject a small extrinsic perturbation in an offline model and check whether the planned path crosses a keep-out region.
A VLA can obscure this error until contact. The safety-relevant quantity is not only pixel error; it is the resulting Cartesian displacement at the end effector and the margin to obstacles.
Calibration belongs in the safety trace
The September 21 NVIDIA safety overview is relevant here only as a reminder that perception error crosses layers. Add calibration age, residual distribution, camera mount change, and frame-transform version to the same preflight record as joint limits. A pose estimate with no uncertainty or provenance cannot safely become a motion target.
MENTAL MODEL / COORDINATES
The same point has different coordinates in different frames.
Rotate the local point (1, 0) counterclockwise into a world frame with the same origin.
x = cos θ
y = sin θ
This example shows only 2D rotation. A real robot also needs consistent translation, 3D frames, units, timestamps, and axis definitions.
SOURCES
01YOUR NOTES