Humanoid robotics | July 2026
Seeing an object is not the same as knowing how to grasp it. A camera image may show a drawer, but the robot still has to estimate depth, choose an approach angle, place its wrist, and apply enough force to pull without slipping. USC researchers recently described CLAMP, a method aimed at this three-dimensional manipulation problem.
The approach uses multiple views, wrist-mounted cameras, and information about actions and language. Those pieces let the robot connect a description such as opening a drawer with the geometry of the scene and the movement of its own hand. The robot is not simply matching a picture. It is trying to relate visual evidence to a possible action.
That connection is also a good explanation of why sensors should be placed near the part doing the work. A camera mounted on the robot's wrist can see how the gripper approaches an object and can update the plan as the hand moves. The view changes, but that change can be useful information rather than a problem.
A classroom team can explore this with a camera attached to a small arm or vehicle. Students can first use a fixed overhead camera to locate a block, then add a camera that moves with the gripper. They can compare which system handles a shifted block or a new background more reliably. The goal is not to copy a research system, but to notice how viewpoint changes affect control.
The recent USC report on CLAMP and 3D manipulation provides the source context. It is an accessible example of how geometry, computer vision, and mechanics meet inside one robot task.

.png)