Abstract
Workspace analysis measures where a robot can place its end effector. For visually guided manipulation, reachability alone is insufficient: a kinematically reachable target may not be visible in the specific pose required to reach it. The robot must then redirect its sensing or move its body to acquire a view, turning a perception limitation into additional motion. Existing humanoids largely inherit this limitation when copying human form factors. We introduce the visible-reachable workspace (VRW), a design-stage measure that conditions visibility on feasible reaching configurations and extends it to concurrent visibility of spatially separated work regions. We apply VRW by building a 31-DoF humanoid with independently actuated RGB-D cameras. On the same robot, camera articulation increases visible-reachable coverage from 38% to 97%. With actuated camera layouts, a second camera raises pairwise coverage from 0.45 to 0.95, while a third changes it only to 0.97. In a controlled two-target reach-and-grasp benchmark, our dual-actuated design reduces mean completion time by 17% and mechanical energy by 19% relative to the same robot with its cameras fixed. Hardware experiments demonstrate simultaneous observation and manipulation of front/back and left/right target pairs without torso reorientation. The results suggest that reachability becomes a more informative design quantity for perception-driven humanoid manipulation when it is evaluated together with the sensing configurations that make the reachable space observable. We will open-source all software and the humanoid hardware design.
Video
Overview
Two moving targets carried by two people on opposite sides. Each camera module tracks one target and the arm on that side follows it, both at once.

Visible-Reachable Workspace
A humanoid can usually reach far more space than it can see. VRW keeps only the reachable targets that can also be observed from a feasible reaching configuration. A pairwise extension, η₂, asks whether a layout can keep concurrent views of a manipulation target and a second, separated region.
Same arms, same body, differing only in whether the camera joints are free. Magenta is visible-reachable; blue is reachable but blind.

Visible-reachable workspace across humanoid platforms.

- Articulation beats count. Actuating the cameras raises visible-reachable coverage from 38% to 97%; a second and third module add under 3%.
- The second module buys concurrency. Pairwise coverage η₂ rises from 0.45 to 0.95 and stays nearly constant as the two work regions separate.
Two-Target Reach-and-Grasp Benchmark
Six simulated scenarios, 900 trials, success-conditional means; lower is better. Act₂ (two actuated cameras) is the adopted configuration.
| Time (s) | Search (s) | Approach (s) | Manipulation (s) | Energy (J) | |
|---|---|---|---|---|---|
| Unitree G1 | 27.5 | 3.8 | 3.8 | 19.9 | 494 |
| Fix₂ | 17.0 | 0.2 | 3.2 | 13.6 | 425 |
| Act₂ (ours) | 14.1 | 0.1 | 2.1 | 11.9 | 346 |
| Fix₁ | 20.5 | 3.3 | 3.3 | 13.9 | 595 |
| Act₁ | 15.5 | 0.6 | 2.7 | 12.2 | 385 |

Real-World Deployment
Four hardware trials play together in each clip.
Left and right, within reach

Front and behind

Hardware

| Specification | Value |
|---|---|
| DoF | 31: 27-DoF body + two 2-DoF camera gimbals, +1 per gripper |
| Mass / height | 36 kg / 1.2 m |
| Arm reach / leg length | 0.46 m / 0.39 m |
| Cameras | 2 × RGB-D, each on its own yaw-pitch gimbal |
| End effectors | Parallel grippers, 350 g each |
| Actuation | Quasi-direct-drive throughout |
| Control | 50 Hz learned whole-body policy, 200 Hz CAN motor loop |
Links
Citation
| |