GOTT connects high-level manipulation intent to robot execution through Reach → Acquire → Move.
The video below shows the full application workflow, from a human demonstration to robot execution.
Reach, acquire, and move take place within the robot-execution stage.
Whole Pipeline: A human demonstration provides the desired object motion for ironing. GOTT reaches a task-relevant contact region, establishes stable contact with its reusable primitive, and tracks the object trajectory to execute the task.
Variable Trajectory Source: The desired object trajectory specifies how the object should move; a reach specification guides the hand toward a task-relevant contact region. These inputs can come from different sources. The videos illustrate trajectories from a generative model, a human demonstration, and keypoint-based planning. Across these sources, the contact-acquisition primitive and tracking backend remain unchanged.
Autonomous Dexterous Manipulation Tasks: GOTT reuses the same contact-acquisition primitive for whiteboard erasing, ironing, pouring, standing objects upright, and tool use. Each task supplies a desired object trajectory and a task-relevant reach specification; the primitive establishes contact before the tracking controller executes the object motion.
Cross-embodiment zero-shot grasping: A single contact-acquisition policy is shared by Sharpa Hand and XHand across Franka and Tianji arms. Select an object to compare grasping across the three arm–hand combinations, without per-embodiment retraining.
Closed-loop policy robustness: The contact-acquisition primitive uses feedback to adjust hand motion as contact evolves. It retries after a failed grasp (left) and responds to human disturbance to maintain stable contact (right).
Zero-shot transfer to Allegro V5: The contact-acquisition primitive transfers to an unseen hand morphology without retraining, supporting pick-and-place (left) and grasping (right).
Training hands: Sharpa Hand, XHand, and Allegro V4. Test hand: Allegro V5, which is excluded from training.
Foundation models and large-scale human data provide rich sources of manipulation intent, but translating this intent into multi-fingered robot behavior remains difficult. Dexterous hands still lack a reusable low-level primitive that reliably establishes contact across tasks and embodiments. We propose GOTT, a reach-acquire-move framework built around a single cross-embodiment contact-acquisition primitive. Given a robot-agnostic object trajectory and a reach specification, GOTT first brings the hand near a task-relevant contact region. The shared closed-loop primitive then establishes stable contact from this approximate initialization, and a pose-conditioned controller tracks the desired object motion. Reach specifications may come from future-aware planning, external models, or human demonstrations, while the primitive and tracking backend remain unchanged. Simulation and real-world experiments show that GOTT is able to establish robust contact across diverse objects, arm-hand platforms, and seen and unseen hand morphologies. It also consistently improves end-to-end task success over open-loop grasp execution.