Encoder-only walking
To our knowledge, the first real-world push-resilient bipedal humanoid locomotion policy using only joint encoders—without IMU or force/torque sensing.
We present blind, whole-body manipulation skills on a Unitree G1 humanoid using only onboard proprioception, without cameras, markers, force-torque, or tactile sensors. Despite this minimal sensing, the trained policies exhibit surprising capability across qualitatively different tasks: push-resilient bipedal walking without IMU feedback, active soccer ball trapping with a foot, seeking and lifting a suitcase by its handle, and mounting a randomly positioned skateboard.
We argue that these capabilities arise from a key underappreciated signal: the way the joint encoder readouts evolve under purposeful compliant contact, effectively forming a whole-body tactile channel. By generating contact-rich motions, the trained policies actively probe the environment; as a result, task-relevant object state (e.g. pose) becomes increasingly decodable from short proprioceptive histories. We expose this information using compact task-specific state estimators trained alongside, but fully separately from, the policies; their prediction errors decrease rapidly after informative contact.
Our results indicate that joint encoder-based proprioception, combined with compliant actuation—now widely available on commercial robots and low-cost motors—is already a strong, practical substrate for whole-body dexterous manipulation and interactive perception, and therefore a natural foundation on which richer sensing can be layered.
Modern robotic dexterity relies on cameras, object trackers, force-torque sensors, or tactile skins. But machine vision is sensitive to occlusion, lighting, color, and changes in scene appearance; dedicated tactile hardware is fragile, expensive, and hard to simulate faithfully at GPU-accelerated scale.
Can blind proprioceptive policies—already so good at locomotion—also do whole-body manipulation?
Under compliant position control, contact-induced loads perturb nominal joint tracking, leaving sparse but informative signatures in encoder histories.
Delete every object-related observation at training time: the manipulation policy receives exactly the proprioceptive inputs used by conventional blind-locomotion policies.
We design and study four skills with increasing levels of interaction complexity:
To our knowledge, the first real-world push-resilient bipedal humanoid locomotion policy using only joint encoders—without IMU or force/torque sensing.
Search for and trap a randomly placed football with one foot.
Locate, orient, and mount a freely moving skateboard.
Find a recessed handle and gently lift a suitcase randomly placed on tables of varying heights.
Our policies learn to sweep, tap, drag, and slide against the world to create the sensations they need in their joint angles. Collectively, they are an existence proof that useful haptic perception can emerge without dedicated tactile sensors.
The policy sweeps with one foot, then taps and traps the ball after first contact.
The policy touches the nose or tail and drags the deck to disambiguate its position and yaw.
The policy touches the table and suitcase, then searches along the top edge until it feels the handle opening; all without toppling the slim object.
Across tasks, estimator error drops after these informative contacts. The policies actively create proprioceptive evidence from which hidden object state can be perceived.
Off-the-shelf robots already possess a latent capacity for haptic perception through their joint encoders.
This capacity should be exploited whenever vision is unreliable—for example, during highly dynamic manipulation or under occlusion—and used as a natural foundation on which richer sensing can be layered.