Abstract:Enabling humanoid robots to acquire human-like locomotion skills through imitation learning is a key approach to building diverse and expressive motor capabilities. Current imitation learning approaches remain fundamentally limited: adversarial methods such as (adversarial motion priors) AMP preserve global motion characteristics but often lose fine-grained details, while precise imitation frameworks like Mimic faithfully reproduce reference motions yet suffer from strong coupling between policy and demonstration, rendering them incapable of responding to real-time control commands. To address this trade-off, we proposed a novel imitation learning architecture through a principled redesign of the observation space and reward structure, explicitly disentangling motion style from task objectives while enabling their effective coordination. The resulting policy operates solely on high-level velocity commands—effectively transforming the original “motion reproducer” into a “velocity commanded controller.” Experiments demonstrate that our method achieves accurate velocity tracking (overall error<0.13 m/s or rad/s) under arbitrary speed commands, while preserving stylistic fidelity: the average dynamic time warping (DTW) distances between generated and reference motions are 0.837 for key joint positions and 1.135 for orientations. These results confirm that our framework successfully reconciles precise task execution with natural motion style reproduction.