速度目标导向的人形机器人风格化运动策略
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

TP 242.6

基金项目:

上海市促进产业高质量发展专项资金—先导产业创新发展项目(RZ-RGZN-01-25-0673)


Velocity-guided stylized locomotion for humanoid robots
Author:
Affiliation:

Fund Project:

undefined

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    通过模仿学习赋予人形机器人拟人化运动技能,是构建多样化、高表现力机器人运动能力的重要途径。当前基于模仿学习的方法仍面临显著局限:以对抗运动先验(adversarial motion priors, AMP)为代表的对抗模仿方法虽能保留整体运动风格,却易丢失精细动作细节;而以Mimic为代表的精确模仿方法虽可高保真复现参考动作,却因策略与参考动作强耦合,难以响应实时控制指令。为此,提出一种新颖的模仿学习训练框架,通过观测空间与奖励机制的协同设计,显式解耦运动风格与任务目标,并实现二者高效协同。所提方法仅依赖高层速度指令驱动,成功将策略从“动作复现器”转变为“速度指令控制器”。实验表明,该策略在任意速度指令下均能实现高精度速度跟踪(整体误差< 0.13 m/s或rad/s),同时关键部位的位置和姿态相对原始参考动作的平均动态时间规整(dynamic time warping, DTW)距离分别为0.837与1.135,充分验证了其在精准响应任务指令的同时,有效保留了参考运动的自然风格特征。

    Abstract:

    Enabling humanoid robots to acquire human-like locomotion skills through imitation learning is a key approach to building diverse and expressive motor capabilities. Current imitation learning approaches remain fundamentally limited: adversarial methods such as (adversarial motion priors) AMP preserve global motion characteristics but often lose fine-grained details, while precise imitation frameworks like Mimic faithfully reproduce reference motions yet suffer from strong coupling between policy and demonstration, rendering them incapable of responding to real-time control commands. To address this trade-off, we proposed a novel imitation learning architecture through a principled redesign of the observation space and reward structure, explicitly disentangling motion style from task objectives while enabling their effective coordination. The resulting policy operates solely on high-level velocity commands—effectively transforming the original “motion reproducer” into a “velocity commanded controller.” Experiments demonstrate that our method achieves accurate velocity tracking (overall error<0.13 m/s or rad/s) under arbitrary speed commands, while preserving stylistic fidelity: the average dynamic time warping (DTW) distances between generated and reference motions are 0.837 for key joint positions and 1.135 for orientations. These results confirm that our framework successfully reconciles precise task execution with natural motion style reproduction.

    参考文献
    相似文献
    引证文献
引用本文

马晓航,梁志远,甘文聪,朱玉迪,李清都.速度目标导向的人形机器人风格化运动策略[J].上海理工大学学报,2026,48(4):407-419.

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-01-30
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-09-15
  • 出版日期:
文章二维码