Is this project an undergraduate, graduate, or faculty project?
Graduate
Project Type
individual
Campus
Daytona Beach
Authors' Class Standing
Gulsum Tuba Cibuk Girgin, Graduate student
Lead Presenter's Name
Gulsum Tuba Cibuk Girgin
Lead Presenter's College
DB College of Engineering
Faculty Mentor Name
Dr. Cagri Kilic
Abstract
Site exploration requires in-situ resource utilization when the physical properties of resources are unknown. Therefore, a generalizable object manipulation method is crucial for extraterrestrial environments. Existing studies develop reinforcement learning policies that enable interaction with objects, in which quadruped robots learn to reach commanded goals with one foot while balancing with the remaining legs. However, in these studies, goal-oriented task execution relies on high-level trajectories provided by human experts, which limits autonomous robotic operations. In this study, we propose a hierarchical DRL in which a high-level pedipulation policy outputs commands for a low-level reach policy, enabling autonomous, smooth and affordable trajectory generation for non-prehensile object manipulation. In the proposed framework, the low-level policy observes its state space as the robot base’s linear and angular velocities, height, projected gravity, joint positions and velocities, previous actions, and end-effector position commands. The action space consists of 12 joint positions. State transitions are calculated by the simulation environment. The reward function is designed to drive the robot’s end effector to reach at the desired position while penalizing excessive base linear velocity, angular velocity, joint torques, accelerations, and action rates. In the high-level policy, the observation space consists of the task-specific goal, force feedback, joint positions, and a segmented point cloud. The action space consists of position commands. The reward includes an object position tracking reward, a foot contact duration reward, and a force feedback rate reward. We employ an actor-critic framework with the PPO algorithm to train both policies. The low-level policy is trained in simulation and transferred on the quadruped robot for real-world validation. Our main contribution is the high-level policy design, which enables smooth manipulation-trajectory generation for objects in exploration sites, supporting sample retrieval, object feature learning, and relocation while incorporating visual features and force feedback.
Did this research project receive funding support (Spark, SURF, Research Abroad, Student Internal Grants, Collaborative, Climbing, or Ignite Grants) from the Office of Undergraduate Research?
No
Included in
Artificial Intelligence and Robotics Commons, Navigation, Guidance, Control and Dynamics Commons
High-Level Trajectory Learning for Non-Prehensile Object Manipulation with Hierarchical Reinforcement Learning
Site exploration requires in-situ resource utilization when the physical properties of resources are unknown. Therefore, a generalizable object manipulation method is crucial for extraterrestrial environments. Existing studies develop reinforcement learning policies that enable interaction with objects, in which quadruped robots learn to reach commanded goals with one foot while balancing with the remaining legs. However, in these studies, goal-oriented task execution relies on high-level trajectories provided by human experts, which limits autonomous robotic operations. In this study, we propose a hierarchical DRL in which a high-level pedipulation policy outputs commands for a low-level reach policy, enabling autonomous, smooth and affordable trajectory generation for non-prehensile object manipulation. In the proposed framework, the low-level policy observes its state space as the robot base’s linear and angular velocities, height, projected gravity, joint positions and velocities, previous actions, and end-effector position commands. The action space consists of 12 joint positions. State transitions are calculated by the simulation environment. The reward function is designed to drive the robot’s end effector to reach at the desired position while penalizing excessive base linear velocity, angular velocity, joint torques, accelerations, and action rates. In the high-level policy, the observation space consists of the task-specific goal, force feedback, joint positions, and a segmented point cloud. The action space consists of position commands. The reward includes an object position tracking reward, a foot contact duration reward, and a force feedback rate reward. We employ an actor-critic framework with the PPO algorithm to train both policies. The low-level policy is trained in simulation and transferred on the quadruped robot for real-world validation. Our main contribution is the high-level policy design, which enables smooth manipulation-trajectory generation for objects in exploration sites, supporting sample retrieval, object feature learning, and relocation while incorporating visual features and force feedback.