Is this project an undergraduate, graduate, or faculty project?

Graduate

Project Type

individual

Campus

Daytona Beach

Authors' Class Standing

Gulsum Tuba Cibuk Girgin, Graduate student

Lead Presenter's Name

Gulsum Tuba Cibuk Girgin

Lead Presenter's College

DB College of Engineering

Faculty Mentor Name

Dr. Cagri Kilic

Abstract

Site exploration requires in-situ resource utilization when the physical properties of resources are unknown. Therefore, a generalizable object manipulation method is crucial for extraterrestrial environments. Existing studies develop reinforcement learning policies that enable interaction with objects, in which quadruped robots learn to reach commanded goals with one foot while balancing with the remaining legs. However, in these studies, goal-oriented task execution relies on high-level trajectories provided by human experts, which limits autonomous robotic operations. In this study, we propose a hierarchical DRL in which a high-level pedipulation policy outputs commands for a low-level reach policy, enabling autonomous, smooth and affordable trajectory generation for non-prehensile object manipulation.   In the proposed framework, the low-level policy observes its state space as the robot base’s linear and angular velocities, height, projected gravity, joint positions and velocities, previous actions, and end-effector position commands. The action space consists of 12 joint positions. State transitions are calculated by the simulation environment. The reward function is designed to drive the robot’s end effector to reach at the desired position while penalizing excessive base linear velocity, angular velocity, joint torques, accelerations, and action rates. In the high-level policy, the observation space consists of the task-specific goal, force feedback, joint positions, and a segmented point cloud. The action space consists of position commands. The reward includes an object position tracking reward, a foot contact duration reward, and a force feedback rate reward.   We employ an actor-critic framework with the PPO algorithm to train both policies. The low-level policy is trained in simulation and transferred on the quadruped robot for real-world validation. Our main contribution is the high-level policy design, which enables smooth manipulation-trajectory generation for objects in exploration sites, supporting sample retrieval, object feature learning, and relocation while incorporating visual features and force feedback.

Did this research project receive funding support (Spark, SURF, Research Abroad, Student Internal Grants, Collaborative, Climbing, or Ignite Grants) from the Office of Undergraduate Research?

No

Share

COinS
 

High-Level Trajectory Learning for Non-Prehensile Object Manipulation with Hierarchical Reinforcement Learning

Site exploration requires in-situ resource utilization when the physical properties of resources are unknown. Therefore, a generalizable object manipulation method is crucial for extraterrestrial environments. Existing studies develop reinforcement learning policies that enable interaction with objects, in which quadruped robots learn to reach commanded goals with one foot while balancing with the remaining legs. However, in these studies, goal-oriented task execution relies on high-level trajectories provided by human experts, which limits autonomous robotic operations. In this study, we propose a hierarchical DRL in which a high-level pedipulation policy outputs commands for a low-level reach policy, enabling autonomous, smooth and affordable trajectory generation for non-prehensile object manipulation.   In the proposed framework, the low-level policy observes its state space as the robot base’s linear and angular velocities, height, projected gravity, joint positions and velocities, previous actions, and end-effector position commands. The action space consists of 12 joint positions. State transitions are calculated by the simulation environment. The reward function is designed to drive the robot’s end effector to reach at the desired position while penalizing excessive base linear velocity, angular velocity, joint torques, accelerations, and action rates. In the high-level policy, the observation space consists of the task-specific goal, force feedback, joint positions, and a segmented point cloud. The action space consists of position commands. The reward includes an object position tracking reward, a foot contact duration reward, and a force feedback rate reward.   We employ an actor-critic framework with the PPO algorithm to train both policies. The low-level policy is trained in simulation and transferred on the quadruped robot for real-world validation. Our main contribution is the high-level policy design, which enables smooth manipulation-trajectory generation for objects in exploration sites, supporting sample retrieval, object feature learning, and relocation while incorporating visual features and force feedback.

 

To view the content in your browser, please download Adobe Reader or, alternately,
you may Download the file to your hard drive.

NOTE: The latest versions of Adobe Reader do not support viewing PDF files within Firefox on Mac OS and if you are using a modern (Intel) Mac, there is no official plugin for viewing PDF files within the browser window.