Heron Systems, Inc. — Department of Defense SBIR Phase II: AF212-D001

Heron Systems, Inc. — SBIR Phase II award from Department of Defense.

Amount
$999,821
Agency
Department of Defense · Air Force
Program / Phase
SBIR · Phase II
Topic
AF212-D001
Solicitation
21.2
NAICS
Place of performance
MD
Period
2022-01-10 → 2023-07-10

Description

Recent demonstrations of super-human performance leveraging reinforcement learning has shown the power of Artificial Intelligence (AI) for solving high dimensional complex problems through long-term decision making. However, most reinforcement learning approaches require hundreds of millions or even billions of training samples to achieve high performing policies. In this paper we propose a novel model based reinforcement learning algorithm call Model Zero that constructs a world model learned from a lower fidelity simulation and transfers the world model into a high fidelity simulation. Model Zero leverages its existing knowledge learned from a low fidelity simulation and continues to learn a more accurate world model after being transferred into the high fidelity simulation. A key advantage of using our model based reinforcement learning approach is that the methods are general and can be applied to various deterministic  high-fidelity physics based environments where training samples are computationally expensive to obtain.