Intelligent Automation, Inc. — Department of Defense SBIR Phase I: A20-061

Intelligent Automation, Inc. — SBIR Phase I award from Department of Defense.

Amount
$111,498
Agency
Department of Defense · Army
Program / Phase
SBIR · Phase I
Topic
A20-061
Solicitation
20.1
NAICS
Place of performance
MD
Period
2020-06-23 → 2021-03-17

Description

The basic Reinforcement Learning algorithms behind recent advances have been around since the 90s. The proliferation of deep learning function approximators; first made popular in computer vision quickly made their way into game theoretic approaches. These high dimensional functions made it possible to learn complex value functions and behavior policies that were not possible before. It is now possible to teach an agent to play Atari games at human level performance from observing pixels, on a desktop CPU in a matter of days. Learned agents in MMO RPGS (e.g.  DOTA and Starcraft) as well as realistic physics-based simulators (e.g. “Hide and Seek”) have demonstrated the potential for the emergence of complex cooperative behavior. During training, complex group behaviors emerged in stages, each more impressive than the last. The ideas behind these advances can be applied to NeuralMMO and transitioned to current and future wargame environment tools to allow military planners to fully realize strategy evolution and how the measures-countermeasure cycle would play out over very long time periods. The nature of the strategies at each transition point has the potential to exceed human-created strategies in the same way AlphaZero developed super-human winning strategies for Chess, Go, and Shogi.