STELLAR SCIENCE LTD. CO. — Department of Defense SBIR Phase I: A22-004
STELLAR SCIENCE LTD. CO. — SBIR Phase I award from Department of Defense.
- Amount
- $111,383
- Agency
- Department of Defense · Army
- Program / Phase
- SBIR · Phase I
- Topic
- A22-004
- Solicitation
- 22.2
- NAICS
- —
- Place of performance
- NM
- Period
- 2023-01-09 → 2023-07-08
Description
Stellar Science proposes a military operation simulator with high-speed, multi-domain capabilities as well as artificial intelligence (AI) and extended reality (XR) interfaces in which AI-enabled command and control (C2) agents will learn by executing simulated multi-domain operations (MDO). The simulator will be based on an existing government-owned modeling, simulation, and analysis (MS&A) tool, the Advanced Framework for Simulation, Integration, and Modeling (AFSIM). AFSIM not only meets the baseline requirements of AI for C2 in MDO via existing add-on capabilities for a bidirectional Python application programming interface (API), much faster-than-real-time simulation speeds, scalability to distributed computing, ability to randomize environment conditions to learn robust policies, adaptability to flexibly model realistic combat characteristics, and existing protocols for user interaction, it exceeds these requirements for reinforcement learning (RL) through the benefits of discrete event simulation (DES) scenario execution, a rich scripting environment that allows for the creation of complex scenario logic, and pre-existing AI implementations such as finite state machines and behavior trees. We will create an OpenAI Gym environment, AfsimEnv, that controls the discrete event incrementation of an AFSIM simulation via Python bindings of AFSIM simulation objects. This environment will enable bidirectional message passing to address interface requirements of the RL state/action/reward cycle and will support MDO policies at multiple echelons including squad, platoon, brigade, division, and corp. We will utilize the AfsimEnv environment to train a policy via RL on an example brigade-level stochastic scenario, an MDO extension of the Tiger Claw scenario baseline. We will also demonstrate the path forward for parallel, distributed data collection for distributed AfsimEnv experience gathering on high-performance computing (HPC) clusters. First, we will demonstrate this capability using community open-source tools, such as Ray’s RLlib package, which can extend any OpenAI Gym environment to highly distributed RL workloads. Second, conditioned on complexities with the DoD Supercomputing Resource Centers (DSRCs), such as inbound networking to DSRC resources and time constraint limitations on GPU-enabled nodes, we will investigate the integration of RLlib with a government-owned tool developed by Stellar Science, Galaxy, that aids in the coordination and deployment of jobs across multiple HPC clusters. Finally, we will invoke existing human-in/on-the-loop (HITL) interactions with AFSIM, to include the distributed wargaming tool, Warlock, as well as standard protocols that interoperate with a running AFSIM simulation, to explore extended reality (XR) scenario interactions as well as augment policy training with expert HITL player actions within the environment.