EPISYS SCIENCE INC — National Aeronautics and Space Administration SBIR Phase I: Z8

EPISYS SCIENCE INC — SBIR Phase I award from National Aeronautics and Space Administration.

Amount
$124,828
Agency
National Aeronautics and Space Administration
Program / Phase
SBIR · Phase I
Topic
Z8
Solicitation
SBIR_21_P1
NAICS
Place of performance
CA
Period
2021-05-12 → 2021-11-19

Description

We believe that increased autonomy and intelligent computing via multiagent AI will quot;further reduce the burden and cost of the ground segment and mission operations in CubeSatsquot; through quot;onboard data processing, autonomous systems, and navigationquot; as quoted in the 2016 Study Reportnbsp;Achieving Science with CubeSats: Thinking Inside the Box. The Decadal Survey identified scientific observation and sample return as a priority. We propose our new TeamAstro as a multiagent AI for spacecraft team coordination and sensing for science return. TeamAstro contributes to increased capabilities on each spacecraft and as a team of spacecraft swarm, increased autonomy in decisions and local actions, and thus minimized communication and control bottleneck with Earth. Our innovation TeamAstro constitutes (1) a derivative of a single DeepMindrsquo;s MuZero for spacecraft application with stringent space domain constraints of orbits, fuel, etc.; (2) enabling heterogeneous single MuZero agents to work together as a team, constituting our new TeamAstro, via attention mechanism and coordination graph learning, while keeping MuZerorsquo;s rules-discovery and Monte-Carlo Tree Search (MCTS) value search, and (3) a design reference Mars mission of TeamAstro for coordinated remote sensing between multiple spacecraft. Based on the goal configuration of the spacecraft swarm and the baseline trajectory in training datasets, TeamAstro will design detailed trajectories considering the observation/control uncertainties and constraints to reconfigure the original swarm configuration to a new one. In this case, the actions would be the thrust firing sequence of each swarm spacecraft at each time, the states would be the spacecraft position/attitude and their velocities, and the reward would be defined based on the achievement of the goal configuration. The training can be performed on the ground beforehand so that the new-optimal policy can be executed on board with a limited computational resource.