KITWARE INC — Department of Defense SBIR Phase I: A20-051

KITWARE INC — SBIR Phase I award from Department of Defense.

Amount
$102,029
Agency
Department of Defense · Army
Program / Phase
SBIR · Phase I
Topic
A20-051
Solicitation
20.1
NAICS
Place of performance
NY
Period
2020-06-19 → 2021-02-28

Description

Complex, chaotic environments are a common feature of the modern battlefield. Threats can exist at various elevations, ranges, and headings and can use civilian crowds to mask their movements and intentions. Maintaining situational awareness in these conditions can be overwhelming. To autonomously attend to all the potential threats within an environment, we propose the Panoptic Guardian-VTA (Video Threat Analysis) processing system. Panoptic Guardian-VT employs a modular processing architecture and hierarchical inference scheme to achieve real-time detection, classification, and tracking of military targets. It will operate in complex, cluttered environments with moving people and vehicles, both day and night. In contrast to existing systems, Panoptic Guardian-VTA identifies objects that not only look like but also move and behave like threats by utilizing state-of-the-art deep-learning models. Deep learning approaches to object and activity recognition are continuously pushing the state of the art. The computer vision community has driven this success by amassing huge, diverse, labeled video datasets by scouring the Internet and video sharing websites. However, this organic approach to data diversity is hard to replicate in the niche domain of target and threat detection from infrared video. Planned data collections are invariably limited to the diversity of the background environment and the total number of unique entities represented. Therefore, training robust detectors in this domain hinges on the development of intelligent data-collection techniques and data-augmentation algorithms to artificially recreate the diversity in the labeled IR video datasets that deep-learning requires. Panoptic Guardian-VTA will create threat detectors in infrared imagery by (1) developing efficient, semi-automated image annotation methods, (2) using image compositing to increase the variety while decreasing the cost of training data collection, and (3) developing advanced deep neural networks architectures based on state-of-the-art methods. Our automated image annotation will use co-boresighted, calibrated RGB and IR cameras.  Modern RGB detectors, trained on large open datasets, will be used to bootstrap IR training data generation, reducing annotation labor costs. Green-screen-like compositing will extract actors from controlled IR videos and place them in scenes in realistic poses based on 3D and semantic knowledge extracted from the scenes. Our advanced architectures will use a multi-stage approach, similar to region proposal networks, to find people, and then in those locations look for weapons and analyze motion. We will combine our various threat cues into a single threat level score that can alert warfighters to danger, reducing mental load and increasing blue force safety and lethality.