X-SCALESOLUTIONS LLC — Department of Energy SBIR Phase II: C53-02a

X-SCALESOLUTIONS LLC — SBIR Phase II award from Department of Energy.

Amount
$1,650,000
Agency
Department of Energy
Program / Phase
SBIR · Phase II
Topic
C53-02a
NAICS
Place of performance
OH
Period
2023-04-03 → 2025-04-02

Description

C53-02a-271265Efficiently parallelizing algebraic solvers such as PETSc would vastly improve the performance of many critical scientific end applications. The major challenge in harnessing GPU systems for such algebraic solvers is simultaneously using all hardware resources via an optimized MPI environment, specifically instruction scheduling (i.e., kernel launches), computation, and communication. Achieving this requires MPI support for streams, efficient non-contiguous memory access, nearest neighbor collectives, communication reduction methods such as message coalescing and on-the-fly compression, and GPU-efficient non-blocking global reductions. We will design and develop SMART-PETSc: a Smart Middleware for Accelerating PETSc on modern high-performance computing hardware. We will work along the following directions: 1) develop optimized methods to overlap kernel launch and execution time to improve parallel hardware utilization, 2) co-design PETSc and MVAPICH using the MPI-T inter- face to support user-specified GPU streams, 3) support in-network communication for non-blocking and neighborhood CPU-/GPU-based collectives, 4) apply on-the-fly compression methods to reduce communication overheads, 5) improve network bandwidth via GPU-aware message coalescing designs 6) enable full-stack observability/explainability through integration with the TAU Performance System via the MPI-T interface, and 7) measure the impact of proposed designs on PETSc end applications. We demonstrated the feasibility of the SMART-PETSc product by implementing optimized datatype processing techniques, intelligent protocol selection and integration with the MPI-T Interface and TAU performance system, and optimized neighborhood collectives. The initial prototype can run successfully on 30 nodes with up to 1,000 MPI processes with the PETSc application. Initial customer engagements with the prototype product have taken place. We aim to build on top of the success of Phase-I to build the complete SMART-PETSc product while primarily focusing on the major technical goals, in-depth Q&A testing, and commercialization. The transformative impact of the proposed SMART- PETSc product will enable many HPC and DL frameworks/applications that use PETSc routines to take advantage of the novel features and capabilities on emerging hardware technologies and reap the additive benefits of co-design and tighter integration between PETSc and MVAPICH libraries. We expect that the solutions SMART-PETSc will provide can reduce the communication overhead by up to 6x. This leads to a significant boost in performance and scalability of many HPC and DL frame- works/ applications that use PETSc APIs. The proposed collaboration will enable DOE Labs to use the SMART-PETSc middleware on upcoming exascale systems and provide a commercial version for use by supercomputing centers and cloud providers worldwide.