VocaliD, Inc. — Department of Health and Human Services SBIR Phase II: NIDCD

VocaliD, Inc. — SBIR Phase II award from Department of Health and Human Services.

Amount
$1,810,965
Agency
Department of Health and Human Services · National Institutes of Health
Program / Phase
SBIR · Phase II
Topic
NIDCD
Solicitation
PA18-591
NAICS
Place of performance
MA
Period
2017-07-01 → 2020-06-30

Description

Our voices are not identicalthey are our identitiesThe human voice is a powerful signal that conveys one s agegendersizeethnicityand personalityamong other attributesYetuntil nowusers of augmentative and alternative communicationAACdevicesscreen reading technologies and other text to speechTTSapplications have relied on a limited set of mass producedgeneric sounding synthetic voicesThis mismatch in vocal identity impacts educational outcomesinfringes on personal safetyand hinders social integrationConventional methods for building a synthetic voice require a voice actor to record an extensive dataset of studio quality recordings which are used to train a computational model and generate the output voiceThe process is time and labor intensive and thus inaccessible to everyday consumers let alone those with speech impairmentVocaliD Inc s award winning technology offers an unprecedented means to build custom crafted synthetic voices that reflect the recipient by combining his her own residual vocalizations with recordings of a matched speaker from our crowdsourced Human VoicebankWe have discovered that even a single vowel contains enough andquot vocal DNAandquotto seed the personalization processVocaliD s custom voices sound like the recipient in agepersonality and vocal identity and have the clarity of everyday talkersHaving made significant progress towards improving the intelligibility and naturalness of our custom voices under Phase IIour voices are within a few percentage points of natural human speech in terms of intelligibility and rated as highly natural sounding by unfamiliar listenersHoweverseveral persistent issues limit the commercial potential of our current methodsFirstour new methods are computationally intensive and thus cannot be utilized on current assistive communication devicesOptimization of the methods to reduce latency and thereby improve usability is criticalAimAnother unintended consequence of advances in clarity and naturalness of our voices is the potential for misappropriationTo counteract thiswe propose developing a multi speaker model to create unique new voices and mask the identity of a given speech donorAimLastalthough the new models are capable of more prosodic variationcurrent methods rely heavily on exemplars in the training dataOur customers indicate a need and desire for greater control of subtle yet meaningful differences in prosodyAimThese additional tasks will further bolster the product and likelihood of commercial success for the AAC market and beyond VocaliD s breakthrough technology powers the first ever custom synthetic voices that are made using only brief samples of recipient vocalizations combined with recordings of matched speaker sfrom our crowdsourced voicebankThis Phase II SBIR Administrative Supplement proposal addresses the challenges of creating a scalable and efficient method for achieving high qualitynatural soundingand controllable personalized voices