Communication Aids for Disabled, Rehabilitation of Speech and Language Disorders, Speech, Alaryngeal, Speech Disorders, Speech Intelligibility, Speech Perception
Conditions
Brief summary
This study will evaluate the ability of MyoVoice to replace natural speech. Referred to generally as an Augmentative and Alternative Communication (AAC) device, MyoVoice uses electrical signals recorded non-invasively from speech muscles (electromyographic, or EMG, signals) to restore communication for those with vocal impairments that resulted from surgical treatment of laryngeal and oropharyngeal cancers.
Detailed description
Over 7.5 million people worldwide are unable to vocalize effectively. Among these individuals are cancer survivors who underwent oropharyngeal/laryngeal surgery and must rely on AAC systems such as text-to-speech applications or artificial voice prostheses as substitutes for their natural voice. Yet most of these devices struggle to convey the expressive attributes of speech (prosody), leading to poor comprehension and a lack of emotional content. The clinical trial will investigate the feasibility of MyoVoice-a novel AAC device that uses surface EMG signals to extract patterns for understanding vocabulary and expressive attributes from articulatory musculature during silently mouthed speech-to effectively restore conversational capabilities for individuals living with vocal impairments due to surgical treatment of laryngeal and oropharyngeal cancers. Patients who underwent a total laryngectomy will be asked to communicate with a conversational partner by silently mouthing words using MyoVoice. The device performance will be evaluated in terms of its ability to accurately and quickly translate articulatory muscle activity into audible speech. MyoVoice will also be compared to that of conventional electrolaryngeal speech aids (i.e., artificial larynx) to evaluate device ease-of-use, functional efficacy, and social reception.
Interventions
Person-centric augmentative and alternative communication (AAC) system comprising a user-specific set of wearable sensors for capturing articulatory muscle activity and mobile software that provides real-time audible speech outputs.
Hand-operated electromechanical device that operates as an artificial larynx to enable a person after laryngectomy to produce speech.
Sponsors
Study design
Eligibility
Inclusion criteria
1. Individuals without Laryngectomy (Controls) Inclusion Criteria: * Primary English speaker * No history of speech, language, cognitive, or hearing disorders * Normal hearing (able to pass a bilateral hearing screening using a threshold of 25 decibels (dB) hearing level (HL) at 125, 250, 500, 1000, 2000, 4000, and 8000 Hz based on the American Speech-Language-Hearing Association) * Capable of signed informed consent
Exclusion criteria
* Inability to understand spoken English or follow simple instructions * History of speech, language, cognitive, or hearing disorders * Inability to provide written informed consent 2. Individuals with Laryngectomy Inclusion Criteria: * At least 6 months S/P total laryngectomy * Primary English speaker * Proficient with an electrolarynx (EL) * Sufficiently available and healthy to comply with multiple test sessions lasting 4-6 hours * Capable of signed informed consent
Design outcomes
Primary
| Measure | Time frame | Description |
|---|---|---|
| Pitch Recognition Accuracy | 30 mins to 2 hours | Percent accuracy between recognized (sEMG) and ground-truth (acoustic) estimates of fundamental frequency (perceptually known as vocal pitch). Accuracy is estimated on a continuous scale from 0 to 100%. |
| Loudness Recognition Accuracy | 30 mins to 2 hours | Percent accuracy between recognized (sEMG) and ground-truth (acoustic) estimates of intensity level (perceptually related to vocal loudness). Accuracy is estimated on a continuous scale from 0 to 100%. |
| Speech Naturalness | 30 mins to 1 hour | Percent naturalness between MyoVoice and electrolarynx speech samples. Speech naturalness is defined based on one's preference for how \[the audio\] sounds in terms of rate, rhythm, intonation and voice quality. Naturalness is estimated on a continuous scale from 0 to 100%. |
Countries
United States
Participant flow
Participants by arm
| Arm | Count |
|---|---|
| Development of Speech Corpora Each participant produces a list of speech tasks that include syllables, phrases, reading passages, and open-ended questions using the MyoVoice system. Myovoice pitch and loudness recognition accuracy of speech is evaluated. Duration: 30 mins to 2 hours. | 10 |
| Experimental and Reference Devices for Communication Each participant listens to speech produced by an experimental (MyoVoice) and reference (electrolarynx) systems that are used to communicate. Participants rate the naturalness of the speech produced from each device (audio files presented in a randomized order). Duration: 30 mins to 1 hour.
Experimental System: MyoVoice: Person-centric AAC system comprising a user-specific set of wearable sensors for capturing articulatory muscle activity and mobile software that provides real-time audible speech outputs.
Reference System: Electrolarynx: Hand-operated electromechanical device that operates as an artificial larynx to enable a person after laryngectomy to produce speech. | 20 |
| Total | 30 |
Baseline characteristics
| Characteristic | Experimental and Reference Devices for Communication | Total | Development of Speech Corpora |
|---|---|---|---|
| Age, Categorical <=18 years | 0 Participants | 0 Participants | 0 Participants |
| Age, Categorical >=65 years | 2 Participants | 2 Participants | 0 Participants |
| Age, Categorical Between 18 and 65 years | 18 Participants | 28 Participants | 10 Participants |
| Age, Continuous | 55 years STANDARD_DEVIATION 7.3 | 42 years STANDARD_DEVIATION 15.58 | 28 years STANDARD_DEVIATION 7.65 |
| Ethnicity (NIH/OMB) Hispanic or Latino | 0 Participants | 1 Participants | 1 Participants |
| Ethnicity (NIH/OMB) Not Hispanic or Latino | 18 Participants | 27 Participants | 9 Participants |
| Ethnicity (NIH/OMB) Unknown or Not Reported | 2 Participants | 2 Participants | 0 Participants |
| Race (NIH/OMB) American Indian or Alaska Native | 0 Participants | 0 Participants | 0 Participants |
| Race (NIH/OMB) Asian | 0 Participants | 0 Participants | 0 Participants |
| Race (NIH/OMB) Black or African American | 0 Participants | 0 Participants | 0 Participants |
| Race (NIH/OMB) More than one race | 1 Participants | 1 Participants | 0 Participants |
| Race (NIH/OMB) Native Hawaiian or Other Pacific Islander | 0 Participants | 0 Participants | 0 Participants |
| Race (NIH/OMB) Unknown or Not Reported | 2 Participants | 2 Participants | 0 Participants |
| Race (NIH/OMB) White | 17 Participants | 27 Participants | 10 Participants |
| Region of Enrollment United States | 20 Participants | 30 Participants | 10 Participants |
| Sex: Female, Male Female | 9 Participants | 14 Participants | 5 Participants |
| Sex: Female, Male Male | 11 Participants | 16 Participants | 5 Participants |
Adverse events
| Event type | EG000 affected / at risk | EG001 affected / at risk |
|---|---|---|
| deaths Total, all-cause mortality | 0 / 10 | 0 / 20 |
| other Total, other adverse events | 0 / 10 | 0 / 20 |
| serious Total, serious adverse events | 0 / 10 | 0 / 20 |
Outcome results
Loudness Recognition Accuracy
Percent accuracy between recognized (sEMG) and ground-truth (acoustic) estimates of intensity level (perceptually related to vocal loudness). Accuracy is estimated on a continuous scale from 0 to 100%.
Time frame: 30 mins to 2 hours
Population: The sEMG and acoustic data collected from participants in the Development of Speech Corpora arm were used to develop and evaluate an algorithm to predict vocal loudness (acoustic correlate: intensity) from sEMG data. Higher accuracy scores equate to better performance.
| Arm | Measure | Value (MEAN) | Dispersion |
|---|---|---|---|
| Development of Speech Corpora | Loudness Recognition Accuracy | 12.1 Percent agreement to reference loudness | Standard Deviation 3 |
Pitch Recognition Accuracy
Percent accuracy between recognized (sEMG) and ground-truth (acoustic) estimates of fundamental frequency (perceptually known as vocal pitch). Accuracy is estimated on a continuous scale from 0 to 100%.
Time frame: 30 mins to 2 hours
Population: The sEMG and acoustic data collected from participants in the Development of Speech Corpora arm were used to develop and evaluate an algorithm to predict vocal pitch (acoustic correlate: fundamental frequency) from sEMG data. Higher accuracy scores equate to better performance.
| Arm | Measure | Value (MEAN) | Dispersion |
|---|---|---|---|
| Development of Speech Corpora | Pitch Recognition Accuracy | 7.70 Percent agreement to reference pitch | Standard Deviation 1.63 |
Speech Naturalness
Percent naturalness between MyoVoice and electrolarynx speech samples. Speech naturalness is defined based on one's preference for how \[the audio\] sounds in terms of rate, rhythm, intonation and voice quality. Naturalness is estimated on a continuous scale from 0 to 100%.
Time frame: 30 mins to 1 hour
Population: Participants in the Experimental and Reference Devices for Communication rated the naturalness of each sample on a visual analog scale from 0 to 100, which 0 anchored at least natural and 100 anchored at most natural. Higher naturalness scores equate to better performance.
| Arm | Measure | Group | Value (MEAN) | Dispersion |
|---|---|---|---|---|
| Experimental and Reference Devices for Communication | Speech Naturalness | MyoVoice | 66.2 Percent of high naturalness ratings | Standard Deviation 22.6 |
| Experimental and Reference Devices for Communication | Speech Naturalness | Electrolarynx | 9.7 Percent of high naturalness ratings | Standard Deviation 19.8 |
| Experimental and Reference Devices for Communication | Speech Naturalness | Overall | 38.0 Percent of high naturalness ratings | Standard Deviation 35.4 |