Speech Sound Disorder
Conditions
Keywords
speech, articulation, motor development
Brief summary
Children with speech sound disorder show diminished accuracy and intelligibility in spoken communication and may thus be perceived as less capable or intelligent than peers, with negative consequences for both socioemotional and socioeconomic outcomes. While most speech errors resolve by the late school-age years, between 2-5% of speakers exhibit residual speech errors (RSE) that persist through adolescence or even adulthood, reflecting about 6 million cases in the US. Both affected children/families and speech-language pathologists (SLPs) have highlighted the critical need for research to identify more effective forms of treatment for children with RSE. In a series of single-case experimental studies, research has found that treatment incorporating technologically enhanced sensory feedback (visual-acoustic biofeedback, ultrasound biofeedback) can improve speech in individuals with RSE who have not responded to previous intervention. A randomized controlled trial (RCT) comparing traditional vs biofeedback-enhanced intervention is the essential next step to inform evidence-based decision-making for this prevalent population. Larger-scale research is also needed to understand heterogeneity across individuals in the magnitude of response to biofeedback treatment. The overall objective of this proposal is to conduct clinical research that will guide the evidence-based management of RSE while also providing novel insights into the sensorimotor underpinnings of speech. The central hypothesis is that biofeedback will yield greater gains in speech accuracy than traditional treatment, and that individual deficit profiles will predict relative response to visual-acoustic vs ultrasound biofeedback. This study will enroll n = 118 children who misarticulate the /r/ sound, the most common type of RSE. This first component of the study will evaluate the efficacy of biofeedback relative to traditional treatment in a well-powered randomized controlled trial. Ultrasound and visual-acoustic biofeedback, which have similar evidence bases, will be represented equally.
Detailed description
Randomized Trial Component: Previous findings suggest that biofeedback interventions can outperform traditional speech therapy for children with RSE, but the research base to date is limited to small-scale studies that do not reach the level of evidence needed to support large-scale changes in practice. The primary objective of the C-RESULTS RCT is to test the working hypothesis that a group of individuals randomly assigned to receive biofeedback-enhanced treatment will show larger and/or faster gains in /r/ production accuracy than an equivalent group receiving the same dose of non-biofeedback treatment. To test this hypothesis, n=110 children will be randomly assigned to receive a standard course of intervention with or without biofeedback. Acoustic and perceptual measures will be used to test for differences in both short-term learning of treated targets (Acquisition) and longer-term carryover of learning to untreated contexts (Generalization). In addition, a survey assessing participants' socio-emotional well-being will be collected from caregivers both pre and post treatment.
Interventions
In ultrasound biofeedback, the elements of traditional treatment (auditory models and verbal descriptions of articulator placement) are enhanced with a real-time ultrasound display of the shape and movements of the tongue. One or two target tongue shapes will be selected for each participant, and a trace of the selected target will be superimposed over the ultrasound screen. Participants will be cued to reshape the tongue to match this target during /r/ production.
In visual-acoustic biofeedback treatment, the elements of traditional treatment (auditory models and verbal descriptions of articulator placement) are enhanced with a dynamic display of the speech signal in the form of the real-time LPC (Linear Predictive Coding) spectrum. Because correct vs incorrect productions of /r/ contrast acoustically in the frequency of the third formant (F3), participants will be cued to make their real-time LPC spectrum match a visual target characterized by a low F3 frequency. They will be encouraged to attend to the visual display while adjusting the placement of their articulators and observing how those adjustments impact F3.
Traditional articulation treatment involves providing auditory models and verbal descriptions of correct articulator placement, then cueing repetitive motor practice. Images and diagrams of the vocal tract will be used as visual aids; however, no real-time visual display of articulatory or acoustic information will be made available.
Sponsors
Study design
Masking description
All perceptual ratings will be obtained from blinded, naive listeners recruited through online crowdsourcing. Following protocols refined in previous published research, binary rating responses will be aggregated over at least 9 unique listeners per token.
Intervention model description
All participants will complete a Dynamic Assessment session (Phase 0) consisting of 2 hours of traditional (non-biofeedback) instruction. Participants will be categorized into high, moderate, and low response groups based on performance in Phase 0, and the response groups will be block randomized to traditional or biofeedback speech treatment. Within the biofeedback condition, individuals will be sub-randomized in equal numbers to receive visual-acoustic or ultrasound treatment. Participants will then complete two phases of speech treatment in their randomly assigned condition. Phase 1 (Acquisition) will consist of high-intensity, highly interactive practice delivered in three 90-minute sessions over one week. Phase 2 (Generalization) will elicit structured practice of /r/ in 16 semiweekly 45-minute sessions over 8 weeks.
Eligibility
Inclusion criteria
* Must be between 9;0 and 15;11 years of age at the time of enrollment. * Must speak English as the dominant language (i.e., must have begun learning English by age 2, per parent report). * Must speak a rhotic dialect of English. * Must pass a pure-tone hearing screening at 20 decibels Hearing Level (HL). * Must pass a brief examination of oral structure and function. * Must exhibit less than thirty percent accuracy, based on trained listener ratings, on a probe list eliciting /r/ in various phonetic contexts at the word level.
Exclusion criteria
* Must not receive a T score more than 1.3 standard deviations (SD) below the mean on the Wechsler Abbreviated Scale of Intelligence-2 (WASI-2) Matrix Reasoning. * Must not receive a standard score below 80 on the Core Language Index of the Clinical Evaluation of Language Fundamentals-5 (CELF-5). * Must not exhibit voice or fluency disorder of a severity judged likely to interfere with the ability to participate in study activities. * Must not have an existing diagnosis of developmental disability or major neurobehavioral syndrome such as cerebral palsy, Down Syndrome, or Autism Spectrum Disorder, or major neural disorder (e.g., epilepsy, agenesis of the corpus callosum) or insult (e.g., traumatic brain injury, stroke, or tumor resection). * Must not show clinically significant signs of apraxia of speech or dysarthria.
Design outcomes
Primary
| Measure | Time frame | Description |
|---|---|---|
| Change in F3-F2 Distance (Hz) Across Sessions, Measured From /r/ Sounds Produced in Syllables or Words During Practice. | Phase I: three 90-min treatment sessions delivered over ~1 week; reported value is the change from Session 1 to Session 3 (slope across sessions) | F3-F2 distance is a number (in Hz) that reflects how close a child's /r/ sound is to a typical adult-like /r/. Smaller numbers indicate more accurate /r/ production; larger numbers indicate a distorted /r/. In typical peers, accurate /r/ is roughly \ 500 Hz, whereas distorted /r/ values are often \>1000 Hz. During Phase I (3 sessions over \ 1 week), children produced /r/ in syllables/words. For this Outcome, we report change across sessions: a single model-based estimate of how much F3-F2 decreased from Session 1 to Session 3 (i.e., the rate of improvement). A more negative change indicates greater improvement. |
Secondary
| Measure | Time frame | Description |
|---|---|---|
| Change From Pre to Post in Percent Correct Ratings by Untrained Listeners, for /r/ Sounds Produced in Word Probes. | Pre (before initiation of treatment) and Post (after the end of all treatment; ~10 weeks later). | The outcome is the percentage of untrained listeners, blinded to time point and treatment condition, who judged each /r/ production as correct from word-probe recordings (0-100%; higher = better). Children completed word probes at Pre and Post. Results in the table summarize change from Pre to Post using a mixed-effects model: specifically, the treatment × time interaction, which estimates the between-group difference in improvement from Pre to Post. A positive change indicates improvement. |
| Impact of Speech Disorder on Social, Emotional, and Academic Well-being (Parent Survey) | Pre (before initiation of treatment) and Post (after completion of all treatment; ~10 weeks later) | Parents completed a questionnaire assessing the impact of their child's speech disorder on social, emotional, and academic well-being. Each item was rated on a 5-point scale (1 = strongly disagree, 3 = neutral, 5 = strongly agree). Scores were averaged across items to yield an overall impact score ranging from 1 to 5, with higher values indicating a greater negative impact. A decrease from Pre to Post indicates improvement. |
Countries
United States
Participant flow
Participants by arm
| Arm | Count |
|---|---|
| Group 1: Traditional Articulation Treatment Traditional articulation treatment | 45 |
| Group 2: Biofeedback--visual-acoustic Biofeedback--visual-acoustic | 32 |
| Group 3: Biofeedback-Ultrasound Biofeedback-ultrasound | 31 |
| Total | 108 |
Withdrawals & dropouts
| Period | Reason | FG000 | FG001 | FG002 |
|---|---|---|---|---|
| Phase 2 (Generalization) | Withdrawal by Subject | 4 | 1 | 1 |
Baseline characteristics
| Characteristic | Total | Group 2: Biofeedback--visual-acoustic | Group 3: Biofeedback-Ultrasound | Group 1: Traditional Articulation Treatment |
|---|---|---|---|---|
| Age, Continuous | 10.8 years STANDARD_DEVIATION 1.57 | 10.86 years STANDARD_DEVIATION 1.87 | 10.71 years STANDARD_DEVIATION 1.3 | 10.72 years STANDARD_DEVIATION 1.54 |
| Ethnicity (NIH/OMB) Hispanic or Latino | 8 Participants | 1 Participants | 3 Participants | 4 Participants |
| Ethnicity (NIH/OMB) Not Hispanic or Latino | 94 Participants | 29 Participants | 25 Participants | 40 Participants |
| Ethnicity (NIH/OMB) Unknown or Not Reported | 6 Participants | 2 Participants | 3 Participants | 1 Participants |
| Race (NIH/OMB) American Indian or Alaska Native | 0 Participants | 0 Participants | 0 Participants | 0 Participants |
| Race (NIH/OMB) Asian | 2 Participants | 0 Participants | 1 Participants | 1 Participants |
| Race (NIH/OMB) Black or African American | 1 Participants | 0 Participants | 1 Participants | 0 Participants |
| Race (NIH/OMB) More than one race | 6 Participants | 0 Participants | 1 Participants | 5 Participants |
| Race (NIH/OMB) Native Hawaiian or Other Pacific Islander | 0 Participants | 0 Participants | 0 Participants | 0 Participants |
| Race (NIH/OMB) Unknown or Not Reported | 5 Participants | 1 Participants | 1 Participants | 3 Participants |
| Race (NIH/OMB) White | 94 Participants | 31 Participants | 27 Participants | 36 Participants |
| Response to dynamic assessment session High initial responders | 50 Participants | 14 Participants | 15 Participants | 21 Participants |
| Response to dynamic assessment session Low initial responders | 58 Participants | 18 Participants | 16 Participants | 24 Participants |
| Sex: Female, Male Female | 40 Participants | 16 Participants | 9 Participants | 15 Participants |
| Sex: Female, Male Male | 68 Participants | 16 Participants | 22 Participants | 30 Participants |
Adverse events
| Event type | EG000 affected / at risk | EG001 affected / at risk | EG002 affected / at risk |
|---|---|---|---|
| deaths Total, all-cause mortality | 0 / 45 | 0 / 32 | 0 / 31 |
| other Total, other adverse events | 0 / 45 | 0 / 32 | 0 / 31 |
| serious Total, serious adverse events | 0 / 45 | 0 / 32 | 0 / 31 |
Outcome results
Change in F3-F2 Distance (Hz) Across Sessions, Measured From /r/ Sounds Produced in Syllables or Words During Practice.
F3-F2 distance is a number (in Hz) that reflects how close a child's /r/ sound is to a typical adult-like /r/. Smaller numbers indicate more accurate /r/ production; larger numbers indicate a distorted /r/. In typical peers, accurate /r/ is roughly \ 500 Hz, whereas distorted /r/ values are often \>1000 Hz. During Phase I (3 sessions over \ 1 week), children produced /r/ in syllables/words. For this Outcome, we report change across sessions: a single model-based estimate of how much F3-F2 decreased from Session 1 to Session 3 (i.e., the rate of improvement). A more negative change indicates greater improvement.
Time frame: Phase I: three 90-min treatment sessions delivered over ~1 week; reported value is the change from Session 1 to Session 3 (slope across sessions)
Population: Three participants from group 2 (visual-acoustic biofeedback) could not be acoustically analyzed due to poor audio quality.
| Arm | Measure | Value (MEAN) | Dispersion |
|---|---|---|---|
| Group 1: Traditional Articulation Treatment | Change in F3-F2 Distance (Hz) Across Sessions, Measured From /r/ Sounds Produced in Syllables or Words During Practice. | 38.2 Hertz / session | Standard Deviation 66 |
| Group 2: Biofeedback--visual-acoustic | Change in F3-F2 Distance (Hz) Across Sessions, Measured From /r/ Sounds Produced in Syllables or Words During Practice. | 66.7 Hertz / session | Standard Deviation 103 |
| Group 3: Biofeedback-Ultrasound | Change in F3-F2 Distance (Hz) Across Sessions, Measured From /r/ Sounds Produced in Syllables or Words During Practice. | 85.2 Hertz / session | Standard Deviation 152 |
Change From Pre to Post in Percent Correct Ratings by Untrained Listeners, for /r/ Sounds Produced in Word Probes.
The outcome is the percentage of untrained listeners, blinded to time point and treatment condition, who judged each /r/ production as correct from word-probe recordings (0-100%; higher = better). Children completed word probes at Pre and Post. Results in the table summarize change from Pre to Post using a mixed-effects model: specifically, the treatment × time interaction, which estimates the between-group difference in improvement from Pre to Post. A positive change indicates improvement.
Time frame: Pre (before initiation of treatment) and Post (after the end of all treatment; ~10 weeks later).
Population: The analysis population differs from the number of participants who completed the Secondary Outcome Assessment visit (Phase 2) because two participants contributed unusable secondary-outcome data (one due to corrupted audio files, and one who completed all treatment sessions but whose number of trials completed did not pass the threshold for inclusion in the analysis).
| Arm | Measure | Value (MEAN) | Dispersion |
|---|---|---|---|
| Group 1: Traditional Articulation Treatment | Change From Pre to Post in Percent Correct Ratings by Untrained Listeners, for /r/ Sounds Produced in Word Probes. | 44.5 Percent words rated correct | Standard Deviation 34 |
| Group 2: Biofeedback--visual-acoustic | Change From Pre to Post in Percent Correct Ratings by Untrained Listeners, for /r/ Sounds Produced in Word Probes. | 43.2 Percent words rated correct | Standard Deviation 33 |
| Group 3: Biofeedback-Ultrasound | Change From Pre to Post in Percent Correct Ratings by Untrained Listeners, for /r/ Sounds Produced in Word Probes. | 49.6 Percent words rated correct | Standard Deviation 34 |
Impact of Speech Disorder on Social, Emotional, and Academic Well-being (Parent Survey)
Parents completed a questionnaire assessing the impact of their child's speech disorder on social, emotional, and academic well-being. Each item was rated on a 5-point scale (1 = strongly disagree, 3 = neutral, 5 = strongly agree). Scores were averaged across items to yield an overall impact score ranging from 1 to 5, with higher values indicating a greater negative impact. A decrease from Pre to Post indicates improvement.
Time frame: Pre (before initiation of treatment) and Post (after completion of all treatment; ~10 weeks later)
Population: The number of participants analyzed for secondary outcome measure #3 is the same as the number of participants who completed Phase 2. The data quality-related exclusions that impacted the secondary outcome measure #2 did not impact secondary outcome measure #3.
| Arm | Measure | Value (MEAN) | Dispersion |
|---|---|---|---|
| Group 1: Traditional Articulation Treatment | Impact of Speech Disorder on Social, Emotional, and Academic Well-being (Parent Survey) | 1.9 Impact score | Standard Deviation 1.1 |
| Group 2: Biofeedback--visual-acoustic | Impact of Speech Disorder on Social, Emotional, and Academic Well-being (Parent Survey) | 1.9 Impact score | Standard Deviation 1.1 |
| Group 3: Biofeedback-Ultrasound | Impact of Speech Disorder on Social, Emotional, and Academic Well-being (Parent Survey) | 2.1 Impact score | Standard Deviation 1.2 |