Skip to content

Video Assisted Speech Technology to Enhance Motor Planning for Speech

Video Assisted Speech Technology to Enhance Functional Language Abilities in Individuals With Autism Spectrum Disorder

Status
Completed
Phases
NA
Study type
Interventional
Source
ClinicalTrials.gov
Registry ID
NCT04764539
Acronym
VAST
Enrollment
6
Registered
2021-02-21
Start date
2019-12-01
Completion date
2020-11-30
Last updated
2023-05-09

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Apraxia of Speech, Autism Spectrum Disorder

Keywords

autism, apraxia, speech therapy, speech pathology, nonverbal

Brief summary

Nearly 3.5 million Americans are diagnosed with Autistic Spectrum Disorder (ASD), a communication disorder that causes skill limitations in the areas of language acquisition, sensory integration, and behavior. This lack of functional language ability limits conversation to its most basic parts, making daily tasks difficult for minimally to non-verbal individuals to achieve. iTherapy is developing the VAST platform, a personalized educational experience for students with ASD by creating a virtual reality-based video-modeling program to stimulate engagement and speech production practice, ultimately providing those with ASD an opportunity to enhance their quality of life by increasing their speech abilities which will enable them to build social networks and handle the events of daily life.

Detailed description

Autism Spectrum Disorder (ASD) is a neurodevelopmental communication disorder resulting in functional language and behavioral delays affecting over 3.5 million Americans. These delays vary with the severity of symptoms that present in ASD but often result in limited speech and increased communication challenges. Alongside linguistic acquisition, oral motor coordination is a crucial part of speech production. Current clinical techniques have shown varying degrees of efficacy in improving functional language proficiency. Most techniques follow a drill-like procedure, where the child is made to repeat various sounds and phrases until they are retained. However, such a process requires potentially over twenty therapy sessions to show improvement which may then only be focused on one aspect of speech. This significantly limits the linguistic and social skills a student will acquire. To improve the efficacy of these therapy sessions, new technology must be developed to provide the most effective educational experience. Video-assisted speech technology (VAST) is a method of using a video of a close-up model of the mouth and speaking simultaneously with it. Rather than present the individual with a static photograph of the initial phoneme, the entire sequence of oral movements can be presented sequentially via video-recorded segments of the orofacial area producing connected speech, combining best practices, video modeling, and literacy with auditory cues to provide unprecedented support the development of vocabulary, word combinations and communication. In this SBIR Phase I proposal, iTherapy will develop a personalized educational experience for students with ASD by creating a virtual reality (VR) based VAST program to stimulate engagement and speech production practice. VR offers several benefits as a therapy technique: overcoming sensory difficulties, more effectively generalizing information, employing visual learning, and providing individualized treatment. As a user moves through the stages of the program, they will be immersed in a proactive environment where they will engross themselves with continuous content. Rather than present the individual with a static photograph of the initial phoneme, the entire sequence of oral movements can be presented sequentially via VR-modelled segments of the orofacial area producing connected speech, combining best practices, video modeling, music therapy, and literacy with auditory cues to provide unprecedented support the development of vocabulary, word combinations and communication. The innovation will be a video series of a realistic VR mouth which will require the use of an app on a tablet or a smartphone, VR goggles, and bone conduction headphones.

Interventions

BEHAVIORALVideo Assisted Speech Therapy (VAST)

Six children with ASD, between the ages of 4 and 8, participated in a 14-sessions-long study that utilized the VR-integrated and the tablet-based VAST application. Three subjects received a 3D VR-integrated, bone conduction VAST prototype, while the remaining group of three received a tablet with a 2D version of the software. Sessions were held twice a week with each lasting approximately 15 minutes (i.e. +/- 5 minutes).

Sponsors

National Institute on Deafness and Other Communication Disorders (NIDCD)
CollaboratorNIH
iTherapy, LLC
Lead SponsorOTHER

Study design

Allocation
RANDOMIZED
Intervention model
PARALLEL
Primary purpose
TREATMENT
Masking
NONE

Intervention model description

Six children with ASD, between the ages of 4 and 8, were recruited to participate in a 12-sessions-long study that utilized the Video-Assisted Speech Therapy (VAST) application. The participants were divided into two groups: one which received the VR-integrated prototype, and one that received a 2D application on a tablet. Each session was approximately 15 minutes long (+/- 5 minutes), occurring twice per week.

Eligibility

Sex/Gender
ALL
Age
4 Years to 8 Years
Healthy volunteers
Yes

Inclusion criteria

* Nonverbal-minimally verbal children (0-5 words) * Diagnosis of Autism Spectrum Disorder

Exclusion criteria

* No history of seizures for participating with VR goggles.

Design outcomes

Primary

MeasureTime frameDescription
Change in Mean Length of Utterance (MLU)Seven weeks--each subject participated in the study twice a week over a 7-week period for a total of 14 sessions. The first and last sessions (session #1 and session #14) were reserved for pre-test and post-test language sample collection and assessment.Participants (aged 4 to 8 years) were given a pre- and post-test 15-minute language sample. MLU was calculated for tests and gain from pre-test to post-test was compared. NOTE: This measure is calculated based on a change in the number of morphemes per utterance during pre-test and post-test language samples. During a five-minute period, two licensed speech-language pathologists (SLP) observed a parent interacting and talking with their child. Parents Both SLPs transcribed the subjects' speech and calculated a mean length of utterance (MLU) for each subject. MLU was calculated by determining how many bound and free morphemes were included within every spoken utterance produced by a subject. The total number of morphemes produced within the 5-minute period were then divided by total number of utterances, which then produced the MLU for each subject. This procedure was use for determining MLU in both the pre- and post-testing procedures.
Change in Percentage of Correctly Transcribed Words Using Automatic Speech RecognitionSeven weeks--each subject participated in the study twice a week over a 7-week period for a total of 14 sessions. The first and last sessions (session #1 and session #14) were reserved for pre-test and post-test language sample collection and assessment.15-minute pre- and post-testing was performed using speech recognition software and transcribed by a licensed speech pathologist. Differences pre and post intervention were compared across group and within groups. NOTE: During our assessment, we used Google's native closed captioning function (a tool which uses machine learning to recognize and transcribe speech) and a third party app, Tactiq Pins, which allows users to keep a transcript of all speaker utterances during a call. We compared our video to the Tactiq Pin transcripts in order to measure any change in the amount of accurately transcribed spoken words between pre-test and post-test language samples. Specific transcription results for each group can be found in the data tables provided.
Change in Articulation AccuracySeven weeks--each subject participated in the study twice a week over a 7-week period for a total of 14 sessions. The first and last sessions (session #1 and session #14) were reserved for pre-test and post-test language sample collection and assessment.Change in % of correct phonemes in each attempted stimulus

Secondary

MeasureTime frameDescription
Parent Perceptions of Communication Changes, Resulting From Study Participation.Seven weeks--each subject participated in the study twice a week over a 7-week period for a total of 14 sessions. The first and last sessions (session #1 and session #14) were reserved for pre-test and post-test language sample collection and assessment.Parent observations -- perceptions of changes in their children's motor-speech, behavioral, and social communication skills after having participated in the study Scale title: Net Positive Changes Score Maximum possible value: 18 Minimum possible value: -2 Higher score is better.
Change in Type-Token RatiosSeven weeks--each subject participated in the study twice a week over a 7-week period for a total of 14 sessions. The first and last sessions (session #1 and session #14) were reserved for pre-test and post-test language sample collection and assessment.A type-token ratio measures the total number of unique words in a given segment of language.
Increase in Response Rate to Treatment StimuliSeven weeks--each subject participated in the study twice a week over a 7-week period for a total of 14 sessions. The first and last sessions (session #1 and session #14) were reserved for pre-test and post-test language sample collection and assessment.The change in response rate measures any significant differences in how often children responded to pre- and post-testing stimuli after having received treatment between the iPad Pro and VR goggles groups. A response is considered a verbal or non-verbal reaction (e.g., eye contact, gestures, vocalizations) to the stimuli presented during the therapy sessions. Higher response rates indicate better engagement and responsiveness to the treatment. The change in response rate is calculated as the value at the post-test time point minus the value at the pre-test time point, with positive numbers representing increases and negative numbers representing decreases in response rate.

Countries

United States

Participant flow

Participants by arm

ArmCount
Stimuli Administered Via 2D Format on an iPad Pro
Three children with ASD, between the ages of 4 and 8, participated in a 14-sessions-long study that utilized the tablet-based Video-Assisted Speech Therapy (VAST) application. Sessions were held twice a week with each lasting approximately 15 minutes (i.e. +/- 5 minutes).
3
Stimuli Administered in 3D Format Via VR Goggles and Bone Conduction Headphones
Three children with ASD, between the ages of 4 and 8, participated in a 14-sessions-long study that utilized a 3D VR-integrated Video-Assisted Speech Therapy (VAST) application, which paired bone conduction audio within a 3D-printed headset prototype. Sessions were held twice a week with each lasting approximately 15 minutes (i.e. +/- 5 minutes).
3
Total6

Baseline characteristics

CharacteristicTotalStimuli Administered Via 2D Format on an iPad ProStimuli Administered in 3D Format Via VR Goggles and Bone Conduction Headphones
Age, Categorical
<=18 years
6 Participants3 Participants3 Participants
Age, Categorical
>=65 years
0 Participants0 Participants0 Participants
Age, Categorical
Between 18 and 65 years
0 Participants0 Participants0 Participants
Age, Continuous5.875 years
STANDARD_DEVIATION 1.818
5.75 years
STANDARD_DEVIATION 2.113
6 years
STANDARD_DEVIATION 1.938
Race (NIH/OMB)
American Indian or Alaska Native
0 Participants0 Participants0 Participants
Race (NIH/OMB)
Asian
3 Participants1 Participants2 Participants
Race (NIH/OMB)
Black or African American
1 Participants1 Participants0 Participants
Race (NIH/OMB)
More than one race
0 Participants0 Participants0 Participants
Race (NIH/OMB)
Native Hawaiian or Other Pacific Islander
0 Participants0 Participants0 Participants
Race (NIH/OMB)
Unknown or Not Reported
0 Participants0 Participants0 Participants
Race (NIH/OMB)
White
2 Participants1 Participants1 Participants
Region of Enrollment
United States
6 participants3 participants3 participants
Sex: Female, Male
Female
5 Participants3 Participants2 Participants
Sex: Female, Male
Male
1 Participants0 Participants1 Participants

Adverse events

Event typeEG000
affected / at risk
EG001
affected / at risk
deaths
Total, all-cause mortality
0 / 30 / 3
other
Total, other adverse events
0 / 30 / 3
serious
Total, serious adverse events
0 / 30 / 3

Outcome results

Primary

Change in Articulation Accuracy

Change in % of correct phonemes in each attempted stimulus

Time frame: Seven weeks--each subject participated in the study twice a week over a 7-week period for a total of 14 sessions. The first and last sessions (session #1 and session #14) were reserved for pre-test and post-test language sample collection and assessment.

Population: All participants for each respective group.

ArmMeasureValue (MEAN)Dispersion
Stimuli Administered Via 2D Format on an iPad ProChange in Articulation Accuracy19.75 percentage of correct phonemesStandard Deviation 28.71
Stimuli Administered in 3D Format Via VR Goggles and Bone Conduction HeadphonesChange in Articulation Accuracy16.24 percentage of correct phonemesStandard Deviation 2.36
Comparison: Articulation Accuracy does not statistically differ for the iPad Pro or VR goggle participant group after having used the VAST system for 2 months.p-value: 0.429695% CI: [-49.68, 42.66]ANOVA
Comparison: Comparison of pre-test and post-test outcome per group.p-value: <0.29495% CI: [-25.7, 65.2]ANOVA
Comparison: Comparison of pre-test and post-test outcome per group.p-value: <0.41995% CI: [-33.86, 66.36]ANOVA
Primary

Change in Mean Length of Utterance (MLU)

Participants (aged 4 to 8 years) were given a pre- and post-test 15-minute language sample. MLU was calculated for tests and gain from pre-test to post-test was compared. NOTE: This measure is calculated based on a change in the number of morphemes per utterance during pre-test and post-test language samples. During a five-minute period, two licensed speech-language pathologists (SLP) observed a parent interacting and talking with their child. Parents Both SLPs transcribed the subjects' speech and calculated a mean length of utterance (MLU) for each subject. MLU was calculated by determining how many bound and free morphemes were included within every spoken utterance produced by a subject. The total number of morphemes produced within the 5-minute period were then divided by total number of utterances, which then produced the MLU for each subject. This procedure was use for determining MLU in both the pre- and post-testing procedures.

Time frame: Seven weeks--each subject participated in the study twice a week over a 7-week period for a total of 14 sessions. The first and last sessions (session #1 and session #14) were reserved for pre-test and post-test language sample collection and assessment.

Population: All participants for each respective group.

ArmMeasureValue (MEAN)Dispersion
Stimuli Administered Via 2D Format on an iPad ProChange in Mean Length of Utterance (MLU)0.5387 Morphemes per utteranceStandard Deviation 0.4119
Stimuli Administered in 3D Format Via VR Goggles and Bone Conduction HeadphonesChange in Mean Length of Utterance (MLU)0.2987 Morphemes per utteranceStandard Deviation 0.4977
Comparison: The Mean Length of Utterance (MLU) will not increase from baseline for the either the iPad Pro or VR goggle participant group after having used the VAST system for 2 months.p-value: <0.55595% CI: [-1.276, 0.796]ANOVA
Comparison: Comparison of pre-test and post-test outcome per group.p-value: <0.212895% CI: [-0.471, 1.548]ANOVA
Comparison: Comparison of pre-test and post-test outcome per group.p-value: <0.53395% CI: [-0.915, 1.511]ANOVA
Primary

Change in Percentage of Correctly Transcribed Words Using Automatic Speech Recognition

15-minute pre- and post-testing was performed using speech recognition software and transcribed by a licensed speech pathologist. Differences pre and post intervention were compared across group and within groups. NOTE: During our assessment, we used Google's native closed captioning function (a tool which uses machine learning to recognize and transcribe speech) and a third party app, Tactiq Pins, which allows users to keep a transcript of all speaker utterances during a call. We compared our video to the Tactiq Pin transcripts in order to measure any change in the amount of accurately transcribed spoken words between pre-test and post-test language samples. Specific transcription results for each group can be found in the data tables provided.

Time frame: Seven weeks--each subject participated in the study twice a week over a 7-week period for a total of 14 sessions. The first and last sessions (session #1 and session #14) were reserved for pre-test and post-test language sample collection and assessment.

Population: All participants for each respective group.

ArmMeasureValue (MEAN)Dispersion
Stimuli Administered Via 2D Format on an iPad ProChange in Percentage of Correctly Transcribed Words Using Automatic Speech Recognition1.388 % of correctly transcribed wordsStandard Deviation 1.966
Stimuli Administered in 3D Format Via VR Goggles and Bone Conduction HeadphonesChange in Percentage of Correctly Transcribed Words Using Automatic Speech Recognition0 % of correctly transcribed wordsStandard Deviation 0
Comparison: The percentage of correctly transcribed words using automatic speech recognition will not increase from baseline for the either the iPad Pro or VR goggle participant group after having used the VAST system for 2 months.p-value: <0.28995% CI: [-1.763, 4.539]ANOVA
Secondary

Change in Type-Token Ratios

A type-token ratio measures the total number of unique words in a given segment of language.

Time frame: Seven weeks--each subject participated in the study twice a week over a 7-week period for a total of 14 sessions. The first and last sessions (session #1 and session #14) were reserved for pre-test and post-test language sample collection and assessment.

Population: All participants for each respective group.

ArmMeasureValue (MEAN)Dispersion
Stimuli Administered Via 2D Format on an iPad ProChange in Type-Token Ratios30.158 number of unique words in a segmentStandard Deviation 24.5
Stimuli Administered in 3D Format Via VR Goggles and Bone Conduction HeadphonesChange in Type-Token Ratios22.26 number of unique words in a segmentStandard Deviation 41.18
p-value: <0.78995% CI: [-68.91, 84.7]ANOVA
Comparison: Comparison of pre-test and post-test outcome per group.p-value: <0.30495% CI: [-40.85, 101.17]ANOVA
Comparison: Comparison of pre-test and post-test outcome per group.p-value: <0.40395% CI: [-43.91, 88.43]ANOVA
Secondary

Increase in Response Rate to Treatment Stimuli

The change in response rate measures any significant differences in how often children responded to pre- and post-testing stimuli after having received treatment between the iPad Pro and VR goggles groups. A response is considered a verbal or non-verbal reaction (e.g., eye contact, gestures, vocalizations) to the stimuli presented during the therapy sessions. Higher response rates indicate better engagement and responsiveness to the treatment. The change in response rate is calculated as the value at the post-test time point minus the value at the pre-test time point, with positive numbers representing increases and negative numbers representing decreases in response rate.

Time frame: Seven weeks--each subject participated in the study twice a week over a 7-week period for a total of 14 sessions. The first and last sessions (session #1 and session #14) were reserved for pre-test and post-test language sample collection and assessment.

Population: All participants for each respective group.

ArmMeasureValue (MEAN)Dispersion
Stimuli Administered Via 2D Format on an iPad ProIncrease in Response Rate to Treatment Stimuli5.67 Number of responsesStandard Deviation 3.4
Stimuli Administered in 3D Format Via VR Goggles and Bone Conduction HeadphonesIncrease in Response Rate to Treatment Stimuli3.33 Number of responsesStandard Deviation 4.03
Comparison: The response to treatment stimuli will not increase from baseline for the either the iPad Pro or VR goggle participant group after having used the VAST system for 2 months.p-value: <0.48595% CI: [-10.79, 6.11]ANOVA
Comparison: Comparison of pre-test and post-test outcome per group.p-value: <0.06595% CI: [-0.569, 11.9]ANOVA
Comparison: Comparison of pre-test and post-test outcome per group.p-value: <0.54995% CI: [-10.85, 17.53]ANOVA
Secondary

Parent Perceptions of Communication Changes, Resulting From Study Participation.

Parent observations -- perceptions of changes in their children's motor-speech, behavioral, and social communication skills after having participated in the study Scale title: Net Positive Changes Score Maximum possible value: 18 Minimum possible value: -2 Higher score is better.

Time frame: Seven weeks--each subject participated in the study twice a week over a 7-week period for a total of 14 sessions. The first and last sessions (session #1 and session #14) were reserved for pre-test and post-test language sample collection and assessment.

Population: All participants for each respective group.

ArmMeasureValue (MEAN)Dispersion
Stimuli Administered Via 2D Format on an iPad ProParent Perceptions of Communication Changes, Resulting From Study Participation.11 Score on a scaleStandard Deviation 2.45
Stimuli Administered in 3D Format Via VR Goggles and Bone Conduction HeadphonesParent Perceptions of Communication Changes, Resulting From Study Participation.9.67 Score on a scaleStandard Deviation 4.18
p-value: <0.65995% CI: [-6.436, 9.096]ANOVA

Source: ClinicalTrials.gov · Data processed: Feb 4, 2026