Skip to content

Flexible Representation of Speech

Flexible Representation of Speech in the Supratemporal Plane

Status
Completed
Phases
NA
Study type
Interventional
Source
ClinicalTrials.gov
Registry ID
NCT05209386
Enrollment
48
Registered
2022-01-26
Start date
2022-05-02
Completion date
2025-04-10
Last updated
2025-09-08

For informational purposes only — not medical advice. Sourced from public registries and may not reflect the latest updates. Terms

Conditions

Epilepsy

Keywords

sEEG, Speech representation, Supratemporal plane

Brief summary

The overarching goal of this exploratory research is to understand the dynamic and flexible nature of speech processing in the human supratemporal plane. The temporal lobe has long been established as a region of interest in the speech perception and processing literature because it contains the auditory cortex. More recently, research has localized the supratemporal plane as an area that exhibits response specificity to acoustic properties of complex auditory signals like speech. The supratemporal plane, comprised of Heschl's gyrus, the planum polare, and the planum temporale, is capable of the rapid spectrotemporal analysis required to map acoustic information to linguistic representation. Neural activity in this area, however, is rarely studied directly because it is difficult to access with non-invasive measures like scalp electroencephalography (EEG). Capitalizing on the unique opportunity to access these areas via routine clinical stereoelectroencephalography (sEEG) in a patient population, this study seeks to understand how cortical responses reflect the diagnosticity of two acoustic-phonetic dimensions of interest and how responses rapidly and flexibly adapt to changes in listening demands. Examining how neural response to voice onset time (VOT) and fundamental frequency (F0) modulates as a function of perceptual weight carried in signaling phoneme categories, and identifying how changes in listening context shift perceptual weight, will provide invaluable data that indicates how speech processing flexibly adapts to short-term acoustic patterns.

Detailed description

The purpose of this study is to understand the dynamic, flexible nature of speech processing as a function of perceptual weight applied to acoustic-phonetic dimensions within varying listening contexts and demands. The specific aims of this study are as follows: 1. To establish the neural response to two acoustic-phonetic dimensions as a function of the perceptual weight they carry when signaling phoneme identity. Aim 1 will specifically evaluate responses to voice onset time (VOT) and fundamental frequency (F0). Data collected will provide a baseline response for participants. 2. To identify how experimental manipulation of listening context impacts perceptual weighting strategies of VOT and F0. Aim 2 will evaluate modulation of neural response to the introduction of noise and the introduction of an accent. A secondary aim of this study is to use control electrodes, which are those placed in clinically necessary regions of the brain but outside of the region of interest for this study (supratemporal plane), to determine if additional regions of the brain are implicated in adaptive plasticity of speech processing. Speech is the primary means by which we convey our needs, wants, and thoughts to others and the ability to process speech is crucial to our everyday functioning, as well as our ability to establish and maintain relationships. Impairments in speech processing have an undeniable negative impact on individuals and society. While habilitative and rehabilitative strategies exist that can improve auditory processing and quality of life, understanding the exact neural mechanism underlying the human brain's ability to process speech would contribute to a more well-defined means by which to target deficits. This study seeks to understand the regions of the brain involved in speech processing, how those regions analyze specific acoustic-phonetic dimensions, and how the system adapts to successfully process speech in different listening contexts. Modern electrophysiological techniques have revolutionized research into activity in the human brain, allowing investigators to identify specific regions or patterns of activity associated with various behaviors and sensory experiences. sEEG recordings, which involve intracerebral measurements of neural activity using depth electrodes, are capable of providing unique access to regions of the brain that are otherwise inaccessible with less invasive measurements. Capitalizing on PI Abel's work with pediatric patients undergoing sEEG recording for localization of seizure foci or language mapping, this study will allow researchers to directly study activity in regions that have already been implicated in the literature as crucial to spectrotemporal analysis of complex acoustic signals like speech. These regions within the supratemporal plane (STP) include Heschl's gyrus, the planum polare, and the planum temporale, all of which are uniquely targeted via sEEG. Existing literature in speech processing has indicated that the mapping of physical input (acoustic signal) to linguistic representation (identification of phonemes or words) is not a static process, but rather highly dependent on listening context. The auditory processing system regularly adapts to changes in signal quality, adverse listening conditions, and short-term deviations from expected and learned regularities in native language input by applying varying importance, or perceptual weight, to specific acoustic-phonetic parameters. This indicates the existence of adaptive plasticity in speech processing, yet existing neurophysiological models do not account for this flexibility in cortical response. Data from pilot EEG and sEEG studies demonstrated that high gamma activity in the STP and behavioral responses were graded by the perceptual weight given to two acoustic-phonetic dimensions, voice onset time (VOT) and fundamental frequency (F0). The proposed study will contribute to existing knowledge by helping to establish a more detailed model of on-line cortical response and adaptation to changing acoustic signals. It is unique in its accounting for the role that perceptual weight of acoustic-phonetic dimensions play in signaling phonemes and making category-based judgments.

Interventions

BEHAVIORALDimension-Based Statistical Learning

Each participant will complete self-paced blocks of stimuli that will first establish a baseline for neural activity and behavioral responses with clear speech, and will then record responses for experimentally manipulated blocks to introduce 1) speech-in-noise and 2) a Canonical-Reverse block to model an accent. Auditory stimuli will be adjusted to a comfortable level for each participant as determined by a calibration process completed by the participant. Each block involves listening to sound via earphones and making a categorical decision between initial consonants (/b/ or /p/) by tapping a button to indicate the word heard by the participant.

Sponsors

National Institutes of Health (NIH)
CollaboratorNIH
Carnegie Mellon University
CollaboratorOTHER
National Institute on Deafness and Other Communication Disorders (NIDCD)
CollaboratorNIH
University of Pittsburgh
Lead SponsorOTHER

Study design

Allocation
NA
Intervention model
SINGLE_GROUP
Primary purpose
BASIC_SCIENCE
Masking
NONE

Eligibility

Sex/Gender
ALL
Age
15 Years to 25 Years
Healthy volunteers
No

Inclusion criteria

* Individuals 15-25 years old * Undergoing sEEG placement in the supratemporal plane for clinically necessary localization of epileptic foci or language mapping * Fluent English speakers * Cognition and speech-language skills within normal limits (as determined by evaluation prior to surgery) * Normal or correct-to-normal visual acuity * Normal hearing acuity in each ear (as determined by audiometric assessment) * No history of autism or ADHD

Exclusion criteria

* Individuals with intellectual disabilities * Abnormal epileptiform activity in the supratemporal plane * Lack of fluent English comprehension/production * Severe language or auditory-specific cognitive dysfunction * History of autism or ADHD

Design outcomes

Primary

MeasureTime frameDescription
Supratemporal Neural Response to Change in Acoustic-Phonetic DimensionsDuring sEEG-EEG recording sessions, up to 3 hours totalNeural activity will be measured via simultaneous EEG-sEEG monitoring in the supratemporal plane as indicated by high-gamma band activity in the electrical signal. Neural activity will be measured as participants listen to acoustic stimuli with gradually manipulated acoustic dimensions, fundamental frequency (F0) and voice onset time (VOT). The data reported here is the number of temporal lobe channels demonstrating significant encoding of change in acoustic dimension (F0 as VOT is held constant).
Behavioral Impact of Change in Acoustic-Phonetic DimensionsDuring sEEG-EEG recording sessions, up to 3 hours totalBehavioral responses in the form of a category judgment will be obtained as participants listen to acoustic stimuli in with gradually varying fundamental frequency (F0) and voice onset time (VOT). Participants will provide a behavioral response by indicating the phoneme perceived at the beginning of stimulus words (/b/ or /p/). Specifically, the outcome is reported as percent of stimuli classified as /p/ over varying F0 with VOT held constant.
Supratemporal Neural Response to Change in Listening ContextDuring sEEG-EEG recording sessions, up to 3 hours totalNeural activity was measured via sEEG monitoring as indicated by high-gamma band activity in the electrical signal. Neural activity will be measured as participants listen to acoustic stimuli in accented speech. The un-transformed voltage represents the difference in electric potential between a specific electrode contact and the reference electrode contact; since we use a common average reference, that means it's the difference between a specific electrode contact and the mean voltage across all electrodes. We then z-score to characterize shifts from baseline activity (where z = 0) at a specific electrode, which is believed to measure changes in voltage due largely to post-synaptic currents. This is a mathematical transformation rather than a published scale or standardized assessment. More extreme z-scores (+ or -) indicate a greater change from baseline local neural activity; in other words, stimuli evoked greater activity in this region.
Behavioral Impact of Change in Listening ContextDuring sEEG-EEG recording sessions, up to 3 hours totalBehavioral responses in the form of a category judgment will be obtained as participants listen to acoustic stimuli in a varied listening context: accented speech. Participants will provide a behavioral response by indicating the phoneme perceived at the beginning of stimulus words (/b/ or /p/). Specifically, the outcome is reported as the mean % of stimuli classified as /p/ with varying VOT and F0 relationships.

Secondary

MeasureTime frameDescription
Neural Response of Non-Regions of Interest to Change in Acoustic-DimensionDuring sEEG-EEG recording sessions, up to 3 hours totalNeural activity will be measured via simultaneous EEG-sEEG monitoring in the supratemporal plane as indicated by high-gamma band activity in the electrical signal. Neural activity will be measured as participants listen to acoustic stimuli with gradually manipulated acoustic dimensions, fundamental frequency (F0) and voice onset time (VOT). The data reported here is the number of temporal lobe channels demonstrating significant encoding of change in acoustic dimension (F0 as VOT is held constant).
Neural Response of Non-Regions of Interest to Change in Listening ContextDuring sEEG-EEG recording sessions, up to 3 hours totalNeural activity was measured via sEEG monitoring as indicated by high-gamma band activity in the electrical signal. Neural activity will be measured as participants listen to acoustic stimuli in accented speech. The un-transformed voltage represents the difference in electric potential between a specific electrode contact and the reference electrode contact; since we use a common average reference, that means it's the difference between a specific electrode contact and the mean voltage across all electrodes. We then z-score to characterize shifts from baseline activity (where z = 0) at a specific electrode, which is believed to measure changes in voltage due largely to post-synaptic currents. This is a mathematical transformation rather than a published scale or standardized assessment. More extreme z-scores (+ or -) indicate a greater change from baseline local neural activity; in other words, stimuli evoked greater activity in this region.

Countries

United States

Participant flow

Recruitment details

Participants were recruited from PI Abel's clinical population of patients undergoing invasive monitoring (sEEG) for drug-resistant epilepsy at Children's Hospital of Pittsburgh between May 2022-May 2025.

Pre-assignment details

As there is only a single experimental group in this study, all enrolled patients were assigned to Arm 1 (Patient Participants).

Participants by arm

ArmCount
Patient Participants
This single-group study will recruit patients through the PI's clinical practice who are undergoing invasive neurophysiological monitoring (sEEG) with clinically necessary placement of electrodes in the supratemporal plane. All participants will complete the same behavioral response paradigms.
16
Total16

Withdrawals & dropouts

PeriodReasonFG000
Overall StudyPeriod of sEEG implant insufficient to collect complete behavioral task24
Overall StudyWithdrawal by Subject8

Baseline characteristics

CharacteristicPatient Participants
Age, Continuous17.62 years
STANDARD_DEVIATION 3.1
Ethnicity (NIH/OMB)
Hispanic or Latino
0 Participants
Ethnicity (NIH/OMB)
Not Hispanic or Latino
14 Participants
Ethnicity (NIH/OMB)
Unknown or Not Reported
2 Participants
Phoneme Categorization81.5 percent of stimuli classified as /p/
STANDARD_DEVIATION 4.1
Race (NIH/OMB)
American Indian or Alaska Native
0 Participants
Race (NIH/OMB)
Asian
0 Participants
Race (NIH/OMB)
Black or African American
0 Participants
Race (NIH/OMB)
More than one race
0 Participants
Race (NIH/OMB)
Native Hawaiian or Other Pacific Islander
0 Participants
Race (NIH/OMB)
Unknown or Not Reported
0 Participants
Race (NIH/OMB)
White
16 Participants
Sex: Female, Male
Female
6 Participants
Sex: Female, Male
Male
10 Participants

Adverse events

Event typeEG000
affected / at risk
deaths
Total, all-cause mortality
0 / 48
other
Total, other adverse events
0 / 48
serious
Total, serious adverse events
0 / 48

Outcome results

Primary

Behavioral Impact of Change in Acoustic-Phonetic Dimensions

Behavioral responses in the form of a category judgment will be obtained as participants listen to acoustic stimuli in with gradually varying fundamental frequency (F0) and voice onset time (VOT). Participants will provide a behavioral response by indicating the phoneme perceived at the beginning of stimulus words (/b/ or /p/). Specifically, the outcome is reported as percent of stimuli classified as /p/ over varying F0 with VOT held constant.

Time frame: During sEEG-EEG recording sessions, up to 3 hours total

Population: 2 participants were excluded because they performed \<50% on catch trials (0 and 40 ms VOTs); 1 participant was excluded because of a data quality issue

ArmMeasureValue (MEAN)Dispersion
Patient ParticipantsBehavioral Impact of Change in Acoustic-Phonetic Dimensions37.4 percent of stimuli classified as /p/Standard Error 6.4
p-value: <0.0001t-test, 2 sided
Primary

Behavioral Impact of Change in Listening Context

Behavioral responses in the form of a category judgment will be obtained as participants listen to acoustic stimuli in a varied listening context: accented speech. Participants will provide a behavioral response by indicating the phoneme perceived at the beginning of stimulus words (/b/ or /p/). Specifically, the outcome is reported as the mean % of stimuli classified as /p/ with varying VOT and F0 relationships.

Time frame: During sEEG-EEG recording sessions, up to 3 hours total

Population: 2 participants were excluded because they performed \<50% on catch trials (0 and 40 ms VOTs); 1 participant was excluded because of a data quality issue

ArmMeasureGroupValue (MEAN)Dispersion
Patient ParticipantsBehavioral Impact of Change in Listening ContextLow F0/Ambiguous VOT54.9 percent of stimuli classified as /p/Standard Error 6.3
Patient ParticipantsBehavioral Impact of Change in Listening ContextHigh F0/Ambiguous VOT30.3 percent of stimuli classified as /p/Standard Error 6
Comparison: The null hypothesis is that there is no effect of change in listening context on behavioral responses of participants.p-value: <0.0001Mixed Models Analysis
Primary

Supratemporal Neural Response to Change in Acoustic-Phonetic Dimensions

Neural activity will be measured via simultaneous EEG-sEEG monitoring in the supratemporal plane as indicated by high-gamma band activity in the electrical signal. Neural activity will be measured as participants listen to acoustic stimuli with gradually manipulated acoustic dimensions, fundamental frequency (F0) and voice onset time (VOT). The data reported here is the number of temporal lobe channels demonstrating significant encoding of change in acoustic dimension (F0 as VOT is held constant).

Time frame: During sEEG-EEG recording sessions, up to 3 hours total

Population: Neural data from 3 participants was contaminated by noise and was thus excluded from analysis.

ArmMeasureValue (COUNT_OF_UNITS)
Patient ParticipantsSupratemporal Neural Response to Change in Acoustic-Phonetic Dimensions17 sEEG Channels
Primary

Supratemporal Neural Response to Change in Listening Context

Neural activity was measured via sEEG monitoring as indicated by high-gamma band activity in the electrical signal. Neural activity will be measured as participants listen to acoustic stimuli in accented speech. The un-transformed voltage represents the difference in electric potential between a specific electrode contact and the reference electrode contact; since we use a common average reference, that means it's the difference between a specific electrode contact and the mean voltage across all electrodes. We then z-score to characterize shifts from baseline activity (where z = 0) at a specific electrode, which is believed to measure changes in voltage due largely to post-synaptic currents. This is a mathematical transformation rather than a published scale or standardized assessment. More extreme z-scores (+ or -) indicate a greater change from baseline local neural activity; in other words, stimuli evoked greater activity in this region.

Time frame: During sEEG-EEG recording sessions, up to 3 hours total

Population: Neural data from 3 participants were contaminated by noise and therefore excluded from analysis. Channels included in this analysis were only those that demonstrated significant encoding of acoustic features (F0).

ArmMeasureValue (MEAN)Dispersion
Patient ParticipantsSupratemporal Neural Response to Change in Listening Context0.55 Difference in z-scored voltageStandard Error 0.12
p-value: <0.0001Mixed Models Analysis
Secondary

Neural Response of Non-Regions of Interest to Change in Acoustic-Dimension

Neural activity will be measured via simultaneous EEG-sEEG monitoring in the supratemporal plane as indicated by high-gamma band activity in the electrical signal. Neural activity will be measured as participants listen to acoustic stimuli with gradually manipulated acoustic dimensions, fundamental frequency (F0) and voice onset time (VOT). The data reported here is the number of temporal lobe channels demonstrating significant encoding of change in acoustic dimension (F0 as VOT is held constant).

Time frame: During sEEG-EEG recording sessions, up to 3 hours total

Population: Neural data from 3 participants was contaminated by noise and was thus excluded from analysis.

ArmMeasureValue (COUNT_OF_UNITS)
Patient ParticipantsNeural Response of Non-Regions of Interest to Change in Acoustic-Dimension4 sEEG Channels
Secondary

Neural Response of Non-Regions of Interest to Change in Listening Context

Neural activity was measured via sEEG monitoring as indicated by high-gamma band activity in the electrical signal. Neural activity will be measured as participants listen to acoustic stimuli in accented speech. The un-transformed voltage represents the difference in electric potential between a specific electrode contact and the reference electrode contact; since we use a common average reference, that means it's the difference between a specific electrode contact and the mean voltage across all electrodes. We then z-score to characterize shifts from baseline activity (where z = 0) at a specific electrode, which is believed to measure changes in voltage due largely to post-synaptic currents. This is a mathematical transformation rather than a published scale or standardized assessment. More extreme z-scores (+ or -) indicate a greater change from baseline local neural activity; in other words, stimuli evoked greater activity in this region.

Time frame: During sEEG-EEG recording sessions, up to 3 hours total

Population: Channels included in this secondary analysis (n = 4 from n = 3 patients) were only those that demonstrated significant encoding of acoustic features (F0).

ArmMeasureValue (MEAN)Dispersion
Patient ParticipantsNeural Response of Non-Regions of Interest to Change in Listening Context-0.11 Difference in z-scored voltageStandard Error 0.064
p-value: 0.48Mixed Models Analysis

Source: ClinicalTrials.gov · Data processed: Feb 4, 2026