$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Ethics statement
All procedures involving human participants were conducted in compliance with institutional and national research guidelines and in accordance with the Declaration of Helsinki. This study was approved by the Institutional Review Board of Hunan Provincial Children's Hospital under approval number KYSQ2023-312 and conducted in accordance with the Declaration of Helsinki. Written informed consent was obtained from all participants' parents or legal guardians prior to enrollment. All methods were performed in compliance with institutional guidelines for human research ethics, and all data were anonymized before analysis.
Study overview and design
This protocol describes a longitudinal observational study designed to evaluate auditory and speech performance in pediatric CI recipients during the first year following implantation. The study integrates four validated behavioral assessment scales, including the Infant-Toddler Meaningful Auditory Integration Scale (IT-MAIS), Categories of Auditory Performance (CAP), Monosyllabic-Trochee-Polysyllabic Word Test (MESP), and Speech Intelligibility Rating (SIR), which have been widely used to evaluate auditory and speech development in pediatric cochlear implant recipients7,30,31,32. Each participant undergoes structured evaluation at six key time points: (1) preoperative baseline (within 1 month before implantation), (2) at CI activation ("switch-on"), (3) 1 month after activation, (4) 3 months after activation, (5) 6 months after activation, and (6) 12 months after activation. At every post-implantation follow-up, standardized auditory and speech tests are conducted by certified audiologists and speech therapists who are blinded to each participant's clinical subtype. Figure 1 provides a schematic representation of the workflow.
Participant recruitment and eligibility
Inclusion criteria
Children aged 1–5 years with bilateral severe-to-profound sensorineural hearing loss who were scheduled for unilateral cochlear implantation were recruited. All candidates met the national criteria for pediatric CI candidacy and demonstrated the following characteristics: 1) No additional developmental, cognitive, or motor disorders aside from hearing loss. 2) Bilateral hearing thresholds differing by <15 dB HL. 3) Normal brain morphology on T1-weighted MRI before surgery. 4) Verified full electrode insertion confirmed by postoperative X-ray imaging and impedance testing. 5) Parental willingness for scheduled follow-up visits for one year.
Exclusion criteria
Children were excluded if they had: 1) previous otologic surgery; 2) progressive or syndromic deafness with neurologic comorbidities; 3) severe visual or motor impairment impeding behavioral testing; or 4) loss to follow-up exceeding two consecutive visits. A total of 64 participants were enrolled, including 32 children with inner-ear malformations and 32 with normal cochlear anatomy, matched by age and sex.
Preoperative assessment
Before cochlear implantation, all participants underwent comprehensive preoperative evaluations to establish baseline auditory and neuroanatomical profiles and to confirm surgical candidacy. The assessment protocol consisted of audiological, electrophysiological, and imaging examinations, supplemented by developmental and behavioral screening.
Audiological evaluation
Pure-tone audiometry (PTA) was performed using a calibrated clinical audiometer (e.g., Interacoustics AC40 or equivalent) with insert earphones in a sound-treated booth following ISO 8253-1 standards. Air-conduction thresholds were measured at octave frequencies from 250 Hz to 8,000 Hz. Auditory brainstem responses (ABR) were recorded using surface electrodes placed at the vertex (Cz), ipsilateral mastoid, and forehead ground. Click stimuli were presented at 90 dB nHL with a repetition rate of 21.1/s, and responses were averaged over 2,000 sweeps. In younger or uncooperative children, ABR and auditory steady-state responses (ASSR) were obtained under natural sleep or mild sedation. Middle-ear function was evaluated using tympanometry to exclude conductive pathology, and distortion product otoacoustic emissions (DPOAE) were recorded to verify outer hair cell dysfunction. Speech audiometry was conducted when feasible to assess speech detection and discrimination thresholds. All measurements were calibrated to international standards (ISO 389-1) and reported in decibels hearing level (dB HL).
Electrophysiological tests
To evaluate auditory pathway integrity, electrically evoked auditory brainstem responses (eABR) were recorded preoperatively using trans-tympanic promontory stimulation. Stimulation intensity was gradually increased in 10 µA increments until a clear wave V was observed. Recording parameters included a band-pass filter of 100–3,000 Hz, analysis time of 15 ms, and averaging over 1,000 sweeps. eABR confirmed the presence of an intact cochlear nerve and provided a physiological basis for electrode mapping.
Imaging assessment
High-resolution temporal bone computed tomography (CT) and magnetic resonance imaging (MRI) were performed to evaluate cochlear anatomy, identify inner-ear malformations, and assess the presence and course of the cochlear nerve prior to implantation. CT imaging was acquired using a multidetector CT scanner with a slice thickness of 0.5–0.6 mm and a reconstruction interval of 0.3 mm to provide detailed visualization of the cochlea, vestibule, and semicircular canals. Images were reconstructed in axial and coronal planes using a high-resolution bone algorithm. MRI examinations were conducted on a 3.0-T scanner using dedicated head coils. The imaging protocol included T1-weighted and T2-weighted sequences with a slice thickness of 1 mm to visualize the internal auditory canal, cochlear nerve, and surrounding neural structures. In addition, diffusion tensor imaging (DTI) was performed with 32 diffusion gradient directions (b = 1,000 s/mm2) to evaluate the integrity of the auditory pathway. DTI tractography was reconstructed using fiber-tracking algorithms to map neural connections from the cochlear nucleus to the primary auditory cortex. These imaging examinations were performed to confirm cochlear nerve integrity and surgical eligibility rather than to serve as outcome variables. All imaging data were independently reviewed by two experienced neuroradiologists who were blinded to clinical information. Discrepancies were resolved by consensus to ensure diagnostic accuracy and consistent classification of cochlear anatomy.
Developmental and behavioral screening
Baseline cognitive and language abilities were evaluated using age-appropriate developmental questionnaires and caregiver interviews. Parental reports on environmental sound awareness, attention, and communication behaviors were recorded to complement objective test results. Children with global developmental delay or neurobehavioral disorders were excluded to minimize confounding effects.
Preoperative counseling and preparation
Before surgery, parents received detailed counseling regarding the implantation procedure, potential benefits, and rehabilitation expectations. Audiologists and speech-language therapists provided preoperative auditory training sessions to facilitate early adaptation after device activation. All data collected during preoperative assessment were recorded as baseline references for longitudinal follow-up.
Cochlear implantation
All cochlear implantations were performed under general anesthesia by experienced otologic surgeons following a standardized surgical protocol. A transmastoid facial recess approach was used to access the round window or cochleostomy site, ensuring complete electrode insertion into the scala tympani while minimizing trauma to the cochlear structures. Mastoidectomy was performed using a high-speed otologic drill with a diamond burr (2–3 mm diameter) under continuous saline irrigation. The facial recess was opened between the facial nerve and chorda tympani nerve under operating microscope magnification (10×–20×).
Surgical procedure
After standard skin preparation and draping, a retroauricular incision was made approximately 2 cm posterior to the auricle. The mastoid cortex was opened, and a facial recess was drilled between the chorda tympani and facial nerve under microscopic guidance. The round window niche was carefully exposed to visualize the cochlear entry point. Depending on anatomical variation, electrode insertion was performed either through the round window membrane or via a limited cochleostomy made anteroinferior to the round window. The electrode array was advanced slowly at approximately 0.5–1 mm/s using microforceps to minimize intracochlear trauma. Intraoperative microscopy was used to confirm correct placement, and the receiver–stimulator package was fixed within a tight periosteal pocket on the skull surface to prevent migration. The wound was closed in layers with absorbable sutures. All patients received a single dose of intravenous antibiotics preoperatively and continued prophylaxis for 48 h postoperatively.
Intraoperative monitoring
To confirm electrode functionality and auditory pathway integrity, intraoperative electrically evoked auditory brainstem responses (eABR) and neural response telemetry (NRT) were performed. Stimulation parameters included a biphasic pulse of 25 µs per phase with an 8 µs interphase gap. Responses were recorded using subdermal electrodes at the vertex and mastoid positions. Latencies and amplitudes of waves III and V were used to verify proper neural activation. Impedance telemetry was also conducted to ensure that all electrode channels were within normal operating limits (1–15 kΩ). Any abnormal responses were investigated immediately to exclude electrode misplacement or partial insertion.
Postoperative care and imaging
All patients underwent routine postoperative care with close monitoring for facial nerve function, vertigo, or wound complications. Computed tomography (CT) was performed within 24 h to confirm full electrode insertion and rule out electrode kinking or migration. Pain and inflammation were managed conservatively, and patients were discharged within 3–5 days post-surgery.
Device activation and initial mapping
Initial activation ("switch-on") was scheduled approximately four weeks after surgery to allow for wound healing and electrode stabilization. Device programming was performed using the manufacturer's clinical fitting software. Threshold (T-level) and comfortable loudness (C-level) were determined for each electrode channel using behavioral responses or objective ECAP measurements. The stimulation rate was initially set at 900 pulses per second, and the dynamic range was adjusted to approximately 40–60 clinical units. Individual electrode threshold (T-level) and comfortable loudness (C-level) levels were established using behavioral feedback and objective measures such as electrically evoked compound action potentials (ECAP). A default frequency allocation table was applied, covering the full speech spectrum (125–8,000 Hz). The dynamic range was fine-tuned over multiple sessions to ensure optimal audibility and comfort. Parents were trained to observe listening behaviors and report auditory responses at home during the early adaptation phase.
Safety and quality assurance
All surgical and programming steps were performed following international standards for cochlear implantation safety. The same surgical team and audiological staff were maintained throughout the study to ensure procedural consistency. Device performance and electrode integrity were rechecked at each follow-up session, and any mapping adjustments were documented for longitudinal analysis.
Longitudinal follow-up evaluations
All participants underwent systematic longitudinal follow-up assessments at 1, 3, 6, and 12 months after initial device activation. Each evaluation was performed in a sound-treated booth under controlled conditions (ambient noise < 30 dB SPL). The same audiologists and speech-language pathologists conducted all sessions to ensure methodological consistency and reduce inter-rater variability. The follow-up battery comprised four validated behavioral scales—IT-MAIS, MESP, CAP, and SIR—designed to quantify auditory perception, speech recognition, and speech intelligibility over time.
Infant-toddler meaningful auditory integration scale (IT-MAIS)
The IT-MAIS was administered as a structured parental interview consisting of 12 items evaluating spontaneous listening behaviors and the integration of auditory experiences in daily life. Each response was rated on a five-point scale (0 = never, 4 = always), yielding a maximum score of 48. Interviews were conducted by trained examiners who followed standardized scripts to avoid bias. Higher scores reflected better auditory awareness and sound discrimination. IT-MAIS was primarily emphasized at the 1- and 3-month sessions to capture early post-activation auditory development.
Categories of auditory performance (CAP)
The CAP scale provided a hierarchical assessment of functional hearing, ranging from 0 (no awareness of sound) to 9 (telephone conversation with a stranger). Each child's CAP level was determined based on direct observation and caregiver reports of everyday listening activities. Assessments were performed at each follow-up point. Longitudinal increases in CAP level were interpreted as indicators of auditory comprehension improvement and speech understanding in real-world environments.
Monosyllabic-Trochee-Polysyllabic word test (MESP)
Speech perception was evaluated using the Mandarin version of the Monosyllabic–Trochee–Polysyllabic Word Test. The test consists of 12 recorded words differing in syllable structure and stress pattern, presented at 65 dB SPL in an open-set format. Children first underwent a short familiarization session with visual cues, followed by auditory-only testing. Correct identifications were scored as 1 point, and total accuracy (%) was calculated. The MESP test was introduced at the 3-month visit and repeated at 6- and 12-month sessions to evaluate progressive speech perception abilities.
Speech intelligibility rating (SIR)
The SIR was used to measure speech production and intelligibility. Ratings were made on a five-point ordinal scale (1 = unintelligible, 5 = intelligible to unfamiliar listeners). Speech samples were collected during spontaneous conversation with caregivers and rated independently by two blinded evaluators. The SIR was introduced once children demonstrated consistent verbal output, typically from the 6-month assessment onward. The 12-month score was considered the key outcome measure for expressive language performance.
Testing environment and standardization
All tests were conducted in a sound-field setup using calibrated free-field loudspeakers positioned 1 meter in front of the participant. The presentation level was fixed at 65 dB SPL, verified before each session with a sound-level meter. Testing was paused immediately if the child displayed fatigue or inattention. Caregivers were allowed to accompany the child to reduce anxiety but instructed not to provide prompts. To ensure reproducibility, examiners underwent inter-rater reliability training, achieving an intraclass correlation coefficient (ICC) > 0.9 across rating sessions.
Data recording and quality control
All raw scores were recorded electronically in the institutional auditory rehabilitation database. Missing data or inconsistencies were cross-checked with clinical records. Progress trajectories were plotted for each individual, and averaged values were calculated for group-level trend analysis. The combination of IT-MAIS, CAP, MESP, and SIR assessments allowed comprehensive tracking of auditory–speech development, from initial sound awareness to functional verbal communication. Detailed scoring procedures and rating criteria for each behavioral scale are provided in Supplementary Tables (Supplementary Table 1, Supplementary Table 2, Supplementary Table 3, and Supplementary Table 4) to facilitate replication of the protocol in other clinical settings.
Statistical analysis
All statistical analyses were performed using SPSS Statistics (version 26.0; RRID:SCR_016479) and GraphPad Prism (version 9.0; RRID:SCR_002798). Data were imported into SPSS, and longitudinal comparisons were performed using the repeated-measures analysis within the General Linear Model menu. Data from all time points were compiled into a standardized database and independently verified by two analysts to ensure accuracy and consistency. A p-value of < 0.05 was considered statistically significant for all two-tailed tests. Continuous variables were expressed as mean ±± standard deviation (SD), and categorical variables as frequencies and percentages. Normality was assessed using the Shapiro-Wilk test. Outliers were screened visually through boxplots and statistically using z-scores; values exceeding ±±3 SD were excluded from parametric analyses. Missing data (< 5%) were handled using mean imputation when missing completely at random (MCAR).
The primary outcome measures were the scores of IT-MAIS, CAP, MESP, and SIR at each follow-up time point (1, 3, 6, and 12 months). Although CAP and SIR are ordinal scales, they were analyzed as continuous variables in parametric models because these instruments contain multiple ordered categories and are widely treated as quasi-continuous measures in cochlear implant outcome research. To ensure robustness, additional nonparametric analyses using Spearman's rank correlation were also performed, yielding results consistent with the parametric findings. Although IT-MAIS, CAP, MESP, and SIR are ordinal scales, they were treated as continuous variables for longitudinal analysis, consistent with prior cochlear implant outcome studies that applied parametric statistical methods to these validated clinical scales. Changes in auditory and speech performance across time were analyzed using repeated-measures analysis of variance (ANOVA) with Greenhouse–Geisser correction when the sphericity assumption (Mauchly's test) was violated. Post hoc pairwise comparisons between consecutive visits were conducted using the Bonferroni adjustment to control for multiple testing. The effect size for each model was reported as partial eta-squared, where 0.01, 0.06, and 0.14 were interpreted as small, medium, and large effects, respectively.
To explore relationships between auditory perception and speech production, Pearson's correlation coefficients (r) were calculated among IT-MAIS, CAP, MESP, and SIR scores at each time point. To confirm the robustness of the findings for ordinal measures, parallel analyses using Spearman's rank correlation were also conducted. Correlation strength was classified as weak (r < 0.3), moderate (0.3 ≤ r < 0.7), or strong (r ≥ 0.7). Additionally, Spearman's rank correlation was applied for ordinal data or non-normally distributed variables. The results were visualized using heatmaps and scatterplots with linear regression trend lines to illustrate temporal coupling between auditory and speech domains.
Stepwise multiple linear regression analysis was conducted to identify early predictors of 12-month SIR outcomes. Independent variables included IT-MAIS and CAP scores at 1 and 3 months, age at implantation, sex, and presence of inner-ear malformation. Variance inflation factor (VIF) values were calculated to check for multicollinearity (VIF < 5 considered acceptable). Model performance was assessed using adjusted R2 and standardized beta coefficients (β).
Predictive accuracy was further evaluated through leave-one-out cross-validation (LOOCV), and residual plots were examined to confirm linearity and homoscedasticity.
Participants were stratified into two subgroups: those with inner-ear malformations and those with normal cochlear anatomy. Independent-samples t tests (or Mann-Whitney U tests for nonparametric data) compared auditory and speech scores between groups at each follow-up interval. The temporal interaction between group and time was evaluated using two-way repeated-measures ANOVA.
CAEP recordings were available for 25 of the 64 participants who successfully completed electrophysiological testing during follow-up visits. Younger children who were unable to maintain stable recording conditions were not included in the electrophysiological analysis. For these participants, latency and amplitude parameters (P1, N1, P2) were correlated with IT-MAIS, CAP, and SIR scores to investigate cortical plasticity patterns. Baseline demographic characteristics between participants with and without CAEP recordings were compared to evaluate potential selection bias. The association between CAEP latency reduction and behavioral improvement was tested using linear mixed-effects models, accounting for repeated measures within subjects.