$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
The study received approval from the Institutional Review Board of Sir Run Run Shaw Hospital, Zhejiang University School of Medicine (2025 Research No. 0660) and was conducted in accordance with the Declaration of Helsinki, Good Clinical Practice guidelines, and Chinese medical research ethics regulations23. All participants provided written informed consent following a comprehensive explanation of the study procedures, potential risks and benefits, data confidentiality measures, and their right to withdraw from the study at any time without penalty. A Data Safety Monitoring Board was established to oversee study conduct and participant safety throughout the data collection period.
Study Design and Setting
This study employed a single-center prospective observational cohort design conducted over a 24-month period from January 2023 to December 2024 at Sir Run Run Shaw Hospital, Zhejiang University School of Medicine, a major tertiary care hospital in Hangzhou, Zhejiang Province, China. Reporting was aligned with the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) framework. The participating institution was selected based on its high surgical volume of more than 12,000 procedures annually, comprehensive perioperative services across multiple surgical specialties, established electronic health record systems with standardized documentation protocols, and demonstrated commitment to nursing research initiatives. During the study period, terminal handling was performed in standardized operating-room units by circulating nurses, scrub nurses, anesthesia personnel, surgeons, and environmental services personnel, following a uniform checklist. The usual staffing model maintained one circulating nurse assigned to each active operating room, with additional scrub or specialty support assigned according to case complexity. Nurse-to-patient exposure was therefore defined at the case or terminal handling episode level rather than as an uninterrupted individual nurse-to-patient ratio. Temporal consistency across the 24-month period was maintained through unchanged institutional terminal handling protocols, quarterly staff refresher training, monthly inter-rater calibration, and continuous audit of checklist adherence. These procedures were intended to improve reproducibility and reduce secular variation in workflow across the observation period.
Participants and Recruitment
The study population comprised registered nurses working in perioperative settings with direct responsibility for terminal handling procedures at Sir Run Run Shaw Hospital, Zhejiang University School of Medicine. Recruitment was conducted through departmental presentations, information sessions, and individual recruitment meetings to ensure representation across shifts and surgical specialties. A total of 232 eligible nurses were approached. Twelve declined participation before consent, 220 provided written informed consent and were enrolled, and 207 completed all required assessments and had complete outcome linkage for the final analytical sample. The 13 enrolled nurses who were not included in the analytical sample consisted of 8 voluntary withdrawals, 3 extended medical leaves, and 2 transfers to non-perioperative units.
Inclusion criteria encompassed: (1) Active employment as a registered nurse in operating room or perioperative services for minimum of six months; (2) Direct involvement in terminal handling procedures as part of regular job responsibilities; (3) Availability of complete performance and outcome data for the entire 24-month study period; (4) Willingness to participate in self-efficacy assessments and follow-up evaluations; (5) Fluency in Mandarin Chinese sufficient for completion of study instruments.
Exclusion criteria included: (1) Temporary, float, or agency nurses without consistent assignment to study units; (2) Nurses in administrative, educational, or research roles without direct clinical responsibilities; (3) Incomplete data or missing key variables exceeding 10% of required measurements; (4) Current participation in other research studies that might influence terminal handling performance; (5) Extended leave of absence (>30 days) during the study period.
Power analysis was performed using G*Power 3.1.9.7 for a two-tailed independent-samples comparison with α = 0.05, 80% power, equal allocation, and an a priori medium effect size of Cohen's d = 0.40. The selected effect size was conservative relative to the pilot observations obtained from 25 perioperative nurses at the same institution and consistent with Cohen's convention for a small-to-medium clinically meaningful difference in behavioral and performance outcomes. This calculation yielded a required total sample of approximately 200 nurses. We targeted enrollment of 220 nurses to allow for attrition and incomplete outcome linkage. The final analytical sample of 207 nurses exceeded the minimum requirement for the primary analyses and was adequate for the planned structural equation model, which contained fewer than 20 freely estimated parameters and therefore maintained more than 10 observations per parameter.
Outcome Measures and Data Collection
The Nursing Professional Self-Efficacy Scale Version 2 (NPSES-2) served as the primary instrument for measuring nurses' self-efficacy beliefs. A Chinese-language version was used after forward translation, back translation, bilingual expert review, and pilot cognitive testing to confirm semantic equivalence in perioperative nursing practice. The validated 7-item instrument was administered through a secure online platform during structured assessment sessions conducted at Sir Run Run Shaw Hospital, Zhejiang University School of Medicine. The scale measures self-efficacy across two theoretically derived domains: care delivery competence (four items) and professional effectiveness (three items). Items are rated on a 5-point Likert scale ranging from 1 (completely disagree) to 5 (completely agree), yielding total scores from 7 to 35. Higher scores indicate stronger self-efficacy beliefs. Incomplete NPSES-2 responses were queried immediately during the assessment session; if an item remained missing, the scale was considered incomplete, and the participant was excluded from analyses requiring the total score.
The instrument demonstrates excellent psychometric properties with Cronbach's α = 0.89 in our sample, and has been extensively validated across diverse nursing populations with established construct validity and test-retest reliability (r = 0.87). Prior to data collection, a pilot study involving 25 nurses at Sir Run Run Shaw Hospital confirmed the instrument's appropriateness for Chinese perioperative nursing contexts and established optimal administration procedures.
Primary outcome measures were extracted from electronic health records, direct observational assessments, standardized performance-monitoring systems, and audit forms developed for this study, according to Centers for Disease Control and Prevention definitions, Chinese hospital infection surveillance standards, and institutional terminal-handling policy. Operational definitions and scoring anchors were prespecified before data collection. The observational checklists covered procedure initiation and completion, environmental cleaning sequence, equipment decontamination, personal protective equipment use, chemical handling, documentation elements, and room readiness verification. Assessors completed standardized training, observed practice sessions, and certification before independent data collection. Checklist domains and representative scoring criteria were added to the study records to improve reproducibility across institutions.
Procedural Completion Time was measured as the elapsed time from the documented start of terminal cleaning, defined as the moment the assigned nurse or terminal handling team began the first post-case cleaning or decontamination action after patient exit, to the final recorded room-readiness confirmation for the subsequent case. Timing was recorded in minutes using synchronized digital systems and standardized time-motion procedures. Inter-observer reliability was established through dual observation of 15% of procedures. Emergency interruptions, unplanned equipment failures, isolation-room conversions, and cases requiring nonstandard environmental decontamination were flagged prospectively and excluded from the primary timing analysis; sensitivity analyses retained these cases to assess robustness.
Protocol Accuracy Score assessed adherence to established terminal cleaning and handling protocols using a comprehensive, standardized 100-point checklist. The checklist was developed by an expert panel comprising perioperative nursing leaders, infection prevention specialists, anesthesiology representatives, and quality management personnel. Items were derived from institutional policy, Chinese Ministry of Health standards, and international infection-prevention guidance, then refined using a two-round consensus process. Item-level content validity was assessed by expert rating, and items with insufficient relevance or clarity were revised before data collection. Scoring was conducted by trained research assistants using structured observation forms, with inter-rater reliability maintained through monthly calibration sessions yielding = 0.91.
Safety Compliance Rate measures the percentage compliance with safety protocols and infection-prevention guidelines using a 25-item safety checklist encompassing personal protective equipment use, chemical handling, sharps and waste management, environmental safety practices, and readiness confirmation. Observations were scheduled across day, evening, night, weekday, and weekend periods using stratified sampling. Nurses were informed that routine audits would occur during the study period, but they were not informed of specific observation times. Photographic documentation was used only for environmental or equipment-readiness verification when appropriate and did not include patient faces, patient identifiers, staff identifiers, or protected health information. Images were stored on encrypted institutional servers and reviewed only by authorized study personnel.
Documentation Completeness evaluated the percentage of required terminal handling documentation elements completed accurately and within the institutional time window. Required elements included terminal cleaning start time, terminal cleaning completion time, disinfectant or decontamination confirmation, equipment-readiness confirmation, room status, responsible nurse identifier, exception notes when applicable, and electronic sign-off. Completeness was defined as the presence of all required elements, whereas accuracy was defined as concordance between the documented element and the audit record or electronic timestamp. Documentation quality was evaluated by independent nurse reviewers who were blinded to participants' self-efficacy scores.
Equipment Readiness Time quantified the time required to prepare and verify equipment functionality for subsequent procedures, measured from completion of cleaning procedures to final equipment check confirmation. Verification was standardized through a specialty-specific readiness checklist that included power-on confirmation, alarm and safety checks, equipment positioning, availability of required disposables, sterilization or decontamination status, and documentation of functional readiness. Device-specific criteria were harmonized by perioperative nursing leaders and biomedical engineering personnel before study initiation. Secondary outcomes incorporated validated measures established by the Institute for Healthcare Improvement and adapted for the Chinese healthcare context.
Surgical Site Infection Rates were determined through a systematic review of infection surveillance databases and patient medical records for infections occurring within 30 days post-procedure, following CDC National Healthcare Safety Network definitions and Chinese hospital infection surveillance standards. Infection data were verified by infection prevention specialists who were blinded to NPSES-2 scores. Because perioperative care is team-based, patient outcomes were not attributed to a single nurse. For case-level linkage, the exposure was assigned from the electronic operating-room staffing record to the primary circulating nurse responsible for terminal handling documentation after the index case. When more than one eligible perioperative nurse was documented for the terminal handling episode, the mean NPSES-2 score of the documented nursing team was used in sensitivity analyses. Models for patient outcomes incorporated available case-context covariates, including surgical specialty, procedure category, operating room, shift, case duration, emergency status, wound category, and patient risk indicators available in the institutional record, such as age, diabetes, body mass index category, and American Society of Anesthesiologists physical status. These procedures were designed to reduce confounding by non-random staff and case assignment, although residual confounding cannot be fully excluded in an observational study.
Patient Satisfaction Scores were derived from standardized patient satisfaction survey items specific to perioperative nursing care, rated on a 1–10 scale, with established psychometric properties adapted for Chinese patients. Surveys were administered after recovery when patients were clinically stable and able to provide meaningful responses. Routine adult practice at the study hospital generally involved preoperative anxiolysis and sedation according to an anesthesiologist's judgment rather than a uniform high-dose amnestic protocol; therefore, patients with documented deep sedation before meaningful interaction with operating-room nurses, postoperative delirium, cognitive inability to respond, or incomplete recall were excluded from the satisfaction linkage. Pediatric patients were excluded from patient satisfaction analyses, and pediatric operating rooms were retained only for nurse-level performance outcomes. Regional anesthesia and spinal anesthesia performed in the operating room were recorded as contextual variables because awake patient-nurse interaction could differ from cases in which blocks were performed outside the operating room.
Nurse Stress Levels were assessed using the Perceived Stress Scale-10 (PSS-10), a validated 10-item instrument measuring subjective stress experiences over the previous month, translated and validated for Chinese healthcare workers. Stress assessments were conducted quarterly during prespecified two-week windows that included weekday and weekend shifts and avoided the first day after extended leave. Assessment timing was recorded relative to shift pattern and high-intensity surgical periods so that temporal workload could be considered in sensitivity analyses.
Team Collaboration Scores used the Collaboration and Satisfaction About Care Decisions (CSACD) instrument, which measured perceived effectiveness of interprofessional collaboration on a 1–5 scale. The instrument was adapted for Chinese perioperative team dynamics through bilingual review, expert assessment of content equivalence, and pilot administration. Team collaboration was assessed using self-report and peer-nomination procedures. Peer nominations were collected anonymously via a secure survey link, reported only in aggregate, and not shared with supervisors, reducing social desirability pressure and hierarchical bias.
Comprehensive demographic, professional, assignment, and contextual data were collected through standardized questionnaires, personnel record reviews, and operating room information systems. Variables included age, gender, educational background, years of nursing experience, years of operating room experience, specialty certifications, work schedule patterns, employment status, surgical specialty assignment, operating-room location, procedure category, shift, case duration, wound category, emergency status, workload intensity, staffing shortages, and participation in continuing education, mentorship, or organizational support programs. These variables were used to explore confounding and moderation related to non-random staff assignment and case complexity.
Data Collection Procedures
Data collection followed a systematic protocol implemented consistently throughout the 24-month observation period at Sir Run Run Shaw Hospital, Zhejiang University School of Medicine. Research staff received standardized training in data collection procedures, with certification achieved through observed practice sessions and inter-rater reliability testing. Electronic data capture systems with built-in range checks and consistency algorithms were employed to minimize data entry errors and ensure data quality.
Participants completed self-efficacy assessments during scheduled assessment sessions held in private conference rooms at the hospital, with research staff available to address questions while maintaining assessment integrity. Performance outcome data were collected continuously throughout the study period using a stratified sampling approach to ensure representation across different times of day, days of the week, and surgical case types.
To address potential observer effects, data collectors remained blinded to participant self-efficacy scores during outcome assessments. A random subset of 20% of observations underwent dual assessment to monitor inter-observer reliability throughout the study period, with drift analysis conducted monthly to maintain measurement consistency. Staffing assignment logs were extracted independently from the operating-room information system and linked after outcome assessment, preserving blinding during direct observation.
Statistical Analysis Plan
Statistical analyses were conducted using SPSS version 29.0 and AMOS version 29.0 for structural equation modeling. All analyses followed a prespecified statistical analysis plan developed before data collection and registered with the Sir Run Run Shaw Hospital research office. The revised analytical description clarifies that NPSES-2 was treated as a continuous exposure for primary inference. Median-based high- and low-self-efficacy groups were retained only to aid clinical interpretation. Reproducibility details were expanded to identify model type, covariate blocks, bootstrap replications, and multiple-comparison denominators.
Data preparation included comprehensive screening for outliers using the interquartile range method and the Mahalanobis distance, assessment of missing-data patterns using Little's MCAR test, and evaluation of distributional assumptions using Shapiro-Wilk tests, Q-Q plots, and histogram examination. Missing data were less than 5% for all variables and were handled using multiple imputation with 20 imputed datasets. The imputation model included NPSES-2 total and subscale scores, age, gender, education, total nursing experience, operating-room experience, specialty certification, shift pattern, surgical specialty, operating room, procedure category, case duration, wound category, emergency status, primary and secondary outcomes, team collaboration score, and stress score.
Between-group comparisons used the sample median NPSES-2 score of 25.1 solely for descriptive contrast, not as the primary inferential approach. Independent-samples t-tests were used for normally distributed continuous variables, Mann-Whitney U tests for non-parametric data, and chi-square tests for categorical variables, with effect sizes calculated using Cohen's conventions. Because dichotomizing a continuous construct can reduce information and create arbitrary boundaries, all regression and mediation analyses modeled NPSES-2 as a continuous variable.
Correlation analyses employed Pearson product-moment correlations for normally distributed variables and Spearman rank correlations for non-parametric data, with 95% confidence intervals calculated using 10,000 bootstrap replications. Multiple regression analyses examined relationships between the continuous NPSES-2 score and outcome measures while controlling for nurse-level covariates and assignment-related case-context covariates. Nurse-level covariates included age, education, total nursing experience, operating-room experience, specialty certification, and shift pattern. Assignment and case-context covariates included surgical specialty, procedure category, operating room, case duration, wound category, emergency status, workload intensity, and staffing shortage indicators when available. Model assumptions were verified through residual plots, leverage statistics, and multicollinearity diagnostics using variance inflation factors.
Structural equation modeling tested the prespecified mediation hypothesis that team collaboration and stress levels partially mediate the association between continuous self-efficacy and terminal handling performance. The theoretical rationale was that higher perceived capability may promote more effective team communication and reduce perceived stress during complex terminal handling tasks, both of which can support more reliable procedural performance. The model included direct paths from NPSES-2 to terminal handling performance, paths from NPSES-2 to team collaboration and stress level, and paths from each mediator to terminal handling performance. Age, education, operating-room experience, specialty certification, shift pattern, surgical specialty distribution, and case-complexity indicators were included as covariates where applicable. Model fit was evaluated using chi-square statistics, comparative fit index (CFI ≥ 0.95), Tucker-Lewis index (TLI ≥ 0.95), root mean square error of approximation (RMSEA ≤ 0.06), and standardized root mean square residual (SRMR ≤ 0.08).
Statistical significance was established at p < 0.05 for primary model parameters. Bonferroni correction was applied within prespecified outcome families. For the five primary terminal handling outcomes, the adjusted threshold was 0.05/5 = 0.010. For the secondary patient and nurse-related outcomes, the adjusted threshold was 0.05/8 = 0.00625. For the experience-stratified descriptive comparisons, the adjusted threshold was 0.05/4 = 0.0125. Effect sizes were interpreted using Cohen's conventions, and clinical interpretation was made cautiously because the study was observational and subject to residual confounding from staff and case assignment.