$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
This study involved human participants and was reviewed and approved by the Ethics Committee on Human Research Protection, Xianda College of Economics and Humanities, Shanghai International Studies University (Approval No. 2025XD1221). All participants signed informed consent forms and completed a 30 min pre-experiment training: EG participants learned to interpret the framework's visual feedback (radar charts, heatmaps), while CG participants received only platform operation training. To assure precision in the data collection, the pre-experiment also involved the 5 min calibration of the eye trackers and physiological sensors.
Study objectives and protocol overview
The objective of this study was to create a closed-loop, end-to-end system that integrates multimodal acquisition, feature extraction, Bayesian fusion, and visual feedback to evaluate engineering skills in an interpretable manner. Figure 1 displays the protocol diagram for the evaluation of engineering skills in health applications based on artificial intelligence-assisted 3D modeling.
Framework overview and hierarchical logic
The visual evaluation framework for engineering skills in AI-assisted 3D modeling health applications proposed in this study is constructed based on the closed-loop logic of "data-driven - feature extraction - evidence fusion - visual feedback". Its core objective is to achieve end-to-end conversion from multi-modal raw data to interpretable skill evaluation results22. To ensure the reliable extraction of multimodal features for evaluating engineering skills, the framework simultaneously captures interaction, physiological, gaze, gesture, and voice signals with millisecond-level temporal synchronization. The framework adopts a hierarchical architecture design, consisting of four core layers: Multi-modal Data Acquisition Layer (L1), AI Feature Extraction Layer (L2), Evidence Fusion and Scoring Engine (L3), and Visual Feedback Layer (L4). Each layer realizes bidirectional interaction through standardized data interfaces, and a framework calibration module (L5) is introduced simultaneously to ensure the dynamic validity of evaluation results. Figure 2 shows the end-to-end pipeline, starting with multimodal data (gaze, kinematics, physiology, voice) acquisition through AI-based feature extraction, Bayesian scoring, and real-time visual and audiovisual feedback. The layers are interactive, with Layer 5 providing continuous calibration. This architecture allows for coordinated processing and validated skill estimation for participants while engaging in live 3-D modelling tasks.
Equation (1) Bayesian Confidence Update Formula:

Where: P(θ) represents the prior probability of skill level θ (based on historical evaluation data); P(D∥θ) denotes the likelihood probability of observing multi-modal feature data D under skill level θ; P(D) is the marginal probability; P(D∥θ) is the posterior probability of skill level θ after fusing data D, i.e., the final confidence score. This framework used a structured Delphi and AHP approach23 to weight indicators. A panel of ten subject matter experts (orthopedic surgeons, biomedical engineers, and experienced medical CAD modelers) participated in three rounds of Delphi rating the importance of an indicator on a 9-point Likert scale until expert scoring reached stability (coefficient of variation <0.15)24. Then, using the converged/distributed scores from the experts, a pair-wise comparison matrix was built and confirmed using the AHP consistency criterion with a consistency ratio (CR) < 0.1. If the CR exceeded 0.1, the matrix was readjusted. The AHP methods' consistency ratios for the pair-wise comparison matrices were between 0.03 and 0.08, indicating great internal logical consistency throughout the multi-level hierarchical approach. The design of each layer in the framework adheres to the "modularization - extensibility" principle. For instance, Layer L1 can expand the physiological data collection dimension by adding biosensors; Layer L2 can adapt to skill evaluation requirements of different health applications by replacing pre-trained models; Layer L4 supports custom visualization dimensions based on user roles. This provides flexibility for subsequent applications in medical 3D modeling education and professional skill certification scenarios.
Multi-modal data acquisition layer
3D interaction event stream
This module collects engineers' operational data from AI-assisted 3D modeling platforms via software plug-ins (16), covering both discrete operation events and continuous parameter data. Discrete events include model selection, vertex adjustment, and AI tool invocation, while continuous parameters involve modeling speed, undo/redo counts, and AI suggestion acceptance rates. Each event is recorded with a 13-digit timestamp and corresponding 3D model coordinate system values (x, y, z). For example, when adjusting the curvature of an orthopedic implant model, the plug-in logs the operation's start and end times, adjusted vertex coordinates, and curvature variations, thereby forming a structured event log (see Table 1 to Table 3 for data fields). These data are subsequently processed using bespoke algorithms to calculate domain-specific geometric, efficiency, and cognitive indicators, enabling fine-grained characterization of modelling behaviours in health-related applications.
Eye movement and gesture signals
Eye movement data is acquired using a desktop eye tracker, focusing on fixation points and saccade parameters in the 3D modeling interface, which reflect the engineer's cognitive attention distribution.
Gesture signals are captured by a depth camera, which extracts 22 upper-limb key joint coordinates to identify typical modeling gestures25. The camera converts 3D gesture coordinate data into normalized vector sequences to mitigate individual body size differences.
Data synchronization for this module relies on the framework's unified hardware trigger (10 Hz timestamp from 3D modeling software), with synchronization error controlled within ±50 ms (calculated via Equation 3-2) to ensure temporal correlation between eye/gesture behaviors and modeling operations 26. Equation (2) presents the Data Synchronization Error Calculation:
(2)
Where: Et = synchronization error between device acquisition time (tdevice) and software timestamp (tsoftware); tdevice = actual acquisition time of eye tracker/depth camera; tsoftware = trigger timestamp from 3D modeling software.
Physiological and voice data
Physiological data is collected via wearable sensors: an Electromyography (EMG) sensor (1000 Hz sampling frequency) 27, attached to the engineer's forearm to record muscle tension during fine operations, reflecting operational stability. Electrocardiography (ECG) sensor (250 Hz sampling frequency)28: It is worn on the chest to collect heart rate variability (HRV) indicators, which characterize cognitive load. Voice data is captured by a directional microphone (44.1 kHz sampling frequency),29, recording engineers' verbal comments during modeling to supplement cognitive state analysis. All physiological and voice data undergo preprocessing before storage in the time-stamped multi-modal database. Figure 3 illustrates the physical experimental hardware configuration. Sensor data from the EMG/ECG wearables, the microphone, the eye tracker, and a depth camera are all connected to a 3-D modelling workstation. Before temporal data is saved in a multimodal database, a synchronization module adjusts all streams (within a margin of ±50 ms). This material specification means reliable, temporally aligned input into the feature extraction and scoring pipeline.
AI Feature extraction layer
Geometric quality indicators
This module extracts quantitative features reflecting the accuracy and rationality of 3D models from the 3D interaction event stream and the final model output, serving as the core basis for evaluating engineers' "modeling precision" skills. To assess structural accuracy, anatomical congruence, and 3D model completeness using deviation, topology, and template-based features metrics, three key indicators are designed, with calculation methods and physical meanings detailed below:
Model deviation rate (MDR)
This measures the geometric deviation between the engineer's modeled result and the standard template. The deviation is calculated by comparing the Euclidean distance of key feature points between the two models. The formula is shown in Equation (3):
(3)
Where: n = number of pre-defined key feature points (set to 50-100 based on model complexity); Pi,engineer = 3D coordinates of the i-th feature point in the engineer's model; Pi,standard = 3D coordinates of the i-th feature point in the standard template; |⋅| = Euclidean distance operator.
A lower MDR indicates higher consistency with the standard model, with a threshold of ≤ 5% for qualified health application models. Modeling competencies were assessed based on time efficiency, redundant movements, gesture stability, and usage patterns of the proposed AI tools during real-time 3D operational scenarios.
Topological consistency (TC)
This evaluates whether the 3D model's topological structure meets medical application requirements. The indicator is calculated by counting the number of topological defects (D) and the total number of topological elements (T, sum of vertices, edges, and faces) in the model, as shown in Equation (4):
(4)
For example, if a knee joint model contains 2 non-manifold edges (D = 2) and 1200 total topological elements (T = 1200), it is 99.83%. A ≥ 98% is required for medical models to avoid manufacturing risks.
Feature completeness (FC)
This assesses whether the engineer has fully implemented all mandatory features of the health application model. The indicator is determined by comparing the number of completed mandatory features (Fcompleted) with the total number of mandatory features (Ftotal) defined in the task requirements:
(5)
Behavioral efficiency indicators
This module processes 3D interaction event streams and gesture signals to extract features reflecting the efficiency of engineers' modeling operations, evaluating their "operation proficiency" skills. Cognitive demands associated with modeling were measured using synchronized physiological, gaze, and EMG signals. These metrics track attention distribution across time and workload fluctuations during the modeling process.
Modeling time efficiency (MTE)
This compares the engineer's actual modeling time (Tactual) with the average time of senior experts for the same task (Texpert), normalized to a 0-100 score:
(6)
For example, if an engineer takes 45 min to complete a task that experts average 30 min on, there is 100 - (45 - 30) x 50 = 75, indicating moderate efficiency.
Redundant operation rate (ROR)
This counts the proportion of unnecessary operations in the modeling process. It is calculated as the ratio of redundant operation count (Oredundant) to total operation count (Ototal):
(7)
Redundant operations are identified by the AI algorithm via rule matching and gesture consistency analysis.
AI tool utilization rate (ATUR)
This measures the engineer's ability to leverage AI-assisted functions to improve efficiency. It is defined as the ratio of AI tool invocation count (IAI) to total tool invocation count (Itotal):
(8)
A higher ATUR indicates better integration of AI tools into the modeling process, which is a key skill for modern medical 3D modeling.
Gesture smoothness (GS)
This evaluates the stability of the engineer's upper limb movements during modeling. The indicator is calculated by the standard deviation of the speed of 22 upper-limb key joints (vj) over time:
(9)
A lower standard deviation (and thus higher GS) indicates smoother gestures, reflecting higher operational proficiency. These indicators are extracted in real time during the modeling process, with the AI algorithm updating the indicator values every 5 min and storing them in the feature database for subsequent fusion.
Cognitive load indicators
This module analyzes eye movement signals and physiological data to quantify the engineer's cognitive load during modeling, evaluating their "cognitive adaptability" skills. Multimodal indicators were weighted using a Delphi-AHP approach and Bayesian evidence updating to output real-time, probabilistic engineering skill scores. Three key indicators are developed:
Fixation dispersion (FD)
This reflects the distribution of the engineer's visual attention. It is calculated as the Euclidean distance between the maximum and minimum fixation point coordinates ( (xmax,ymax) and (xmin,ymin) ) in the 3D modeling interface over a 2 min window:
(10)
A larger FD indicates scattered attention, often associated with high cognitive load.
Heart rate variability (HRV) - SDNN
SDNN (Standard Deviation of Normal-to-Normal Intervals) is extracted from ECG data to measure the fluctuation of the engineer's heart rate. A lower SDNN indicates higher sympathetic nerve activity, reflecting increased cognitive load. The AI algorithm first filters ECG noise (via wavelet transform) and then calculates SDNN using Equation (11):
(11)
Where: N = number of normal heartbeats in a 5-minute window; RRk = time interval between the k-th and (k + 1)-th normal heartbeat;
= average RR interval in the window.
Electromyography (EMG) - Root mean square (RMS)
RMS of forearm EMG signals reflects muscle tension: higher RMS indicates increased muscle stiffness due to cognitive stress. The calculation formula is:
(12)
Where: m = number of EMG signal samples in a 1 min window, and EMG(t) = amplitude of the EMG signal at time t. A summary diagram has been created to assist with conceptual clarity and facilitate replication of the research. It illustrates the hierarchical relationship among the ten evaluation indicators and the three main constructs behind the proposed framework: geometric quality, operational efficiency, and cognitive adaptability. This diagram depicts how multi-modal input signals are converted into measurable indicators that are subsequently fused through a Bayesian approach to create a composite engineering skill score. This overview conveys a compacted version of the evaluation logic and increases the methodological transparency of this study formatively and for future educational and regulatory direction. Figure 4 depicts the suggested engineering skill evaluation framework's hierarchical structure, which organizes ten multimodal indicators into three fundamental constructs: geometric quality, operational efficiency, and cognitive flexibility. These factors contribute to the overall skill outcome used to evaluate AI-assisted 3D modeling performance.
Evidence fusion and scoring engine
Indicator weight assignment (Delphi + AHP)
This module addresses the problem of differentiated importance among multi-dimensional features by combining the Delphi method (expert consensus) and the Analytic Hierarchy Process (AHP) to assign weights to the ten extracted indicators (three geometric quality indicators, four behavioral efficiency indicators, and three cognitive load indicators), and to represent these complex multimodal indicators in visual and interpretable formats, such as radar charts, heatmaps, and annotated playbacks for biomedical-relevant diagnostic inference.
Delphi consensus stage
A panel of ten experts assessed the importance of each indicator using a 9-point scale (1 = "least important", 9 = "most important") across three rounds of anonymous questionnaires. Following each round, statistical feedback-comprising mean scores and coefficients of variation (CV)-was provided to the panel to facilitate score adjustment. Convergence was defined as a CV ≤ 0.15 for each indicator. Notably, consensus for indicators such as the Model Deviation Rate (MDR) and AI Tool Utilization Rate (ATUR) was typically achieved by the second round, attributed to their direct clinical relevance.
To ensure technical robustness, the protocol standardized multi-modal data alignment (synchronization tolerance ≤ 50 ms; 30 Hz unified trigger clock) and indicator calculation methods. For instance, MDR was derived from 50-100 anatomical key points, while the Topological Consistency (TC) threshold was set at >98% for defect-free elements. These thresholds were justified against clinical tolerance standards, such as a < 5% deviation for porous bone and a < 3° error for vascular bifurcation angles. Furthermore, the methodology incorporated Bayesian scoring with Gaussian likelihood modeling (using calibrated σ ranges for distinct skill tiers) and utilized a CNN-LSTM feature extractor trained with a batch size of 32, a learning rate of 1 x 10-4, and 120 epochs.
AHP Hierarchy establishment
A three-level hierarchy is constructed: Target Layer (A): Engineering skill comprehensive evaluation; Criterion Layer (B): 3 first-level indicators (B1: Geometric Quality, B2: Behavioral Efficiency, B3: Cognitive Load); Indicator Layer (C): 10 second-level indicators (C1: MDR, C2: TC, ..., C10: EMG RMS).
Pairwise comparison and consistency check
Based on Delphi consensus results, a pairwise comparison matrix is constructed for the Criterion Layer and Indicator Layer. Taking the Criterion Layer as an example, the matrix is defined as:
(13)
Where: aij is the importance ratio of Bj to Bj. The weight vector of the Criterion Layer (WB = [wB1,wB2,wB3]) ) is calculated via the eigenvector method. To avoid logical contradictions, the Consistency Ratio (CR) is checked using Equation (14):
(14)
Where: ( λmax = maximum eigenvalue of the matrix, = number of indicators), and is the random consistency index (for n = 3, RI = 0.58). The result is acceptable if CR < 0.1.
Through this process, typical weights are obtained: (Geometric Quality), (Behavioral Efficiency), (Cognitive Load); the weight of MDR (C1) in the Indicator Layer is usually the highest (~0.25), while EMG RMS (C10) is the lowest (~0.05).
Bayesian confidence update
This module fuses the weighted indicators to generate the final skill confidence score, addressing the "uncertainty of single-feature evaluation" by updating prior probabilities with real-time multi-modal data. The process is based on Equation (1) (Bayesian formula) and is implemented in two steps:
Step 1: Prior probability initialization (P(θ))
The skill level is divided into 5 grades: (Novice), (Beginner), (Intermediate), (Advanced), (Expert). The prior probability is initialized using historical data: if the system has evaluated 500 engineers, and 100 are classified as θ3, then P(θ3) = 100/500 = 0.2. For new users without historical data, a uniform prior (P(θi) = 0.2 for all i) is adopted.
Step 2: Posterior probability calculation (P(θ∥D))
The multi-modal feature data is the weighted sum of the 10 normalized indicators :
(15)
Where: wCk is the weight of the k-th indicator (from Section 3.4.1), and Cknorm is the normalized value of the k-th indicator ([0,1] range).
The likelihood probability is modeled using a Gaussian distribution, assuming that the feature data of engineers at skill level follows (mean μi, variance
). These parameters are calibrated using 200 labeled samples.
Finally, the posterior probability is calculated via Equation (3-1), and the skill grade corresponding to the maximum is the real-time evaluation result. For example, if an engineer's D = 0.72, and (higher than other grades), they are classified as "Advanced".
To visualize the confidence update process, Figure 5 shows how known equal skill probabilities are updated into a posterior distribution using Bayesian inference and real-time signals as evidence. After observing the task, the model updates to strengthen the evidence towards the skill grade with the largest posterior probability. Validation tests indicated strong convergence stability of the expert skill grades, and the Bayesian scores strongly correlated with expert ratings, supporting the trusted use of Bayesian updating in the workflow.
The output of this layer is a comprehensive skill score (0-100, mapped from the maximum posterior probability) and a confidence interval, which is transmitted to the Visual Feedback Layer for intuitive presentation.
Visual feedback layer
Real-time radar chart
This module converts the 10 normalized indicators and 3 three-level criterion scores into a radar chart, enabling intuitive comparison of engineers' skill strengths and weaknesses. The radar chart is designed with 7 axes (3 for first-level criteria: B1-B3; 4 for core second-level indicators: C1=MDR, C4=ATUR, C7=FD, C9=SDNN-selected for their high weight and interpretability), with each axis range normalized to [0,100] (mapped from the [0,1] normalized values via Equation 3-7). To continually improve measurement accuracy through expert annotations and active-learning updates at the inception of the framework, a health-modelling context and growing understanding are incorporated.
Equation (16) Indicator Value Mapping for Radar Chart:
(16)
Where: Vradar = value displayed on the radar chart axis; Cknorm = normalized value of the k-th indicator ([0,1] range).
The radar chart updates in real time (every 5 min) and includes two reference benchmarks: a standard benchmark line (dashed line, Vradar = 80, representing the "Advanced" skill level threshold) and an expert average line (dotted line, calculated from 50 senior engineers' historical data). Figure 6 shows the engineer's behaviour across a composite of key indicators: geometric quality, cognitive load, tool-use efficiency, and fixation dispersion, compared to the corresponding expert benchmarks. As part of the feedback layer, true scores, using fused metrics, were translated into an interpretable profile. User studies indicated its successful implementation for quickly finding competence gaps, while promoting a structured, data-driven calculation for improvement during modelling tasks.
Temporal heatmap
This module visualizes the dynamic change trend of skill indicators over the entire modeling process, helping identify time periods where engineers face difficulties. The heatmap uses a 2D grid: the x-axis represents time intervals (10-minute segments, totaling 12 segments for a 2-hour task), and the y-axis represents 5 key indicators (C1=MDR, C4=ATUR, C6=GS, C7=FD, C9=SDNN). The color intensity of each grid cell corresponds to the indicator's normalized value, with a color scale from blue (low value, poor performance: Cknorm < 0.5) to red (high value, good performance: Cknorm < 0.8), as defined in Equation (17):
Equation (3-8) presents Color Intensity Calculation for Heatmap:
(17)
Where: Icolor = RGB color intensity value (0-255, 0=blue, 255=red); min(Cknorm) and max(Cknorm) = minimum and maximum normalized values of the indicator across all time intervals.
Figure 7 illustrates time-resolved changes in cognitive load, fixation dispersion, gaze stability, and decision-latency metrics in 10 min segments. Note that the highlighted regions indicate high cognitive demand or workflow disruptions. This time-view is generated directly from coalescent multimodal streams, supported by annotation by experts, illuminating its critical role in engaging with modelling fatigue, attentional shifts, and task-specific difficulties.
D playback annotation
This module overlays skill evaluation annotations on the 3D model's operation playback, enabling step-by-step analysis of the engineer's modeling behavior. The playback speed is adjustable (0.5×-2× real time) and synchronizes with the multi-modal data timestamp. Three types of annotations are embedded: Geometric Error Annotations: Red wireframes highlight model regions where MDR (C1) exceeds 5% (qualified threshold), with text labels showing the deviation value. Efficiency Warning Annotations: Yellow exclamation marks appear at the location of redundant operations, with pop-up windows showing the operation type. Cognitive Load Annotations: Green/red status bars in the corner indicate real-time cognitive load (based on C9: SDNN and C10: EMG RMS), with red meaning "High Load" (SDNN < 50 ms) and green meaning "Normal Load" (SDNN > 100 ms).
Framework calibration and iteration
Expert annotation interface
This module provides a human-in-the-loop interface for senior experts to correct and label the framework's evaluation results, addressing the "drift of evaluation accuracy" caused by changes in health application scenarios. The interface is designed with three core functions: result correction, indicator reweighting, and annotation storage.
Result correction function
Experts first review the visual feedback results and the framework's comprehensive skill score (0-100). If the score deviates from expert judgment, the expert can manually adjust the score and record the reason. The correction magnitude is quantified via Equation (18):
(18)
Where: Scorrected = final corrected score; Sframework = initial score output by the framework; Sexpert = expert's manual score; α = correction weight (set to 0.8 by default, ensuring expert judgment dominates while retaining framework rationality).
Indicator reweighting function
For scenario-specific calibration, experts can adjust the weight of individual indicators via a sliding bar (0-0.5 range). The interface automatically recalculates the consistency ratio (CR) after adjustment (using Equation 3-6) and alerts experts if (illogical weight distribution), ensuring the adjusted weights remain mathematically valid.
Annotation storage function
All corrected scores, adjusted weights, and expert comments are stored in an annotated database with time stamps and scenario tags. This database serves as the training data for the subsequent online active learning module, with a typical data schema shown in Table 2.
Online active learning update
This module uses the annotated data to iteratively optimize the framework's core algorithms (AI feature extraction models in L2 and Bayesian scoring engine in L3), enabling the framework to adapt to new health application scenarios without full retraining. The update process follows a query-by-committee (QBC) strategy, divided into three stages:
Stage 1: Uncertainty calculation
Three parallel "committee models" are deployed (all based on the same architecture as the L2 feature extraction model but trained on different historical datasets):
Model A: Trained on orthopedic implant modeling data;
Model B: Trained on organ reconstruction data;
Model C: Trained on dental implant modeling data.
For each unlabeled sample (engineer's modeling data), the committee models output three predicted skill scores. The uncertainty of the sample is calculated via the coefficient of variation (CV) of the three scores:
(19)
Where: SA,SB,SC = scores predicted by the three committee models; std (⋅)= standard deviation;
= average of the three scores.
A higher CV indicates greater disagreement among the committee models (higher uncertainty), meaning the sample is more valuable for annotation.
Stage 2: Sample selection
The framework selects the top 10% of samples with the highest CV and sends them to the expert annotation interface for labeling. This active selection strategy reduces the number of samples requiring expert annotation (compared to random selection) by 40-60%, improving calibration efficiency.
Stage 3: Model retraining and update
The newly annotated samples are used to fine-tune the framework's algorithms:
For L2 feature extraction models: The model is retrained with a learning rate of 1e-5 (10% of the initial training rate) using the annotated geometric error data, updating only the last two layers to avoid overfitting.
For L3 Bayesian scoring engine: The Gaussian distribution parameters (
) for each skill level are updated using the annotated score data, with the update formula shown in Equation (20):
(20)
Where:
= updated mean of skill level θi;
= original mean before update; S-θi = average corrected score of annotated samples labeled as θi; β = forgetting factor (set to 0.7, balancing historical data and new annotations).
The updated framework is deployed online immediately after retraining, with a performance check before deployment: if the average absolute error between the framework's new output and expert annotations is ≤3 (reduced by ≥20% from the previous version), the update is confirmed; otherwise, the retraining process is repeated with additional annotated samples.
This iterative calibration mechanism ensures the framework's evaluation accuracy remains above 90% even as health application 3D modeling technologies evolve, maintaining long-term adaptability.