$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
This study was conducted as a case series using representative results from four children to demonstrate protocol feasibility rather than statistical inference. Two children typically developing (male 4 years old, 100 cm, 14.5 kg; male 10 years old, 144 cm, 40.8 kg) and two children with cerebral palsy, Gross Motor Function Classification System III (female 4 years old, 93 cm, 13.8 kg; male 4 years old, 102 cm, 16.7 kg) participated in this study. Parents/legal guardians and children provided informed consent/assent prior to participation in accordance with our Institutional Review Board-approved study protocol (IRB-FY2024-33).
NOTE: While this protocol describes procedures for both 3D marker-based and 2D markerless motion capture systems, the focus of this study is on the implementation of the 2D markerless approach, which represents the core clinical and practical application of this work. The core implementation of the current protocol requires a marker-based motion capture system (minimum 8 cameras) for 3D kinematic data acquisition, and a single RGB camera (minimum resolution 1920 × 1080 pixels, 50 Hz) positioned perpendicular to the sagittal plane for 2D markerless capture. The method is suitable for controlled clinical or laboratory environments with adequate lighting (> 500 lux) and minimal background clutter. To improve robustness in non-ideal environments, the vailá toolbox integrates pre-processing tools (e.g., drawbox masking, video crop/resize) that allow users to isolate the participant from background distractions or adjust suboptimal camera framing. Key limitations include reduced accuracy for the single-camera, markerless approach for children exhibiting significant out-of-plane movements, as this approach cannot capture depth information or movements outside the camera's field of view. Users should ensure participants remain within the predefined capture volume and maintain sagittal plane alignment during task execution to optimize markerless tracking performance.
1. Pre-test checks
- Power on and warm up the infrared motion capture camera system and the video camera. The infrared motion capture system requires approximately 30-60 min of warm-up to reach thermal stabilization before calibration and data collection. Verify the sampling rate at 100Hz (infrared cameras for marker-based system) and 50Hz (video camera for markerless system).
- Clear and/or cover reflective objects from the capture volume.
2. Camera placement
- Marker-based system (3D camera system)
- Position cameras approximately 1.5-2.5 m from the center of the movement area, mounted at heights ranging from 0.5-2.5 m. Ensure that the cameras are distributed evenly around the capture volume to optimize marker visibility and minimize occlusion during the sit-to-stand task. Alternate elevations of camera position to improve triangulation accuracy and reduce line-of-sight interference.
- Each point within the target volume (where the child will perform the task) must be visible to at least two cameras at all times, following standard recommendations for optimal 3D data capture. Define the capture volume based on the anthropometrics of pediatric participants and the task being evaluated.
NOTE: The working space was centered at approximately 0.80 m above the floor.
- Use the marker-based system's software tools to refine the three main optical settings (zoom, aperture, focus) to ensure low calibration residuals and high marker tracking consistency across trials.
- Markerless system (2D camera)
- Place the video camera laterally at 3.0 m from the participant (or closer, when possible, as long as the whole body of the subject can be seen during the entire task, from seating to standing), centered at the pelvis height when the participant is standing, orthogonal to the sagittal plane.
- Level the optical axis (no roll) and align the image midline with the hip joint center region.
- Configure the camera with fixed imaging parameters for all recordings.
- Set the acquisition region to an area of interest of 1920 × 1200 px to define the effective resolution. Use a 20 mm focal lens and keep the aperture (enough contrast between the environment and participant) constant throughout the experiment.
- Set the frame rate to 50 Hz and fix the shutter duration at 5 ms. Apply the following hardware settings and do not modify them during data collection: gain 9.998, gamma 0.4, black level 0, saturation 1.
- Disable all automatic exposure and white-balance adjustments to prevent temporal drift in image appearance.
NOTE: To adjust the focus, a sign can be used with a short phrase written in font Calibri, style bold, size 48, placed where the participant will perform the task. Within the camera's field of view, the letters should be sharp, with no visible blur.
- When using any lens with noticeable distortion, prepare a Charuco/Checkerboard18,19for calibration/undistortion (optional for strictly sagittal, stationary tasks).
NOTE: Both marker-based and markerless video camera systems used are shown in Figure 1.

Figure 1: Close-up of the two camera systems used in this study. The 3D kinematic system's infrared camera (left) was used in which 12 cameras made up the system for the marker-based approach, and one video camera (right) was used for the markerless approach. Please click here to view a larger version of this figure.
3. Synchronization
NOTE: This step is only needed when using different systems concurrently. Synchronization between systems is not necessary when using only a single kinematic (2D or 3D) system.
- Connect the LockLab (or equivalent control/synchronization unit) to send a transistor-transistor logic (TTL) trigger to both systems at the start of each trial.
- If hardware triggers are unavailable, use a visible light-emitting diode (LED) flash (preferable method) or an audible clap captured by both systems for later alignment.
4. Capture volume calibration (marker-based system)
- Visually inspect cameras' field of view (both infrared and video cameras) to remove unnecessary objects from each view. In the marker-based system's software, mask reflective objects that could not be removed from the camera's view (e.g., other cameras).
- Perform wand calibration where the task will be performed. Move the wand throughout the full capture volume to ensure the entire space is covered. Accept calibration only if spatial accuracy ≤ 1 mm (mean residual).
NOTE: Verify that at least three cameras simultaneously view the calibration wand during this procedure. Move the wand through the capture volume at a speed that approximates the motion expected in the experimental task: For high-velocity tasks, move the wand rapidly; for slower movements, sweep the wand more slowly. At the end of the calibration pass, check the software indicators for any cameras that have not captured sufficient views of the wand and perform additional targeted wand movements within those cameras' fields of view before concluding the calibration. Residual errors after wand calibration typically fall between 0.3 and 0.8 mm when the wand pass is smooth, the capture volume is uniformly illuminated, and all cameras have unobstructed views.
- Set the volume origin by placing the calibration wand at the designated origin location within the defined testing space. Define the floor plane according to the manufacturer's instructions.
5. Marker placement (marker-based system)
- Place retroreflective markers according to the Plug-in Gait20,21,22 model (Figure 2) by a trained examiner.
- Verify visibility of bilateral hip, knee, and trunk markers during a brief test movement prior to collecting static trial (required for 3D marker-based systems).

Figure 2: Plug-in Gait model adapted for the present study. The standard marker set was applied to define lower limb, pelvis, trunk, and head segments. Upper limb markers were removed to focus exclusively on lower-body and trunk kinematics. This adaptation preserves the integrity of the original Plug-in Gait full-body model while excluding arm motion from the analysis. Please click here to view a larger version of this figure.
6. Chair and environment
- Use an armless chair and adjust the seat height so that the hip and knee angles are ~90°, with the feet flat on the floor. Record the seat height (cm) for each participant.
- Provide uniform lighting; avoid backlight and occlusions in the sagittal view. Ensure a uniform, high-contrast background.
7. Participant instruction and safety
- Confirm no identifiable information is included with recording files (only non-identifiable alphanumeric codes for each participant).
- For sit-to-stand trials, ensure participants are aware of proper instructions and protocol.
- Instruct typically developing (TD) participants to keep their hands on their hips.
- For children with cerebral palsy (GMFCS III), provide manual support as needed to achieve an upright stance without compromising safety.
NOTE: While pediatric patients may require manual support to ensure safety and task completion, care must be taken to minimize occlusion of body segments or markers, as this can compromise data quality.
- Cue the movement verbally ("up") and allow a self-selected comfortable speed during sit-to-stand trials.
8. Trial acquisition
- Before capturing the dynamic trials, capture a static calibration trial with the participant standing in a neutral position with feet shoulder-width apart and arms at 90° of abduction.
- Ensure all markers are visible to at least two cameras; verify in the marker-based system's acquisition software that ≥ 95% of markers are detected across all cameras. Record for 2-3 s. One 3D perspective frame must have all markers.
- In the marker-based system's software, apply the Plug-in Gait static model to define joint centers and segment coordinate systems. Verify proper model fit by inspecting the 3D skeleton visualization and confirming anatomically plausible joint center positions.
NOTE: This static trial provides the biomechanical model parameters required for subsequent 3D dynamic trial processing. In this step, one will need to input the participant's anthropometric measures (e.g., weight, height, leg length, etc.) in the marker-based system's software. When manual support is needed for pediatric patients during the static trial, care must be taken to avoid marker occlusion, which can compromise data quality.
- Record three sit-to-stand23 trials per participant. Resting period between trials can be as needed per participant, as this is not a physically challenging task.
- Use a standardized file-naming convention including group (e.g., TD, CP), participant alphanumeric ID, and trial number.
- In the marker-based system's software, reconstruct marker trajectories using the calibrated camera system.
- Label each marker according to the Plug-in Gait marker set nomenclature. Use automatic gap-filling (maximum gap size: 10 frames) for brief occlusions.
- Verify labeling accuracy by visually inspecting the skeleton in 3D perspective. Once all markers are correctly labeled and trajectories are continuous, export the trial as a .c3d file for analysis in the vailá toolbox.
- Save optical data as .c3d and video frames or stream as lossless or high-bitrate compressed files.
9. Data processing
NOTE: The vailá toolbox (v0.10.21)24 is cross-platform (Windows, macOS, Linux) and requires Python 3.12 or newer. The markerless analysis modules (e.g., MediaPipe25) are optimized to run efficiently on standard CPUs, ensuring accessibility in clinical environments without specialized hardware. A CUDA-enabled GPU is optional and only utilized if the user selects alternative high-speed engines (e.g., YOLO-Pose) for faster processing. For detailed execution commands, step-by-step guides, and full documentation, users are directed to the toolbox's online resources and integrated help files, available at: https://github.com/vaila-multimodaltoolbox/vaila/tree/main/docs/modules/markerless-analysis.
- Marker-based approach: In the marker-based system's software, process the static and dynamic trials. Apply the Dynamic Plug-in Gait model to the labeled marker trajectories (see step 7.4) to compute the joint angle time-series.
- Export the processed trial data as a .c3d file. Ensure that the computed joint angle time-series (e.g., hip, knee, ankle angles) are included in the export.
- Import the processed .c3d file into the vailá toolbox using the vaila.load_c3d() function. The toolbox automatically extracts the pre-calculated joint angle time-series data from the file.
- Apply the 4th-order zero-lag Butterworth low-pass filter (6 Hz cutoff) to the joint angle data using the integrated filtering functions (as described in step 9.2.4 below) to ensure consistent processing with the markerless data.
- Subsequent analyses (e.g., cyclogram generation, coordination metrics computation) are then performed within vailá.
- Markerless approach: Load MediaPipe Pose estimation module (v0.10.21) within vailá. Set the following parameters: static_image_mode = False, model_complexity = 1-2 (use level 2 when GPU/CPU resources allow higher accuracy), smooth_landmarks = True, min_detection_confidence = 0.6, min_tracking_confidence = 0.6.
NOTE: Alternative pose estimation frameworks, such as OpenPose, can be used here.
- Process each video frame to extract the MediaPipe 33 landmarks (Figure 3). Discard frames with detection confidence below the threshold or apply interpolation for missing data.
NOTE: MediaPipe detects 33 body landmarks, including face, torso, and limb keypoints (Figure 3). For the current protocol, only the trunk and lower-body landmarks relevant to sit-to-stand analysis was used: Shoulder [landmarks 11 (left shoulder) and 12 (right shoulder)], Hip [landmarks 23 (left hip) and 24 (right hip)], Knee [landmarks 25 (left knee) and 26 (right knee)], and Ankle [landmarks 27 (left ankle) and 28 (right ankle)]. These landmarks are automatically detected from the RGB video, and they are estimated from visual features (rather than retroreflective marker positions used in the Plug-in Gait model).
- Confirm all key landmarks (shoulder, hip, knee, ankle) are detected with confidence ≥ 0.7.
- (Optional) If lens distortion is relevant, undistort frames using calibration parameters (camera matrix and distortion coefficients) before running MediaPipe.
- YOLO-Pose: In parallel, load the YOLO-Pose module integrated in vailá.
- Configure the model with default weights for human skeleton detection (keypoint set consistent with COCO format). Each frame is processed to detect skeletal keypoints using the YOLO backbone, optimized for robust performance under motion blur, occlusion, or complex backgrounds.
- Apply temporal smoothing filters to reduce frame-to-frame jitter and align detected keypoints with the coordinate system used for MediaPipe.
- For the temporal smoothing filter, apply a low-pass 4th-order Butterworth filter with a cutoff frequency of 6 Hz (sit-to-stand movements have dominant frequency content below 5 Hz) to all joint angle time-series data, both marker-based and markerless, to remove high-frequency noise while preserving the movement signal.
- Implement this filter by using the scipy.signal.butter() and scipy.signal.filtfilt() functions in Python, using zero-phase forward-backward filtering to avoid temporal shifts.
- Data handling in vailá: Export landmark trajectories from both pipelines (MediaPipe and YOLO-Pose) in a standardized .csv or .json format. Store metadata (frame number, timestamp, detection confidence) together with keypoint coordinates. This ensures comparability across methods and reproducibility of analyses.
- Coordinate definitions and joint angles: Define sagittal-plane hip flexion and knee flexion angles with positive values for flexion.
- For markerless data, derive joint centers/segments from MediaPipe landmarks (e.g., hip: pelvis proxy; knee: femur-tibia; trunk: shoulder-hip segment) and compute sagittal angles using a consistent joint coordinate system.
NOTE: Sagittal-plane joint angles are computed geometrically from the 2D landmark coordinates. The vailá toolbox calculates the angle (θ) between two segment vectors (vector a and vector b) using the arccosine of their dot product: θ = arccos((a · b) / (|a| |b|)). Hip Flexion is defined as the angle between the 'trunk' vector (vector connecting MediaPipe landmarks 11/12 [shoulder] and 23/24 [hip]) and the 'thigh' vector (vector connecting 23/24 [hip] and 25/26 [knee]). Knee Flexion is defined as the angle between the 'thigh' vector (vector connecting 23/24 [hip] and 25/26 [knee]) and the 'shank' vector (vector connecting 25/26 [knee] and 27/28 [ankle]).
- Apply the same sign convention and filtering (if needed) to both approaches to facilitate comparison.
- Temporal synchronization: Align data from both approaches using the LockLab trigger time stamp; otherwise, align the first visible flash/clap frame across data streams.
- Resample time series, if necessary, to a common time base before comparisons.
- Define movement cycle:
- Define the sit-to-stand cycle from onset of forward trunk lean (positive trunk angular velocity threshold) to stable upright stance for children typically developing (trunk angular velocity ~0 and knee full extension) and, for children with cerebral palsy, stable upright stance and highest position of cervical marker (marker-based system).
- For children with cerebral palsy, who may not achieve full extension (i.e., trunk aligned with pelvis, hip at 0° flexion, knees fully extended) as observed with children typically developing, define the end of the cycle as the most upright posture achieved consistent with the static standing trial, and confirmed by visual inspection to ensure a stable stance before the next trial.
- Do not time-normalize for the representative plots; report the actual task duration per trial. (Time-normalized analyses can be provided as supplementary, if desired.)
- Cyclogram computation and metrics: Plot hip (x-axis) vs. knee (y-axis) angles for each trial to generate hip-knee cyclograms.
- Overlay trials and approaches (marker-based vs. markerless) for visual comparison.
- Compute optional summary metrics: hip/knee ranges of motion, loop area, loop orientation (principal axis angle), and path length.
- Export processed time series and metrics as .csv for statistical analysis.

Figure 3: MediaPipe's 33 landmarks. For this study, we utilized only the landmarks pertinent to the sit-to-stand task, specifically landmarks 11 and 12 (shoulders), 23 and 24 (hips), 25 and 26 (knees), and 27 and 28 (ankles). Please click here to view a larger version of this figure.