$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
Kinematic assessment of the drinking task has been recommended for use in stroke rehabilitation research6. However, practical concerns have been cited as barriers to the uptake of kinematics in rehabilitation research, and these concerns include the cost and technical requirements for motion capture technologies conventionally used to measure human biomechanics15. These concerns subsequently impede clinical uptake and translation of kinematic assessment to clinical practice. Alternatively, a more practical approach to extract kinematics by computer vision can be realized using a single stereo camera, a personal computer, and items commonly found in an outpatient clinic setting16. We present a protocol designed for clinicians that employs a web-based guide to record images and video that are compatible with a computer vision workflow to extract kinematic metrics of the drinking task. The long-term goal is to increase accessibility to these metrics, which have been previously developed and recommended for use in stroke rehabilitation research5,6,12,13,14.
Video synchronization is important for dual-camera approaches to computer vision. Dual-camera setups mimic human eyes, which simultaneously perceive a scene in the real world from slightly different vantage points to allow depth perception. In the case of computer vision, appropriate synchronization is then critical to the accuracy of 3-dimensional information (e.g., position of the human wrist in a Cartesian coordinate system) determined from two 2-dimensional sources (e.g., pixel location of the human wrist in a photo). In our experience, dual cameras can provide reasonable results when synchronized to within 1 ms16. Several factors may reduce synchronization, including unreliable camera frame rates, limited data transmission rates in cables, software inefficiencies, and limited hardware specifications (e.g., low RAM or low-tier processors). To alleviate synchronization issues, this protocol utilizes an off-the-shelf stereo camera, which benefits from a smaller physical footprint and is designed as a plug-and-play device (e.g., single USB connection to a standard personal computer) to automatically capture synchronized images from two built-in image sensors.
Camera characteristics require careful consideration for any computer vision application. In theory, higher resolution images and higher frame rate recordings of human movement offer more data for detecting anatomic key points with greater spatiotemporal accuracy. However, an inevitable balance is necessary between the amount of data and the computing resources available to process the data. The described protocol is based on a stereo camera that records video with 1280 x 720 pixel resolution at a frame rate of 60 Hz, which is comparable to prior studies applying computer vision for drinking task kinematics16,18. For recording these videos, we have observed reliable success using a higher-tier consumer-grade computer (e.g., 3.2 GHz CPU with 128 GB RAM) and mixed success using a laptop computer (e.g., 2.6 GHz CPU with 16 GB RAM). When applying lower-tier computers, a typical issue is missing video frames, which often presents as an output video with a duration less than the intended 20 s video recording. Of note, while data computing is a common culprit behind missing video frames, data transmission is also of importance. In the healthcare setting, for example, security settings associated with computer access ports (i.e., USB drives) might result in data transmission delays that interfere with reliable video recording.
Line-of-sight interference will challenge key steps in a computer vision workflow, including calibration and human pose estimation. Calibration relies on digital images of a standard-sized checkerboard pattern, which allows computer vision to map image pixels to the physical dimensions of the real world. Clear visualization of the checkerboard pattern improves the success rate of this mapping. Visual interference may arise due to a number of reasons including but not limited to the following: (i) lighting issues (e.g., bright glare from a nearby window or dark shadows); (ii) excessively fast movements of the checkerboard (e.g., image blurring due to prolonged camera exposure time); or (iii) inappropriate intrinsic camera properties (e.g., incorrect focal length or severe lens distortion). Human pose estimation leverages artificial intelligence to automatically annotate anatomic landmarks from images of a single person (single-pose estimation) or multiple people (multi-pose estimation). Clear visualization of the human is then critical to the success rate of pose estimation. Some visual interference may be tolerated depending on computer vision parameters (e.g., AI architecture, prediction thresholds). However, to gauge the suitability of a video for human pose estimation, common pitfalls to be wary of include the following: (i) bystanders in view in case of single-pose estimation; (ii) body segments intermittently outside the camera field-of-view; and (iii) self-occlusions due to overlapping body segments.
Decoupling data collection from the computer vision workflow itself offers both advantages and drawbacks. A notable drawback is the delay in clinically relevant information available to clinicians and patients to inform decision-making. A potential advantage is reduced interference with the clinician-patient interaction. For example, assuming calibration is performed prior to a patient's appointment, video recording 10 repetitions of the drinking task as described in prior literature12 would ideally require less than 4 min of video recording time plus any time required to orient patients who are naive to the protocol. Furthermore, this decoupling may boost clinician acceptance by dividing the responsibility for information. Within this separation of duties, the clinician takes ownership of when and why video assessments occur, and data scientists/engineers take ownership of refining the computer vision workflow that is responsible for accurate/reliable clinical metrics (e.g., drinking task kinematics). Indeed, this can be likened to the current reality of many blood-based (e.g., complete metabolic panels) and image-based tests (e.g., MRI studies) routinely ordered by clinicians. Lastly, this decoupling may be advantageous as the field of AI continues to advance. The video data acquired today could undergo repeat analysis by more advanced computer vision solutions of tomorrow, which may extract more kinematic metrics with greater speed, accuracy, and reliability.
Our goal is to further advance access to upper limb kinematics among the rehabilitation community, which includes clinicians in outpatient settings and foreseeably individuals in the home setting (e.g., telerehabilitation). This protocol and the web-based app should be viewed as an early stage of development, which has so far considered a very focused activity (i.e., the drinking task) and a relatively focused group of stakeholders (i.e., clinician users and subjects who can readily perform the drinking task). Future versions will be necessary to accommodate a more diverse user group, more diverse activities, and more diverse subject populations (e.g., individuals with severe deficits requiring activity accommodation). Additionally, future development will include important checkpoints before widespread implementation, such as integration into electronic medical records and maintenance of patient privacy and confidentiality. These future versions will also likely incorporate new technology and refine currently available technology. For example, the current protocol assumes a computer vision workflow based on binocular vision (e.g., stereo cameras). A logical next development is a protocol featuring monocular vision (e.g., more basic cameras), which may be achievable by utilizing AI trained on binocular datasets. Of note, the current protocol yields image and video data compatible with current state-of-the-art computer vision solutions, which are already being superseded in performance by newer solutions18. The web-based protocol described here may allow clinicians to implement into their practice a foundation for computer vision, which is likely to advance dramatically as AI technology permeates the healthcare industry.