Review Article

Unsupervised Digital Cognitive and Functional Assessment and Training: Is There a Potential Impact of The Lack of Human Contact?

DOI:

10.3791/70431

April 21st, 2026

In This Article

Summary

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Digital strategies enable unsupervised remote assessment and training, but several considerations are needed to ensure the validity of the collected data and the usefulness of the interventions.

Abstract

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Advances in technology allow for the possibility of digital delivery of cognitive and functional assessment and training. There are two broad classes of digital strategies: those that were initially developed for digital delivery and those that were migrated from earlier traditional paper and pencil assessments. There has also been considerable research on supervised vs unsupervised assessment and training. Concurrent digital developments, such as artificial intelligence (AI), have enabled the rapid generation of parallel assessment forms and the translation of assessment materials for international use, suggesting the possibility of worldwide, widely accessible deployment. Translation does not ensure validation, and many assessment instruments have been translated but not normed or validated post-translation. In this review, we examine the current state of digital assessment and training efforts, using examples from widely deployed in-person assessments that have undergone translation and digital migration, and comparing them with digital strategies that were never delivered on paper or in person. Comments are provided on the utility of unsupervised performance and its implications for individuals self- administering digital assessments or training programs. There are several areas that require careful thought to deliver digital strategies remotely that are accessible and personalized, including psychometric integrity, linguistic validity, the social support provided by an in-person human tester, and whether that support can be quantified and supplemented with digital alternatives.

Introduction

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The digital revolution has changed the landscape of cognitive and functional assessment and training in aging-related and psychiatric conditions. A large proportion of assessments performed today rely fully or partially on digital delivery. Digital assessments can be migrated from legacy standard neuropsychological tests administered by a human tester1,2 or fully native to a digital platform3,4,5. Some strategies include delivery on mobile devices6,7,8, which can be accompanied by other remote assessments to capture momentary moods, locations, and environmental conditions9.

Systematic digital migration allows for comparison of digital versions of procedures to legacy versions. For some tasks, extensive validation information has been collected, with systematic comparisons to legacy procedures. An example is the Brief Assessment of Cognition in Schizophrenia (BACS), wherein an extensively normed10 legacy test was migrated to digital delivery11, with a comparative normative study12. A similar example was the digital migration of the Loewenstein-Acevedo Scales for Semantic Interference and learning (LASSI-L)13, which was migrated to an Avatar-guided, remotely administrable14 version with three digital forms15.

An advantage of digital tasks is that rapid, generally automated updates, including testing stimuli, are highly feasible. This includes language translations with Artificial Intelligence (AI) strategies, which can rapidly produce alternate versions. Also facilitated by AI strategies is the development of parallel forms required for repeated testing, particularly of memory or problem-solving, to avoid unacceptable practice effects16, even in populations with dementia17. One of our goals in this review is to evaluate the quality of previous digital migrations and AI-adaptations, including the data on equivalence after digital migration, linguistic translation, and development of parallel forms.

A further development in digital assessment is the remote or unsupervised self-administration of digital assessment or training procedures. While this has been standard practice for years for commercial computerized cognitive training (CCT)18, some studies have suggested that unsupervised self-administration of cognitive testing can also be feasible19,20,21. Unsupervised self-administration can be completed at home on an internet-connected digital device or in a testing center, where multiple participants take the tests under minimal supervision. Feasibility studies have been generally short-term, although some were the same duration or longer durations compared to standard clinical trials. We will review the general state of unsupervised assessments and training and comment on areas with challenges or where more information is needed.

A critical issue in unsupervised assessment and training is that of cross-cultural adaptation and feasibility. In addition to language differences, there are substantial economic differences, likely impacting digital access and literacy, including experience with mobile phones, tablet or laptop computers, and at-home internet access22. The worldwide digital divide is no longer age-related but economic. The motivation for participation in online remote digital activities may vary considerably across participants. Feasibility metrics for unsupervised assessment have included the consent rate, assessment completion, data validity, and the ability to manage the technology. Most clinical unsupervised assessments target individuals with some expected impairments who may be given considerable instruction20,21, including in-person, regarding log-in and completion of assessments or training. Feasibility data from Individuals who sign up for remote studies in which all registration and training are performed fully independently19,23 may not be generalizable to individuals with lower levels of digital skills, lower confidence in their abilities, and lower general cognitive performance. This would include individuals who suspect that they have early signs of dementia and perform an unsupervised assessment in a clinic.

In a study of the transition of two digital assessment systems from device-resident computer presentation to cloud-based delivery24, there were systematic advantages in touchscreen data collection speed compared to a computer mouse. There are common digital literacy challenges in functional capacity assessments: some participants have minimal keyboard skills and need to be taught how to type a capital letter for a password25. In a multi-site study of cognition and functional capacity in the United states26, performance on a digital functional capacity assessment varied by more than 1.0 SD across geographic locations, tracking regional educational attainment and income. We will evaluate the feasibility and validity of unsupervised assessment and training using all available metrics.

As described below, the major challenge of unsupervised technology delivery is not the failure of the digitally derived results to converge with standard measures. The major challenge is the broad feasibility of the technology needed for self-administered assessments and training, as well as overly variable performance in the absence of human supervision. There is an intrinsically social interaction during standard cognitive tasks that provides supervision and assistance to participants who are having difficulty. An important question also addressed is whether such support can be delivered digitally, through AI-based interactions that simulate participant-tester interactions in legacy human-delivered assessment procedures.

Review and Perspective

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Below is a review of the data on the issues introduced above. They include standards for validating performance-based tests, applicable to both standard and digital assessments; standards for the digital migration of legacy assessments; and comparisons of the validation process for digital-native tasks with those that were migrated from legacy paper/human-delivery formats, including alternative forms, translation, and cross-cultural adaptation. There is also a wide-ranging discussion of the feasibility of unsupervised administration compared to human-delivered strategies, considering the overall feasibility of these strategies and the role of the human tester in traditional assessment and training interventions. Several functions are performed in these assessment situations, some of which could, in theory, be delivered with AI-informed assistance. Supplemental FIle 1 presents the issues considered in this review.

General principles in the validation of cognitive and functional assessment
When developing a well-validated digital assessment or treatment strategy, several considerations arise. A prototypical example of this process was the collaborative, large-scale outcomes assessment development program, the Measurement and Treatment Research to Improve Cognition in Schizophrenia (MATRICS) initiative, which conducted a validation study of a large selection of legacy, tester-delivered assessments. Based on discussions at a consensus meeting27, test selection in this context was guided by five core considerations: reliability across repeated administrations, appropriateness for longitudinal use, association with functional outcomes, sensitivity to treatment-related change, and overall feasibility and tolerability. Suitability for repeated administration was defined as either minimal performance improvement due to practice or the availability of equivalent alternative forms. Although the resulting MATRICS Consensus Cognitive battery (MCCB)28 selected a set of tests meeting these criteria, the battery has proven challenging to administer remotely29. A similar development process was used for a similarly focused assessment, the Brief Assessment of Cognition in Schizophrenia (BACS)3. The BACS was developed with similar strategies and the six tester-delivered subtests in the BACS were examined for all these characteristics, including in a large validation study with healthy controls (n = 404)8.

The paper BACS was widely validated internationally, with multiple studies examining its use (e.g.,30,31,32,33,34,35). Interestingly, the MCCB has also been translated multiple times, with the developer’s website36 listing 45 different translations. Also noted on the website was that only ~10 of the 45 versions have official language norms, despite their widespread use in international clinical trials. This is an issue addressed later.

The digital migration of the BACS into the BAC-App was systematically conducted in a comparison study11, including participants with schizophrenia and healthy controls, all of whom were tested with both the legacy and digital versions. Prior to initiation of the study, alternate forms were developed for several BAC-App subtests, including Verbal Memory (seven versions), Symbol Coding (eight versions), and Tower of London (eight versions). Thus, the digital migration of the task was designed to enable direct comparison of the two versions and to validate parallel forms. A subsequent large-scale validation study (N = 55312) with healthy controls and a small number of participants with subjective cognitive decline was conducted, yielding normative information on performance on the BAC-App. Thus, normative data were collected with the legacy version and then repeated with the digital version, including an evaluation of the equivalent forms (although the data have not been published to date). It is important to note that the resulting BAC-app is not intended for unsupervised administration; a tester must be present.

A similar digital migration effort was conducted for a legacy paper assessment, the Loewenstein-Acevedo Scales for Semantic Interference and learning (LASSI-L)13,14. This assessment targets the detection of proactive semantic interference, which has been shown to predict faster cognitive decline and the presence of ADRD biomarkers even in cases with minimal additional impairment15. Extensive normative data were available for the LASSI-L, and a digital version (LASSI-D) was subsequently developed for unsupervised assessment. In that development process, in evaluating these instruments, emphasis was placed on reliability across administrations, suitability for repeated use, relevance to functional outcomes, and overall feasibility and tolerability. An advancement over the BACS migration was the concurrent validation of Spanish and English versions of the LASSI-D, as opposed to similar versions of the LASSI-L, and the development of the migrated version using AI-informed strategies to increase the feasibility and validity of unsupervised self-administration. Specifically, a human avatar performs the functions of a human tester, delivering instructions, correcting errors, and providing encouragement to optimize participant performance. Qualitative data collected from participants indicated high levels of satisfaction with avatar-delivered assessments, with preference scores higher for these than for human-delivered assessments.

Digital-native assessments should also adhere to these same reliability and validity principles, and some programs have. For instance, the development of two different digital functional capacity measures, the Virtual Reality Functional Capacity Assessment tool (VRFCAT)3,37 and the Functional Skills Assessment and Training Program (FUNSAT)25,38 adhered to the same five principles in validation studies. Both VRFCAT and FUNSAT underwent a secondary digital migration to cloud-based delivery24, and both tests have alternative forms with demonstrated similarity. The migration study of the FUNSAT also considered the performance of Spanish and English speakers across three alternative forms.

Another feature of both VRFCAT and FUNSAT was an extensive effort at convergent validation, with comparisons to both traditional and digital assessment strategies performed. For instance, the correlation of the VRFCAT with negative symptoms in schizophrenia was examined39, as was its correlation with other paper-and-pencil functional capacity indices40. A qualitative study was performed to assess its acceptability to the target assessment population41. For the FUNSAT, studies of its factor structure42, sensitivity to treatment43, correlations with the BAC-app44, and other functional capacity measures45, and correlations with self-reported functioning have also been reported46.

Technology of unsupervised assessments
Paper-and-pencil neuropsychological assessment is the defining case of a supervised assessment, which includes social interaction, supervision, guidance, and suggestions to avoid problematic responding. As noted above, some digitally administered cognitive and functional measures are designed to be administered by an examiner. Examples include the Identical-Pairs CPT47 from the MCCB, the BAC-app, and the VRFCAT. Other assessments are designed such that they could be administered unsupervised without the presence of an examiner, because the assessment procedure is automated and runs on its own4. Some of these programs have shown vulnerability to missing48 or invalid scores49 on some subtests, although the proportion of cases with invalid scores has been low. These scores may reflect poor test-taker engagement, which can be corrected in some cases with examiner feedback. With current digital technology, interruptions in responding during testing can be detected, and the process can be temporarily suspended while the participant is reminded to continue. For instance, the FUNSAT assessment and training modules pause the procedure and display a pop-up window if no response is received within 15 s. This built-in detection process is not present in all of the strategies mentioned above.

Data on the feasibility of unsupervised assessments
There are many studies of unsupervised assessments of cognitive performance in serious mental illness and ADRD. A large-scale ADRD review20 identified 28 recent research reports and 23 different tools for unsupervised assessments, completed either at an office visit or remotely. Acceptability data were quite good, with informed consent and participation rates over 85%. Many participants reported that they would be willing to participate again.

Adherence to assessments varied considerably, ranging from 63% in an 8-week study to 94% in another 8-week study. Rates of unusable data were low, but certain assessment strategies led to higher rates of challenges in the completion of the assessments. In home-based augmented reality settings, technical limitations led to the loss of roughly 21% of data, and device incompatibility resulted in nearly 32% of participants being unable to complete the remote assessments.

Psychometric considerations, which are commonly carefully managed in traditional assessment strategies, were not as well considered as they could have been. For instance, only 2/23 tasks reported on the convergence of parallel forms and test-retest reliability were in the low end of the acceptable range. However, as some of these assessments were delivered remotely and ad hoc, environmental factors could influence performance, which is an important consideration. Thus, what appears to be low test-retest reliability could actually reflect sensitivity to environmental influences and within-participant variance associated with changing environmental factors9,50. Convergent validity with standard cognitive assessments was consistently acceptable, with correlations ranging from r = 0.50 to 0.70.

The review of remote cognitive assessments in schizophrenia21 was similarly sized (34 articles) and raised some of the same issues, although only a subset of the studies actually performed truly unsupervised assessments. Most remotely administered assessments did not mention parallel forms, and there were essentially no reports of norms for remote delivery compared to in-person assessments. In fact, only one of 34 studies examined normative performance on remote vs in person assessments versus supervised assessments. A further finding that corresponded with the Polk et al. study was that the inability to complete assessments, due to internet connectivity issues, reports of technology breakdowns, and other technological considerations, was substantial and present in nearly all unsupervised studies.

An issue that arose in an attempt to remotely administer the BAC-app was participants' concern about data integrity, as they were tested in person and remotely with the BAC-app51. In that study, participants, on average, performed substantially better on a word-list learning test administered remotely, suggesting that, as the stimuli were presented, they wrote or recorded them. Adding to the list of considerations for unsupervised assessments, some type of security strategy to maintain data integrity seems required.

Remotely delivered digital cognitive and skills training
There are multiple studies that have been conducted regarding training cognitive and functional skills on a remote basis in schizophrenia52 and mild cognitive impairment53. In addition, cognitive training has been added as a combined intervention to augment the efficacy of other interventions, such as app-based negative symptoms treatments in schizophrenia54and digital functional skills training in MCI25. The general findings have been that the rate of adherence and the training gains seen are quite similar at home and in the office, for populations ranging from schizophrenia to MCI and ADRD. Adherence is never perfect, and training gains are affected by individual differences characteristics55,56. One factor that may lead to adequate gains and adherence is that the cognitive training procedures, whether conducted in the office or remotely, are essentially identical. The issues mentioned above about technical challenges and inability to launch the programs are clearly the same for remote training as for remote assessment. Studies in which participants received in-person training to facilitate launching and operating the software unsupervised have reported the best remote adherence (e.g.,53).

Worldwide utilization of digital assessments and treatments
As noted above, there have been multiple translations of legacy paper assessments, as well as multiple international efforts in the digital domain. For instance, digital assessments of dyslexia57 and mathematical skills58 have been developed in Spain, although it is unclear whether these tasks have been translated into other languages. In China, a strategy was developed to modify the Trail-Making Test to include additional motor skills measurement59 and another study in China developed a dual-task interference procedure to examine the processing of complex sentences60. In Singapore, a unique cross-linguistic training program involving bilingual individuals trained in English and one other Asian language, with the aim of improving cognitive performance61. All these studies reported successful development efforts and acceptability, although only one allowed for a direct comparison of the usefulness of the tests across language within cultural groups.

The role of local validation of translated tests
The most widely used traditional in-person assessment strategies, such as the MCCB and the BAC, have been extensively validated in different languages and cultural groups as mentioned above. In contrast, the NIH Toolbox, a wide-ranging assessment aimed at individuals ages 3–85, is offered only in English and administered on a single in-person digital platform62. There is considerably less regional and global information available on the validation of digitally migrated versions of cognitive and functional assessment tasks and training programs. For example, in the ERUDITE63study of luvadaxistat, the BAC-app was administered in eight different languages, and there has been only one published study64 validating a different language form (i.e., Spanish) of that assessment post-digital migration. Some direct-to-digital assessments, such as the VRFCAT, have been used in translated variants in studies with more than 30 different languages65. Again, no substantial evidence base exists for the use of the VRFCAT across languages, although validation of the similarity of the parallel forms was presented when the original English-language version of the assessment was developed. Validation information collected with legacy versions of digitally migrated tasks is likely to be quite relevant if the migration did not change any features of the test, but there are a number of languages where outcome measures, legacy and digital migrations, have not been studied for their psychometric characteristics. However, as noted in two different reviews20,21 for ADRD and schizophrenia, there has been remarkably little attention paid to characteristics of tasks on an international basis after modification and digital migration.

The challenges in these two studies were that the VRFCAT showed greater variability than expected, and a substantial number of participants had extremely low scores. It is possible that the translations had adverse effects on performance. It is just as likely that challenges with inexperienced testers led to poor supervision of participant engagement during the assessments. Despite the obvious potential for tester challenges, it is also probably premature to think that a linguistically validated translation would guarantee a psychometrically similar instrument. During the process of systematic digital migration, the paper and pencil legacy instrument should be examined for the generation of similar scores as the digitized version, although, as noted by the two reviews above, this is not documented to be common. The strategy described above for the digital migration of the LASSI-D is not susceptible to these challenges, as it includes a systematic psychometric comparison of each of the three digital parallel forms with the results of the human-administered legacy version across both Spanish and English. It is, however, a very detailed process and would be extremely challenging to implement in a timely manner for a 30-language clinical trial.

What is the role of the human tester?
The data suggests that a major challenge in unsupervised digital assessment and training is enabling participants to make the technology work and access the system. Other challenges, such as participants’ reports of technological incompatibility, may also stem from their lack of familiarity with digital strategies and their inability to operate technology. Thus, in unsupervised assessment paradigms, personalized instructions prior to the initiation of assessments may be a helpful, or even requisite, strategy.

Some technology-based assessments are primarily administered during office visits, including the majority of tasks designed for remote assessment in schizophrenia. For instance, the VRFCAT requires in-person supervision of assessments, as does the BAC-App. VRFCAT has an introduction component administered prior to the formal assessment, during which the human tester answers any questions and redirects any inadequate effort, while the main assessment section runs continuously. The examiner is directed to supervise the participant and give them guidance if they fail to persist in their efforts. However, a recent VRFCAT study found that some participants had many “forced progressions,” meaning they did not solve a testing objective within 300 seconds. Since some of the objectives are a single brief task (“Pick up the wallet”), failures in examiner supervision could be suspected. The BAC-app technology administers all stimuli and collects all data but still requires the human tester to supervise performance and advance the subtasks after participant completion. The Cogstate system does not require supervision of participants’ performance, but there have been multiple reports of invalid performance on certain components of the system and reduced sensitivity to impairments in schizophrenia66 that could arise from inadequate sustained performance and erratic scores on the part of participants, which could be detected and corrected with in-person supervision.

Is it possible to use digital supervision?
There is considerable current interest in Chatbots or other AI-delivered digital psychotherapeutic interventions. Among the challenges of these interventions is the requirement for bidirectional communication, in which the AI therapist must understand the client's communication in real time and respond appropriately. A different AI strategy was used in developing the LASSI-D. A human avatar was created with AI and served as a test technician, providing all instructions and feedback that would be issued by a human tester in a standard testing situation. The avatar was also taught to answer a predetermined set of FAQs, which were triggered by the examine touching a “?” icon on the screen. There was no attempt to communicate interactively at that stage. Acceptability of the procedure was quite high (95% completion of testing sessions) in a largely minority, monolingual Spanish-speaking population, with very high completion rates (over 96% of assessment sessions) and qualitative statements of preferences for avatar testers over human testers.

A possible avenue for expansion of AI guidance for unsupervised assessment and training would be to develop a training and assistance system designed to help participants set up and originate their unsupervised assessments or training for the first time. Although still in development, FUNSAT is being modified to include an initial instructional module with human avatar guidance to train potential participants to sign up for training, log in to the website, and initiate their training efforts. These strategies are more likely to be delivered consistently and accurately than chatbot therapeutic interactions, because avatar communication would be based on pre-scripted answers to questions and the delivery of consistent instructions.

Other digital strategies
The current review focuses on performance-based, intentional cognitive assessments. There is an impressive array of passive digital phenotyping measures that are now collected and can be related to unsupervised assessment results. These include geolocation67, physical activity68, social context69, interactions with devices70, and speech and language characteristics71. Although these are not necessarily measures of cognition, they reflect potential environmental factors that may also influence cognitive performance. For instance, mood state inferred from facial or vocal responses, social context, and background distraction can correlate with cognitive performance, particularly during unsupervised assessments. Some of the challenges associated with unsupervised assessment may be detectable and reducible through the concurrent use of passive eC digital phenotyping. As these measures do not require intentional effort to collect the data, they do not add burden to assessments.

Conclusions

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Digital assessment and treatment of cognitive and functional deficits are highly promising because they enable scaling the process and reaching participants who might not have access otherwise. Technological developments enable easy translation of stimuli, the creation of parallel assessment forms, and remote delivery. Most data suggest that the results of unsupervised assessment and training are highly convergent with those of legacy strategies, in terms of global performance indices. However, there are challenges in both participant engagement and the technical features of test development. Many participants are unable to manage technology unsupervised. Although technology makes it easy to translate stimuli and create parallel forms, there has not been adequate attention to the psychometric similarity of resulting translations or new forms. The systematic approach to psychometric evaluation of test quality is not excused because technology makes it easier to generate test stimuli.

Some pharmacological treatment studies using digital outcomes assessments have failed to identify treatment effects, and those failures could partially relate to different versions of the tests, in terms of parallel forms or translated stimuli, not converging with the original version. These strategies, either due to tester issues or challenges with the newly developed forms, have also been found to have the potential to yield overly variable test performance and missing data.

A supplementary approach has been to use AI-guided avatar testing and training assistants. Such approaches have considerable promise, but themselves require validation before full deployment. In cases where such strategies are used, a stepped approach would be required: the assistant would initially deliver fully scripted instructions and feedback, then be trained with large-language models to understand participant questions and interact with participants.

The social context of assessment and intervention is important to sustain technology-related efforts. Human testers can adapt to participant challenges and facilitate engagement with training, while providing technical assistance. This interaction will need to be integrated into digital approaches to unsupervised cognitive assessment and training to broaden the reach of these strategies. Participants who are most in need of these assessments and intervention strategies are also most likely to need facilitated assistance while self-administering cognitive and functional skills assessment and training tasks.

Supplemental File 1: Issues associated with developing valid performance-based assessments: Focus on validity, feasibility, and tolerability of digitally delivered assessments and training. Please click here to download this file.

Disclosures

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

Dr. Harvey has received consulting fees or travel reimbursements from Alkermes, BMS (Karuna Therapeutics), Boehringer Ingelheim, Kynexis, Minerva Neurosciences, Neurocrine Biosciences, and Recognify. He has a research grant from Intracellular Therapeutics (Now part of Johnson and Johnson). He receives royalties from the Brief Assessment of Cognition in Schizophrenia (Owned by Clario, formerly WCG Endpoint Solutions, Inc, formerly Verasci, and contained in the MCCB). He is the chief scientific officer of i-Function, Inc., and a scientific consultant to EMA Wellness, Inc. Mr. Kallestrup is the co-founder and chief executive officer of i-Function, Inc.

Acknowledgements

Loading...
$$\rightleftharpoonup{xx}$$ $$\longleftharp{xx}$$, $$\longrightharp{xx}$$,

The research supporting the development of the FUNSAT and the LASSI-D was funded by United States National Institute on Aging (NIA) grants R44 AG074818-02, R43 AG057238-02, and R44 AG057238-04, with Peter Kallestrup as PI.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

Digital Cognitive AssessmentFunctional AssessmentUnsupervised AssessmentDigital TrainingArtificial IntelligenceAssessment TranslationPsychometric IntegrityLinguistic ValidityHuman Tester SupportSelf Administered Assessment

Related Articles