$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
The representative results of this study include the usability questionnaire results, eye tracking data analysis, NASA-TLX scale data, fNIRS data analysis, and dynamic cognitive load changes. For the usability questionnaire results, eye tracking data analysis, NASA-TLX scale data and fNIRS data analysis, normality tests, and differences tests were conducted. For dynamic cognitive load changes, this study selected fNIRS and eye tracking data from a single participant to demonstrate the validity of the multimodal measurement.
Usability questionnaire results
None of the items in the usability questionnaire followed normal distribution (Table 2). The reliability of AR and the website in the usability questionnaire were tested, and the Cronbach's alpha score was considered acceptable (Cronbach's alpha = 0.974).
Table 2: Normality test of the usability questionnaire. None of the items in the usability questionnaire followed normal distribution. The data were analyzed using the Wilcoxon signed rank test. Please click here to download this Table.
The median difference scores of the usability questionnaire between AR and the website are shown in Table 3. The data distributions of the AR and website conditions are shown in Figure 3. A significant difference was observed between the AR and website conditions, with the median scores for AR being higher than those for the website. The results showed that the participants had a better user experience in the AR condition than in the website condition.
Table 3: Median difference scores of usability questionnaire between the AR and website conditions. The median scores for AR were significantly higher than those for the website. Please click here to download this Table.

Figure 3: Data distribution of the usability questionnaire. A schematic illustration of data distribution of the usability questionnaire. Please click here to view a larger version of this figure.
Eye tracking data analysis
All eye tracking indicators were tested for normality, and the results are presented in Table 4. In Tasks 1 and 3, only fixation frequency followed normal distribution, while all the other indicators did not follow normal distribution. In Task 2, fixation frequency and saccade frequency followed normal distribution, but the rest of the indicators did not follow normal distribution. For Task 4, only saccade frequency followed normal distribution. When comparing differences between the AR and website conditions, the data were reported separately based on normality/non-normality. Table 5 shows differences in eye-tracking indicators between the AR and website conditions in Task 1 (quality of water). There were significant differences in all eye-tracking indicators between the AR and website conditions, p < 0.001. The average fixation duration was significantly longer in the AR condition than in the website condition (MedianAR = 586.85, interquartile range (IQR) = 482.55-714.6; Medianwebsite = 398.05, IQR = 362.775-445.275). The other indicators were significantly lower in the AR condition than in the website condition.
Table 4: Normality test of eye-tracking indicators. The eye-tracking data that follow a normal distribution were analyzed using a paired samples t-test, and the eye-tracking data that did not follow a normal distribution were analyzed using the Wilcoxon signed rank test. Please click here to download this Table.
Table 5: Differences in the eye-tracking indicators between AR and the website in Task 1. There were significant differences in all eye-tracking indicators between AR and website conditions, p < 0.001. The average fixation duration was significantly longer in AR than in the website condition (MedianAR = 586.85,interquartile range (IQR) = 482.55-714.6; Medianwebsite = 398.05,IQR = 362.775-445.275). The other indicators were significantly lower in AR than in the website condition. Please click here to download this Table.
Table 6 shows the differences in the eye-tracking indicators between the AR and website conditions in Task 2 (storage temperature). All eye tracking indicators showed significant differences between AR and website conditions, p < 0.001. The average fixation duration was significantly longer in the AR condition than in the website condition (MedianAR = 477.2,IQR = 398.675-596.575; Medianwebsite = 397.1,IQR = 353.35-451.075). The other indicators were significantly lower in the AR condition than in the website condition.
Table 6: Differences in the eye-tracking indicators between AR and the website in Task 2. There were significant differences in all eye tracking indicators between the AR and website conditions, p < 0.001. The average fixation duration was significantly longer in AR than in the website condition (MedianAR = 477.2, IQR = 398.675-596.575; Medianwebsite = 397.1, IQR = 353.35-451.075). The other indicators were significantly lower in AR than in the website condition. Please click here to download this Table.
Table 7 shows the differences in the eye-tracking indicators between the AR and website conditions in Task 3 (matching diet). There were significant differences in all eye-tracking indicators between the AR and website conditions, p < 0.001. The average fixation duration was significantly longer in the AR condition than in the website condition (MedianAR = 420.45,IQR = 352.275-467.8; Medianwebsite = 360.6, IQR = 295-399.075). The other indicators were significantly lower in the AR condition than in the website condition.
Table 7: Differences in the eye-tracking indicators between AR and the website in Task 3. There were significant differences in all eye-tracking indicators between the AR and website conditions, p < 0.001. The average fixation duration was significantly longer in AR than in the website condition (MedianAR =420.45,IQR = 352.275-467.8; Medianwebsite = 360.6,IQR = 295-399.075). The other indicators were significantly lower in AR than in the website condition. Please click here to download this Table.
Table 8 shows differences in the eye-tracking indicators between the AR and website conditions in Task 4 (price per liter). There were significant differences in all eye-tracking indicators between the AR and website conditions, p < 0.001. The average fixation duration was significantly longer in the AR condition than in the website condition (MedianAR = 495.25,IQR = 404.8-628.65; Medianwebsite = 263.1, IQR = 235.45-326.2). However, the other indicators were significantly lower in the AR condition than in the website condition.
Table 8: Differences in the eye-tracking indicators between AR and website in Task 4. There were significant differences in all eye-tracking indicators between the AR and website conditions, p < 0.001. The average fixation duration was significantly longer in AR than in the website condition (MedianAR =495.25,IQR = 404.8-628.65; Medianwebsite = 263.1,IQR = 235.45-326.2). The other indicators were significantly lower in AR than in the website condition. Please click here to download this Table.
For the visual search tasks, lower eye-tracking indicators were associated with a higher efficiency of information search (except for the average fixation duration). Taken together, the eye-tracking data demonstrated that the participants had higher information search efficiency when using AR than when using the website.
NASA-TLX scale data
None of the items of the NASA-TLX scale followed a normal distribution (Table 9). The Cronbach's alpha score was considered acceptable (Cronbach's alpha = 0.924).
Table 9: Normality test of NASA-TLX scale. None of the items of the NASA-TLX scale followed a normal distribution. The data were analyzed using the Wilcoxon signed rank test. Please click here to download this Table.
The median difference scores of the NASA-TLX scale between AR and website conditions are presented in Table 10. Data distributions of AR and website conditions are shown in Figure 4. A significant difference was observed between AR and website conditions. The NASA-TLX scale scores of the AR condition were lower than those of the website condition, indicating that the AR technique led to a lower cognitive load than that by the website.
Table 10: Median difference scores of NASA-TLX scale between AR and the website. The NASA-TLX scale scores of the AR condition were significantly lower than those of the website condition. Please click here to download this Table.

Figure 4: Data distribution of the NASA-TLX scale. A schematic illustration of data distribution of the NASA-TLX scale. Please click here to view a larger version of this figure.
fNIRS data analysis
The mean O2Hb values were tested for normality, and the results are presented in Table 11. When comparing differences between the AR and website conditions, data were reported separately based on normality/non-normality. Differences in mean O2Hb between the AR and website conditions are presented in Table 12. There were significant differences between the two conditions when participants performed Task 1 (adjusted p = 0.002), Task 3 (adjusted p = 0.007), and Task 4 (adjusted p < 0.001). The mean O2Hb of the tasks performed in the AR condition was significantly lower than in the website condition (Task 1: MeanAR = -1.012, SDAR = 0.472, Meanwebsite = 0.63, SDwebsite = 0.529; Task 3: MeanAR = -0.386, SDAR = 0.493, Meanwebsite = 1.12, SDwebsite = 0.554; Task 4: MeanAR = -0.46, SDAR = 0.467, Meanwebsite = 2.27, SDwebsite = 0.576). While performing Task 2, differences between the AR and website conditions did not reach a significant level (adjusted p = 0.154 > 0.05). These results indicate that participants had a lower cognitive load when using the AR technique than when using the website.
Table 11: Normality test of mean O2Hb. The fNIRS data that follow a normal distribution were analyzed using a paired samples t-test, and the fNIRS data that did not follow a normal distribution were analyzed using the Wilcoxon signed rank test. Please click here to download this Table.
Table 12: Differences in mean O2Hb between AR and website. There were significant differences between the two conditions when participants performed Task 1 (adjusted p = 0.002), Task 3 (adjusted p = 0.007), and Task 4 (adjusted p < 0.001). The mean O2Hb of the tasks performed in the AR condition was significantly lower than in the website condition (Task 1: MeanAR = -1.012, SDAR = 0.472, Meanwebsite = 0.63, SDwebsite = 0.529; Task 3: MeanAR = -0.386, SDAR = 0.493, Meanwebsite = 1.12, SDwebsite = 0.554; Task 4: MeanAR = -0.46, SDAR = 0.467, Meanwebsite = 2.27, SDwebsite = 0.576). While performing Task 2, differences between AR and website conditions did not reach a significant level (adjusted p = 0.154 > 0.05). Please click here to download this Table.
Dynamic cognitive load changes
Figure 5 shows the changes in O2Hb concentration when a participant performed Task 4 in the website condition. At point 1, the participant had trouble calculating the price per liter. The intense searching process induced an increase in O2Hb concentration, which indicated an increase in the instantaneous load. When the participant received a cue, the O2Hb concentration dropped to point 2, and the instantaneous load reached a valley value at that moment. The participant then started working hard to calculate the price per liter and wanted to complete the task as soon as possible. In this context, the O2Hb concentration continued to increase and reached a maximum (point 3). In summary, the multimodal measurement of eye tracking and fNIRS could effectively measure dynamic changes in cognitive load while interacting with information systems and can also examine individual differences in consumer behavior.

Figure 5: fNIRS instantaneous load. A schematic illustration of the dynamic cognitive load changes using fNIRS instantaneous load. Please click here to view a larger version of this figure.
Taking the four tasks of AR integrated into IoT as examples, this study combined NeuroIS approaches with subjective evaluation methods. The experimental results suggested that: (1) for the usability questionnaire, participants had a better subjective evaluation in the AR condition than in the website condition (Table 3 and Figure 3); (2) for the eye tracking data, participants had higher information search efficiency when using AR than when using the website (Table 5, Table 6, Table 7, and Table 8); (3) for the NASA-TLX scale data and fNIRS data, the AR technique led to lower cognitive load than that with the website (Table 10 and Table 12); and (4) for the dynamic cognitive load, the multimodal measurement of eye tracking and fNIRS could effectively measure dynamic changes of cognitive load while interacting with information systems, and could also examine the individual differences in consumer behavior (Figure 5). By comparing the differences in the neuroimaging data, physiological data, and self-reported data using the usability questionnaire and NASA-TLX scale between the AR and website conditions, the AR technique could promote efficiency for information search and reduce cognitive load during the shopping process. Thus, as an emerging retail technology, AR could effectively enhance the user experience of consumers and in turn may increase their purchase intention.
Supplemental Figure 1: Screenshot of information displayed on the AR application used in the study. Please click here to download this File.
Supplemental Figure 2: Screenshot of information displayed on the website used in the study. Please click here to download this File.