Research Article

Explainable AI-Supported Cyber-Physical Collaboration for Sustainable Manufacturing in Industry 5.0

1 views

DOI:

10.3791/72379

September 1st, 2026

 , 

Corresponding Authors: Tao Ni <proftaoni@gmail.com>

In This Article

Summary

This study presents an Explainable AI-enabled Industry 5.0 cyber-physical collaboration framework that uses a CNN-transformer and SHAP for industrial acoustic anomaly detection on the MIMII dataset. Edge-computing integration supports interpretable, human-centric predictive maintenance.

Abstract

In this study, we present an innovative approach to the sustainable manufacturing of industrial parts using an explainable artificial intelligence (XAI)- based cyber-physical collaboration system for Industry 5.0. Current cyber-physical human systems (CPHSs) have been found to integrate AI only to a limited extent and often lack explainability. Consequently, there is a need to improve their scalability and flexibility, in keeping with the tenets of Industry 5.0: resilience, long-term viability, and human-centricity. To address these shortcomings, we introduce an architecture that integrates deep learning and explainable AI into the Cyber-Physical System (CPS) ecosystem to enable smart, explainable decision-making. Specifically, we use the Malfunctioning Industrial Machine Investigation and Inspection (MIMII) dataset for machine fault detection from acoustic signals. Audio signals undergo spectral transformation to generate spectrogram images, which are then processed by a convolutional neural network (CNN) for spatial feature extraction and a transformer encoder for temporal features. To improve transparency and facilitate human-in-the-loop cooperation, the Shapley additive explanations (SHAP)-based approach helps operators understand the system's decision-making process by highlighting input features that influence model predictions. In addition, the framework is applied in an edge-computing-based cyber-physical system to achieve low latency and low energy consumption while providing timely responses, thereby helping ensure a sustainable production environment. Experimental findings show that the proposed CNN-Transformer architecture achieves 98.5% accuracy and outperforms existing CPHS systems while maintaining low latency. The use of explainable AI in such a cyber-physical cooperation system greatly increases operator confidence and enables seamless human-machine collaboration.

Introduction

The transition from Industry 4.0, which emphasized technology, mechanization, and communication, to the new Industry 5.0 model, which combines technological advancement with adaptability, sustainable development, and human-centeredness, has influenced the dynamic evolution of industrial systems1,2. The European Commission defines Industry 5.0 as a paradigm focused not only on productivity and efficiency but also on the three pillars of human-centricity, sustainability, and resilience. It is important to note that recent research has pointed out the role of AI systems, human-machine cooperation, and sustainable manufacturing as the main drivers of Industry 5.0 environments3,4. Traditional maintenance methods are no longer sufficient in this evolving industrial environment. To enable resilience in complex, unstable contexts, maintenance must evolve into a more adaptable and systemic form5. Since 2011, Industry 4.0 has been the dominant model for industrial transformation in Europe, encouraging the digital transformation of businesses to enhance operations and maximize available resources4. Increasing efficiency without eliminating human labor in production, while protecting workers’ health and safety, is therefore imperative. Because of this, businesses and academics have begun defining the tenets of the Industry 5.0 paradigm5,6,7, partly motivated by the Japanese notion of Society 5.08,9,10,11. Despite the increased emphasis on sustainable maintenance in recent literature, most existing systematic evaluations remain fragmented or domain-specific, lacking a comprehensive framework that integrates resilience, sustainability, and AI. For example, the authors examine sustainable maintenance practices for specific systems or strategies, including sustainable total productive maintenance (STPM), whereas the researchers talk about incorporating sustainability themes into maintenance techniques12,13,14. Author conducted a comprehensive assessment of sustainability adoption criteria and highlighted technological factors that enable manufacturing to have a positive sustainable impact15. Nevertheless, within the Industry 5.0 paradigm, sustainability, resilience, and AI-based decision-making are not fully integrated in any of these works. As a result, separate maintenance methods frequently address sustainability. SHAP stands out among other explainable artificial intelligence (XAI) techniques due to its theoretically sound feature attribution method16,17,18.

Data-driven decision-making is enabled by AI, but accurate models often require large amounts of data. Conventional machine learning (ML) techniques, such as decision trees, logistic regression, and linear regression, frequently perform poorly in highly nonlinear situations because they assume relatively simple data distributions19. By learning intricate feature representations from massive datasets, deep neural networks (DNNs) get around these restrictions20. Deeper networks often perform better than shallow structures on complex decision-making tasks, according to studies21, but their increasing depth and parameter complexity make the model less interpretable22.

A Cyber-Physical System (CPS) is a crucial enabling technology for Industry 5.0 manufacturing because it combines distributed intelligence, embedded computers, physical processes, and communication. To enable flexible, intelligent, and real-time industrial operations, modern CPSs are rapidly integrating edge computing, edge AI, and green Internet of Things (G-IoT)23,24. Sustainability and resilience have been highlighted in recent studies on Maintenance 4.0 and smart maintenance. While earlier works have examined the significance of Maintenance 4.0 technology in sustainable manufacturing and suggested modeling frameworks for sustainable maintenance, bibliometric and systematic studies have assessed smart maintenance technologies25,26,27,28. Further research has investigated the integration of sustainable maintenance into Industry 4.0 environments and sustainability-oriented production systems29,30,31,32. The Operator 4.0 concept33, human-centric Industry 5.0 frameworks34, and socioeconomic perspectives35 have all garnered significant attention to human-centric maintenance. Additionally, advancements in digital maintenance36, predictive and prescriptive maintenance37,38, and sustainable maintenance strategies39,40 have been emphasized in recent studies. Despite these advancements, research remains dispersed, and few studies offer a cohesive paradigm that integrates explainable AI, sustainability, resilience, and human-centered cyber-physical cooperation. Cyber-Physical Systems (CPSs) have not attained full autonomy despite their growing deployment, and the Industry 5.0 paradigm does not favor it. Industry 5.0, on the other hand, places greater emphasis on human supervision, cooperation, and decision-making. By incorporating human thinking and cooperation into cyber-physical processes, cyber-human systems such as human-in-the-loop CPSs (HiLCPSs)41, Human-Centered CPSs (HCPSs)42,43, and Cyber-physical human systems (CPHSs)44extend traditional CPSs. This integration improves adaptability, transparency, safety, and resilience by enabling real-time collaboration between intelligent systems and human operators45,46.

Industrial audio anomaly detection has been greatly enhanced by new architectures, including Vision Transformer (ViT), Audio Spectrogram Transformer (AST), Conformer, and self-supervised learning; however, their integration with explainable cyber-physical systems for Industry 5.0 remains limited. Without integrating explainability, edge deployment, and human-in-the-loop decision support into a single framework, current research primarily focuses on classification accuracy or on post hoc XAI-based interpretability. This study proposes a novel Explainable AI-enabled CNN-Transformer cyber-physical collaboration architecture that combines edge computing, temporal dependency modeling, SHAP-based explainability, spatial feature learning, and human-centric decision support to enable sustainable Industry 5.0 manufacturing and overcome these constraints.

Thus, there remains a substantial research deficit. Current CNN/Transformer-based acoustic anomaly detection techniques favor classification performance but lack cyber-physical cooperation and explainability. While Industry 5.0 frameworks offer conceptual human-centric models that do not integrate explainable deep learning, edge intelligence, and operator-assisted decision-making into a single system, XAI-based predictive maintenance techniques increase transparency but operate independently of CPHS architectures. Therefore, a single Industry 5.0 Cyber-physical human system that integrates precise auditory anomaly detection, SHAP-based explainability, edge-enabled deployment, and human-in-the-loop cooperation is still lacking.

This paper proposes a novel Explainable AI-enabled CNN-Transformer cyber-physical collaboration paradigm for Industry 5.0 to overcome these constraints. The suggested architecture combines edge computing deployment, CNN-based spatial feature extraction, Transformer-based temporal dependency learning, SHAP-based explainability, and human-in-the-loop decision support into a single CPHS framework. The suggested framework offers an integrated solution that concurrently enhances anomaly detection performance, model transparency, operational efficiency, and human-centric collaboration for sustainable manufacturing, in contrast to current methods that handle these elements independently.

This study's main goal is to create an Explainable AI-enabled Cyber-physical human system (CPHS) for sustainable Industry 5.0 manufacturing by combining SHAP-based explainability, edge computing, and human-in-the-loop decision support with a hybrid CNN-Transformer-based acoustic anomaly detection approach. In particular, the suggested framework aims to improve anomaly detection accuracy while providing clear, understandable forecasts, facilitating effective edge deployment, and encouraging cooperative decision-making between intelligent systems and human operators. By concurrently achieving high detection performance, explainability, low-latency operation, and human-centric cyber-physical collaboration, this integrated architecture addresses the shortcomings of current acoustic anomaly detection, XAI-based predictive maintenance, and CPHS frameworks.

The suggested architecture presents a unified Explainable AI-enabled Cyber-physical human system (CPHS) for Industry 5.0, in contrast to current acoustic anomaly detection techniques, which primarily focus on classification accuracy. Within a single architecture, it combines CNN-based spatial feature extraction, Transformer-based temporal modeling, SHAP-based explainability, edge computing, and human-in-the-loop decision support. The suggested architecture addresses key issues of transparency, operator assistance, and cyber-physical cooperation by fusing precise anomaly detection with comprehensible decision-making and real-time edge deployment. The main scientific contribution of this work is the system-level integration, which sets the suggested approach apart from previous research that focuses on anomaly detection, explainable AI, or Industry 5.0 frameworks alone.

Protocol

The suggested methodology uses a system-based Industry 5.0 Cyber-physical human system (CPHS) approach that combines explainable deep learning with a cyber-physical collaboration framework enabled by edge computing to support sustainable manufacturing processes. First, the problem is addressed by recognizing certain limitations of current CPHS solutions, such as a lack of AI-powered intelligence, explainability, and human-centric collaboration. To address the challenges, the proposed solution leverages the MIMII dataset for industrial machine condition monitoring using acoustic signal processing.

The proposed framework begins by feeding the generated spectrograms to a CNN model that extracts spatial features by recognizing significant frequency patterns associated with different machine conditions. The extracted spatial features will be converted into sequences and fed into a transformer encoder module to capture temporal dependencies via self-attention. The transformed feature representation is subsequently passed to a fully connected classifier module to classify machine states as normal or abnormal. In addition, to ensure transparency and facilitate human-centric decision-making, the SHAP explainability approach is employed to explain predictions by selecting significant time-frequency regions in the model output. The designed model will be implemented in an edge-enabled CPS setting, enabling real-time inference with low computational latency and energy expenditure. In addition, a human-in-the-loop scheme will be integrated to enable human operators to validate using explainability results. Performance is validated using accuracy, precision, recall, and F1-score, as well as other system-level metrics such as latency and energy efficiency.

The proposed framework is an XAI-driven CPHS mechanism for sustainable manufacturing in the context of Industry 5.0, and its workflow is presented as a pipeline from data acquisition to intelligent decision-making and execution, as shown in Figure 1. In the initial stage of the data acquisition procedure, sound signals are collected from machines in industries such as motors, pumps, and bearings using datasets including MIMII. Data signals thus obtained are forwarded to the next stage, referred to as data preprocessing, where they undergo filtering, normalization, and segmentation. The segmented data signals are finally converted to a spectrogram using the short-time Fourier transform (STFT) at the feature representation layer.

Then, the AI model layer uses the spectrogram images to extract features via the CNN module, identifying spatial patterns and anomalies. Then, the extracted features are reshaped and fed into the transformer encoder, which learns temporal dependencies and long-range relationships via self-attention mechanisms. The resulting feature vectors are classified as normal or abnormal using the classification layer. To ensure transparency in the prediction process, the explainability layer uses SHAP to highlight significant time-frequency regions in the spectrogram.

Additionally, there is a human-in-the-loop process in which operators interpret and provide feedback on the model's outcomes. In the deployment layer, the model runs on an edge-based CPHS infrastructure, ensuring fast inference and low power consumption, while optionally interfacing with cloud systems for storage requirements or advanced analytics. Altogether, the framework ensures accurate fault detection, explainability, and efficient real-time operation in line with the requirements of Industry 5.0.

Problem definition and requirement analysis

Cyber-physical human systems (CPHS) in Industry 5.0 settings promise intelligence, sustainability, and human-centeredness. Still, contemporary frameworks for building CPHS are hindered by a lack of advanced AI-enabled capabilities, explainability, and adequate support for human-machine collaboration. Indeed, conventional CPHS are based on rule-based or uninterpretable models, which impede adaptation to changing industry conditions and limit the operators' confidence in automation. Moreover, conventional CPHS cannot effectively analyze highly sophisticated data, such as industrial acoustic signals containing both spatial (i.e., frequency) and time-dependent components, which are important for predictive maintenance. Hence, it is necessary to create a framework that can detect machine anomalies with high precision, provide explanations, and operate at the edge, thereby enabling CPS.

Here x(t) is the acoustic input signal, S(f,t) is the spectrogram representation of the signal, figure-protocol-1 is the label and θ is the parameter set of the CNN-Transformer architecture. Symbols used in subsequent equations are introduced right after their first appearance in order to preserve mathematical correctness.

From a mathematical point of view, the task can be formalized as a supervised learning problem. Let the industrial acoustic signal be denoted by x(t). The STFT of x(t) is computed according to the formula:

figure-protocol-2   (1)

where ω(.) is the window function and f is the frequency. The generated spectrogram figure-protocol-3 serves as the input to the deep learning algorithm. The goal is to train the mapping function

figure-protocol-4    (2)

where figure-protocol-5 is the class label (either normal or abnormal), and where S is the spectrogram input, y is the predicted output label and θ is the set of learnable parameters of the CNN-Transformer architecture. The optimization problem is formulated as the minimization of the classification loss:

figure-protocol-6  (3)

subject to the constraints of minimum latency figure-protocol-7, energy efficiency figure-protocol-8, and interpretability, wherein the explanation ∅(SHAP) captures the importance of features in the prediction process,

figure-protocol-9   (4)

where figure-protocol-10 is the contribution of each individual feature. In line with the foregoing, the design requirements include the following: (i) reliable anomaly detection based on hybrid deep learning, (ii) interpretable outputs for human intervention, and (iii) implementation in an edge-enhanced CPS framework.

Dataset acquisition and preprocessing

In this regard, to assess the proposed Explainable AI-based CPHS framework, the MIMII Dataset13 has been selected. This dataset is a widely used benchmark for detecting anomalies in industrial machinery using acoustic signals. This database contains sound recordings of various machines, such as slide rails, fans, pumps, and valves, under both normal and faulty operating conditions. For each machine, recordings vary with different settings and noise environments. This makes it suitable for training and evaluating models under real-time industrial conditions.

Whereas the MIMII dataset is widely used in unsupervised anomaly detection studies that rely solely on normal samples during training, it includes explicit annotations of both normal and anomalous machine states. The problem is posed as a supervised binary classification task to test the ability of the proposed Explainable AI-based CNN-Transformer architecture to separate healthy from faulty machine states. Specifically, the official MIMII labels are used, assigning class label 0 to normal samples and class label 1 to anomalous samples.

Table 1 shows the Dataset Description Details. Approximately 12,000 recordings of acoustic signals are considered in the current research, including both normal and faulty machinery operations. As per the requirements, all recordings in this study have been captured at 16 kHz and are 10 s long. Although the dataset contains approximately 12,000 recordings, a class-balanced subset of 2,000 test recordings was selected solely to visualize the confusion matrix. All reported performance metrics were computed using the complete test set (approximately 2,400 recordings). Furthermore, training and test samples are separated at an 80:20 ratio for model training. This database is well-suited for Industry 5.0 applications because it includes real-time noise conditions. The recording distribution used in this study is given in Table 2. From this dataset, approximately 12,000 recordings covering all machine types and signal-to-noise ratio (SNR) conditions were selected for the supervised binary classification experiments using an 80:20 train/test split, shows the industrial machine types, corresponding machine IDs, total numbers of normal and anomalous acoustic recordings, SNR conditions, and the experimental train/test partition employed for supervised binary classification.

The experiment was carried out using the MIMII (Malfunctioning Industrial Machine Investigation and Inspection) dataset version 1.0, which comprises four types of industrial machines: fans, pumps, valves, and slide rails. Data were collected from different machine IDs (ID00–ID06, depending on the machine type) and conditions. The MIMII database comprises data acquired at signal-to-noise ratios (SNRs) of −6 dB, 0 dB, and +6 dB. Normal data corresponds to the healthy condition of the industrial machine, while anomalous data corresponds to an anomalous condition. The classes were labeled based on the database annotation, with normal data labeled class 0 and anomalous data labeled class 1.

A fair evaluation procedure required that the split between train and test sets occur at the original audio-recording level prior to any splitting, augmentation, or spectrogram creation steps. As a result, audio segments from the same recording could not appear in both the train and test sets. For the evaluation experiment, recordings with different signal-to-noise ratios (−6 dB, 0 dB, and +6 dB) were used, which provided an opportunity to evaluate the suggested framework under various industrial conditions.

To prevent data leakage, the train/test split was performed using the original audio file as the basis, prior to any segmentation or spectrogram generation. Segments from the same 10 s audio recording were assigned exclusively to either the training or the testing set, with no segments present in both sets. Afterward, spectrograms were generated and augmented separately for each dataset. This way, the problem of data leakage is avoided, thereby preventing an overly optimistic assessment of performance.

Dataset preprocessing

The unprocessed acoustic signals from the MIMII dataset are then processed to improve signal quality and enable effective feature extraction. The input signal is denoted by x(t), where t is time. The first process is noise filtering, whereby background noises are filtered using a filtering function such that:

figure-protocol-11   (5)

where ∗ represents the convolution process and h(t) is the filter's impulse response. This process enhances the clarity of the acoustic signal and minimizes the effects of noise often prevalent in industrial environments. Signal normalization is used to ensure that the amplitudes of all audio signals are comparable before feeding them into the deep learning model. The normalized signal, denoted as xn(t), is calculated as:

figure-protocol-12   (6)

where μ and σ are the mean value and standard deviation of the input signal, respectively. All input signals are guaranteed to have zero means and unit variances thanks to this procedure. Finally, for processing, the normalized input signal is divided into multiple tiny pieces. Given that the length of the entire signal is T, it is broken down into N overlapping frames, each having a length L using a sliding window function:

figure-protocol-13   (7)

where H is the hop size. The spectrogram can be defined by the formula below:

figure-protocol-14   (8)

The spectrograms obtained in this way are fed into the CNN-transformer-based framework to extract spatial and temporal features for successful anomaly detection. Figure 2 shows the preprocessing pipeline.

Figure 2 illustrates a vertical preprocessing pipeline that may be employed for acoustic data acquired in industry. Firstly, the raw audio signals are extracted from the MIMII dataset and fed into the noise-removal stage to reduce noise. This is followed by normalizing the signal to remove amplitude inconsistencies, ensuring the signal is consistent. The signal is later split into smaller overlapping sections. Following segmentation, STFT transforms will be applied to the data to convert time-domain signals into time-frequency signals. Then, the power spectrum is calculated as part of generating the signal's spectrogram. Preprocessed spectrograms will thus be the output generated in the process and will be used in the proposed CNN-transformer model.

Spectrogram generation

After data normalization and segmentation, acoustic signals are converted into a time-frequency representation using the Short-Time Fourier Transform (STFT). The resulting spectrogram provides a pictorial representation of changes in frequency composition over the period of analysis. In industrial settings, machine faults usually manifest as frequency patterns, making spectrograms more useful for detecting machine anomalies.

The resulting processed data is therefore converted into spectrogram images to distinguish between normal and abnormal instances within the machines. The audio data will be converted into multiple spectrograms to enable the learning model to analyze local frequency components and temporal changes within the machine process. These images are used as input data for the CNN-transformer model, where the CNN handles spatial (frequency) patterns while Transformers handle temporal relationships.

Algorithm 1: Spectrogram Generation from Acoustic Signals

Input: Raw audio signal x(t)

Output: Spectrogram S(t, f)

Step 1: Load audio signal  x(t)

Step 2: Apply noise filteringfigure-protocol-15

Step 3: Normalize signal xn(t)

Step 4: Segment signal into frames:

for k = 1 to N do

figure-protocol-16

end for

Step 5: Apply a window function w(t) to each segment

Step 6: Compute STFT:

figure-protocol-17  (9)

Step 7: Compute magnitude:

figure-protocol-18   (10)

Step 8: Store spectrogram S(t, f)

Step 9: Return spectrogram dataset 

The process of generating spectrograms starts with acquiring the original audio signal, followed by preprocessing steps such as noise reduction and signal normalization to ensure signal integrity and quality. After that, the normalized audio is sliced into multiple overlapping windows to capture the temporal characteristics of both long- and short-term periods. To prevent spectral leakage, each of the slices is multiplied by the window function before being passed through the Short-Time Fourier Transform. Afterward, the transformed signal is represented as a complex-valued function in the frequency domain, and its magnitudes are squared. Thus, the energy of the frequencies over the given time is represented by the spectrograms. Finally, they are stored as image-like data and fed to the neural network for training.

Spatial feature extraction using CNN

After the process of creating spectrograms, the output time-frequency images are taken as inputs and are given to a Convolutional Neural Network (CNN). Assuming the input spectrogram as figure-protocol-19, where H is the frequency domain and W is the time domain. The purpose of using the CNN here is to learn discriminative spatial features from the input spectrogram. These discriminative spatial features will help differentiate between normal and abnormal machine behavior in an industrial environment.

A basic operation performed by the Convolutional Neural Network (CNN) is convolution. Convolution with a given filter figure-protocol-20 is given as:

figure-protocol-21   (11)

The operation described above enables the neural network to identify local spatial features, such as edges, frequencies, and harmonic content, in the spectrogram. This is done through multiple filter operations that generate various feature representations, producing several feature maps. Convolution is followed by a non-linear activation function, as shown below:

figure-protocol-22  (12)

This non-linearity enables the network to capture and represent complex relationships among variables. The next stage in a convolutional neural network involves the application of pooling operations. The common technique of max pooling is expressed mathematically as follows:

figure-protocol-23   (13)

where Ω refers to the pooling window. Pooling helps to achieve translation invariance and retain only the relevant information while discarding redundancy. From low-level variables such as frequencies to complex machine behavior, higher-level features can be extracted hierarchically using a convolutional neural network. Consequently, feature maps can be used to depict the network's ultimate output:

figure-protocol-24   (14)

where represents the learnable parameters of the CNN. The feature maps contain rich spatial information and are transformed into sequences to serve as inputs to the transformer encoder in the next phase. The above-described spatial feature extraction process is crucial for enabling the model to detect minute anomalies in industrial acoustic signals.

Figure 3 demonstrates how spatial features are extracted horizontally from the spectrogram input using a CNN model. It starts with the input spectrogram, which encodes time-frequency properties of the audio signal in a pictorial form. The input spectrogram is then passed through a convolutional layer, in which several types of filters extract spatial features such as frequency bands, edges, and textures, thereby producing feature maps. Thereafter, the extracted feature maps undergo a nonlinear activation function, rectified linear unit (ReLU), allowing the neural network to detect dependencies in the input data. The feature maps are then passed through a pooling layer, where the map size is reduced while retaining its most important characteristics. This process of convolutional, activation, and pooling layers repeats itself multiple times, creating a hierarchy of feature representation that begins with simple features and evolves toward more sophisticated ones related to machine operations.

Temporal modeling using a transformer encoder

Once the spatial features have been extracted using the CNN, the next step is to derive high-level feature maps that capture frequency-based patterns learned from the spectrogram. The resulting feature maps are then flattened and reshaped into sequences to capture temporal information in the acoustic signal. Assuming the feature maps are represented by figure-protocol-25, where H', W', and C are height, width, and channels, respectively, the following sequence is derived:

figure-protocol-26    (15)

where T represents the sequence length, and d is the feature dimension. The Transformer then uses the above sequence to model long-term dependencies using a self-attention mechanism. The transformer model, in contrast to conventional recurrent models, processes each element of the sequence in parallel and uses attention weights to determine each piece's relative relevance. The way the attention mechanism works is by using Query Q, Key K, and Value V vectors generated for each input xt according to:

figure-protocol-27   (16)

whereas WQ, WK, WV are trainable weights. The similarities among the elements are determined through:

figure-protocol-28   (17)

dk is important vectors' dimension. This procedure makes it easier for the model to learn dependencies from distant time steps in the signal and focus just on the pertinent time segments. To increase the network's capacity for learning, multihead attention is used, which refers to the process whereby several heads of attention learn from the input in parallel:

figure-protocol-29  (18)

Each head attends to distinct features of the temporal input sequence. This is followed by feed-forward neural networks and residual connections:

figure-protocol-30   (19)

The resulting encoded sequence captures both the spatial and temporal features of the input. This sequence is sent to the classifier layer to detect anomalies.

Figure 4 shows the overall process of Temporal modeling using a transformer Encoder on features extracted from the CNN. First, the feature maps extracted from the CNN are used; they contain spatial information from the spectrogram. The feature maps are flattened into a sequence, with each item corresponding to a timestamp. After that, the sequence passes through a linear transformation that maps the data into vectors of a specific dimension. As transformer models lack the ability to capture temporal dependencies, positional encoding is also added to the embedding vectors to preserve the sequence's temporal structure. Finally, the sequence passes through a transformer encoder, consisting of a stack of multiple layers, in which a multi-head self-attention mechanism is employed to enable the model to learn connections between timestamps in the sequence. The feed-forward neural network further improves these embeddings, while residual connections ensure that gradients propagate efficiently. The process is repeated across the stacked transformer encoder layers, which capture progressively more complex temporal interactions. In the end, the sequence embeddings are combined into a fixed-length representation via pooling (e.g., average pooling or the classification token (CLS)), yielding a global temporal feature representation.

Classification layer

Following temporal encoding with the transformer encoder, the outputs will be a sequence of encoded vectors, each with features related to both the spatial and temporal aspects of the input signal. The first step toward classification from the encoded outputs is to aggregate all those features into a single vector. This is achieved either by applying mean pooling to the sequence of feature vectors or by employing the so-called "classification token". Suppose the output of the Transformer is figure-protocol-31, where each figure-protocol-32. The aggregated feature h is computed as:

figure-protocol-33   (20)

Then, the aggregated feature vector h represents the input in a compact form and serves as the classifier's input. The classification part comprises one or more dense layers, which transform the learned representation into the output space. A dense layer implements a transformation that is defined as follows:

figure-protocol-34   (21)

where o is the output (logits), b is the bias vector, and W is the weight matrix. The latter are the class scores in an unnormalized form. To incorporate nonlinearity into the model, activation functions such as ReLU are often used. Finally, a softmax activation function is applied to the logits to obtain a probability for each class:

figure-protocol-35   (22)

where C represents the number of classes, (C = 2; normal and abnormal). The class prediction will be based on the maximum probability value. During training, the optimal values of the model’s parameters are obtained by minimizing the cross-entropy loss between the actual and predicted labels, thereby achieving accurate classification of machine states. In summary, the classification layer serves an essential function in transforming the extracted temporal features using the Transformer architecture into practical decisions. Through proper discrimination between healthy and faulty machine states, anomaly detection and intelligent decision-making become possible within the Industry 5.0 paradigm.

Figure 5 shows how the classification layer works, which comes after the transformer Encoder. The process starts with the encoded sequence output, in which each vector contains temporal information learned from the input. These vectors are combined into a single vector, known as the global feature vector. The next step is to pass the global feature vector through several dense neural network layers. In each dense layer, matrix multiplication operations are performed to create linear combinations of the inputs. After that, a non-linear activation function, such as ReLU, can be applied to increase the flexibility in learning decision boundaries. At the end of the last dense layer, there are logit scores for each class. These logit scores can be used on the Softmax function to calculate probability scores for each class. These probability scores are used to determine whether the input belongs to Class 0 (Normal) or Class 1 (Abnormal).

Explainability using SHAP

After the model classifies input signals, predictions are generated based on whether the machine operates normally or abnormally. Nonetheless, within the Industry 5.0 environment, it is not enough to produce predictions; the system must also provide transparency and explainability to support human decision-making. The following framework uses SHAP, a game-theory-based model that quantifies the importance of each feature in determining the prediction. In doing so, SHAP assigns weights to features, allowing one to evaluate the relative importance of different sections of the spectrogram input to the overall prediction.

The mathematical formulation of SHAP is based on a decomposition of the model f(x):

figure-protocol-36   (23)

where ∅0 is the baseline value (mean model output), ∅i is the effect of the i-th feature, and M is the number of input features. For the purpose of this study, the features are the time-frequency bins from the spectrogram. SHAP values are positive if they contribute to the probability of the abnormal state and negative otherwise. This formulation guarantees a unique, locally consistent explanation for each prediction. The SHAP method can be applied to a trained CNN-transformer model to provide an explanation based on feature importance. The visual representation will show the most important time-frequency bins in the spectrogram that affect the decision-making process. For instance, abnormal behavior of the equipment may result in certain high-energy frequency bands that the SHAP framework identifies as important contributors to the decision process.

In addition, SHAP can foster cooperation between people and machines in a cyber-physical setting. This is because the predictions can be explained using intuitive visuals, allowing the operators and decision-making systems to verify predictions, analyze the problem, and take appropriate steps. This not only increases trust but also makes AI-based manufacturing processes reliable and accountable. By integrating deep learning algorithms that meet current performance standards with SHAP, which provides explainability, the proposed architecture fulfills the requirements of Industry 5.0.

Algorithm 2: SHAP-based Explainability for Model Prediction

Input: Trained model f(x), input spectrogram S

Output: SHAP values figure-protocol-37 and explanation map

Step 1: Input spectrogram S into a trained model

Step 2: Compute prediction: y = f(S)

Step 3: Initialize a deep learning-oriented SHAP explainer

Step 4: Compute SHAP values:
figure-protocol-38

Step 5: For each feature i in S:

Calculate contribution figure-protocol-39

Step 6: Generate explanation map:

Highlight regions with high figure-protocol-40

Step 7: Visualize feature importance (heatmap)

Step 8: Return SHAP values and explanation map

The SHAP-based explainability algorithm starts by obtaining a preprocessed spectrogram image as input and feeding it into the trained CNN-transformer architecture to make predictions (e.g., normal or abnormal machine condition). After making predictions using the trained architecture, the SHAP method is used to interpret the model. In particular, the SHAP values (ϕ) are computed for each input feature to find out which time-frequency components of the input contribute most to the prediction. To calculate these values, the model outputs for inputs with and without certain features must be compared.

After calculating the SHAP values for each feature, its influence on the prediction can be evaluated by the sign and magnitude of the SHAP value. For example, features with higher SHAP values are more likely to influence the outcome, and a positive or negative sign indicates that the feature pushes the prediction toward the normal/abnormal class, respectively. Based on the results, an explanation map can be made in the form of a heatmap.

Lastly, the algorithm outputs both the SHAP values and the visualizations, providing insight into how humans can analyze the model's decision. The technique makes the process more transparent and allows humans to engage in decision-making, thus ensuring that both the algorithm's predictions and its reliability are achieved in the AI-driven industry. Representative test samples were selected from accurately diagnosed anomalous instances across various machine types to improve the reliability of the explainability analysis. While global SHAP feature-importance analysis was performed by aggregating SHAP values across the test dataset to identify consistently influential time-frequency locations, local SHAP explanations were generated for these representative cases. Without requiring expert annotation, this integrated local and global analysis offers a reliable interpretation of the model predictions.

Edge-enabled CPS Integration

This proposed solution is used within the edge-enabled CPS, which aims to achieve efficiency, sustainability, and real-time industrial operations. Once the model has been trained, the CNN-transformer will be deployed via an edge computing approach. The edge computing layer in question is near the industrial Internet of Things (IIoT) devices, including the sensors and machinery. Rather than transmitting all collected data to the cloud, the edge device performs inference on the incoming acoustic signal without first transferring it there. Therefore, the system can detect machinery abnormalities almost instantly and take the necessary steps before the situation deteriorates further. Another advantage of this edge-enabled approach is improved latency and efficiency since data is analyzed near the point of generation. In other words, unlike in the cloud approach, where inference time Tinf

figure-protocol-41   (24)

The latency, in general, will be considerably reduced through a decrease in Tcommunication. Another benefit is reduced bandwidth consumption and power expended when transmitting data continuously. This system is therefore scalable and sustainable for large industrial plants with numerous connected devices. In addition, the CPS enables interaction with the physical layer (machines), the AI (computational layer), and the human layer.

Human-in-the-loop collaboration

A key component of the suggested Industry 5.0 CPS, which includes human operators actively involved in the decision-making process alongside AI models, is human-in-the-loop collaboration. Namely, according to the designed framework, the system does not become an autonomous black box but instead allows operators to interact with it and observe explainable prediction results in the form of SHAP explanations. The predictions in question refer to the outcomes (normal/abnormal), accompanied by visualizations of parts of the spectrograms used to make them.

Such an interaction helps operators better understand exactly how the prediction is formed and which factors contribute to it. Human involvement implies verifying the system's results and allowing operators to override them when necessary, based on their professional judgment. Collaboration with human operators makes the system more trustworthy, efficient, safe, and reliable, helping prevent errors.

Results

The presented Explainable AI-integrated cyber-physical collaboration system was experimentally tested on approximately 12,000 audio recordings from the MIMII Dataset, which contains both normal and abnormal machine conditions. The training dataset consisted of about 9,600 data records, while the test dataset included roughly 2,400 audio samples. Based on the findings, the suggested CNN-transformer model achieved a classification accuracy of 98.5% on the test data set, 0.97–0.98. It relates to the findings in Table 3, the confusion matrix, and the statistical stability analysis. Integrating the transformer encoder enhances the architecture's temporal modeling capabilities, while the CNN extracts discriminative spatial features from the input spectrogram.

In addition to classification performance, other system-related characteristics were assessed. Firstly, inference latency for an edge-based implementation did not exceed 20 ms, enabling real-time anomaly detection. Secondly, the presented system helped improve energy efficiency by reducing communication requirements. Finally, the proposed Explainable AI framework provides SHAP interpretability, making the model easier to analyze from a domain expert's perspective. The MIMII data was converted to spectrograms via Short-Time Fourier Transform (STFT), using a window size of 1024, a hop size of 512, and a Fast Fourier Transform FFT size of 1024. The resulting log Mel-spectrograms were resized to 224 × 224 pixels and then z-score normalized. Data augmentation methods, including time masking, frequency masking, and Gaussian noise addition, were used to increase the robustness of the trained model. The architecture of the CNN consisted of four convolutional layers with 32, 64, 128, and 256 filters, respectively, and included batch normalization, ReLU, and max pooling layers. The extracted feature maps were transformed into sequences and used as inputs to the transformer encoder, which consisted of four encoder layers, eight attention heads, a 256-dimensional embedding, and a 1024-dimensional feed-forward dimension. A dropout probability of 0.3 was applied to minimize overfitting. The Adam optimizer with a learning rate of 0.0001 and a weight decay of 1 x 10⁻5 was used for training. The number of epochs and batch size in training were 100 and 32, respectively. The random seed was set to 42 to maintain experiment reproducibility. The following methods were used to verify the statistical validity of the results: all experiments were run independently 10 times with different random seeds and data shuffles. The average values and standard deviations of Accuracy, Precision, Recall, F1-score, area under the receiver operating characteristic curve (ROC-AUC), and area under the precision–recall curve (PR-AUC) were calculated. Moreover, 95% confidence intervals were calculated to estimate performance uncertainty. Finally, a paired t-test was done comparing the results of our CNN-transformer approach with the best-performing baseline model. The early stopping method was applied using 10 epochs as the patience parameter. Table 4 presents the Environmental Setup and Hyperparameter Configuration, and Table 5 provides the details of the Simulation Environment.

The evaluation experiment in the current work was conducted using the MIMII dataset, which includes various machine operating conditions and acoustic fault scenarios. Although the results from the proposed framework validate the model, further validation across various industry datasets and manufacturing environments is needed.

The proposed framework was implemented on an edge-computing platform equipped with a multicore processor, 16 GB system memory, and a graphics processing unit (GPU)-enabled edge accelerator to support real-time inference and explainability. Industrial acoustic sensors transmitted audio streams to the edge gateway via the IIoT communication layer. Edge nodes were used to perform tasks such as spectrogram creation, anomaly inference, and SHAP explanation generation, while the cloud server was used to store long-term data and update the model. Such a setup not only reduces the amount of communication required but also allows real-time interaction between sensing devices, AI modules, and human operators in cyber-physical systems.

The CNN-transformer approach introduced above is compared to several existing methods, including conventional ML models such as support vector machine (SVM) and random forest, CNN models without the transformer, and long short-term memory (LSTM)- based methods. While traditional ML models do not achieve satisfactory results due to a lack of capacity to process spatial and temporal correlations in data, CNN models can efficiently handle spatial features, while LSTM models can handle temporal aspects of acoustic signals. Nonetheless, the former are inferior to the latter in terms of computing cost and parallel processing capability.

To ensure a fair comparison between the proposed CNN-transformer model and the baseline models, the former was trained and evaluated using the same MIMII dataset split (80% training and 20% testing), preprocessing steps, spectrogram creation process, and evaluation measures. In the SVM method, the radial basis function (RBF) kernel was used, with C = 10 and γ = 0.01. The random forest method included 200 trees, with a maximum tree depth of 20. As for CNN, it had four convolution layers (32, 64, 128, and 256 filters) followed by fully connected layers without a Transformer module. Lastly, for the LSTM method, two stacked LSTM layers were used with 128 hidden units, and a dropout rate of 0.3.

On the contrary, the use of the CNN-transformer framework enables the best results by efficiently leveraging both the spatial and temporal capabilities of acoustic signals. Besides, the use of edge computing allows for reduced computing costs, improved latency, and energy savings. In addition, the explainability of predictions implemented using SHAP can be regarded as another advantage that sets the proposed method apart from existing techniques.

Despite recent achievements in the development of novel models for acoustic anomaly detection, such as Vision Transformer (ViT), Audio Spectrogram Transformer (AST), Conformer, and self-supervised learning-based models, an experimental comparison of these architectures was not conducted in the present study. In the future, these models are planned to be used as benchmarks for evaluating the proposed framework.

Metrics, including accuracy, precision, recall, and F1-score, are used to evaluate the suggested model's performance. While precision measures the ratio of true positives to the total number of true positives and false positives, accuracy assesses the overall correctness of predictions. The recall provides an estimate of the model's capacity to forecast anomalies in a particular system. This is crucial because failing to forecast a defect could lead to system breakdowns. The following is how these metrics are calculated mathematically:

figure-results-1   (25)

figure-results-2   (26)

figure-results-3   (27)

figure-results-4   (28)

Apart from these, additional metrics related to system performance, such as latency, energy utilization, and scalability, are considered when evaluating the real-world performance of the suggested CPS model based on edge computing.

Apart from standard performance measurement techniques, there are many other measures that may help evaluate the proposed solution from a more holistic perspective, especially given its applicability to real-world imbalanced and streaming data case studies. For example, the algorithm's capacity to distinguish between the two groups based on different thresholds is measured by the area under the receiver operating characteristic curve (ROC-AUC). Given its robustness and separability, the ROC-AUC value ought to be high. Additionally, the precision-recall Area under Curve (PR-AUC) metric is important for the anomaly detection case study since the abnormal classes are rare.

An important measure to consider is the Matthews Correlation Coefficient (MCC). It considers all four confusion matrix items and provides an accurate evaluation, even in the presence of class imbalance. This measure is calculated as follows:

figure-results-5   (29)

The values of MCC range from -1 to +1, with +1 indicating perfect prediction. The other metric that could be applied is the Specificity (True Negative Rate), which would be useful in an industrial application for assessing whether there are any false positive predictions:

figure-results-6    (30)

To make sure that there was no inconsistency between the confusion matrix and performance metrics, all the evaluation criteria were computed based on the final test dataset. The notations of true positive (TP), true negative (TN), false positive (FP), and false negative (FN) are used throughout the following equations. Accuracy, Precision, Recall, Specificity, F1-score, and MCC were calculated based on Eqs. (25)–(30). Additionally, ROC-AUC and PR-AUC were computed from the CNN-transformer model's probability outputs via independent threshold evaluation.

The key metrics to consider at the system level include latency, throughput, and computational cost. Latency is the time required to make a prediction. Throughput is the amount of data samples processed per second. Computation cost, which represents energy efficiency, is important in ensuring computational efficiency. In addition, explainability metrics, such as SHAP consistency and stability of feature importance, could also be considered when evaluating the model.

It is clear from the results in Table 3 for the MIMII Dataset that the proposed CNN-transformer model outperforms all other machine learning and deep learning methods. The performance of the SVM and random forest models is acceptable, albeit limited, with accuracies of 88.6% and 91.2%, respectively. This could be explained by these models' inability to extract complex temporal and spatial features from acoustic signals. However, LSTM outperforms SVM and random forest, achieving 94.5% accuracy by modeling time series data. However, the suggested CNN-transformer approach demonstrates the best performance, with improved accuracy (98.5%), precision (98.06%), recall (99.02%), and F1-score (98.54%). Thus, the suggested algorithm can accurately classify normal and abnormal machinery conditions. The improved specificity (97.96%) demonstrates the approach's effectiveness in reducing false positives, which is crucial for the industry. It should be noted that the performance of CNN-transformer can be explained by the combination of CNN and Transformer, where the former extracts discriminative features and the latter captures temporal relations in long sequences.

The ROC curve in Figure 6 shows the comparative performance of SVM, random forest, LSTM, and proposed CNN-transformer approaches concerning their classification abilities regarding the MIMII dataset. As shown in Figure 6, the highest Area Under Curve (AUC) was observed for the suggested CNN-transformer, which secured 0.994. It is evident that the proposed method provides better discriminative performance for distinguishing between normal and abnormal categories. On the other hand, LSTM yields slightly lower results than the suggested CNN-transformer approach; thus, its AUC is about 0.97. Both random forest and SVM have relatively low AUCs, around 0.958 and 0.913, respectively. These are attributed to both approaches' inability to learn complex acoustic patterns from the given data.

In anomaly detection, where abnormal samples are few, the precision-recall (PR) curve provides a better analysis. In the PR curve, the CNN-transformer model achieves the highest Average Precision (AP), around 0.97–0.98, suggesting superior ability to detect abnormal instances with fewer false alarms, as shown in Figure 7. Conversely, the LSTM classifier ranks second, with an AP ranging from 0.92 to 0.94. On the other hand, the random forest and SVM classifiers are ranked third and fourth, respectively, with APs of 0.88 and 0.84. The proposed hybrid classifier achieves higher accuracy across all recall levels, which is important for practical industrial applications, as they require both high precision and low false alarms.

The CNN-transformer model has been compared against existing models on the MIMII Dataset using the advanced performance metrics shown in Table 6. The ROC-AUC value indicates the model's discriminative capability between the two classes across various thresholds. The SVM and random forest classifiers achieve ROC-AUC values of 0.913 and 0.958, respectively, while the LSTM model attains a ROC-AUC of 0.970, and the CNN secured 0.98. The proposed CNN-transformer framework achieves the highest ROC-AUC of 0.994, demonstrating superior discrimination between normal and abnormal machine conditions. Similarly, the PR-AUC values indicate that the proposed model can effectively handle class imbalance, an essential factor in anomaly detection. The proposed CNN-transformer achieves the highest PR-AUC (Average Precision) of 0.982, followed by CNN (0.950), LSTM (0.912), random forest (0.891), and SVM (0.765), indicating superior precision-recall performance for industrial acoustic anomaly detection. Furthermore, the Matthews Correlation Coefficient (MCC), which measures the balance of the model's performance, is indicated by a relatively high value of 0.97 for the proposed model. However, earlier methods such as SVM (0.73) and random forest (0.78) had lower predictive accuracy.

To test the stability of the presented approach, the same experiment was conducted 10 times with varying initialization seeds and data-shuffling schemes. The results in Table 7 for the CNN-transformer architecture were as follows: accuracy = 98.50% ± 0.19%, precision = 98.06% ± 0.23%, recall = 99.02% ± 0.22%, and F1-score = 98.54% ± 0.24%. The relatively small standard deviation indicates the model's stability. The 95% confidence intervals were computed for all evaluation measures. It was determined that the confidence interval for the accuracy metric was [98.1%, 98.9%], indicating that the model consistently performs well. Narrow confidence intervals were also determined for ROC-AUC, PR-AUC, and MCC. The paired t-test comparing the proposed CNN-transformer model and the best-performing baseline model (LSTM) yielded a p value < 0.05, indicating that the observed performance gain is significant and not due to chance.

To test the robustness of the proposed CNN-transformer architecture, the experiments were repeated 10 times with different random initialization seeds and shuffling configurations, as shown in Figure 8. The average accuracy, precision, recall, and F1-score obtained by the model were 98.50% ± 0.19%, 98.06% ± 0.23%, 99.02% ± 0.22%, and 98.54% ± 0.24%, respectively. From the box plot analysis, there are minimal variations in the model's performance metrics, indicating its stability and robustness. Thus, the proposed approach exhibits high stability across multiple experiments, with no significant variation, thereby indicating the model's good generalization capabilities.

Additionally, a statistical test was performed based on the output of multiple runs of the experiment. The small standard deviations across all evaluation measures suggest little susceptibility to changes in randomness or training. This confirms that the proposed CNN-transformer model can deliver stable and consistent performance in real-life cyber-physical manufacturing systems.

The SHAP heatmap gives a sense of the importance of time-frequency regions in the spectrogram in predicting whether the machine condition is normal or abnormal. SHAP values indicate the contribution of each predictor to the final prediction.

In the SHAP value heatmap shown in Figure 9A, positive SHAP values (e.g., between +0.6 and +1.2) represent contribution toward the abnormality, whereas negative SHAP values (i.e., between −0.3 and −0.8) represent contribution to the normal category. The most intense SHAP score was found in the middle frequency range (40–80 frequency bins) at time frames 50–100. These regions are the key areas where critical fault signals are detected in acoustic measurements. Furthermore, the low-frequency band (fewer than 20 bins) shows SHAP scores of almost zero (−0.1 to +0.1), indicating a negligible impact on the model's decision-making. It indicates that the model ignores the background noise irrelevant to fault prediction. Temporal consistency of high SHAP scores within specific time bands also highlights the Transformer module's ability to capture long-range dependencies in acoustic signals. In summary, the SHAP heatmap clearly shows that the CNN-transformer model focuses on unique time-frequency bands with high discriminative power. It enhances the model's interpretability and helps the human operator visually detect faults.

It is shown that the SHAP method is used solely as an interpretability technique and is not part of the CNN-transformer model's training process. The classification accuracy is determined solely by the parameters learned by the CNN and Transformer components. SHAP is used after the model inference step to interpret the predictions by identifying the most influential time-frequency windows that drive predictions toward either normal or abnormal categories. Thus, SHAP improves human transparency and understanding without impacting classification accuracy. The SHAP values provided by the proposed framework enable a clear visualization that will help maintenance operators verify model predictions by comparing detected regions with actual machine behavior and maintenance records. Thus, it will be easier for people to collaborate with machines during predictive maintenance activities.

To evaluate the level of explainability, a global SHAP analysis of feature importance was conducted, as illustrated in Figure 9B. The results show that some mid-frequency components play a significant role in anomaly detection. Moreover, a case study illustrated in Figure 9C shows that SHAP correctly identifies fault-relevant regions in abnormal machine spectrograms. The study involved a qualitative case study utilizing abnormal data samples for pumps and valves from the MIMII dataset. In terms of faulty pumps, SHAP identified frequency ranges of 2–4 kHz associated with cavitation and wear. In the case of faulty valves, SHAP activation was dominant in areas with repeated harmonic signals indicative of leakage. Across representative test samples, the SHAP visualizations reliably identified discriminative time-frequency zones associated with anomalous machine conditions. By highlighting the spectral regions that most significantly contribute to each prediction, these explanations shed light on how the model makes decisions. The SHAP analysis is offered as an interpretability tool for understanding model behavior, rather than as expert-validated fault verification, because formal validation by maintenance professionals was outside the scope of this study.

The confusion matrix highlights how well the proposed CNN-transformer model performs classification on the MIMII dataset shown in Figure 10. Overall, 960 normal samples were correctly identified as normal (True Negatives), whereas 20 normal samples were falsely identified as abnormal (False Positives). Furthermore, out of all samples, 1010 abnormal samples were detected as such (True Positives), while 10 abnormal samples were identified as normal (False Negatives).

Thus, high detection performance was achieved, especially for abnormal samples, which are crucial in predictive maintenance. As shown in Figure 10, a False Negative error is observed very rarely (10 samples) – this is vital for any industry to ensure its safety and reliability. On the other hand, the False Positive error (20 samples) occurs slightly more often but remains at a relatively low level. The confusion matrix shown in Figure 10 was generated using a balanced subset of 2,000 test recordings (equal numbers of normal and anomalous samples) to provide a clear visualization of the classifier's performance. All reported evaluation metrics, including accuracy, precision, recall, specificity, and F1-score, were computed using the complete test set (approximately 2,400 recordings).

Taking into account the confusion matrix depicted in Figure 10 and in Table 8 (TN = 960, FP = 20, TP = 1010, FN = 10), the following calculations can be done: Accuracy = (TP + TN) / (TP + TN + FP + FN) = 98.50%, Precision = TP/(TP + FP) = 98.06%, Recall = TP/(TP + FN) = 99.02%, Specificity = TN / (TN + FP) = 97.96%, F1-score = 98.54%, and MCC = 0.97. Moreover, the ROC-AUC and PR-AUC values were 0.994 and 0.982, respectively. The confusion matrix shown in Figure 10 was generated using a balanced representative subset of 2,000 recordings from the test data for visualization purposes only. All quantitative performance metrics reported in this study were calculated using the complete test set (approximately 2,400 recordings). Overall, the results demonstrate that the proposed CNN-transformer architecture achieves 98.5% classification accuracy, in line with the results reported in Table 8 and the ablation analysis.

To verify the effectiveness of the edge-based implementation, additional measurements were conducted. The average inference latency of the proposed model is 18.7 ms/sample, and it can process approximately 53 samples per second, as shown in Table 9. Power consumption during inference was 8.9 W, yielding an estimated energy consumption of 0.167 J/prediction. To evaluate the sustainability impact of the proposed approach in Table 10, communication overhead, energy consumption, and resource usage have been considered. Since the inference process occurs locally on edge nodes, only summary diagnostic data is transmitted to the cloud environment. The experimental findings indicated a 73% decrease in communication compared with traditional cloud-based CPS systems. In addition, local inference decreased energy consumption by 32%, while communication time decreased from 75 ms to 20 ms.

Ablation study

The MIMII dataset is subjected to ablation research to evaluate the impact of each component of the explainable AI-based CPHS architecture. This ablation study investigates the effects of using a CNN to learn spatial features, a transformer encoder to learn temporal features, and the SHAP explainability algorithm itself. Several modifications of the proposed model are considered, namely: (i) a CNN-only version, (ii) a Transformer-only version, (iii) the model with CNN and Transformer combined but no SHAP algorithm, and (iv) the whole proposed model with CNN and Transformer combined with SHAP. For evaluation, common metrics such as accuracy, precision, recall, and F1-score are employed.

These ablation study results demonstrate the significance of each element in the framework outlined in Table 11. The CNN-only model achieves 93.8% accuracy, indicating that spatial feature extraction from spectrograms is effective, yet it cannot capture temporal dependencies in acoustic signals. The same applies to the Transformer-only model, which achieves a slightly lower accuracy (92.6%) because it extracts temporal features while neglecting spatial information. Finally, the proposed CNN-transformer model obtained an accuracy of 98.5%, precision of 98.06%, recall of 99.02%, and F1 score of 98.54%, in line with the results of the main experiment. Since SHAP is used only for post hoc interpretation and not for model training, the classification accuracy is the same as that of the trained CNN-transformer model. SHAP is then used to create feature attribution explanations. The improved recall (99.02%) indicates that the model is effective at detecting abnormal conditions in the machines, ensuring no faulty parts are missed, which is imperative for implementing a predictive maintenance system.

DATA AVAILABILITY:

The Malfunctioning Industrial Machine Investigation and Inspection (MIMII) dataset used in this study is publicly available from the official Zenodo repository. The experiments were conducted using the publicly available audio recordings and annotations provided by the dataset authors. The dataset can be accessed at https://zenodo.org/records/3384388. The implementation code accompanying this study is publicly available through the Zenodo repository at https://zenodo.org/records/21405292. The repository includes the implementation code for the proposed framework, including the data preprocessing pipeline, model architecture, training workflow, evaluation scripts, configuration files, and supporting utilities to facilitate understanding and independent reproduction of the proposed methodology. The MIMII dataset is not redistributed with the implementation code and should be obtained separately from its official repository.

figure-results-7
Figure 1: Architecture of proposed work. Architectural design of explainable AI-driven cyber-physical human system combining preprocessing, CNN-based spatial learning, transformer-based temporal modeling, SHAP explainability, and edge computing for sustainable manufacturing. Please click here to view a larger version of this figure.

figure-results-8
Figure 2: Pipeline of preprocessing. Acoustic signal preprocessing pipeline consisting of noise removal, normalization, segmentation, STFT transformation, and spectrogram extraction. Please click here to view a larger version of this figure.

figure-results-9
Figure 3: Spatial feature extraction using CNN. Architecture for spatial feature extraction by CNNs using spectrograms through the process of convolution, activation, and pooling. Please click here to view a larger version of this figure.

figure-results-10
Figure 4: Temporal modeling using a transformer encoder. Transformer encoder architecture to capture temporal dependencies and long-range dependencies through spectrogram sequence representation. Please click here to view a larger version of this figure.

figure-results-11
Figure 5: Classification layer. A classifier layer that maps the encoded temporal features to probability values for predicting the machine condition. Please click here to view a larger version of this figure.

figure-results-12
Figure 6: Result of the ROC curve. Comparison of ROC curves for SVM, random forest, LSTM, CNN, and proposed CNN-transformer algorithms on the MIMII dataset. Please click here to view a larger version of this figure.

figure-results-13
Figure 7: Result of the precision-recall curve. Precision-recall curve comparison representing the anomaly detection efficiency of various machine learning and deep learning methods. Please click here to view a larger version of this figure.

figure-results-14
Figure 8: Multiple-run performance stability analysis. Box plot showing the stability in performance of Accuracy, Precision, Recall, and F1-score results obtained by ten independent experiments conducted with the proposed CNN-transformer architecture. The small variations indicate stability and robustness of the model. Please click here to view a larger version of this figure.

figure-results-15
Figure 9: SHAP-based explainability analysis of the proposed anomaly detection framework. (A) SHAP-based explainability analysis. Heatmap using the SHAP method to show the most relevant time-frequency zones that lead to anomalous behavior identification. (B) Feature importance analysis using the global SHAP approach for the identification of the most influential features. (C) Visualization of SHAP explainability for an abnormal machine. Please click here to view a larger version of this figure.

figure-results-16
Figure 10: Confusion matrix for MIMII dataset. The confusion matrix represents the classification performance of the proposed CNN-transformer algorithm. Please click here to view a larger version of this figure.

ParameterDescription
Dataset NameMIMII (Malfunctioning Industrial Machine Investigation and Inspection) Dataset
Machine TypesValves, Pumps, Fans, Slide Rails
Total Samples UsedApproximately 12,000 audio recordings
ClassesNormal, Abnormal
Sampling Rate16 kHz
Duration per Sample10 seconds
Audio FormatWAV
Feature RepresentationSTFT-based Spectrogram Images
Preprocessing StepsNoise Filtering, Normalization, Segmentation, STFT Transformation
Label AssignmentNormal recordings labeled as Class 0; abnormal recordings labeled as Class 1 according to MIMII annotations
Train-Test Split80% Training (≈9,600 samples), 20% Testing (≈2,400 samples)
Data Partition StrategyAudio-file-level split performed before segmentation to prevent data leakage
Industrial Noise ConditionsRealistic industrial acoustic environments included in MIMII recordings
Application DomainPredictive Maintenance and Machine Anomaly Detection
Target Industry 5.0 ObjectiveExplainable, Human-Centric, and Sustainable Fault Detection

Table 1: Dataset description. Description of the MIMII dataset with respect to machine type, sampling distribution, signal characteristics, feature representation, and application aspects.

Machine TypeMachine IDsNormal RecordingsAnomalous RecordingsSNR Conditions (dB)Experimental Split
FanID00–ID068,4001,600−12 to −980% Training / 20% Testing
GearboxID00–ID067,1241,147−17 to −1580% Training / 20% Testing
PumpID00–ID067,2001,800−18 to −980% Training / 20% Testing
Slide RailID00–ID057,0001,800−14 to −1280% Training / 20% Testing
ValveID00–ID057,0001,800−12 to −980% Training / 20% Testing

Table 2: Description of the MIMII dataset used in this study. Industrial machine types, machine IDs, numbers of normal and anomalous recordings, SNR conditions, and the train/test partition used for supervised binary classification.

ModelAccuracy (%)Precision (%)Recall (%)F1-score (%)Specificity (%)
SVM88.686.985.486.190.2
Random Forest (RF)91.289.888.589.192.6
LSTM94.593.292.89395.1
CNN96.395.494.995.196.8
CNN–Transformer (Proposed)98.598.0699.0298.5497.96

Table 3: Classification result for MIMII dataset. Comparison of machine learning and deep learning classifiers' performance using accuracy, precision, recall, F1 score, and specificity.

ParameterValue
DatasetMIMII Dataset
Total Samples~12,000
Train-Test Split80% / 20%
Input TypeSpectrogram Images
CNN Layers3–5 Convolution Layers
Transformer Layers2–4 Encoder Layers
Batch Size32
Learning Rate0.001
OptimizerAdam
Epochs100
HardwareEdge Device + GPU (for training)
STFT Window Length1024
Hop Size512
FFT Size1024
Spectrogram Resolution224 × 224
NormalizationZ-score
Data AugmentationTime Masking, Frequency Masking, Gaussian Noise
CNN Layers4
CNN Filters32, 64, 128, 256
Transformer Layers4
Attention Heads8
Embedding Dimension256
Feed-Forward Dimension1024
Dropout0.3
Random Seed42
Early Stopping Patience10

Table 4: Environmental setup and hyperparameter configuration. Parameters of training and experiments used to implement the proposed CNN-transformer architecture.

ComponentSpecification
Programming LanguagePython
FrameworkTensorFlow / PyTorch
ProcessorIntel i7 / Ryzen 7
GPUNVIDIA GTX 1660 / RTX Series
Edge DeviceRaspberry Pi / NVIDIA Jetson
VisualizationMatplotlib, SHAP
Edge DeviceNVIDIA Jetson Xavier NX
CPU6-core ARM v8.2
RAM16 GB
Operating SystemUbuntu 20.04
Deep Learning FrameworkTensorFlow/PyTorch
Dataset Size12,000 MIMII samples
ConnectivityEthernet/Wi-Fi
Model Parameters8.4 Million
Model Size32.7 MB
FrameworkPyTorch + TensorRT
Batch Size1
QuantizationFP16
PruningNo
Average Latency17.8 ms
Median Latency17.2 ms
95th Percentile Latency19.6 ms
Throughput56 samples/s
Memory Usage1.4 GB
Power Consumption11.3 W

Table 5: Simulation environment. Hardware and software were used to train, deploy, and evaluate the proposed model.

ModelROC-AUCPR-AUCMCC
SVM0.9130.7650.73
Random Forest (RF)0.9580.8910.78
LSTM0.970.9120.86
CNN0.980.950.9
CNN–Transformer (Proposed)0.9940.9820.97

Table 6: Comparison of experimental results for ROC-AUC, PR-AUC, and MCC. Comparison between various models based on ROC-AUC, PR-AUC, and Matthews Correlation Coefficient measures.

MetricMean (%) + Std(%)95% Confidence
Accuracy98.50 ± 0.19[98.1, 98.9]
Precision98.06 ± 0.23[98.0, 98.2]
Recall99.22 ± 0.22[98.8, 99.4]
F1-score98.54 ± 0.24[98.1, 98.9]
ROC-AUC0.9840.004
PR-AUC0.9820.005
MCC0.970.007

Table 7: Result of multiple-run statistical analysis. Robustness statistics for the proposed CNN-transformer architecture across 10 experiments, including the mean, standard deviation, 95% confidence interval, and significance for various performance measures.

MetricValue
Accuracy98.50%
Precision98.06%
Recall99.02%
Specificity97.96%
F1-score98.54%
MCC0.97
ROC-AUC0.994
PR-AUC0.982

Table 8: Final test-set of confusion matrix. The test set confusion matrix for the proposed architecture, showing the classification of machines' normal and abnormal states from which the evaluation metrics were derived.

MetricCloud DeploymentEdge Deployment
Inference Latency (ms)85.618.7
Throughput (samples/s)2253
Network Usage (MB/min)12018
Energy per Prediction (J)0.520.17
Real-Time CapabilityModerateHigh

Table 9: Result of inference latency. The latency performance of the edge deployment of the Explainable AI-enabled CPHS framework, including processing time, throughput, memory consumption, and other real-time characteristics.

MetricCloud-based CPSProposed Edge-enabled CPHSImprovement
Data Transmission per Hour100%35%65% Reduction
Communication Latency75 ms20 ms73.3% Reduction
Energy Consumption100%68%32% Reduction
Estimated Carbon Emission100%70%30% Reduction

Table 10: Sustainability assessment of edge-enabled CPHS. Measurements of communication overhead, energy consumption, resource utilization, and latency are used to assess the sustainability of the edge-enabled CPHS.

Model VariantAccuracy (%)Precision (%)Recall (%)F1-Score (%)
CNN only93.892.594.293.3
Transformer only92.691.293.592.3
CNN + Transformer 98.598.0699.0298.54

Table 11: Result of ablation study. Result of ablation study highlighting the effect of CNN, Transformer, and SHAP modules on the performance of the proposed model.

Discussion

The experimental results demonstrate that the proposed CNN-transformer architecture outperforms individual algorithms by effectively integrating spatial and temporal feature learning. According to the ablation study, although the CNN extracts frequency-based discriminative features from spectrograms, the transformer encoder improves model performance by capturing long-range temporal dependencies in acoustic signals. Moreover, the confusion matrix illustrates very few false negatives, implying that the proposed model can effectively identify machine anomalies. Such architecture has led to improved classification results, with an accuracy of 98.5% and a recall rate of 99.02%. This result is crucial in industrial fault detection systems, as the model should not miss any faulty machines. Experimental results have shown that the proposed XAI-enabled CNN–Transformer framework successfully combines high anomaly detection accuracy, interpretability, and real-time edge implementation. The combination of the two modules enables simultaneous extraction of spatial and temporal features from acoustic data, while SHAP explanations enhance explainability and reliability. These results align with the principles of Industry 5.0, which emphasize human-centered decision-making, sustainability, and synergy between the cyber and physical domains. However, a more thorough evaluation with other industrial datasets and practical deployment is required to estimate the generalizability of the proposed approach.

The implementation of SHAP for explainability not only improves the system's efficiency but also significantly enhances understanding of the model's underlying logic. It focuses especially on the time-frequency bands that are important for these choices. This is crucial in Industry 5.0 scenarios where people and AI systems need to trust each other. It is evident from the SHAP heatmaps that the model has recognized certain frequency bands associated with faults, thereby ensuring alignment with reality. The novelty of this study is not merely based on the use of CNN-transformer or SHAP alone, both of which have been mentioned in the existing scientific literature. Rather, it is due to their integrated application within the structure of Industry 5.0 cyber-physical collaboration that they simultaneously solve explainability, sustainability, edge intelligence, and human-centric decision support problems. Although the suggested framework draws on ideas from Industry 5.0, such as human-centric collaboration, explainability, resilience, and sustainable edge computing, the current research evaluates its technical performance solely using anomaly detection accuracy, latency, energy consumption, and interpretability. The direct evaluation of human-centric aspects of operator trust, cognitive load, decision quality, and acceptance was out of the scope of the current research design. Future research will include human-in-the-loop experiments to measure the effectiveness of collaborative decision-making in the proposed framework. From a theoretical standpoint, this research makes a valuable contribution to the expansion of hybrid deep learning models through showing the effectiveness of combining convolutional neural networks and Transformer structures for processing industrial spatiotemporal data. Indeed, the paper enhances the CPS and CPHS approaches by implementing the explainable AI concept throughout the decision-making process. Moreover, the implementation of the SHAP technique as a tool for measuring the contributions of various time-frequency features gives evidence to the theoretical basis of explainable AI-based systems. Overall, such an approach paves the way for a new paradigm in which deep learning models not only become more accurate but also can be interpreted by humans, consistent with the idea of Industry 5.0.

In practical terms, the proposed system is an excellent solution for real-time industrial anomaly detection and predictive maintenance. Using CPHS architectures at the edge, the system enables timely processing of acoustic signals and is well-suited for implementation in smart manufacturing facilities. The high recall rate (99.02%) reduces the probability of missed faults, thereby increasing operational safety and preventing unexpected machine breakdowns. Moreover, the use of SHAP-based explanations allows humans to understand model results, which is especially important in critical situations where human involvement in the loop is necessary. The findings from the experimental deployment clearly demonstrate the benefits of edge computing over traditional cloud computing. This is because inference done at the edge of the network minimizes latency times and improves efficiency through higher throughput. In addition, reducing energy consumption per prediction will help achieve sustainable manufacturing goals in line with Industry 5.0 standards. Despite these encouraging results, however, the outlined approach has several limitations. First, the framework is mainly validated on the MIMII dataset; although it is fairly extensive, it might fail to cover all possible scenarios that may be encountered in an actual industrial setting and across different kinds of machines. The second limitation is that using SHAP for explainability increases computational complexity, which may be an issue when deploying the framework in highly constrained edge settings. Lastly, the existing model is limited to acoustic signals alone, whereas industrial systems require multimodal processing, including, for example, vibration and temperature.

Even though the use of the MIMII dataset ensures real-world fault scenarios in industrial acoustics, validation on other benchmark datasets, such as ToyADMOS, DCASE industrial anomaly detection data, IMS bearing data, and CWRU bearing fault data, can further improve the generalizability of the proposed framework. In future work, we plan to conduct cross-domain validations in machine types, sensing methods, and manufacturing settings to ensure the generalization of the Explainable AI-based cyber-physical collaboration framework. The other limitation of the current research is the lack of comparative experiments on recently developed acoustic anomaly detection systems, such as Vision Transformers, Audio Spectrogram Transformers, Conformers, and novel self-supervised learning techniques. Although the CNN-transformer model has demonstrated strong results and interpretability, further research will be conducted to benchmark it against these advanced technologies. Another drawback of the current research is that the evaluation of Industry 5.0 features, such as resilience and sustainability, was conducted indirectly via system-level metrics, including inference latency and energy consumption. A more thorough assessment of resilience to communication and sensor failures, and of sustainability in terms of carbon footprint and resource utilization, was not performed in the current research. These factors should be evaluated in future studies.

However, it should be noted that the human-in-the-loop collaboration paradigm described in this research is theoretical and is intended to illustrate the use of explainable AI outputs for decision-making in Industry 5.0 cyber-physical systems assisted by the operator. The experimental validation with human operators the analysis of usability, response time, and expert evaluation fall outside the scope of this research. In this work, a new Explainable AI-driven cyber-physical collaboration model is presented for sustainable manufacturing in Industry 5.0 settings. In the developed system, a hybrid CNN-transformer architecture is used to efficiently learn the spatial and temporal properties of acoustic spectrograms, thereby enabling the detection of anomalies in industrial machines. The proposed model operates in four steps: preprocessing acoustic data into spectrograms, learning spatial properties with a CNN, learning temporal dependencies with a transformer encoder, and classification. Results obtained with the proposed method on the MIMII dataset demonstrate that the suggested model outperformed the baseline models, achieving 98.5% accuracy and 99.02% recall rates. Moreover, incorporating explainability enhances the system's clarity. Although the proposed framework for Explainable AI-assisted cyber-physical collaboration performed well on the MIMII benchmark dataset, the research is currently constrained to experimental validation. The framework still needs to be evaluated through large-scale production in an industrial environment where machines are heterogeneous, conditions vary, and the deployment is long-term. Additionally, comparisons with architectures such as ViT, AST, Conformer, and self-supervised learning models should be included in future work. In other words, while the framework demonstrates that integrating explainable AI, cyber-physical collaboration, and human-in-the-loop collaboration is feasible, further validation is needed before deploying it at scale in industry. Future work will further investigate the proposed framework across various fault-diagnosis benchmarking datasets for industry and cross-domain manufacturing, demonstrating its robustness, transferability, and applicability within sustainable manufacturing ecosystems. Future validation will cover operator-centered experiments, resilience benchmarking under disruptive industrial conditions, and sustainability-related life-cycle evaluation.

For reproducibility, this work provides details on data preprocessing, the splitting strategy, spectrogram generation settings, the CNN-transformer model architecture, the optimizer, hyperparameters, and the metrics used to measure the model’s performance. The random seeds were set before experimentation. To ensure replicability, the experiments were conducted on the publicly available MIMII (Malfunctioning Industrial Machine Investigation and Inspection) Dataset. The link to the dataset is as follows: DOI: 10.5281/zenodo.3384388. The implementation code accompanying this study has been made publicly available through the Zenodo repository: https://zenodo.org/records/21405292. The audio signals were sampled at 16 kHz and converted into log-Mel spectrograms using the STFT with a window size of 1024, a hop length of 512, and an FFT size of 1024. The proposed CNN-transformer architecture had four convolutional layers (with 32, 64, 128, and 256 filters), followed by a transformer encoder with four layers, eight attention heads, an embedding dimension of 256, and a feed-forward dimension of 1024. The training process involved Adam optimization with a learning rate of 0.0001 and a batch size of 32, and 100 epochs. Algorithms 1 and 2 describe the data preprocessing stage and the explainability procedure. The full algorithm, including implementation details, is presented.

Disclosures

An AI tool (ChatGPT) was used only to assist with editing the language and improving the grammar, clarity, and readability. AI tools were not used to generate scientific data, perform experiments, analyze results, interpret findings, or draw scientific conclusions. All scientific content, data analysis, interpretation, and the final version of the manuscript were independently reviewed, verified, and approved by the authors, who take full responsibility for the accuracy and integrity of the work. The authors have no conflicts of interest to disclose.

Acknowledgements

FUNDING: This work was supported by the Key Project on Higher Education Teaching Reform in Shaanxi Province (Project No. 23BG053); the Scientific Research Foundation for High-Level Talents of Shaanxi Polytechnic University (Project No. BSJ-2023-08); and the Scientific Research Project of Shaanxi Polytechnic University (Project No. 2024YKYB-006). The authors sincerely thank Shaanxi Polytechnic University, Xianyang, China, and Xi’an Traffic Engineering University, Xi’an, China, for providing the facilities and institutional support that enabled this research.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Social Support ScalePublished academic scale; originally developed by Park and revised/supplemented by Kim25-item version; 5-point Likert response format; no commercial catalog numberMeasures perceived social support among pre-service early childhood teachers across Emotional Support, Informational Support, Material Support, and Evaluative Support. Overall Cronbach’s alpha = 0.96.
Teacher Self-Efficacy ScalePublished academic scale; originally developed by Enochs and Riggs; Chinese-adapted study version25-item version; 5-point Likert response format; no commercial catalog numberAssesses teacher efficacy in the Chinese pre-service early childhood education context across General Teacher Self-Efficacy and Personal Teacher Self-Efficacy. Overall Cronbach’s alpha = 0.96.
Teacher-Child Interaction ScalePublished academic scale; developed by Lee and reviewed by the study expert panel for Chinese use30-item version; 5-point Likert response format; no commercial catalog numberAssesses self-reported teacher-child interaction competence rather than observed classroom interaction quality. Includes Emotional Interaction, Verbal Interaction, and Behavioral Interaction; overall Cronbach’s alpha = 0.94.
Demographic Information FormResearch team, Sichuan University of Arts and Science / Daegu University collaborationCustom study questionnaire section; no commercial catalog numberCollected participant characteristics including gender, educational background, professional category, internship location, internship duration, internship type, and kindergarten type.
Participant Information and Informed Consent ScriptResearch team; approved by Daegu University Research Ethics CommitteeEthics approval number: XXX; local approval record to verifyUsed before questionnaire distribution to explain study purpose and survey content and obtain informed consent from pre-service early childhood education teacher participants.
Pilot Survey Questionnaire PackageResearch team and expert panelPilot version used June 1–5, 2024; no commercial catalog numberPilot-tested the clarity and appropriateness of the Professional Perceptions Scale, Social Support Scale, Teacher Self-Efficacy Scale, Teacher-Child Interaction Scale, and demographic items in 50 pre-service early childhood education teachers.
Expert Panel Review ProcedureResearch team; expert panel of two early childhood education professors, one preschool director, three internship supervisors, and the research teamConsensus-based content review; no commercial model numberUsed to revise scale items for ambiguity, appropriateness, semantic clarity, developmental suitability, and cultural relevance before the formal survey.
Printed On-Site Questionnaire FormsResearch team; local institutional survey suppliesPaper survey forms; no commercial catalog numberUsed for on-site questionnaire administration. In the final sample, 336 valid on-site questionnaires were obtained before combining with online questionnaires.
WeChatTencentWeChat mobile application; version should be verified from the device/app record used during June–July 2024Used as an online distribution channel for questionnaire administration; 27 completed online questionnaires were collected through the online/WeChat route.
Telephone and In-Person Recruitment ProcedureResearch teamStandardized recruitment procedure; no commercial catalog numberUsed to contact institutions/participants, explain the study purpose and survey content, and support informed consent before questionnaire administration.
SPSS StatisticsIBMVersion 27.0Used for descriptive statistics, Pearson correlation coefficients, Cronbach’s alpha reliability analysis, normality checks using skewness and kurtosis, and Harman’s single-factor test for common method bias.
SPSS AmosIBMVersion 26.0Used for confirmatory factor analysis, structural equation modeling, multi-group measurement invariance testing, bootstrap testing of indirect effects, and alternative model comparisons.
Microsoft ExcelMicrosoftMicrosoft 365 Excel or locally installed Excel version; verify exact local build before submissionRecommended/used as the workbook environment for organizing coded survey data, data dictionary, supplementary tables, and the JoVE Table of Materials. Exact local version should be verified by the authors.
Structural Equation Modeling Fit-Index CriteriaPublished methodological guidance; Schreiber et al. and Kline references cited in the manuscriptNo commercial catalog number; criteria applied in the statistical analysis planDefines and reports model-fit indices including chi-square statistic, chi-square divided by degrees of freedom, Tucker-Lewis Index, Comparative Fit Index, Adjusted Goodness-of-Fit Index, Incremental Fit Index, Root Mean Square Error of Approximation, and Standardized Root Mean Square Residual.
Bootstrap Mediation Testing ProcedureSPSS Amos / research team statistical analysis plan5,000 bootstrap samples; 95% bias-corrected confidence intervalsUsed to test indirect associations through teacher efficacy. Indirect effects were considered significant when the 95% bias-corrected confidence interval did not include zero.
Measurement Invariance Testing ProcedureSPSS Amos / research team statistical analysis planMulti-group confirmatory factor analysis; ΔCFI < .010 and ΔRMSEA < .015 criteriaUsed to evaluate configural, metric, and scalar invariance across gender, program type, internship duration, and institution before pooled SEM estimates were interpreted.

References

  1. Fraga-Lamas P, Barros D, Lopes SI, Fernández-Caramés TM. Mist and edge computing cyber-physical human-centered systems for Industry 5.0: A cost-effective IoT thermal imaging safety system. Sensors (Basel). 2022;22(21):8500.
  2. Breque M, De Nul L, Petridis A. Industry 5.0: Towards a sustainable, human-centric and resilient European industry. Brussels: European Commission; 2021:https://research-and-innovation.ec.europa.eu/knowledge-publications-tools-and-data/publications/all-publications/industry-50-towards-sustainable-human-centric-and-resilient-european-industry_en.
  3. Sariişik G, Demir S. Industry 5.0: A human-centric paradigm for sustainable and resilient industrial transformation. Journal of Social Perspective Studies. 2025;2(2):50–66.
  4. Liu X et al. Evaluating the interactions of multi-dimensional value for sustainable product-service systems with a grey DEMATEL-ANP approach. J Manuf Syst. 2021;60:449–458.
  5. Di Nardo M et al. The maintenance in Industry 4.0: Assistance and implementation [conference paper]. Presented at: 32nd European Safety and Reliability Conference; Dublin, Ireland; 2022. p. 2286–2292. https://hdl.handle.net/20.500.12607/48222.
  6. Kagermann H, Lukas WD, Wahlster W. Industrie 4.0: Mit dem Internet der Dinge auf dem Weg zur 4. industriellen Revolution. VDI Nachrichten. 2011;(13):2–3. https://www.dfki.de/fileadmin/user_upload/DFKI/Medien/News_Media/Presse/Presse-Highlights/vdinach2011a13-ind4.0-Internet-Dinge.pdf.
  7. Council for Science, Technology and Innovation. Report on the Fifth Science and Technology Basic Plan. Tokyo: Cabinet Office, Government of Japan; 2015:https://www8.cao.go.jp/cstp/kihonkeikaku/5basicplan_en.pdf.
  8. Nahavandi S. Industry 5.0—A human-centric solution. Sustainability (Basel). 2019;11(16):4371.
  9. Paschek D, Mocan A, Draghici A. Industry 5.0—The expected impact of the next industrial revolution [conference paper]. Presented at: MakeLearn and TIIM International Conference; Piran, Slovenia; 2019. p. 125–132. https://toknowpress.net/ISBN/978-961-6914-25-3/papers/ML19-017.pdf.
  10. Maddikunta PKR et al. Industry 5.0: A survey on enabling technologies and potential applications. J Ind Inf Integr. 2022;26:100257.
  11. Longo F, Padovano A, Umbrello S. Value-oriented and ethical technology engineering in Industry 5.0: A human-centric perspective for the design of the factory of the future. Appl Sci (Basel). 2020;10(12):4182.
  12. Gong Y, Chung YA, Glass J. AST: Audio Spectrogram Transformer [conference paper]. Presented at: Interspeech 2021; Brno, Czech Republic; 2021. p. 571–575. https://www.isca-archive.org/interspeech_2021/gong21b_interspeech.html.
  13. Purohit H et al. MIMII dataset: Sound dataset for malfunctioning industrial machine investigation and inspection [conference paper]. Presented at: 4th Workshop on Detection and Classification of Acoustic Scenes and Events; New York, NY; 2019. p. 209–213. https://dcase.community/documents/workshop2019/proceedings/DCASE2019Workshop_Purohit_21.pdf.
  14. Hallioui A et al. A review of sustainable total productive maintenance. Sustainability (Basel). 2023;15(16):12362.
  15. Vasić S et al. Identification of criteria for enabling the adoption of sustainable maintenance practice: An umbrella review. Sustainability (Basel). 2024;16(2):767.
  16. Barredo Arrieta A et al. Explainable artificial intelligence: Concepts, taxonomies, opportunities, and challenges toward responsible AI. Inf Fusion. 2020;58:82–115.
  17. Lundberg SM, Lee SI. A unified approach to interpreting model predictions [conference paper]. Presented at: 31st International Conference on Neural Information Processing Systems; Long Beach, CA; 2017. p. 4765–4774. https://papers.nips.cc/paper/7062-a-unified-approach-to-interpreting-model-predictions.
  18. Du M, Liu N, Hu X. Techniques for interpretable machine learning. Commun ACM. 2020;63(1):68–77.
  19. Sarker IH. Machine learning: Algorithms, real-world applications and research directions. SN Comput Sci. 2021;2(3):160.
  20. Saxe A, Nelli S, Summerfield C. If deep learning is the answer, what is the question? Nat Rev Neurosci. 2021;22(1):55–67.
  21. Piccialli F et al. A survey on deep learning in medicine: Why, how and when? Inf Fusion. 2021;66:111–137.
  22. Li Z et al. A survey of convolutional neural networks: Analysis, applications, and prospects. IEEE Trans Neural Netw Learn Syst. 2022;33(12):6999–7019.
  23. Gil M, Albert M, Fons J, Pelechano V. Engineering human-in-the-loop interactions in cyber-physical systems. Inf Softw Technol. 2020;126:106349.
  24. Krugh M, Mears L. A complementary cyber-human systems framework for Industry 4.0 cyber-physical systems. Manuf Lett. 2018;15:89–92.
  25. Jasiulewicz-Kaczmarek M et al. Recent advances in smart and sustainable maintenance [invited-session proposal]. Presented at: 10th IFAC Conference on Manufacturing Modelling, Management and Control; Nantes, France; 2022. https://hub.imt-atlantique.fr/mim2022/wp-content/uploads/2021/12/7a519.pdf.
  26. Jasiulewicz-Kaczmarek M, Gola A. Maintenance 4.0 technologies for sustainable manufacturing—An overview. IFAC-PapersOnLine. 2019;52(10):91–96.
  27. Saihi A, Ben-Daya M, As’ad RA. Maintenance and sustainability: A systematic review of modeling-based literature. J Qual Maint Eng. 2023;29(1):155–187.
  28. Bastas A. Sustainable manufacturing technologies: A systematic review of the latest trends and themes. Sustainability (Basel). 2021;13(8):4271.
  29. Franciosi C, Iung B, Miranda S, Riemma S. Maintenance for sustainability in the Industry 4.0 context: A scoping literature review. IFAC-PapersOnLine. 2018;51(11):903–908.
  30. Franciosi C et al. Measuring maintenance impacts on the sustainability of manufacturing industries: From a systematic literature review to a framework proposal. J Clean Prod. 2020;260:121065.
  31. Vrignat P, Kratz F, Avila M. Sustainable manufacturing, maintenance policies, prognostics and health management: A literature review. Reliab Eng Syst Saf. 2022;218:108140.
  32. Durán O, Aguilar J, Capaldo A, Arata A. Fleet resilience: Evaluating maintenance strategies in critical equipment. Appl Sci (Basel). 2021;11(1):38.
  33. Ciccarelli M, Papetti A, Germani M. Exploring how new industrial paradigms affect the workforce: A literature review of Operator 4.0. J Manuf Syst. 2023;70:464–483.
  34. Madzik P et al. Human-centricity in Industry 5.0—Revealing hidden research topics by unsupervised topic modeling using latent Dirichlet allocation. Eur J Innov Manag. 2025;28(1):113–138.
  35. Farivar F, Klarin A, Kanani-Moghadam V. Industry 5.0: A socio-technical system perspective on human agency and institutional legitimacy. J Manag Organ. 2026;32(4):1168–1189.
  36. Karki BR, Porras J. Digitalization for sustainable maintenance services: A systematic literature review. Digit Bus. 2021;1(2):100011.
  37. Santiago RAdF et al. Data-driven models applied to predictive and prescriptive maintenance of wind turbines: A systematic review of approaches based on failure detection, diagnosis, and prognosis. Energies (Basel). 2024;17(5):1010.
  38. Carvalho JN, Silva FR, Nascimento EGS. Challenges of the biopharmaceutical industry in the application of prescriptive maintenance in the Industry 4.0 context: A comprehensive literature review. Sensors (Basel). 2024;24(22):7163.
  39. Santos ACdJ, Cavalcante CAV, Wu S. Maintenance policies and models: A bibliometric and literature review of strategies for reuse and remanufacturing. Reliab Eng Syst Saf. 2023;231:108983.
  40. Orošnjak M, Brkljač N, Ristić K. Fostering cleaner production through the adoption of sustainable maintenance: An umbrella review with a questionnaire-based survey analysis. Clean Prod Lett. 2025;8:100095.
  41. Ji Z, Zhou Y, Wang B, Zang J. Human–cyber–physical systems in the context of new-generation intelligent manufacturing. Engineering (Beijing). 2019;5(4):624–636.
  42. Wang B et al. Toward human-centric smart manufacturing: A human–cyber–physical systems perspective. J Manuf Syst. 2022;63:471–490.
  43. Jiao J, Zhou F, Gebraeel NZ, Duffy V. Towards augmenting cyber-physical-human collaborative cognition for human–automation interaction in complex manufacturing and operational environments. Int J Prod Res. 2020;58(16):5089–5111.
  44. Yilma BA, Panetto H, Naudet Y. Systemic formalisation of cyber-physical-social systems: A systematic literature review. Comput Ind. 2021;129:103458.
  45. Zhou Y, Yu FR, Chen J, Kuo YH. Cyber-physical-social systems: A state-of-the-art survey, challenges and opportunities. IEEE Commun Surv Tutor. 2020;22(1):3

Reprints and Permissions

Tags

Deep LearningMachine Fault DetectionSpectrogram ImagesConvolutional Neural NetworkTransformer EncoderHuman Machine Collaboration