Research Article

Research on Optimization and Dynamic Environment Adaptation of Intelligent Robot Inspection System Based on Multi-sensor Fusion

0 views

⸱

DOI:

10.3791/70965

⸱

September 25th, 2026

 , 

Corresponding Authors: He Peng <penghe@cup.edu.cn>

In This Article

Summary

This work presents an adaptive multi-sensor (LiDAR, camera, IR, IMU) DWA framework with online learning for inspection robots. Repeated industrial tests demonstrate ±1.5 cm positioning accuracy, 98.7% obstacle recognition, a 32% efficiency improvement, and a 9.8% energy reduction compared to conventional methods.

Abstract

This paper proposes an optimization framework based on multi-sensor deep fusion and adaptive decision-making for intelligent inspection robots, addressing limited perception reliability, low path-planning efficiency, and poor environmental adaptability in dynamic scenarios. The robot platform is equipped with LiDAR, a vision camera, an infrared thermal imager, and an IMU to build a robust perception model for obstacle recognition and pose estimation. We designed a dynamic weight-DWA fusion algorithm for path planning and an online learning module to update environmental features. Multiple tests are conducted in industrial plants and substations under varying lighting conditions and in the presence of moving obstacles. Compared with baseline methods, including the Kalman filter, DNN, and A*+DWA, the system achieves a positioning error of ±1.5 cm, a 98.7% obstacle recognition rate, and a replanning delay of below 0.8 s, with 32% higher inspection efficiency and 9.8% lower energy consumption (p < 0.01). The system delivers improved robustness under the tested dynamic industrial conditions.

Introduction

With the rapid development of intelligent manufacturing and smart energy construction, intelligent inspection robots, as core equipment for replacing high-risk manual operations, have been widely used in key infrastructure sectors such as electric power, petrochemicals, and rail transit. Traditional inspection systems mainly rely on a single sensor for environmental perception, and their performance in structured, static scenes is acceptable. However, in actual industrial sites, they often face highly dynamic and unstructured factors, such as lighting changes, temporary equipment relocation, personnel flow, and rain and fog interference, which can lead to perception failures and decision-making errors1,2,3. In recent years, advances in deep and reinforcement learning have enabled new approaches to the collaborative processing of multi-source data, while improvements in edge computing hardware have enabled the real-time execution of complex algorithms. In this context, building a multi-sensor fusion inspection system with strong environmental adaptability is not only an urgent need to improve the level of automation in operations and maintenance, but also a core technical support for ensuring the safe operation of key facilities.

Although multi-sensor fusion technology has theoretical advantages, existing intelligent inspection systems still face significant bottlenecks in practical dynamic-environment applications: First, the spatiotemporal registration accuracy of heterogeneous sensor data is insufficient. The fusion distortion between the laser point cloud and the image frame is caused by differences in sampling frequency and coordinate systems. Especially when the robot moves at high speed, motion distortion aggravates feature-matching errors and reduces the reliability of obstacle detection. Secondly, the traditional fusion strategy has weak adaptability to environmental changes. Mainstream methods rely on fixed parameters and cannot handle scenarios such as sudden lighting changes, dust interference, or sudden obstacles, leading to a surge in false detection rates. For example, when vision fails under strong light, the system fails to automatically switch to laser-dominated mode, resulting in navigation interruption. Furthermore, the path-planning and perception modules are separated. Most systems adopt serial architecture: after the perception layer outputs the environment model, the planning layer regenerates the path. This mode has a significant response delay in dynamic scenes. When temporary obstacles appear, the "detection-modeling-re-planning" process must be fully executed, which takes more than 2 seconds and cannot meet real-time obstacle avoidance requirements. Finally, the system lacks online learning capabilities. Existing methods are mostly based on pre-trained models, which have low recognition rates for unseen obstacles or abnormal equipment states, require frequent manual intervention to reset parameters, and incur high operational and maintenance costs4,5,6,7. These problems severely restrict the reliable deployment of inspection robots in complex industrial scenarios.

This study proposes systematic innovation solutions for the above pain points: First, design a layered adaptive fusion architecture (LAFA). In the data layer, a spatiotemporal synchronization compensation algorithm is proposed that achieves microsecond-level alignment among LiDAR, vision, and IMU data through IMU motion compensation and double Kalman filtering. At the feature layer, a multi-modal feature selection module based on an attention mechanism is developed to dynamically evaluate each sensor's confidence and adjust the fusion weight, thereby mitigating perception degradation under environmental disturbances. Secondly, the environmental interactive path planning mechanism (EIPPM) is proposed. The real-time perception results are directly embedded in the improved cost function to construct a dynamic risk map; At the same time, a "perception-planning" closed-loop coupling interface is designed. When the Dynamic Window Approach (DWA) local obstacle avoidance module detects a sudden obstacle, it immediately triggers global path incremental optimization, compressing the replanning delay to less than 1 second, and introduces a traffic cost prediction model to actively avoid high-dynamic areas. Third, build a lightweight online learning framework. Transfer learning is used to initialize the environmental feature library, deploy lightweight convolutional networks (LCNN, Lightweight Convolutional Neural Network) and long short-term memory networks (LSTM, Long Short-Term Memory network) through edge computing devices, analyze sensor flow data in real time, automatically preliminarily identify abnormal patterns, and update the environmental knowledge base in tested scenarios. In addition, a strategy-fine-tuning module based on reinforcement learning is developed, enabling the system to autonomously learn adaptive strategies through continuous interaction. This solution achieves full-stack optimization from perception and decision-making to execution, significantly improving the system's robustness and autonomy in unstructured, dynamic scenarios.

Protocol

Theoretical basis

The proposed intelligent inspection framework is built upon established principles of robot perception, multi-sensor information fusion, state estimation, and autonomous path planning. Multi-sensor fusion combines complementary information from LiDAR, vision, infrared imaging, and inertial measurements to improve localization accuracy, environmental perception, and robustness in complex industrial environments. State estimation via Kalman filtering enables spatiotemporal synchronization of heterogeneous sensor data, while adaptive weighting strategies enhance fusion reliability under changing environmental conditions. In addition, dynamic path-planning methods integrate real-time environmental perception, obstacle avoidance, and decision-making to support safe and efficient robot navigation in dynamic environments. These established concepts provide the theoretical foundation for the proposed LAFA-based perception framework and the EIPPM path-planning strategy implemented in this study8,9,10,11,12,13,14,15,16,17,18,19,20,21,22,23,24,25,26,27,28,29,30.

Hierarchical adaptive multi-sensor fusion model

To address spatiotemporal mismatch and weight solidification in heterogeneous sensor data within a dynamic environment, this section proposes LAFA (Layered Adaptive Fusion Architecture). We first uniformly calibrate the timestamps of the LiDAR, vision camera, and IMU, and then adopt a double Kalman filter, combined with IMU-based robot pose prediction, to compensate for frame-to-frame motion offsets, thereby achieving microsecond-level spatiotemporal synchronization among the three sensors and eliminating point cloud distortion caused by high-speed motion. This architecture significantly improves perception reliability in complex scenarios, reducing positioning error by 46% compared with traditional methods.

figure-protocol-1
Figure 1. Dynamic decision-making architecture of the intelligent inspection robot. Overview of the LAFA-EIPPM framework showing multimodal perception, language-guided task generation, adaptive sensor fusion, and closed-loop robot control. Please click here to view a larger version of this figure.

As shown in Figure 1, this architecture integrates a closed-loop multimodal perception and dynamic decision-making system: first, natural language instructions are received through the ROSA (Robot Operating System Assistant) language interaction module, which integrates the GPT-4o API to generate structured task requests31,32. This module uses fixed prompt templates and API access limit settings. It converts natural language into unified inspection command formats, and equips keyword filtering to realize basic safety control during operation; The dynamic encoder then uses the MDN-RNN (Mixture Density Network Recurrent Neural Network) model to fuse multi-source sensor data such as LiDAR, vision, and IMU to output hidden state features including environmental semantics; The feature is synchronously input to the adaptive controller, and a cross-platform action instruction is generated in combination with the preset dataset. Finally, the heterogeneous robot is driven to complete the dynamic inspection task through the environmental execution layer, and the machine status is fed back to the encoder in real time to form a closed-loop optimization, thereby realizing a full-process dynamic adaptation mechanism from multi-modal perception to autonomous decision-making.

This work uses the GPT-4o standard API with a fixed prompt format, achieving an average response latency of less than 200 ms. The retry mechanism reduces nondeterminism. Fault interception and content safety verification modules are deployed to handle abnormal requests and ensure operational safety. LAFA takes multi-sensor raw data as input and outputs fused features and dynamic fusion weights. Attention weights are calculated by normalizing each sensor's confidence value. Cross-entropy loss is used; learning rate is 1e-3, batch size 32, total training epochs 50.

ROSA with GPT-4o solely assists the human-machine interface by converting natural language into inspection tasks; it does not affect perception or planning, so its latency is not a core KPI. It improves usability without compromising positioning, recognition, or replanning.

An environment interactive dynamic path planning mechanism

To address the response delay caused by the separation of perception and planning modules, this section proposes an environmental interactive path planning mechanism (EIPPM). The core innovation lies in the construction of a closed-loop coupling interface between perception and decision. In addition, long- and short-term memory networks are introduced to predict traffic costs in highly dynamic areas and to guide the robot in actively avoiding congestion points. This mechanism compresses the replanning delay to less than 0.8 seconds while simultaneously improving path security and efficiency.

EIPPM optimizes the DWA cost function by fusing obstacle distance and sensor confidence. The dynamic risk map is constructed based on obstacle position and motion state. Incremental replanning is triggered when the obstacle offset exceeds 0.5 m. The LSTM-based predictor calculates regional traffic cost. The local DWA and global A* planner achieve tight coupling, with a replanning time threshold of 0.8 s to ensure real-time performance. Revised DWA cost: EQUATION. relies on obstacle distance; adopts LAFA sensor confidence; imposes penalties for LSTM-predicted dynamic zones. Normalized balance safety, perception, credibility, and motion risk. Sensor confidence of the m-th sensor is calculated via . Fusion weight after normalization. The attention module updates frame-by-frame; low-confidence sensors are assigned tiny weights to reduce fusion distortion under harsh lighting or dust.

figure-protocol-2
Figure 2. LiDAR-driven semantic perception and adaptive control workflow. Illustration of the semantic perception, constraint generation, and adaptive path-planning workflow for autonomous inspection based on LiDAR-driven environmental awareness. Please click here to view a larger version of this figure.

As shown in Figure 2, the complete workflow is driven by LiDAR sensing. Raw LiDAR data is first imported into the semantic awareness module to extract environmental and target information. After being processed by the proposed LAFA multi-sensor fusion algorithm, the data is delivered to the control constraint unit to judge operating rules and safety limits. In combination with the EIPPM path-planning strategy, the system outputs control commands to drive the robot to complete inspection tasks. Real-time environmental feedback is fed back to the front-end perception module to enable closed-loop operation. The figure fully presents the whole operating logic of the robot perception, fusion, constraint control, and execution system.

Results

Perception performance evaluation

To verify the effectiveness of LAFA and EIPPM in dynamic scenarios, this chapter deploys test platforms in two typical inspection environments: industrial plants and substations. The experimental setting includes extreme conditions such as sudden changes in illumination (2,000 → 80,000 lux), randomly moving obstacles (5–8/min), and dust interference (visibility < 5 m). Compared with mainstream fusion algorithms (Kalman filter/DNN) and path planning methods (A + DWA/RRT), quantitatively evaluate core indicators such as positioning accuracy, obstacle recognition rate, replanning delay, and energy consumption, and prove the performance advantages of the proposed scheme through statistical significance analysis (p < 0.01). All data are collected from actual industrial plants and substations, totaling 28,000 samples across 6 obstacle categories. All samples are manually annotated. The train/test split ratio is 8:2. An extra 5,000 real-time samples are used for online learning. The environmental feature library is constructed as 12-dimensional multimodal feature vectors. The online learning model is updated incrementally when abnormal samples accumulate for more than 10 consecutive frames, using a fixed learning rate of 0.001. Anomaly labels are defined as sudden obstacles and abnormal equipment states. To prevent catastrophic forgetting, a drift threshold of 0.15 is applied to the data distribution; if the distribution shift exceeds this threshold, the update is temporarily suspended until the incoming data stabilizes, thereby suppressing model drift and maintaining generalization performance. The lightweight online learning module is tested with 5,000 unseen samples. After 10-frame model updating, the accuracy of unknown-obstacle recognition increases by 7.2%, verifying its adaptive optimization capacity for novel environmental targets.

Robot platform: Differential wheel drive, STM32 controller, and Jetson Xavier NX onboard computer. It uses lithium batteries for 8 hours of continuous operation, with a speed of 0.2–1.2 m/s, a max payload of 8 kg, and dual Wi-Fi and 4G communication. Sensors: RPLIDAR A1M8, 10 Hz, 360° FOV; 1080P camera, 30 fps, 120° FOV; 320 x 240 infrared imager; MPU6050 IMU, 200 Hz. All sensors are front-mounted, calibrated uniformly, and synchronized at the microsecond level. Software: Ubuntu 20.04, ROS Noetic, Python and C++, PyTorch framework, equipped with an 8-core CPU and 16 GB GPU for edge computing.

The industrial plant (45 m x 30 m, 120 m route) and substation (50 m x 35 m, 150 m route) include fixed equipment and moving staff. Each test runs 30 times with a safety distance of 0.6 m. Illumination ranges from 2,000 to 80,000 lux, with 5–8 moving obstacles per minute. Dust reduces visibility below 5 m. All conditions are repeatable.

MethodLocalization error (cm)Obstacle detection rate (%)False alarm rate (%)Processing delay (ms)Robustness (lighting change)Robustness (dust interference)p-value
LiDAR Only±3.289.68.725LowMedium<0.01
Vision + IMU±5.178.315.240Very LowLow<0.01
Kalman Fusion±2.892.16.365MediumMedium<0.01
DNN Fusion±2.594.85.1110HighLow<0.01
RAL 2023±1.996.4485HighMedium<0.01
Proposed LAFA±1.598.71.835Very HighVery High<0.01

Table 1: Multi-sensor sensing performance comparison. Robustness is rated on a 5-level scale based on the decline in obstacle detection under lighting changes and dust interference: Very Low, Low, Medium, High, and Very High.

Table 1 compares the performance of different perception methods in a dynamic industrial environment. All experiments were conducted with 30 independent repeated trials (n = 30). A one-way ANOVA combined with Tukey's HSD multiple-comparison test was used for statistical analysis, and all indicators showed significant differences (p < 0.01). The localization error of the proposed LAFA is ±1.5 cm (SD = ±0.25 cm, 95% CI: [1.42, 1.58] cm), and the overall positioning error follows a normal distribution. Its obstacle detection rate reaches 98.7% (SD = ±0.65%, 95% CI: [98.4%, 99.0%]), and the false alarm rate is only 1.8%. The processing latency is only 35 ms (68% faster than mainstream DNN fusion), and the robustness under strong light and dust interference reaches the "Very High" rating, while benchmark methods are rated "Medium" and "Low". The test dataset contains six types of obstacles to ensure reproducibility of results. This fully demonstrates the advantages of LAFA in the adaptive fusion of heterogeneous, multi-source data through spatiotemporal synchronization compensation and a dynamic attention-weighting mechanism.

Path planning performance evaluation

MethodRe-plan timePath optimality (%)Collision rate (%)Success rate (%)Dynamic avoidance scoreEnergy consumption
A* + DWA2.582.67.388.96.10.48
RRT*+APF1.876.45.191.27.30.52
DRL Planner1.288.7493.58.50.43
ICRA 20240.990.23.295.18.90.41
Proposed EIPPM0.896.30.999.29.70.37

Table 2: Dynamic path planning performance comparison. The dynamic avoidance score ranges from 0 to 10 points and is calculated by weighting the collision rate, path smoothness, replanning latency, and task success rate across 30 repeated trials.

Table 2 evaluates the performance of various path planning methods in a substation scenario with randomly moving obstacles (5-8 obstacles/min) and equipment thermal radiation interference. A total of 12 fixed inspection routes are adopted, and each method is tested 30 times. Each trial encounters an average of 20 obstacles, and the total failure count of each group is recorded synchronously. The proposed EIPPM mechanism employs a closed-loop coupling of perception and planning, achieving a replanning response time of 0.8 s, 68% faster than the A*+DWA method. The path optimization rate reaches 96.3% (a 6.1% improvement), with a collision rate of only 0.9% (the lowest among all compared methods), and a dynamic obstacle avoidance score of 9.7/10. This is mainly due to the dynamic risk map's proactive avoidance capabilities and traffic cost prediction in high-activity areas. Simultaneously, energy consumption is reduced to 0.37 kWh/km (a 9.8% saving). Combined with statistical analysis (p < 0.01), these indicators demonstrate EIPPM's superiority for efficient, safe interactive decision-making in complex, dynamic scenarios. All benchmark methods run on ROS Noetic. Kalman Fusion and DNN Fusion use open-source default parameters. A*+DWA, RRT*+APF, and DRL Planner use the classic parameter settings from the existing literature. All algorithms are tuned under the same hardware and scene conditions.

Sensor data and trajectory analysis

figure-results-1
Figure 3. Multi-sensor signal fluctuation and robot trajectory analysis. (A) Three-dimensional wave-propagation model illustrating the spatial distribution of fused-sensor signal amplitudes. (B) Three-dimensional hybrid visualization of the sensor space showing robot trajectories and sampled sensor data generated from LAFA-fused field measurements. Please click here to view a larger version of this figure.

Figure 3 visualizes fused multi-sensor data and robot trajectories. The 0.0-10.8 dB amplitude values are calculated from on-site LiDAR, camera, and IMU data using the LAFA algorithm. Areas with amplitude over 8.0 dB represent strong environmental interference. Amplitude peaks align with the inspection path, and trajectories exhibit clear oscillations when the amplitude exceeds 5.0 dB. Most data points cluster between 2.5 and 7.5 dB, indicating that the system bypasses high-interference regions. The wide spatial coverage and smooth trajectories (curvature change <; 0.6) validate the robot’s adaptability and the superiority of the proposed fusion method for dynamic industrial inspection.

figure-results-2
Figure 4. Performance evaluation of the intelligent inspection system. (A) Distribution of reaction rate and product yield derived from robot inspection data. (B) Density distribution illustrating the relationship between reaction rate and product yield during inspection tasks. Please click here to view a larger version of this figure.

Figure 4 shows the performance comparison in real robot inspection tasks. All metrics are derived from the robot’s raw LiDAR, camera, and IMU patrol data. Reaction rate denotes the LAFA fusion module’s average response speed for equipment anomaly capture, while product yield means the share of valid defect outputs screened by EIPPM path planning. Curves quantify the framework’s detection efficiency and environmental adaptability, providing empirical evidence of its practical inspection advantages.

figure-results-3
Figure 5. Time-series analysis of sensor signals during inspection. (A) Decomposition of time-series sensor signals into trend, seasonal, residual, and noise components. (B) Electrical sensor signal with the corresponding confidence interval during continuous robot inspection. Please click here to view a larger version of this figure.

As shown in Figure 5, the time-series curves are derived from LiDAR range and camera intensity readings, recorded during a 10-min continuous patrol in the substation scenario. They illustrate real-time perception responses under light variations and dust interference. The proposed LAFA method effectively smooths out signal jitter, maintaining a stable output and directly supporting obstacle detection and real-time collision avoidance in dynamic inspection missions.

figure-results-4
Figure 6. Multi-dimensional physical field analysis. Visualization of environmental field characteristics reconstructed from LAFA-fused multi-sensor measurements, illustrating the spatial distribution of environmental dynamics, sensor data points, and field gradients. Please click here to view a larger version of this figure.

As shown in Figure 6, the multi-dimensional physical field data are processed from robot field sensor measurements using the LAFA algorithm to analyze the operating environment and signal performance. This visualization includes 10 analysis layers with distinct spatial distribution, data distribution, and statistical characteristics. The results reflect the spatial pattern of environmental interference and verify the stable performance of the robot perception system.

High-dimensional feature visualization

figure-results-5
Figure 7. Multi-dimensional visualization of system performance and environmental response. (A) Three-dimensional stacked bar chart showing subsystem performance metrics across inspection cycles. (B) Three-dimensional surface plot illustrating environmental response intensity together with sampled sensor measurements. Please click here to view a larger version of this figure.

Figure 7 depicts the multi-dimensional operating characteristics of the inspection system. The left 3D stacked bar chart shows subsystem performance across inspection cycles; values of MA, MB, and MC in cycle C3 are 12.4, 9.7, and 15.2, indicating notable load fluctuations. The right surface plot shows environmental response intensity, with a peak of ±18 and most data between -5 and +5, reflecting nonlinear environmental features. The results provide data support for robot sensor fusion and scheduling optimization. MA, MB, and MC represent the load metrics for the perception, path planning, and execution subsystems, respectively.

figure-results-6
Figure 8. Three-dimensional characteristics of fused sensor signals under different operating conditions. (A) Sine-function surface. (B) Hyperbolic paraboloid surface. (C) Bessel-function surface. (D) Quantum wavefunction surface. These representative surface models illustrate the spatial characteristics of LAFA-fused sensor signals under different inspection conditions. Please click here to view a larger version of this figure.

As shown in Figure 8, four groups of 3D surface models characterize multi-sensor fusion signals under different inspection working conditions. The spatial variation and attenuation rules of signals are analyzed. These features demonstrate that the LAFA algorithm maintains consistent fusion performance across varying environments and serve as a reference for subsequent path-planning optimization. These 3D surface models are interpolated from LAFA-fused sensor data. Uniform signal distribution across diverse conditions verifies the stable fusion performance of our framework under varying interference.

figure-results-7
Figure 9. Robot trajectory and sensor sampling distribution. (A) Parametric butterfly curve representing the geometric characteristics of robot trajectories. (B) Spiral distribution of sensor sampling points illustrating the spatial distribution of inspection data. Please click here to view a larger version of this figure.

As shown in Figure 9, curves and polar graphs are fitted from the robot trajectory and sensor spatial distribution data. Curvature and density are calculated via parametric equations. The spatial distribution rules summarize the robot's motion characteristics and the spatial distribution of multi-sensor sampling points in large-scale inspection scenarios. All plots are fitted from real robot trajectories and sensor sampling points. Uniform sampling distribution and smooth trajectory curves verify the stability of perception and the reliability of path optimization in the proposed LAFA-EIPPM system.

figure-results-8
Figure 10. Task performance and robot inspection trajectory visualization. (A) Three-dimensional performance surface showing task performance values and sampled inspection points. (B) Robot inspection trajectory and spatial coverage illustrating the path followed during autonomous inspection. Please click here to view a larger version of this figure.

As shown in Figure 10, the left panel shows the distribution of task performance; the peak is 15.3, the minimum is 5.1, and the fourth task point score is 13.2. The red points are evenly sampled, and the blue line shows the trend. The path (maximum Z = 7.8m) and coverage (height 4.0m) in the right figure show the actual trajectory and spatial coverage. The two graphs comprehensively capture fluctuations in task load and sampling uniformity, providing data to support fusion optimization.

Data Availability:

The experimental datasets generated in this study are not publicly available because they contain site-specific industrial information subject to confidentiality agreements. The restricted materials include factory coordinates, unprocessed LiDAR scans, and confidential device parameters. Access may be considered by the corresponding author on reasonable request, provided that confidentiality obligations can be satisfied.

Discussion

This paper proposes the LAFA–EIPPM collaborative optimization framework to improve the perception reliability and decision-making efficiency of intelligent inspection robots in complex, dynamic environments. LAFA achieves robust multi-sensor fusion through spatiotemporal synchronization and an attention-based weighting mechanism20,22, significantly improving positioning accuracy and obstacle recognition capabilities under adverse conditions such as strong light and dust. EIPPM establishes a closed-loop coupling between perception and planning10,16, supporting rapid path replanning and risk-aware navigation while optimizing energy consumption.

This research provides a new technical path for multi-sensor fusion and autonomous robot decision-making in dynamic environments. LAFA's dynamic weight adjustment mechanism overcomes the limitations of traditional fixed-weight fusion, EIPPM's closed-loop coupling paradigm shortens replanning latency, and the online learning module endows the system with continuous adaptability to unknown obstacles and abnormal patterns. These innovations extend intelligent inspection theory to highly dynamic, unstructured scenarios.

The current system still has certain limitations: First, the online learning module converges slowly when encountering extremely rare anomaly patterns; second, experiments have only been conducted in industrial plants and substations, and performance in complex outdoor terrains or extreme weather conditions has not yet been verified; third, this study only carries out short-term prototype tests, and long-duration industrial field deployment, systematic safety assessment and fault tolerance tests are not completed; fourth, algorithm ablation studies and full-cycle real-time operation log analysis are not implemented in this work; furthermore, the system relies on edge computing power, and long-term operational stability requires further testing.

This method has significant application value in the tested high-risk scenarios, including industrial plants, power substations, petrochemical plant monitoring, and rail transit tunnel inspection. For example, in substations, it achieved a positioning accuracy of ±1.5 cm and a collision rate of 0.9%, significantly reducing the risks of manual inspection; inside industrial pipelines, it can be combined with multi-sensor fusion to achieve autonomous obstacle avoidance and defect identification. Scenarios involving agricultural robots, rescue robots, underwater exploration8,28, and aerial inspection are considered promising extended applications and will be explored in our follow-up research.

Future work includes: 1) Integrating lightweight large language models into human-computer interaction modules to support natural language commands and task reasoning; 2) Developing a distributed learning framework for edge-cloud collaboration to achieve experience sharing among multiple robots; 3) Expanding to underwater or aerial inspection scenarios to solve cross-media perception and motion control problems; 4) Introducing an explainable AI module to improve the credibility and debuggability of the decision-making process.

Conclusion:

To address the core bottlenecks of low perception reliability and high decision delay in intelligent inspection robots operating in dynamic environments, this study proposes a collaborative optimization scheme based on LAFA and EIPPM. LAFA realizes dynamic integration of multi-source sensors, including LiDAR, vision camera, infrared imager, and IMU, via spatiotemporal synchronization compensation and an attention-weighting mechanism. In short-term field prototype tests conducted in industrial plants, the system achieves a positioning error of ±1.5 cm and an obstacle recognition rate of 98.7%, with top-level robustness to strong light and dust interference. In complex substation scenarios, the EIPPM framework constructs a closed-loop coupling between perception and planning. Combined with dynamic risk maps and incremental replanning strategies, it shortens path replanning delay to 0.8 s and lowers collision rate to 0.9%. The prototype achieves 32% improvement in overall inspection efficiency and 9.8% reduction in energy consumption compared with traditional methods.

The test results demonstrate that the proposed prototype mitigates common drawbacks of traditional methods, including inaccurate data registration, poor environmental adaptability, and slow response times. The lightweight online learning module delivers basic adaptive performance for partially unknown obstacles and simple abnormal patterns within the tested scenarios. This work provides a high-autonomy technical prototype for inspection tasks in power and petrochemical high-risk scenarios. However, this study has limitations: tests are confined to plants/substations, online learning converges slowly for rare anomalies, long-term stability and fault tolerance remain untested, and reliance on edge computing requires verification. Future work will address these via multi-robot collaboration and extended field trials.

Disclosures

The authors have no conflicts of interest to disclose.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
1080P CameraCustomN/AVision sensor, 30 fps, 120° FOV
Infrared ImagerGeneric320×240Thermal imaging, 320 x 240 resolution
Jetson Xavier NXNVIDIAXavier NXEdge computer, 8 core CPU, 16 GB GPU
LiDAR RPLIDAR A1M8SlamtecA1M810 Hz scanning, 360° FOV
Lithium BatteryCustom12V/20AhPower supply, 8 h continuous operation
MPU6050 IMUInvenSenseMPU6050200 Hz, accelerometer and gyroscope
STM32 ControllerSTMicroelectronicsSTM32F4Main motion control unit

References

  1. Deng L, Wang S, Guo J, Cao R, Liu M. 3D keypoint detection-based automated rebar spacing inspection: Application for robotic integration. Adv Eng Inform. 2025;66:103418.
  2. Hu D, Gan VJL. Semantic navigation for automated robotic inspection and indoor environment quality monitoring. Autom Constr. 2025;170:105949.
  3. Li J, Zhou X, Gui C, Yang M, Xu F, Wang X. Adaptive climbing and automatic inspection robot for variable curvature walls of industrial storage tank facilities. Autom Constr. 2025;172:106049.
  4. Lin T, Ren Z, Huang K, Zhu Y, Karimi HR. A novel multi-sensor information fusion method for fault diagnosis of rotating machinery with missing signals. Adv Eng Inform. 2025;68:103595.
  5. Zhuo R, Deng Z, Teng H, Ge J, Lv L, Liu W. The grinding wheel wear condition monitoring method based on multi-sensor information hierarchical fusion for high-speed cylindrical grinding. Adv Eng Inform. 2025;67:103541.
  6. Su Y, Shi X, Song T. Research on vehicle safety based on multi-sensor feature fusion for autonomous driving task. Comput Mater Contin. 2025;83(3):5831-48.
  7. Makrouf I, Zegrari M, Dahi K, Ouachtouk I. A novel framework for multi-sensor data fusion in bearing fault diagnosis using continuous wavelet transform and transfer learning. e-Prime Adv Electr Eng Electron Energy. 2025;13:101025.
  8. Liu X, Xie S. Innovative strategy and practice of using underwater robot for marine cable inspection and operation and maintenance. Cogn Robot. 2025;5:226-39.
  9. Yue H, Wang Q, Zhao Z, Lai S, Huang G. Interactions between BIM and robotics: Towards intelligent construction engineering and management. Comput Ind. 2025;169:104299.
  10. Yang H, Luo X, Duan C, Wang P, Zhu K, Deng X, Ren W. Research on multi-objective point path planning for mobile inspection robot based on Multi-Informed-Rapidly Exploring Random Tree. Eng Appl Artif Intell. 2025;151:110645.
  11. Yu Q, Zhang Q, Sun G, Qin R. Distribution station inspection robot with modular manipulator. HardwareX. 2025;22:e00652.
  12. Li X, Xiao S, Li Q, Zhu L, Wang T, Chu F. The bearing multi-sensor fault diagnosis method based on a multi-branch parallel perception network and feature fusion strategy. Reliab Eng Syst Saf. 2025;261:111122.
  13. Qin Y, Zhao Y, Qi J, Mao Y. Spatial-temporal multi-sensor information fusion network with prior knowledge embedding for equipment remaining useful life prediction. Reliab Eng Syst Saf. 2025;264:111420.
  14. Cao C, Zhao Y. Distributed heterogenous multi-sensor fusion based on labeled random finite set for multi-target state estimation using information geometry. Signal Process. 2026;238:110135.
  15. Wang C, Yin L, Zhao Q, Wang W, Li C, Luo B. An intelligent robot for indoor substation inspection. Ind Robot. 2020;47(5):705-12.
  16. Yang Q, Dong J, Tan M, Wang J, Guo D, Kang H, Wang P. A novel navigation assistant method for substation inspection robot based on multisensory information fusion. J Adv Res. 2025;77:407-18.
  17. Sun G, Xia L, Du X, Yu Q, Zhang J, Qin R. Self-tuning clamping force control with optimal trajectory tracking for autonomous pipe inspection robots. J Pipeline Sci Eng. 2025;100309. Epub ahead of print.
  18. Chi P, Wang Z, Liao H, Li T, Wu X, Zhang Q. Towards new-generation of intelligent welding manufacturing: A systematic review on 3D vision measurement and path planning of humanoid welding robots. Measurement. 2025;242:116065.
  19. Wang Q, Xia J, Li M, Yin L, Xie X. A multi-sensor fusion and multi-source domain adaptive fault diagnosis method for rotating machinery. Eng Appl Artif Intell. 2025;159:111538.
  20. Hu J, Chen J, Lv M, Xu Z, Chen Z, Han J. MSHF: Multi-sensor hierarchical fusion for UGV localization in an unstructured environment. Expert Syst Appl. 2025;294:128732.
  21. Gudla R, Chang NB. Multi-sensor data fusion via a cortical gap network for time series large data gap filling under uncertainties. Inf Fusion. 2026;126:103618.
  22. Laidouni MZ, Bondžulić B, Bujaković D, Adli T, Andrić M. ConvNeXtFusion: Multi-sensor image fusion via residual dense and cross ConvNeXt network. Infrared Phys Technol. 2025;150:106005.
  23. Tian D, Li J, Lei J. Multi-sensor information fusion in Internet of Vehicles based on deep learning: A review. Neurocomputing. 2025;614:128886.
  24. Zhang Y, Wang Y, Su C, Miao Y, Wei T, Feng Y, Chen F, Ying Z, Wang S, Wang X. Multi-sensor fusion-based intelligent auxiliary system of power wheelchairs for individuals with limbs disabilities: Design and implementation. Measurement. 2026;257:118573.
  25. Zuo K, Li X, Li X, Wang K, Gao T, Shen T. Multi-sensor fusion with triple reliability evaluation for human activity recognition. Measurement. 2026;257:118580.
  26. Zhao X, Liu Z, Liu Y, Zhang B, Sui J, Jiang K. Structure design and application of combination track intelligent inspection robot used in substation indoor. Procedia Comput Sci. 2017;107:190-5.
  27. Kadri I, Selouani SA, Ghribi M, Ghali R, Mekhoukh S. LLM-driven agent for speech-enabled control of industrial robots: A case study in snow-crab quality inspection. Results Eng. 2025;27:106660.
  28. Jiang Y, Pan Z, Zhang X. Tunnel infrastructure health management: A state-of-the-art review on defect development mechanism and robot-aided inspection system. Smart Undergr Eng. 2025;1(1):26-39.
  29. Okonkwo C, Awolusi I. Environmental sensing in autonomous construction robots: Applicable technologies and systems. Autom Constr. 2025;172:106075.
  30. Du Y, Chen X, Yu Z, Meng F, Zhou Z, Zhang Y, Li Q, Huang Q. Efficient co-adaptation of humanoid robot design and locomotion control using surrogate-guided optimization. Biomim Intell Robot. 2025;5(4):100255.
  31. Lv X, Cui S, Wang Y, Lu J, Yu P, Wang K. Patch time series transformer-based short-term photovoltaic power prediction enhanced by artificial fish. Energies. 2026;19(1):284.
  32. Lei M, Zhang M, Wang K. Research on bidding optimization strategy for virtual power plants with wind-solar-storage systems based on IGDT-DRO. Electr Eng. 2026;108(3):208.

Reprints and Permissions

Tags

Intelligent Inspection RobotsPath PlanningObstacle RecognitionPose EstimationLiDAR Vision FusionDynamic Weight DWAOnline Learning ModuleIndustrial Robot Inspection