A subscription to JoVE is required to view this content. Sign in or start your free trial.

Method Article

Real-time Production Line Safety Monitoring Using Deep Learning-based Object Detection and Feature Enhancement

74 views

⸱

DOI:

10.3791/70968

⸱

August 21st, 2026

In This Article

Summary

This paper proposes a production line safety monitoring method integrating You Only Look Once version 11 (YOLOv11) with convolutional neural networks. It achieved mean average precision at 0.5 intersection-over-union threshold (mAP@0.5) of 0.91, mAP@0.5:0.95 of 0.82, and 120 frames per second on a graphics processing unit, outperforming evaluated baseline models.

Abstract

With the deepening of Industry 4.0, the automation and intelligence levels of production lines have significantly improved, placing higher demands on the real-time performance, accuracy, and safety of monitoring. Traditional monitoring systems, relying on manual inspections or simple threshold-based decisions, generally suffer from slow response times, high false alarm rates, and limited intelligence. Therefore, this paper proposes a production-line safety monitoring system that integrates You Only Look Once version 11 (YOLOv11) with a convolutional neural network (CNN). First, the acquired images were preprocessed. Then, YOLOv11 was used to identify workers, equipment, and potential hazards in real time. Next, an enhanced CNN network with multi-scale feature fusion and attention mechanisms was introduced to improve the feature extraction capabilities for small and occluded targets. Finally, the detection results were fused with the CNN-enhanced features to assess safety status. Experiments were conducted using a self-built production-line safety dataset, employing the stochastic gradient descent (SGD) optimizer with momentum 0.9, an initial learning rate of 0.01, weight decay of 0.0005, cosine-annealed learning rate adjustment, a batch size of 32, and training for 200 epochs. mAP@0.5 and detection speed in frames per second (FPS) were used as evaluation metrics for comparison with the evaluated baseline algorithms. The results show that the proposed system achieved an mAP@0.5 of 0.91 and a detection speed of 120 FPS. It also demonstrated robust performance under complex conditions such as shading and varying lighting, supporting its effectiveness and practical application potential in production line safety monitoring.

Introduction

In modern industrial production, safety is the cornerstone of ensuring production efficiency and personnel well-being. The production line environment typically involves high-speed mechanical equipment, complex technological processes, and collaborative operations that involve multiple tasks. Any small potential safety hazard can lead to serious production accidents, resulting in significant economic losses and even fatalities. Therefore, it is crucial to establish an efficient, intelligent, and real-time safety monitoring system to enhance safety in industrial production1,2.

Tradit....

Access restricted. Please log in or start a trial to view this content.

Protocol

YOLOv11 and CNN algorithm

System overall architecture
The production line safety monitoring system proposed in this paper adopts a hybrid architecture that combines YOLOv11 with an improved CNN to achieve high-precision, high-speed real-time safety monitoring. As shown in Figure 1, the system mainly consists of an input preprocessing module, a YOLOv11 object detection module, a CNN feature enhancement module, and a fusion decision module. The input image is first normalized and augmented by the preprocessing module to improve the model's generalization ability. Sub....

Access restricted. Please log in or start a trial to view this content.

Results

To verify the effectiveness of the proposed YOLOv11 and CNN fusion model in production line safety monitoring, a dataset covering various safety hazard scenarios was constructed. This dataset contains 12,000 production line scene images, divided into training, validation, and test sets in a 7:1.5:1.5 ratio, with a fixed random seed of 42. It covers various working conditions, including standard lighting, low lighting, partial occlusion, heavy occlusion, motion blur, and complex backgrounds. The confidence threshold for i.......

Access restricted. Please log in or start a trial to view this content.

Discussion

Experimental results show that the proposed YOLOv11–CNN fusion model improves detection accuracy and real-time performance for production line safety monitoring. Leveraging the efficient single-stage detection of YOLOv11 and the feature enhancement of the CNN module, the system identified workers, equipment, and hazardous areas under complex lighting and occlusion conditions. Multi-scale feature fusion and attention mechanisms enhanced the ability to focus on key safety elements, providing interpretable visual evid.......

Access restricted. Please log in or start a trial to view this content.

Disclosures

The authors have no conflicts of interest to declare.

Acknowledgements

This work was sponsored in part by Taizhou University Scientific Research Foundation for Advanced Talents "Research on key technologies of precision detection for intelligent manufacturing" (TZXY2020QDJJ006).

....

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
COCO API   N/AAssessment tool used for target detection accuracy evaluation and statistical analysis.
CUDA 11.4NVIDIAN/AParallel computing platform used with the PyTorch 1.12 framework.
Edge TPU   N/ADeployment platform used for latency and power consumption testing; reported latency was 18 ms and power consumption was 15 W.
Jetson Xavier edge computing deviceNVIDIA   Edge computing device used for throughput testing; reported processing speed was 70 FPS with 99th percentile latency within 50 ms.
NVIDIA Tesla V100 GPUNVIDIA   GPU used in the training/evaluation server.
PyTorch 1.12   N/ADeep learning framework used for model training and evaluation.
scikit-learn    N/AAssessment/statistical analysis tool used for model evaluation.
Self-built production line safety datasetN/AN/ADataset of 12,000 production line scene images in VOC format, covering standard lighting, low lighting, partial occlusion, heavy occlusion, motion blur, and complex backgrounds; six annotation categories: workers, safety helmets, robotic arms, hazardous areas, tools, and vehicles.
Training server       Server equipped with an NVIDIA Tesla V100 GPU and Ubuntu 20.04 operating system.
Ubuntu 20.04 operating system     N/AOperating system used on the training/evaluation server.
VOC annotation formatN/AN/AAnnotation format used for the self-built production line safety dataset.
YOLOv11 object detection modelUltralyticsN/AObject detection model used to detect workers, equipment, and potential hazards in real time.

References

  1. Ramadan MNA, Ali MAH, Jaber H, Alkhedher M. Blockchain-secured IoT-federated learning for industrial air pollution monitoring: a mechanistic approach to exposure prediction and environmental safety. Ecotoxicol Environ Saf. 2025;300:118442.
  2. Ahn J, et al. SafeFac: video-based smart safety monitoring for preventing industrial work accidents. Expert Syst Appl. 2023;215:119397.
  3. Natha S, et al. A scalable and generalized deep ensemble model for road anomaly detection in surveillance videos. Comput Mater Contin. 2024;81(3):3707-3729.
  4. Diallo AR, Homri L, Dantan JY. Reducing false alarms in fault detection: a comparative analysis between conformal prediction an....

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Tags

Real-Time MonitoringDeep Learning DetectionYOLOv11Convolutional Neural NetworkMulti-Scale Feature FusionAttention MechanismIndustrial Automation