This study presents a lightweight U-Net variant with improved segmentation accuracy using hybrid convolution and channel attention, achieving state-of-the-art results on MoNuseg and GlaS with fewer parameters.
A subscription to JoVE is required to view this content. Sign in or start your free trial.
Method Article
This study presents a lightweight U-Net variant with improved segmentation accuracy using hybrid convolution and channel attention, achieving state-of-the-art results on MoNuseg and GlaS with fewer parameters.
Artificial neural network-based computer image processing technology has rapidly developed in recent years and found widespread applications across multiple fields. In medical image processing, UNet and its variants have shown great success in lesion detection, cell segmentation, and polyp segmentation tasks. This research presents a modified U-shaped network with reduced network parameters achieved by decreasing the network depth and increasing the network channels to enhance the model's learning ability. To counteract the reduction in learning capacity caused by depth compression, a hybrid-channel convolutional module is introduced to replace the original network's convolutional module. The model introduces a channel attention mechanism between the layer and the point convolution layer to improve the practical channel feature extraction ability. The article also concludes that using mixed-depth convolution can effectively solve the size span of the segmentation target, making it difficult for one model to be fit for multiple data sets. The proposed model achieves state-of-the-art results on two popular public datasets, MoNuseg and GlaS, with a mean dice increase of 1.0% and 1.37%, respectively. The total parameters of the modified model are reduced to 1.71M, representing a 38.6x reduction compared to UCtransNet, with 65.6M parameters.
Medical image segmentation is critical in various clinical applications, including diagnosis, treatment planning, and disease monitoring. Traditional image processing methods based on feature engineering algorithms are no longer suitable for handling the large volumes of data generated in modern healthcare environments. Fortunately, the advancement of artificial neural network technology and improvements in computer hardware capabilities have made it possible to train and deploy large-scale neural network models. As a result, computer vision processing algorithms based on machine learning models are continuously being developed and applied in various fields.
Access restricted. Please log in or start a trial to view this content.
This section describes the proposed MixKNet architecture for medical image segmentation. The MixKNet model extends the UNet architecture, incorporating attention mechanisms and mixed depthwise convolution layers. The software used in this study is listed in the Table of Materials.
1. Overall architecture
Access restricted. Please log in or start a trial to view this content.
Comparison with state-of-the-art methods
The MixKNet architecture has been evaluated on two medical image segmentation datasets, GlaS and MoNuSeg, and compared with state-of-the-art methods in the field. The evaluation has been conducted by reporting the dice and IoU scores of the proposed method and several comparable models, as shown in Table 3. The results indicate that the MixKNet architecture outperforms the other models, achieving the highest me.......
Access restricted. Please log in or start a trial to view this content.
This study presents MixKNet, a novel modification of the UNet1 architecture for medical image segmentation, integrating mixed-depth convolution and a channel attention mechanism to enhance segmentation performance while significantly reducing computational complexity. The hybrid channel convolution combines multiple convolutional kernels operating on separate channel groups. This design allows the network to capture diverse spatial and contextual features at different receptive field scales within.......
Access restricted. Please log in or start a trial to view this content.
The authors declare no competing financial interests or other conflicts of interest.
The second batch of Ningbo City's 2023 Social Welfare Research Projects: Research and Application of Key Technologies for Telecom Network Fraud Identification Based on Ontology and NLP (2023S169).
....Access restricted. Please log in or start a trial to view this content.
| Name | Company | Catalog Number | Comments |
|---|---|---|---|
| Linux OS (Ubuntu 22.04) | Various (Open Source Community) | N/A(Version 22.04 LTS) | Open-source operating system for development https://releases.ubuntu.com/22.04/ |
| NVIDIA RTX 3090 GPU | NVIDIA | 3090 | High-performance GPU for deep learning https://www.nvidia.com/en-us/geforce/graphics-cards/30-series/rtx-3090/ |
| Python | Python Software Foundation | N/A (Version ≥ 3.8) | Programming language for AI/ML development https://www.python.org/ |
| PyTorch | Meta AI | N/A (Version 1.13) | Open-source machine learning framework https://pytorch.org/ |
| Visual Studio Code (VS Code) | Microsoft | N/A | Lightweight and powerful code editor https://code.visualstudio.com/ |
Access restricted. Please log in or start a trial to view this content.
This article has been published
Video Coming Soon