HCTSpeckle and Lightweight YUV Transformer with ViT Encoder-Sandwich Decoder Network For Underwater Microplastic High Resolution Image Segmentation
Автор: Badugu Vimala Victoria, Kamil Reza Khondakar, Ravi Kumar Suggala
Журнал: International Journal of Image, Graphics and Signal Processing @ijigsp
Статья в выпуске: 4 vol.18, 2026 года.
Бесплатный доступ
Microplastics are tiny particles made of plastic that are a significant source of pollution in the sea, and are dangerous to both the environment and human health. Nevertheless, the existing detection techniques are not robust and general in dynamic underwater settings with varying microplastic shapes and sizes. To overcome these issues, a high-resolution underwater microplastic segmentation framework was implemented that integrates HCTSpeckle and Vision Transformer Encoder–Sandwich Decoder Network. Initially, underwater sensors record continuous visual images. The captured images are pre-processed with HCTSpeckle, a CNN-Transformer denoising network, for removing speckle noise, and it retains important structural information through the use of hybrid convolution-transformer blocks and double residual interactions. The denoised images were contrasted with the Lightweight YUV Transformer-based Network, in which multistage squeeze-and-excitation fusion is used to improve the visibility of object boundaries. The enhanced images are subjected to a hybrid Vision Transformer- Sandwich Decoder to produce precise underwater microplastic segmentation, where the ViT encoder captures high-level features that are globally correlated without distorting positioning information in space by patch embeddings and self-attention mechanisms. They are decoded using a Sandwich Decoder Network, which learns both local and global dependencies. Also, ranking and region-based pooling priorities fine edges and microplastic structures, whereas the pixel-wise segmentation head precisely categorizes the identified microplastics into fiber, film, pellet, and fragment types. The proposed approach attains the pixel accuracy of 0.97 and specificity of 0.98 with a Dice Coefficient of 0.934, which contains the effective segmentation result of underwater microplastic images. These findings ensure the framework efficiently identifies and categorizes various types of microplastics in diverse underwater sceneries.
Marine Pollution, Microplastic Segmentation, Hctspeckle Denoising, Lightweight YUV Transformer-Based Network, Vision Transformer, Sandwich Decoder Network
Короткий адрес: https://sciup.org/15020562
IDR: 15020562 | DOI: 10.5815/ijigsp.2026.04.04