Hybrid ViT-UNet Framework for Accurate River Segmentation and Buffer Zone Mapping in High-Resolution Satellite Imagery

Автор: T. Satyanarayana Murthy, K. Gangadhara Rao, Swathi Sowmya Bavirathi

Журнал: International Journal of Engineering and Manufacturing @ijem

Статья в выпуске: 4 vol.16, 2026 года.

Бесплатный доступ

Precise mapping of water bodies is crucial for flood monitoring, disaster risk and response reduction, as well as sustainable water resource management. In this paper, we introduce a deep learning model for effective segmentation of rivers, lakes, and reservoirs from high-resolution Gaofen-2 satellite images. Leveraging the Five-Billion-Pixels dataset-more than 5 billion annotated pixels for 24 land cover classes—our approach solves the problem of segmenting water bodies on various terrains and environmental conditions. The proposed U-Net and ViT-UNet models, with the former employing Vision Transformers to enhance global context perception. For enhancing generalization, the dataset is augmented using Albumentations and flipping, rotation, and scaling transformations. Hybrid loss functions of Dice Loss, Binary Cross-Entropy, and Focal Loss are employed to handle class imbalance, especially for slender river segments. The ViT-UNet model attained 98.8% pixel accuracy, which mirrors its ability to preserve fine detail and large-scale spatial pattern. Mixed-precision training and the AdamW optimizer has enhanced the computational efficiency. Further, demonstrates the potential of transformer-based segmentation models for remote sensing achieved accuracy of 98% for environmental risk management and decision support in disaster-prone areas.

Segmentation, Remote Sensing, Vision Transformers, Flood Prediction, Semantic Segmentation

Короткий адрес: https://sciup.org/15020577

IDR: 15020577   |   DOI: 10.5815/ijem.2026.04.05