Обработка изображений, распознавание образов. Рубрика в журнале - Компьютерная оптика
Gradient-based technique for image structural analysis and applications
Статья научная
This paper is devoted to application of gradients field characteristics in selected problems of image intellectual analysis and processing. To analyse the properties and structure of an image several approaches and models based on the use of the gradients field characteristics, are proposed. In this paper, models based on Weibull distribution are considered, an image dominant direction estimation algorithm using the parameters of scattering ellipse of gradients field components is proposed, and a similarity measure of two images with arbitrary dimensions and orientation is proposed. Some examples of applications of these models for estimation of blur and structuredness of an image, for the quality assessment of resizing and rotating algorithms, as well as for detection of a specified object on the image delivered by an unmanned aerial vehicle, are given.
Бесплатно
Head model reconstruction and animation method using color image with depth information
Статья научная
The article presents a method for reconstructing and animating a digital model of a human head from a single RGBD image, a color RGB image with depth information. An approach is proposed for optimizing the parametric FLAME model using a point cloud of a face corresponding to a single RGBD image. The results of experimental studies have shown that the proposed optimization approach makes it possible to obtain a head model with more prominent features of the original face compared to optimization approaches using RGB images or the same approaches generalized to RGBD images.
Бесплатно
High-performance discrete wavelet transform for JPEG XS standard
Статья научная
This paper presents a high-speed method for forward and inverse discrete wavelet transform (DWT) intended for the JPEG XS image compression standard. Unlike state-of-the-art approaches, which process pixels sequentially, the proposed algorithm employs the Winograd method to compute groups of 2-5 pixels in parallel within a single clock cycle. We determine the minimum fractional bit-widths required for fixed-point arithmetic to ensure reconstructed image quality with a peak signal-to-noise ratio (PSNR) of at least 40 dB. Hardware modeling using the OpenLane environment demonstrates that the proposed method increases throughput by up to 109% for forward DWT and up to 144% for inverse DWT compared to state-of-the-art techniques. The optimal configurations are 3-pixel fragments for forward and 4-pixel fragments for inverse transforms. The proposed DWT approach is recommended for real-time systems where processing speed is critical, particularly in medical imaging and satellite data processing.
Бесплатно
High-speed recursive-separable image processing filters with variable scanning aperture sizes
Статья научная
In the process of development of computer technologies, the number of areas of their application naturally grows and, along with it, the complexity of the tasks to be solved, which entails the need for new research. Similar tasks include digital filtering of images in the field of medical technologies and active-pulse television measuring systems. There are many methods and algorithms of digital filtering designed to solve the problem of improving the quality; algorithms that can improve the quality of images while reducing computational costs are widely used. High demands, which are made due to the constant growth in the size of the generated images, as well as the requirement for modern television systems, is real-time operation. When solving practical problems, it is required to use different filter aperture sizes, which provide an increase in quality and preservation of image details. The solution of these problems was the reason for the emergence of adaptive filters that are able to change the parameters in the process of processing the received data, while not spending additional time on processing with an increase in the size of the aperture. The paper presents the principles of constructing adaptive image processing filters, which, by obtaining an input parameter indicating the required dimension of a multi-element aperture, are able to implement the construction of the required aperture. The Laplacian “Truncated Pyramid” filter and the “double pyramid” Laplacian were modified. A feature of these filters is the oddness of the multi-element aperture, so the coefficient used to build the mask is always set to odd. When using these filters, it is possible to use two coefficients that are responsible for increasing the filtration efficiency, since, in their original form, the Laplacian filters have a sum of coefficients equal to zero. The experiment shows a comparison with high-dimensional filters that work when using classical two-dimensional convolution. The next stage of the presented research will be the application of parallel computing techniques, which will increase the speed of the developed filters.
Бесплатно
Image compression and encryption based on wavelet transform and chaos
Статья научная
With the rapid development of network technology, more and more digital images are transmitted on the network, and gradually become one important means for people to access the information. The security problem of the image information data increasingly highlights and has become one problem to be attended. The current image encryption algorithm basically focuses on the simple encryption in the frequency domain or airspace domain, and related methods also have some shortcomings. Based on the characteristics of wavelet transform, this paper puts forward the image compression and encryption based on the wavelet transform and chaos by combining the advantages of chaotic mapping. This method introduces the chaos and wavelet transform into the digital image encryption algorithm, and transforms the image from the spatial domain to the frequency domain of wavelet transform, and adds the hybrid noise to the high frequency part of the wavelet transform, thus achieving the purpose of the image degradation and improving the encryption security by combining the encryption approaches in the spatial domain and frequency domain based on the chaotic sequence and the excellent characteristics of wavelet transform...
Бесплатно
Improvements of programing methods for finding reference lines on X-ray images
Статья научная
The paper gives an overview of the algorithms developed to obtain reference lines and angles on X-ray images. These geometrical characteristics are used in the medical analysis of human joints. We propose the algorithm’s modifications based on the analysis of numerous X-ray images. These modifications allowed obtaining a great increase in calculation speed and the improvement of final results quality given by the corresponding application. They also lead to a significant reduction of manual tuning of the program, arising only in the rare cases when the properties of given images differ significantly from the mean ones.
Бесплатно
Статья научная
The purpose of this work is to design a system for microscopic bacterial images classification that can be generalized to new data. In the course of work, a dataset containing 23 bacterial species was collected. We use a strain-wise method for dividing the dataset into training and test sets. Such splitting (in contrast to random division) allows evaluating the performance of classifiers on new strains in the case of intra-species visual variability of bacteria. We propose a “Multi-headed” ResNet (ResNet-MH) for the analysis of microscopic images of bacterial colonies. This approach forces the neural network to analyze features of different resolutions, such as the shape of individual bacterial cells and the shape and number of bacterial clusters during training. Our network achieves the 41.6% accuracy species-wise and 64.06% accuracy genera-wise. The proposed method of dataset splitting guarantees generalization to new unseen strains, whereas random splitting into training and test sets leads to overfitting of the system (accuracy is over 90%). For the 10 visually strain-wise stable species, the accuracy of the proposed system reaches 83.6% species-wise.
Бесплатно
Статья научная
The purpose of this work is software implementation of the temporal frame interpolation, the formation of selection criteria and the choice of a suitable neural network model based on the obtained practical data. And also, evaluation of its efficiency for eliminating the interframe shift effect of dynamic objects on the depth maps of multi-area active-pulse television measuring systems in order to improve the accuracy of map building. As initial data for the experiments, static frames were recorded while moving the test rig along the X and Z axes. The static frames are images of the test rig, averaged 100 times, at a distance of 13 meters, which moved along an automated linear guide with a step of 1 mm. As a result of the work, an assessment of the interframe shift effect influence on space depth maps of multi-area active-pulse television measuring systems containing dynamic objects was made. The implementation and testing of the temporal frame interpolation algorithm for suppressing the interframe shift effect of dynamic objects on depth maps was also performed. The algorithm was implemented using Python and the PyCharm IDE with SciPy, NumPy, OpenCV, PyTorch, Threading and other libraries. Numerical values of the RMSE, PSNR, and SSIM metrics were obtained before and after eliminating the effect of interframe shift of dynamic objects on depth maps. The use of the temporal frame interpolation algorithm allows more accurate measurement of distance to moving object in the field of view of multi-area active-pulse television measuring systems.
Бесплатно
Integrating landscape ecological risk with ecosystem services in the Republic of Tatarstan, Russia
Статья научная
It is a novel approach to linking landscape ecological risk (LER) and ecosystem services (ESs) for environmental management and sustainable development, since it enables real-time decision-making. This study used 12 natural factors relevant to LER and 11 ESs factors to analyze spatiotemporal changes and establish a relationship between them in Tatarstan, Russia, for the years 2010, 2015, and 2020. The statistical tests (Global Moran's I, Getis-Ord Gi*), analysis of habitat vulnerability, and ecological loss in the ArcGIS platform reveal a consistent variance in factor clustering and pattern as well as the impact of governmental policies in the studied area. According to analysis findings, 2015 had the best ecological conditions of the three years because 44.79 % of the research area had decreased landscape ecological risk, which increased ecosystem services. Additionally, the results show that both maps have significant spatial disparities and that LER and ESs are negatively impacted by high human-socioeconomic activity. The integration of LER and ESs through the overlap of both maps provides a significant amount of spatial information for mapping, monitoring, management, and the protection of the fragile environment for sustainable landscape development and management.
Бесплатно
Interpretable graph methods for determining nanoparticles ordering in electron microscopy images
Статья научная
An important step in determining the properties of carbon materials is the analysis of images from a scanning electron microscope (SEM). These images show the material surface after the application of metal nanoparticles. The order of these nanoparticles is a key characteristic that affects the material properties. We have previously proposed an approach to formalize the order features based on the identification of lines by nanoparticles in the SEM image. This paper proposes a novel approach to line allocation that is based on the concept of constructing a minimum spanning forest. Additionally, it introduces a set of novel ordering functions that are derived from this approach. The experimental study demonstrates that the combination of these new and previously extracted features improves the recognition quality of SEM images with ordered and disordered nanoparticles arrangements. This approach allows us to gain a better understanding of the nanoparticles arrangement and their effect on the material properties.
Бесплатно
Localization of mobile robot in prior 3D lidar maps using stereo image sequence
Статья научная
The paper studies the real-time stereo image-based localization of a vehicle in a prior 3D LiDAR map. A novel localization approach for mobile ground robot, which successfully combines conventional computer vision techniques, neural network based image analysis and numerical optimization, is proposed. It includes matching a noisy depth image and visible point cloud based on the modified Nelder-Mead optimization method. Deep neural network for image semantic segmentation is used to eliminate dynamic obstacles. The visible point cloud is extracted using a 3D mesh map representation. The proposed approach is evaluated on the KITTI dataset and a custom dataset collected from a ClearPath Husky mobile robot. It shows a stable absolute translation error of about 0.11 – 0.13 m. and a rotation error of 0.42 – 0.62 deg. The standard deviation of the obtained absolute metrics for our method is the smallest among other state-of-the-art approaches. Thus, our approach provides more stability in the estimated pose. It is achieved primarily through the use of multiple data frames during the optimization step and dynamic obstacles elimination on depth image. The method’s performance is demonstrated on different hardware platforms, including energy-efficient Nvidia Jetson Xavier AGX. With parallel code implementation, we achieve an input stereo image processing speed of 14 frames per second on Xavier AGX.
Бесплатно
MIDV-2020: a comprehensive benchmark dataset for identity document analysis
Статья научная
Identity documents recognition is an important sub-field of document analysis, which deals with tasks of robust document detection, type identification, text fields recognition, as well as identity fraud prevention and document authenticity validation given photos, scans, or video frames of an identity document capture. Significant amount of research has been published on this topic in recent years, however a chief difficulty for such research is scarcity of datasets, due to the subject matter being protected by security requirements. A few datasets of identity documents which are available lack diversity of document types, capturing conditions, or variability of document field values. In this paper, we present a dataset MIDV-2020 which consists of 1000 video clips, 2000 scanned images, and 1000 photos of 1000 unique mock identity documents, each with unique text field values and unique artificially generated faces, with rich annotation. The dataset contains 72409 annotated images in total, making it the largest publicly available identity document dataset to the date of publication. We describe the structure of the dataset, its content and annotations, and present baseline experimental results to serve as a basis for future research. For the task of document location and identification content-independent, feature-based, and semantic segmentation-based methods were evaluated. For the task of document text field recognition, the Tesseract system was evaluated on field and character levels with grouping by field alphabets and document types. For the task of face detection, the performance of Multi Task Cascaded Convolutional Neural Networks-based method was evaluated separately for different types of image input modes. The baseline evaluations show that the existing methods of identity document analysis have a lot of room for improvement given modern challenges. We believe that the proposed dataset will prove invaluable for advancement of the field of document analysis and recognition.
Бесплатно
MIDV-500: a dataset for identity document analysis and recognition on mobile devices in video stream
Статья научная
A lot of research has been devoted to identity documents analysis and recognition on mobile devices. However, no publicly available datasets designed for this particular problem currently exist. There are a few datasets which are useful for associated subtasks but in order to facilitate a more comprehensive scientific and technical approach to identity document recognition more specialized datasets are required. In this paper we present a Mobile Identity Document Video dataset (MIDV-500) consisting of 500 video clips for 50 different identity document types with ground truth which allows to perform research in a wide scope of document analysis problems. The paper presents characteristics of the dataset and evaluation results for existing methods of face detection, text line recognition, and document fields data extraction. Since an important feature of identity documents is their sensitiveness as they contain personal data, all source document images used in MIDV-500 are either in public domain or distributed under public copyright licenses. The main goal of this paper is to present a dataset. However, in addition and as a baseline, we present evaluation results for existing methods for face detection, text line recognition, and document data extraction, using the presented dataset.
Бесплатно
Method for removing haze from images, captured under a wide range of lighting conditions
Статья научная
The presence of haze on images degrades the quality of perception and automatic analysis of scenes. One of the most popular methods of haze removal is the dark channel prior method, which is based on the Koschmieder atmospheric scattering model. However, its underlying assumptions are not met for nighttime, since localized light sources make a significant, if not the main, contribution to lighting. We propose here to use the degree of belonging of an image element to a localized light source, determined based on a one-class classifier, as a value that characterizes the confidence of the corresponding element of the estimated transmission map during its rectifi-cation based on the gamma-normal model, which makes it possible to increase the accuracy of dehazing when processing images, captured in low-light or nighttime conditions.
Бесплатно
Статья научная
Modified versions of the Wiener filter that replace the mean with the median (MMWF) can better reduce Gaussian noise in images than the classical Wiener filter (WF). However, performance gradually decreases as the noise variance increases. To overcome these limitations, we propose a modified Wiener filter (MADNWF) that replaces the mean with the scaled median absolute deviation (MADN). Similar to MMWF, our modification to the Wiener filter uses a local kernel to average over each pixel. Instead of using the median alone, we replace it with MADN. We used four public datasets for the experiment: Set12, the Tampere17 noise-free dataset, TID2008, and the BSD68 dataset. The first step of this study was to generate 'degraded' images. To achieve this, we added a random amount of zero-mean Gaussian noise with noise variances ranging from 10 to 90, in increments of 10, to every image in the dataset. For comparison with the WF and MMWF, we then applied the proposed MADNWF to reduce noise in degraded images. We evaluated the resulting performance by comparing the reduced noise images with the original images using the Structural Similarity Index (SSIM) and Peak Signal-to-Noise Ratio (PSNR) metrics. Experimental results demonstrate that MADNWF consistently outperforms the other methods. At a low noise level (variance = 10), MMWF with a 3x3 filter slightly outperforms in fine-structural preservation, achieving a peak SSIM of 0.6676. However, as the noise intensity increases (variances from 20 to 90), MADNWF achieves complete dominance across all datasets. MADNWF achieves the highest PSNR, up to 33.27 dB at low noise levels, representing an improvement of up to 0.28 dB over WF. Under extreme noise conditions (variance = 90), the MADNWF 7x7 configuration achieves superior image cleanliness (up to 28.91 dB), whereas the MADNWF 5x5 setting serves as the optimal trade-off for structure preservation, outperforming MMWF with an SSIM margin increase of up to 0.0276. In conclusion, our results demonstrate that MADNWF helps reduce the noise distribution in natural images.
Бесплатно
Multispectral optoelectronic device for controlling an autonomous mobile platform
Статья научная
The paper substantiates the use of multispectral optoelectronic sensors intended to solve the problem of improving the positioning accuracy of autonomous mobile platforms. A mathematical model of the developed device operation has been suggested in the paper. Its distinctive feature is the cooperative processing of signals obtained from sensors operating in ultraviolet, visible, and infrared ranges and lidar. It reduces the computational complexity of detecting dynamic and stationary objects within the field of view of the device by processing data on the diffuse reflectivity of materials. The paper presents the functional organization of a multispectral optoelectronic device that makes it possible to detect and classify working scene objects with less time spending as compared to analogs. In the course of experimental research, the validity of the mathematical model was evaluated and there were obtained empirical data by means of the proposed hardware and software test stand. The accuracy evaluation of the detected object, at a distance of up to 100m inclusive, is within 0.95. At a distance of more than 100 m, it decreases. This is due to the operating range of a lidar. Error in determining spatial coordinates is of exponential character and it also increases sharply at a distance close to 100 m.
Бесплатно
Mutual modality learning for video action classification
Статья научная
The construction of models for video action classification progresses rapidly. However, the performance of those models can still be easily improved by ensembling with the same models trained on different modalities (e.g. Optical flow). Unfortunately, it is computationally expensive to use several modalities during inference. Recent works examine the ways to integrate advantages of multi-modality into a single RGB-model. Yet, there is still room for improvement. In this paper, we explore various methods to embed the ensemble power into a single model. We show that proper initialization, as well as mutual modality learning, enhances single-modality models. As a result, we achieve state-of-the-art results in the Something-Something-v2 benchmark.
Бесплатно
New information technology for phytocenoses regional monitoring using remote sensing data
Статья научная
A new information technology for plant communities monitoring using remote sensing data, oriented for application at the regional level, is proposed. The technology is based on maintaining a base of reference polygons, accumulating data on the boundaries of specific plant communities and related semantic information. This database provides a source of verified and up-to-date information for solving problems of rational nature management. To expand the database of reference polygons, two algorithms for finding new ones are presented: a reliable algorithm (analyzing several growing seasons) and an urgent algorithm (based on the current growing season). The advantage of the proposed system is the integration of data storage, processing, and analysis, which enables the automation of the creation of new and the monitoring of existing polygons, as well as the solution of a wide range of problems based on remote sensing data and artificial intelligence algorithms. Practical tasks in studying of phytocenoses in the Samara Region, implemented using the proposed monitoring technology, are considered: updating forest inventory data, monitoring the status of a rare reintroduced species (Paeonia Tenuifolia), searching for new reference polygons across a vast territory of several steppe protected areas, and searching for areas of presence of an invasive plant species (Elaeagnus angustifolia L).
Бесплатно
Noise reduction and mammography image segmentation optimization with novel QIMFT-SSA method
Статья научная
Breast cancer is one of the most dreaded diseases that affects women worldwide and has led to many deaths. Early detection of breast masses prolongs life expectancy in women and hence the development of an automated system for breast masses supports radiologists for accurate diagnosis. In fact, providing an optimal approach with the highest speed and more accuracy is an approach provided by computer-aided design techniques to determine the exact area of breast tumors to use a decision support management system as an assistant to physicians. This study proposes an optimal approach to noise reduction in mammographic images and to identify salt and pepper, Gaussian, Poisson and impact noises to determine the exact mass detection operation after these noise reduction. It therefore offers a method for noise reduction operations called Quantum Inverse MFT Filtering and a method for precision mass segmentation called the Optimal Social Spider Algorithm (SSA) in mammographic images. The hybrid approach called QIMFT-SSA is evaluated in terms of criteria compared to previous methods such as peak Signal-to-Noise Ratio (PSNR) and Mean-Squared Error (MSE) in noise reduction and accuracy of detection for mass area recognition. The proposed method presents more performance of noise reduction and segmentation in comparison to state-of-arts methods. supported the work.
Бесплатно
Novel approach of simplification detected contours on X-ray medical images
Статья научная
This paper gives description of a method for simplifying the number of points representing detected contours of the bones on digital X-ray images. Such simplification permits simplify way for correction the location of these points in the cases, if the analyzed image has poor quality, and to reduces the time of analysis it to get the reference lines and angles for diagnosis purposes of the area under investigation.
Бесплатно