Статьи журнала - Компьютерная оптика

Все статьи: 2629

High-speed recursive-separable image processing filters

High-speed recursive-separable image processing filters

Kamenskiy Andrey Victorovich

Статья научная

The development of modern technologies in the field of image formation leads to an increase in the size of the generated images, as a result the question of reducing the processing computational costs arises, and this is an important factor in the creation of real-time systems. The study provides a description of high-speed recursive-separable filters for improving the quality of images, which, due to the peculiarities of their implementation, can reduce the number of computational operations required for the image processing process. This type of filters is obtained from two-dimensional linear digital filters, which are modified by applying recursive and separable properties to them. The MATLAB environment computing method for implementation of these filters is described. An extensive performance research of the developed filters has been carried out at various sizes of the test image and on various experimental installations. The comparison with the classical two-dimensional convolution method of the developed filters is demonstrated, and it shows the time gain required for the image processing. The results obtained can be applied in biomedical image processing systems or in vision systems working in heavy weather conditions.

Бесплатно

High-speed recursive-separable image processing filters with variable scanning aperture sizes

High-speed recursive-separable image processing filters with variable scanning aperture sizes

Kamenskiy A.V., Kuryachiy M.I., Krasnoperova A.S., Ilyin Yu.V., Akaeva T.M., Boyarkin S.E.

Статья научная

In the process of development of computer technologies, the number of areas of their application naturally grows and, along with it, the complexity of the tasks to be solved, which entails the need for new research. Similar tasks include digital filtering of images in the field of medical technologies and active-pulse television measuring systems. There are many methods and algorithms of digital filtering designed to solve the problem of improving the quality; algorithms that can improve the quality of images while reducing computational costs are widely used. High demands, which are made due to the constant growth in the size of the generated images, as well as the requirement for modern television systems, is real-time operation. When solving practical problems, it is required to use different filter aperture sizes, which provide an increase in quality and preservation of image details. The solution of these problems was the reason for the emergence of adaptive filters that are able to change the parameters in the process of processing the received data, while not spending additional time on processing with an increase in the size of the aperture. The paper presents the principles of constructing adaptive image processing filters, which, by obtaining an input parameter indicating the required dimension of a multi-element aperture, are able to implement the construction of the required aperture. The Laplacian “Truncated Pyramid” filter and the “double pyramid” Laplacian were modified. A feature of these filters is the oddness of the multi-element aperture, so the coefficient used to build the mask is always set to odd. When using these filters, it is possible to use two coefficients that are responsible for increasing the filtration efficiency, since, in their original form, the Laplacian filters have a sum of coefficients equal to zero. The experiment shows a comparison with high-dimensional filters that work when using classical two-dimensional convolution. The next stage of the presented research will be the application of parallel computing techniques, which will increase the speed of the developed filters.

Бесплатно

Human Action Recognition Based on The Skeletal Pairwise Dissimilarity

Human Action Recognition Based on The Skeletal Pairwise Dissimilarity

Surkov E.E., Seredin O.S., Kopylov A.V.

Статья научная

The main idea of the paper is to apply the principles of featureless pattern recognition to human activity recognition problem. The article presents the human figure representing approach based on pairwise dissimilarity function of skeletal models and a set of reference objects, also known as a basic assembly. The paper includes a basic assembly analysis and we propose the method for selecting the least-correlated basic objects. The video sequence proposed for analysis of human activity within frames is represented as an activity map. The activity map is a result of computing the pairwise dissimilarity function between skeletal models from the video sequence and the basic assembly of skeletons. The paper conducts frame-by-frame annotation of activities in the TST Fall Detection v2 database, such as standing, sitting, lying, walking, falling, post-fall lying, grasp, ungrasp. A convolutional neural network based on the ResNetV2 with the SE-block is proposed to solve the activity recognition problem. SE-block allows to detect inter-channel dependencies and selecting the most important features. Additionally, we prepare a data for training, determine an optimal hyperparameters of the neural network model. Experimental results of human activity recognition on the TST Fall Detection v2 database using the Leave-one-person-out procedure are provided. Furthermore, the paper presents a frame-by-frame assessment of the quality of human activity recognition, achieving an accuracy exceeding 83%.

Бесплатно

Hybrid Tamm-cavity modes in photonic crystal with resonant nanocomposite defect layer

Hybrid Tamm-cavity modes in photonic crystal with resonant nanocomposite defect layer

Vetrov Stepan Yakovlevich, Avdeeva Anastasia Yurievna, Pyatnov Maxim Vladimirovich, Timofeev Ivan Vladimirovich

Статья научная

Hybrid optical modes in a one-dimensional photonic crystal with a resonant nanocomposite defect bounded by a metallic layer are studied. The nanocomposite consists of spherical metallic constituents, that are distributed in a dielectric matrix. Transmittance, reflectance, and absorbance spectra of this structure, which is shined by light with normal incidence, are calculated. The possibility of control of the hybrid modes spectral characteristics by changing the thickness of the layer adjacent to the metal, the number of layers, and the nanocomposite filling factor is shown.

Бесплатно

Hyperspectral image segmentation using dimensionality reduction and classical segmentation approaches

Hyperspectral image segmentation using dimensionality reduction and classical segmentation approaches

Myasnikov Evgeny Valerevich

Статья научная

Unsupervised segmentation of hyperspectral satellite images is a challenging task due to the nature of such images. In this paper, we address this task using the following three-step procedure. First, we reduce the dimensionality of the hyperspectral images. Then, we apply one of classical segmentation algorithms (segmentation via clustering, region growing, or watershed transform). Finally, to overcome the problem of over-segmentation, we use a region merging procedure based on priority queues. To find the parameters of the algorithms and to compare the segmentation approaches, we use known measures of the segmentation quality (global consistency error and rand index) and well-known hyperspectral images.

Бесплатно

Hyperspectral remote sensing data compression and protection

Hyperspectral remote sensing data compression and protection

Gashnikov Mikhael Valeryevich, Glumov Nikolay Ivanovich, Kuznetsov Andrey Vladimirovich, Mitekin Vitaly Anatolyevich, Myasnikov Vladislav Valerievich, Sergeyev Vladislav Victorovich

Статья научная

In this paper, we consider methods for hyperspectral image processing, required in systems of image formation, storage, and transmission and aimed at solving problems of data compression and protection. A modification of the digital image compression method based on a hierarchical grid interpolation is proposed. Methods of active (on the basis of digital watermarking) and passive (on the basis of artificial image distortion detection) data protection against unauthorized dissemination are developed and investigated.

Бесплатно

Identification of parameters of discrete stochastic systems with unknown inputs based on an information filtering algorithm

Identification of parameters of discrete stochastic systems with unknown inputs based on an information filtering algorithm

A.V. Tsyganov

Статья научная

The paper considers the problem of parameter identification of discrete linear stochastic systems in the state space. An identification criterion is proposed for systems with unknown inputs based on the information version of the Gillijns and De Moor algorithm. We apply this criterion to identify the diffusion coefficient of a one-dimensional diffusion model with unknown boundary conditions of the first kind. The results of computer modeling validate the presented approach.

Бесплатно

Illustration visual communication based on computer vision image retrieval algorithm

Illustration visual communication based on computer vision image retrieval algorithm

Zhang H.Z.

Статья научная

In illustration design, good visual communication can make the audience resonate. Computer vision image retrieval algorithm provides important support and assistance for the visual communication of illustration. However, the traditional image retrieval algorithm has problems of subjectivity and inaccuracy in complex image classification. Therefore, this paper optimizes the feature extraction module of convolutional neural network and fuses hash algorithm to improve the efficiency and speed of image retrieval. The experimental results show that the accuracy of the improved convolutional neural network is 82.7 %, which is more than 6 percentage points higher than the traditional algorithm model. The recall rate of the volume neural network model improved by hashing algorithm is 94.1 %. Research is of great significance to the visual communication of illustration design, which helps designers to find relevant materials more accurately, improve the artistic quality and ornamental value of their works, and promote the innovation and development of illustration design.

Бесплатно

Image compression and encryption based on wavelet transform and chaos

Image compression and encryption based on wavelet transform and chaos

Gao Haibo, Zeng Wenjuan

Статья научная

With the rapid development of network technology, more and more digital images are transmitted on the network, and gradually become one important means for people to access the information. The security problem of the image information data increasingly highlights and has become one problem to be attended. The current image encryption algorithm basically focuses on the simple encryption in the frequency domain or airspace domain, and related methods also have some shortcomings. Based on the characteristics of wavelet transform, this paper puts forward the image compression and encryption based on the wavelet transform and chaos by combining the advantages of chaotic mapping. This method introduces the chaos and wavelet transform into the digital image encryption algorithm, and transforms the image from the spatial domain to the frequency domain of wavelet transform, and adds the hybrid noise to the high frequency part of the wavelet transform, thus achieving the purpose of the image degradation and improving the encryption security by combining the encryption approaches in the spatial domain and frequency domain based on the chaotic sequence and the excellent characteristics of wavelet transform...

Бесплатно

Image compression using discrete orthogonal transforms with the «Noise-like» basis functions

Image compression using discrete orthogonal transforms with the «Noise-like» basis functions

Chernov V.M., Dmitriyev A.G.

Статья научная

The generalization of the discrete orthogonal transforms with the basis functions generated in a pseudorandom way is the subject of the article. The examples of such transforms application in the field of videoinformation coding in the channels with the high level of «seldom» noise are also given.

Бесплатно

Improve Image Text Descriptions using Large Language Models

Improve Image Text Descriptions using Large Language Models

N. Andriyanov, A. Kim

Статья научная

In this paper, a multi-stage approach to improve text queries (prompts) for image generation is proposed and comprehensively investigated. First, the GPT-2 model, pre-trained on 18,000 raw query-quality query pairs from the Lexica.art platform, automatically expands and refines the original prompts to produce more detailed and semantically accurate images when generated by the diffusion network Stable Diffusion 1.5. Next, the Image Captioning task (BLIP2) and four large language models (DeepSeek, Grok, ChatGPT, YandexGPT) are used to compare the quality of signature expansion, demonstrating different stylistic strategies for augmenting initial descriptions. The proposed "Captioning → Prompt Enhancer (Mistral) → Stable Diffusion+LoRA" Pipeline additionally includes a pre-training of Mistral's own model on BLEU, METEOR and CIDEr metrics, providing a steady increase in the quality of textual descriptions (BLEU from 0.12 to 0.435 after 500 epochs) and a significant reduction in the FID metric for image generation (from 0.3482 to 0.1873). In the final stage, the LoRA modules embedded in UNet and the Stable Diffusion text encoder allow efficient learning of the generation of previously "unknown" objects (rare or fictional), reducing FID to 0.172 at rank = 64. An expert survey (134 respondents) confirmed the visual preference of images generated by optimized queries, demonstrating the potential of the proposed technique to improve the quality of multimodal systems.

Бесплатно

Improvements of programing methods for finding reference lines on X-ray images

Improvements of programing methods for finding reference lines on X-ray images

Al-Temimi Ammar Mudheher Sadeq, Pilidi Vladimir Stavrovich

Статья научная

The paper gives an overview of the algorithms developed to obtain reference lines and angles on X-ray images. These geometrical characteristics are used in the medical analysis of human joints. We propose the algorithm’s modifications based on the analysis of numerous X-ray images. These modifications allowed obtaining a great increase in calculation speed and the improvement of final results quality given by the corresponding application. They also lead to a significant reduction of manual tuning of the program, arising only in the rare cases when the properties of given images differ significantly from the mean ones.

Бесплатно

Improving Data Matrix mobile recognition via fast Hough transform and adaptive grid extractors

Improving Data Matrix mobile recognition via fast Hough transform and adaptive grid extractors

Rybakova E.O., Limonova E.E., Bezmaternykh P.V.

Статья научная

The Data Matrix is a barcode symbology originally designed for industrial needs. Today, its symbols are increasingly found on everyday products such as pharmaceutical packaging, electronic components, food labels, and clothing tags. This widespread usage presents a challenge: reading Data Matrix symbols from images captured by mobile cameras in uncontrolled environments. The reading process mainly consists of three steps, namely barcode localization, segmentation and decoding. In this work, we focus on the precise localization and segmentation of Data Matrix barcodes. We introduce a new method that involves the localization of the finder pattern using fast Hough transform and subsequent iterative segmentation to extract the encoded message. Our approach demonstrates superior localization quality, as measured by the mean Intersection over Union metric (0.889), and achieves better recognition accuracy (0.903) compared to open–source solutions for reading Data Matrix barcodes, such as libdmtx (0.665), ZXing (0.569), and ZXing–cpp (0.858). Our method requires only 35 milliseconds for computations on an ARM device, enabling real–time processing. It is significantly faster than libdmtx (10 seconds), ZXing (610 milliseconds), although it is slightly slower than ZXing–cpp (6.65 milliseconds).

Бесплатно

Improving generalization in classification novel bacterial strains: a multi-headed resnet approach for microscopic image classification

Improving generalization in classification novel bacterial strains: a multi-headed resnet approach for microscopic image classification

Yachnaya V.O., Mikhalkova M.A., Malashin R.O., Lutsiv V.R., Kraeva L.A., Khamdulayeva G.N., Nazarov V.E., Chelibanov V.P.

Статья научная

The purpose of this work is to design a system for microscopic bacterial images classification that can be generalized to new data. In the course of work, a dataset containing 23 bacterial species was collected. We use a strain-wise method for dividing the dataset into training and test sets. Such splitting (in contrast to random division) allows evaluating the performance of classifiers on new strains in the case of intra-species visual variability of bacteria. We propose a “Multi-headed” ResNet (ResNet-MH) for the analysis of microscopic images of bacterial colonies. This approach forces the neural network to analyze features of different resolutions, such as the shape of individual bacterial cells and the shape and number of bacterial clusters during training. Our network achieves the 41.6% accuracy species-wise and 64.06% accuracy genera-wise. The proposed method of dataset splitting guarantees generalization to new unseen strains, whereas random splitting into training and test sets leads to overfitting of the system (accuracy is over 90%). For the 10 visually strain-wise stable species, the accuracy of the proposed system reaches 83.6% species-wise.

Бесплатно

Improving plot-level growing stock volume estimation using machine learning and remote sensing data fusion

Improving plot-level growing stock volume estimation using machine learning and remote sensing data fusion

Mirpulatov I., Kedrov A., Illarionova S.

Статья научная

Forest characteristics estimation is a vital task for ecological monitoring and forest management. Forest owners make decisions based on timber type and its quality. It usually requires field based observations and measurements that is time- and labor-intensive especially in remote and vast areas. Remote sensing technologies aim at solving the challenge of large area monitoring by rapid data acquisition. To automate the data analysis process, machine learning (ML) algorithms are widely applied, particularly in forestry tasks. As ground truth values for ML models training, forest inventory data are usually leveraged. Commonly it involves individual forest stand measurements that are less precise than sample plots. In this study, we delve into ML-based solution development to create spatial-distributed maps with volume stock using sample plot measurements as reference data. The proposed pipeline includes medium-resolution freely available Sentinel-2 data. The experiments are conducted in the Perm region, Russia, and show a high capacity of ML application for forest volume stock estimation based on multispectral satellite observations. Gradient boosting achieves the highest quality with MAPE equal to 30.5%. In future, the proposed solution can be used by forest owners and integrated in advanced systems for ecological monitoring.

Бесплатно

Improving the quality of building space depths maps using multi-area active-pulse television measuring systems in dynamic scenes

Improving the quality of building space depths maps using multi-area active-pulse television measuring systems in dynamic scenes

Zabuga S.A., Kapustin V.V., Musikhin I.D.

Статья научная

The purpose of this work is software implementation of the temporal frame interpolation, the formation of selection criteria and the choice of a suitable neural network model based on the obtained practical data. And also, evaluation of its efficiency for eliminating the interframe shift effect of dynamic objects on the depth maps of multi-area active-pulse television measuring systems in order to improve the accuracy of map building. As initial data for the experiments, static frames were recorded while moving the test rig along the X and Z axes. The static frames are images of the test rig, averaged 100 times, at a distance of 13 meters, which moved along an automated linear guide with a step of 1 mm. As a result of the work, an assessment of the interframe shift effect influence on space depth maps of multi-area active-pulse television measuring systems containing dynamic objects was made. The implementation and testing of the temporal frame interpolation algorithm for suppressing the interframe shift effect of dynamic objects on depth maps was also performed. The algorithm was implemented using Python and the PyCharm IDE with SciPy, NumPy, OpenCV, PyTorch, Threading and other libraries. Numerical values of the RMSE, PSNR, and SSIM metrics were obtained before and after eliminating the effect of interframe shift of dynamic objects on depth maps. The use of the temporal frame interpolation algorithm allows more accurate measurement of distance to moving object in the field of view of multi-area active-pulse television measuring systems.

Бесплатно

Indexing of computer optics in the emerging sources citation index database

Indexing of computer optics in the emerging sources citation index database

Stafeev Sergey S.

Ред. заметка

Inclusion of the journal Computer Optics in the Emerging Sources Citation Index database is described in this editorial.

Бесплатно

Innovative Integration of Residual Networks for Enhanced In-loop Filtering in VVC Using Deep Convolutional Neural Networks

Innovative Integration of Residual Networks for Enhanced In-loop Filtering in VVC Using Deep Convolutional Neural Networks

Ibraheem M.K.I., Dvorkovich A.V., Al-Temimi A.M.S.

Статья научная

This paper explores the integration of Residual Networks (ResNets) into the in-loop filtering (ILF) process of the Versatile Video Coding (VVC) standard, aiming to enhance video compression efficiency and video quality through the application of Deep Convolutional Neural Networks (DCNNs). The study introduces a novel architecture, the Residual Deep Convolutional Neural Network (RDCNN), designed to replace conventional VVC in-loop filtering modules, including Deblocking Filter (DBF), Sample Adaptive Offset (SAO), and Adaptive Loop Filter (ALF). By leveraging the Rate Distortion Optimization (RDO) technique, the RDCNN model is applied to every coding unit (CU) to optimize the balance between video quality and bitrate. The proposed methodology involves offline training with specific parameters using the TensorFlow-GPU platform, followed by feature extraction and prediction of optimal filtering decisions for each video frame during the encoding process. The results demonstrate the effectiveness of the proposed RDCNN in significantly reducing the bitrate while maintaining high visual quality, outperforming existing methods in terms of compression efficiency and peak signal-to-noise ratio (PSNR) values across various video files (YUV color space). Specifically, the RDCNN achieved a YUV PSNR of 41.2 dB and a BD-rate reduction of – 2.43% for the Y component, – 6.96% for the U component, and – 9.43% for the V component. These results underscore the potential of deep learning techniques, particularly ResNets, in addressing the complexities of video compression and enhancing the VVC standard. The evaluation across various YUV video files, including Stefan_cif, Soccer, Mobile, Harbour, Crew, and Bus, revealed consistently higher average YUV PSNR values compared to both VTM 22.2 and other related methods. This indicates not only improved compression efficiency but also enhanced visual quality, crucial for diverse video processing tasks.

Бесплатно

Insight into plasmonics: resurrection of modern-day science (invited)

Insight into plasmonics: resurrection of modern-day science (invited)

Butt M.A.

Статья научная

Plasmonics is a field of research and technology that focuses on the interaction between light and free electrons in a metal structure called plasmon. The study of plasmonics has gained significant attention in recent years due to its potential for several applications and its ability to manipulate light at nanoscale dimensions. Plasmonics enables the control of light at the nanoscale, far beyond the diffraction limit of conventional optics. This allows for the development of new devices and technologies with enhanced performance and functionality. In this paper, recent advances in plasmonics in medicine, agriculture, agriculture, environmental monitoring, lasers and solar energy harvesting are reviewed. Despite these promising prospects, plasmonic devices must overcome obstacles such as significant energy losses, complicated production processes, and the need for better material characteristics. Plasmonics will continue to advance because of ongoing work in nanotechnology, material science, and engineering, which will make it a more significant field with a wide range of usages in the future. In the end, the advantages and the limitations related to the realization of plasmonic devices in the real world are discussed.

Бесплатно

Integrated fiber-based transverse mode converter

Integrated fiber-based transverse mode converter

Gavrilov Andrey Vadimovich, Pavelyev Vladimir Sergeevich

Статья научная

A transverse mode converter based on a binary microrelief implemented directly on the end-face of a few-mode fiber was numerically investigated. The results of numerical simulation demonstrated the converter to form LP-11 and LP-21 modes with high efficiency, providing a more-than 92 % mode purity. Transformations of modes excited by a fiber microbending were also numerically investigated. The excited beams were shown to save their mode purity even in a strong bending as the arising parasitical modes were mostly unguided by the fiber. The resulting beam power and mode content were also demonstrated to depend on the beam and bending mutual orientation for beams with strong rotational symmetry.

Бесплатно

Журнал