Статьи журнала - Компьютерная оптика
Все статьи: 2609
Статья научная
Laser light modes are beams in whose cross-section the complex amplitude is described by eigenfunctions of the operator of light propagation in the waveguide medium. The fundamental properties of modes are their orthogonality and their ability to retain their structure during propagation for example in a lenslike medium, in free space or a Fourier stage. Novel Diffractive Optical Elements (DOEs) of MODAN-type [1] open up new promising potentialities of solving the tasks of generation, transformation, superposition and subsequent separation again of different laser modes. Now we present new results obtained by synthesis and investigation of beams consisting of more than one twodimensional Gaussian laser modes with the same value of propagation constant (invariant multimode beams) formed by DOEs. The exploitation of these phenomena could enhance the fiber optical system transfer capacity without pulse enlargement.
Бесплатно
Inverse scattering transform algorithm for the Manakov system
Статья научная
A numerical algorithm is described for solving the inverse spectral scattering problem associated with the Manakov model of the vector nonlinear Schrödinger equation. This model of wave processes simultaneously considers dispersion, nonlinearity and polarization effects. It is in demand in nonlinear physical optics and is especially perspective for describing optical radiation propagation through the fiber communication lines. In the presented algorithm, the solution to the inverse scattering problem based on the inversion of a set of nested matrices of the discretized system of Gelfand-Levitan-Marchenko integral equations, using a block version of the Levinson-type Toeplitz bordering algorithm. Numerical tests carried out by comparing calculations with known exact analytical solutions confirm the stability and second order of accuracy of the proposed algorithm. We also give an example of the algorithm application to simulate the collision of a differently polarized pair of Manakov optical vector solitons.
Бесплатно
Статья научная
The focusing properties of phase correcting Fresnel lenses with small values of focal length - to - diameter (F/D) and with focal lengths of two wavelengths or less are investigated. For these lenses, the paraxial approximation for the Rayleigh resolution criterion is no longer valid. For Fresnel lenses designed with F/DF ≤ λ, spatial resolutions of less than 0.5λ are possible, which is finer than what can typically be achieved for conventional (paraxial) designs. The spot beams in these cases are not quite axially symmetrical due to the presence of anti-symmetric field components, which vanish for larger values of F/D.
Бесплатно
Laser beam characterization by means of diffractive optical correlation filters
Статья научная
Analyzing of amplitude-phase characteristics of laser beam is topical in experimental physics and in a great number of laser applications, such as, for example, laser material treatment. The task of analyzing the amplitude-phase beam structure may be treated as that of analyzing the modal composition, if this is thought of as both analyzing individual modal powers and intermode phase shifts. In this paper the problem is tackled using a special diffractive optical element (DOE), called MODAN, matched to a group of laser radiation modes and their special combinations. The experimental results reported indicate that such an approach shows promise. Key words: laser beam, Gaussian modes, intermode power distribution, intermode phase shifts.
Бесплатно
Laser generation thresholds of the cholesteric liquid crystal layer
Статья научная
The laser threshold of the eigenmodes in cholesteric liquid crystal (CLC) cells are calculated. The influence of gain on light localization was investigated. The influence of absorption and gain on the light energy density in the CLC layer both at isotropic and anisotropic absorption and gain were investigated for the first time. The calculated threshold values were compared with analytical expression for laser thresholds obtained under the condition ImK<<1/d, where d is the CLC layer thickness, and K is the resonance wave vector.
Бесплатно
Lightweight neural network-based pipeline for barcode image preprocessing
Статья научная
Barcode scanning greatly benefited from deep learning research, as well as the image processing stages included in its workflow. These stages commonly handle pre-processing tasks like localizing barcode symbols in the input image, identifying their type, and normalizing the found regions. They are especially important when there is no a priori knowledge of input image capturing conditions. Thus, a case of multiple barcode recognition within a unique image drastically differs from a single barcode processing in video stream via smartphone. We assess how accuracy of these stages affects the accuracy of the whole barcode scanning as its best and propose a lightweight neural network-based pipeline implementing tasks listed above. To perform this assessment and evaluate the performance of the proposed pipeline elements, we conduct a series of experiments using the set of popular open source scanners, including OpenCV, WeChat, ZBar, ZXing and ZXing-cpp over the SE-barcode and Dubska datasets. These experiments reveal how the proposed pipeline can be configured for optimum speed and accuracy performance depending on the objective and the chosen scanner.
Бесплатно
Localization of mobile robot in prior 3D lidar maps using stereo image sequence
Статья научная
The paper studies the real-time stereo image-based localization of a vehicle in a prior 3D LiDAR map. A novel localization approach for mobile ground robot, which successfully combines conventional computer vision techniques, neural network based image analysis and numerical optimization, is proposed. It includes matching a noisy depth image and visible point cloud based on the modified Nelder-Mead optimization method. Deep neural network for image semantic segmentation is used to eliminate dynamic obstacles. The visible point cloud is extracted using a 3D mesh map representation. The proposed approach is evaluated on the KITTI dataset and a custom dataset collected from a ClearPath Husky mobile robot. It shows a stable absolute translation error of about 0.11 – 0.13 m. and a rotation error of 0.42 – 0.62 deg. The standard deviation of the obtained absolute metrics for our method is the smallest among other state-of-the-art approaches. Thus, our approach provides more stability in the estimated pose. It is achieved primarily through the use of multiple data frames during the optimization step and dynamic obstacles elimination on depth image. The method’s performance is demonstrated on different hardware platforms, including energy-efficient Nvidia Jetson Xavier AGX. With parallel code implementation, we achieve an input stereo image processing speed of 14 frames per second on Xavier AGX.
Бесплатно
Losses and orbital part of the poynting vector of air-core modes in hollow-core fibers
Статья научная
In our earlier works, we investigated a relationship between the formation of vortices in the transverse component of the Poynting vector of core modes and the regimes of strong localization of these modes in solid core micro-structured optical fibers. In this paper, we consider the behavior of the orbital part of the Poynting vector of fundamental and high-order modes in hollow-core fibers, and make comparisons with similar fundamental core mode behavior in solid core micro- structured optical fibers. We then demonstrated the impact of the “negative” curvature of the core-cladding boundary of a hollow-core fiber on the behavior of the orbital part of the Poynting vector of the air-core modes.
Бесплатно
MIDV-2020: a comprehensive benchmark dataset for identity document analysis
Статья научная
Identity documents recognition is an important sub-field of document analysis, which deals with tasks of robust document detection, type identification, text fields recognition, as well as identity fraud prevention and document authenticity validation given photos, scans, or video frames of an identity document capture. Significant amount of research has been published on this topic in recent years, however a chief difficulty for such research is scarcity of datasets, due to the subject matter being protected by security requirements. A few datasets of identity documents which are available lack diversity of document types, capturing conditions, or variability of document field values. In this paper, we present a dataset MIDV-2020 which consists of 1000 video clips, 2000 scanned images, and 1000 photos of 1000 unique mock identity documents, each with unique text field values and unique artificially generated faces, with rich annotation. The dataset contains 72409 annotated images in total, making it the largest publicly available identity document dataset to the date of publication. We describe the structure of the dataset, its content and annotations, and present baseline experimental results to serve as a basis for future research. For the task of document location and identification content-independent, feature-based, and semantic segmentation-based methods were evaluated. For the task of document text field recognition, the Tesseract system was evaluated on field and character levels with grouping by field alphabets and document types. For the task of face detection, the performance of Multi Task Cascaded Convolutional Neural Networks-based method was evaluated separately for different types of image input modes. The baseline evaluations show that the existing methods of identity document analysis have a lot of room for improvement given modern challenges. We believe that the proposed dataset will prove invaluable for advancement of the field of document analysis and recognition.
Бесплатно
MIDV-500: a dataset for identity document analysis and recognition on mobile devices in video stream
Статья научная
A lot of research has been devoted to identity documents analysis and recognition on mobile devices. However, no publicly available datasets designed for this particular problem currently exist. There are a few datasets which are useful for associated subtasks but in order to facilitate a more comprehensive scientific and technical approach to identity document recognition more specialized datasets are required. In this paper we present a Mobile Identity Document Video dataset (MIDV-500) consisting of 500 video clips for 50 different identity document types with ground truth which allows to perform research in a wide scope of document analysis problems. The paper presents characteristics of the dataset and evaluation results for existing methods of face detection, text line recognition, and document fields data extraction. Since an important feature of identity documents is their sensitiveness as they contain personal data, all source document images used in MIDV-500 are either in public domain or distributed under public copyright licenses. The main goal of this paper is to present a dataset. However, in addition and as a baseline, we present evaluation results for existing methods for face detection, text line recognition, and document data extraction, using the presented dataset.
Бесплатно
MIDV-DM: A Document-Oriented Dataset for Image Manipulation Detection and Localization
Статья научная
As the scope of application of document recognition systems in business processes increases, so does the number of attacks on these systems. One form of such attacks could involve software for manipulating a digital image of a document. The development of methods for image manipulation detection and localization is complicated with the fact that available datasets neither contain images of documents nor lack diversity in capture conditions and document types. Furthermore, these datasets do not cover the range of possible kinds of manipulations that occur under natural conditions. In this paper, we introduce MIDV-DM – a publicly available benchmark designed for the development and testing of methods aimed at detecting and localizing manipulations in identity document images. It contains images subjected to eight types of manipulations, which we have conceptually categorized based on our analysis of over 2000 real-world fraud attempts. In total, MIDV-DM contains 1000 original document images from the public MIDV-2020 dataset and 8000 automatically created manipulated images based on them, along with the ground truth masks and annotations. The paper also describes the process of obtaining baseline quality based on the IML-ViT model. The authors believe that MIDV-DM will open new opportunities for researchers to advance technologies for document image authenticity analysis.
Бесплатно
MIMO communication system capacity in random visible light channel
Статья научная
Being a promising one, optical information transmission standard expands capabilities of communication systems in the conditions of heavy frequency band load. Optical communication system efficiency in a room can be improved by multi-antenna systems. The aim of this paper is a theoretical study of MIMO Li-Fi communication system capacity. The calculation of ergodic capacity is performed for MIMO optical communication system in terms of various scenarios of light propagation. Receiving and transmitting system is modeled in the form of receivers and transmitters randomly placed in a room with randomly oriented light-emitting and photo diodes. A matrix of channel parameters is modeled using corresponding probability density functions and additive Gaussian noise at receiver inputs. The paper also considers various scenarios of optical signal propagation and their influence on optical channel capacity. The comparison of various methods of power distribution between original modes of MIMO optical communication system as well as their influence on capacity is carried out. Optimal power distribution between MIMO system eigenmodes is determined by maximum capacity criterion.
Бесплатно
Статья научная
An automatic speech recognition system has the possibility of enhancing the standard of living for persons with disabilities by solving issues such as dysarthria, stuttering, and other speech defects. In this paper, we introduce a voice assistant using hyperkinetic dysarthria (HD) defect speeches. It contains the data preprocessing steps and the development of a novel convolutional recurrent network (CRN) model that is built depending on the convolutional neural networks and recurrent neural networks. We implemented data preprocessing methods, including filtering, down-sampling, and splitting, to prevent overfitting and decrease processing power as well as time. In addition, the technique of Mel Frequency Cepstral Coefficients (MFCC) has been utilized to extract speech characteristics. The proposed model is trained to recognize HD speech disorders using a dataset including 2000 Russian speeches. The experimental results demonstrate that the proposed method obtains a character error rate (CER) of 14.76 %. It indicates that approximately 85 % of characters are able to correctly recognize on the test dataset. We have created a telegram bot that utilizes our trained model to help people with hyperkinetic dysarthria speech disorder. This bot is capable of providing assistance independently, without the need for any third-party assistance.
Бесплатно
Many heads but one brain: fusionbrain - a single multimodal multitask architecture and a competition
Статья научная
Supporting the current trend in the AI community, we present the AI Journey 2021 Challenge called FusionBrain, the first competition which is targeted to make a universal architecture which could process different modalities (in this case, images, texts, and code) and solve multiple tasks for vision and language. The FusionBrain Challenge combines the following specific tasks: Code2code Translation, Handwritten Text recognition, Zero-shot Object Detection, and Visual Question Answering. We have created datasets for each task to test the participants' submissions on it. Moreover, we have collected and made publicly available a new handwritten dataset in both English and Russian, which consists of 94,128 pairs of images and texts. We also propose a multimodal and multitask architecture - a baseline solution, in the centre of which is a frozen foundation model and which has been trained in Fusion mode along with Single-task mode. The proposed Fusion approach proves to be competitive and more energy-efficient compared to the task-specific one.
Бесплатно
Many-parameter m-complementary Golay sequences and transforms
Статья научная
In this paper, we develop the family of Golay–Rudin–Shapiro (GRS) m-complementary many-parameter sequences and many-parameter Golay transforms. The approach is based on a new gen-eralized iteration generating construction, associated with n unitary many-parameter transforms and n arbitrary groups of given fixed order. We are going to use multi-parameter Golay transform in Intelligent-OFDM-TCS instead of discrete Fourier transform in order to find out optimal values of parameters optimized PARP, BER, SER, anti-eavesdropping and anti-jamming effects.
Бесплатно
Mapping and evaluating urban density patterns in Moscow, Russia
Статья научная
The defense of the notion of ‘compact city’ as a strategy to reduce urban sprawl to support greater utilization of existing infrastructure and services in more compact areas and to improve the connectivity of employment hubs is actively discussed in urban research. Using the urban residential density as a surrogate measure for urban compactness, this paper empirically examines a cadaster database that contains details of every property with a view of capturing changes in urban residential density patterns across Moscow using geospatial techniques. The policy of densification in chase of a more compact city has produced mixed results. Findings of this study signal that the urban densities across the buffer zones around Moscow city are significantly different. The Landsat images from 1995, 2005 and 2016 are classified based on the maximum likelihood to expand the land use/cover maps and identify the land cover. Then, the area coverage for all the land use/cover types at different points in time is combined with the distance from the city center. After that, urbanization densities from the city center toward the outskirts for every 1-km distance from 1 to 60 km are calculated. The city density on the distance of 1 to 35 km is found to be very high in the years 1995 to 2016. As usual, the population, traffic conditions, industrialization and government policy are the major factors that influenced the urban expansion.
Бесплатно
Статья научная
The relaxation of a three-level atom interacting with a photon heat bath and an external stochastic field is investigated. For the reduced density matrix, a master equation averaged over stochastic process realizations is derived. An exact solution is obtained and the radiation line shapes are calculated.
Бесплатно
Статья научная
The paper studies entangled states of two qubits interacting with each other and with an electromagnetic field. The state of the qubits is determined by a statistical density matrix. The degree of entanglement of the state is characterized by the Peres-Gorodeckii (PG) parameter. The statistical density matrix and its evolution are determined in the energy representation within the framework of the path integral formalism. The obtained equations determine the dependence of the PG parameter on the parameters of qubit dipole-dipole interaction and the acting electromagnetic field. The results of numerical calculations are presented in graphs for the PG parameter. It is shown that it is possible to choose parameters corresponding to qubit states with a high degree of entanglement (0.99).
Бесплатно
Method for removing haze from images, captured under a wide range of lighting conditions
Статья научная
The presence of haze on images degrades the quality of perception and automatic analysis of scenes. One of the most popular methods of haze removal is the dark channel prior method, which is based on the Koschmieder atmospheric scattering model. However, its underlying assumptions are not met for nighttime, since localized light sources make a significant, if not the main, contribution to lighting. We propose here to use the degree of belonging of an image element to a localized light source, determined based on a one-class classifier, as a value that characterizes the confidence of the corresponding element of the estimated transmission map during its rectifi-cation based on the gamma-normal model, which makes it possible to increase the accuracy of dehazing when processing images, captured in low-light or nighttime conditions.
Бесплатно