Численные методы и анализ данных. Рубрика в журнале - Компьютерная оптика
Many heads but one brain: fusionbrain - a single multimodal multitask architecture and a competition
Статья научная
Supporting the current trend in the AI community, we present the AI Journey 2021 Challenge called FusionBrain, the first competition which is targeted to make a universal architecture which could process different modalities (in this case, images, texts, and code) and solve multiple tasks for vision and language. The FusionBrain Challenge combines the following specific tasks: Code2code Translation, Handwritten Text recognition, Zero-shot Object Detection, and Visual Question Answering. We have created datasets for each task to test the participants' submissions on it. Moreover, we have collected and made publicly available a new handwritten dataset in both English and Russian, which consists of 94,128 pairs of images and texts. We also propose a multimodal and multitask architecture - a baseline solution, in the centre of which is a frozen foundation model and which has been trained in Fusion mode along with Single-task mode. The proposed Fusion approach proves to be competitive and more energy-efficient compared to the task-specific one.
Бесплатно
Many-parameter m-complementary Golay sequences and transforms
Статья научная
In this paper, we develop the family of Golay–Rudin–Shapiro (GRS) m-complementary many-parameter sequences and many-parameter Golay transforms. The approach is based on a new gen-eralized iteration generating construction, associated with n unitary many-parameter transforms and n arbitrary groups of given fixed order. We are going to use multi-parameter Golay transform in Intelligent-OFDM-TCS instead of discrete Fourier transform in order to find out optimal values of parameters optimized PARP, BER, SER, anti-eavesdropping and anti-jamming effects.
Бесплатно
Статья научная
An original approach to solving difficult time-consuming problems of registration and analysis of random point images is described. The approach is based on the development and application of high-performance specialized computer algebra systems. Three software packages have been created specifically for carrying out equivalent analytical transformations on a computer. The first software system is designed to calculate formulas describing the volumes of convex polyhedra with parametrically specified boundaries in n -dimensional space. The second system is based on the calculation of multidimensional integral expressions by the method of cyclic differentiation of the integral with respect to the parameter. The third system is based on the accelerated implementation of complex combinatorial-recursive transformations on a computer. Another distinctive feature of the work is the extension of the classical Catalan numbers to the multidimensional case (they were required to solve a number of intermediate probabilistic-combinatorial problems). The implementation of the above computer algebra software systems on a multi-core cluster of Novosibirsk State University, together with the direct use of the explicit form of generalized Catalan numbers, allowed the authors to obtain several new previously unknown probabilistic formulas and dependencies required for solving problems in the field of analysis of random point images.
Бесплатно
Статья научная
Beam divergence is one of the instrument resolution parameters in neutron computed tomography. In pinhole geometry, due to the finite size of the source, geometric unsharpness affects the transmission images and therefore influences the reconstructed data. In this paper, we propose an approach for deterministic simulation of this effect for a voxelized 3D object. The idea behind the proposed approach is to use multiple point sources at a pinhole position and collect transmission images from each of them. The implementation was done using the ASTRA toolbox by calculating cone beam projections from each point source. This approach was applied to a porous phantom. Artifacts associated with beam divergence were identified in the reconstructed data. The influence of beam divergence on the segmentation of pores by binarization of the reconstructed data has been considered.
Бесплатно
Multigrammatical modelling of neural networks
Статья научная
This paper is dedicated to the proposed techniques of modelling artificial neural networks (NNs) by application of the multigrammatical framework. Multigrammatical representations of feed-forward and recurrent NNs are described. Application of multiset metagrammars to modelling deep learning of NNs of the aforementioned classes is considered. Possible developments of the announced approach are discussed.
Бесплатно
Network community partition based on intelligent clustering algorithm
Статья научная
The division of network community is an important part of network research. Based on the clustering algorithm, this study analyzed the partition method of network community. Firstly, the classic Louvain clustering algorithm was introduced, and then it was improved based on the node similarity to get better partition results. Finally, experiments were carried out on the random network and the real network. The results showed that the improved clustering algorithm was faster than GN and KL algorithms, the community had larger modularity, and the purity was closer to 1. The experimental results show the effectiveness of the proposed method and make some contributions to the reliable community division.
Бесплатно
Neural network task specialization via domain constraining
Статья научная
This paper introduces a concept of neural network specialization via task-specific domain constraining, aimed at enhancing network performance on data subspace in which the network operates. The study presents experiments on training specialists for image classification and object detection tasks. The results demonstrate that specialization can enhance a generalist's accuracy even without additional data or changing training regimes -- solely by constraining class label space in which the network performs. Theoretical and experimental analyses indicate that effective specialization requires modifying traditional fine-tuning methods and constraining data space to semantically coherent subsets. The specialist extraction phase before tuning the network is proposed for maximal performance gains. We also provide analysis of the evolution of the feature space during specialization. This study paves way to future research for developing more advanced dynamically configurable image analysis systems, where computations depend on the specific input. Additionally, the proposed methods can help improve system performance in scenarios where certain data domains should be excluded from consideration of the generalist network.
Бесплатно
Статья научная
This research addresses the problem of automatic object detection in images under limited-visibility conditions, where objects are partially occluded, the background is complex, and lighting and viewpoints vary widely. The proposed approach combines pretraining on a programmatically generated synthetic dataset of 18,000 images - produced using the Visualization Toolkit (VTK) library - with fine-tuning on a compact real-image dataset of 2,000 annotated photographs (500 per class). Six deep neural network architectures - Faster R-CNN ResNet-50 FPN, SSD MobileNet V3, YOLOv11n, EfficientDet-D7, DETR-DC5, and CenterNet- were evaluated across three training regimes: synthetic-only, real-only, and combined (90% synthetic / 10% real). Hybrid training yielded the most substantial improvements: YOLOv11n achieved mAP@0.5 = 0.91 and mAP@0.75 = 0.86 (Precision = 0.89, Recall = 0.90, F1 = 0.89, 82 FPS), compared to 0.79 (synthetic-only) and 0.78 (real-only), representing a gain of up to +15 percentage points in mAP@0.5. EfficientDet-D7 reached mAP@0.5 = 0.87 and mAP@0.75 = 0.81, while CenterNet achieved mAP@0.5 = 0.88 at 35 FPS. Robustness analysis under simulated occlusion demonstrated that hybrid-trained models maintain reliable detection even under severe conditions: YOLOv11n retained mAP@0.5 = 0.78 at 50% occlusion and mAP@0.5 = 0.65 at 25% object visibility, while the degradation in mAP under 75% occlusion did not exceed 20% of the baseline level. The results confirm the viability of synthetic data as a standalone pretraining resource and validate the proposed hybrid pipeline for applications in autonomous driving, video surveillance, and industrial inspection.
Бесплатно
Point cloud registration based on global compatibility feature
Статья научная
In this paper, we present a point cloud registration method that utilizes a global point cloud compatibility feature. We introduce an evaluation technique called global compatibility, which helps distinguish between correct and incorrect feature point pairs by calculating the corresponding compatibility weights. To begin, we employ a spectral matching technique to select reliable seed points, allowing us to construct a consistent point set in the vicinity of these seed points. We then design a consistent filter to eliminate outliers from the obtained set. Our approach includes proposing optimal weight matching based on the characteristics of each compatible point set, alongside spectral matching for decomposing the constructed multiple compatible point sets. We assign smaller weights for points affected by larger noise, which aids in generating the corresponding rigid transformation. Ultimately, we select the best transformation as the final result. Notably, our method does not require retrieving all features from the entire point set, and it effectively removes discrete points, thereby constructing a more efficient and robust consistent point set. Experimental results demonstrate that our method performs very well on both indoor and outdoor datasets, as well as on datasets with low overlap.
Бесплатно
Recognition of biosignals with nonlinear properties by approximate entropy parameters
Статья научная
More and more attention is being paid to the development of methods for the objective analysis of biosignals for computer medical systems. The search for new non-standard methods is aimed at improving the reliability of diagnostics and expanding the areas of their practical application. In this paper, methods for recognizing biomedical signals by the degree of severity of their nonlinear components are considered. An approach based on the use of approximate entropy closely related to Kolmogorov entropy ( K -entropy) is used. Its parameters can be used to detect dynamic irregularities associated with nonlinear properties of signals. The algorithm for calculating this characteristic is considered in detail. Based on model experiments, its main properties are analyzed. It is shown that the entropy of a finite sequence, calculated in accordance with a multistep procedure, can give an erroneous estimate of the degree of regularity of the signal. A procedure for correcting the approximate entropy is proposed, which expands the area of analysis of this function for estimating nonlinearity. It has been established that the transition to adjusted entropy makes it possible to increase the reliability of the detection of chaotic components. A set of entropy parameters is proposed for constructing recognition procedures. Examples of solving the problems of detecting atrial fibrillation by the parameters of the nonlinearity of the rhythmogram, as well as assessing the depth of anesthesia by the electroencephalogram (EEG) are given. Experiments conducted on real recordings of electrocardiogram (ECG) and EEG signals have shown the high efficiency of the proposed algorithms. The proposed methods and algorithms can be used in the development of systems for monitoring ECG of cardiological patients, as well as monitoring the depth of anesthesia by EEG during surgical operations.
Бесплатно
Renewed empirical formulas of Weibull distribution parameters estimates
Статья научная
The empirical formulas proposed in the literature for estimating the parameters of a two-parameter Weibull distribution, obtained using the equations of the moment method, are considered. It is noted that the formulas used to estimate the shape parameter take the form of various types of dependences on the coefficient of variation of the distribution. By modeling the empirical formulas selected for analysis, a comparative analysis of their errors relative to accurate numerical solutions of the moment method equations was carried out. A renewed empirical formula for the shape parameter is proposed. An approach to estimating the scale parameter is proposed, in which the empirical formula of the latter is reduced to the product of the standard deviation of the distribution by a power function of the coefficient of variation with an exponent equal to – 1.027. The results of applying the updated empirical formulas to numerical data obtained by modeling a random sample from the Weibull distribution are presented. It is shown that the accuracy of the proposed empirical formulas is quite high.
Бесплатно
Research on an image color restoration method for old art films
Статья научная
The preservation and restoration of old art films have certain practical value. Focusing on the color restoration of old art films, this paper introduced the same mapping loss on the basis of a cycle-consistent generative adversarial network, which enables the algorithm to capture more image details and achieve better transfer results. The old art films Roman Holiday and Tracks in the Snowy Forest were used as experimental data to verify the color restoration effect of the improved cycle-consistent generative adversarial network algorithm. It was found that compared with the generative adversarial network and cycle-consistent generative adversarial network algorithms, the improved cycle-consistent generative adversarial network algorithm was superior. It achieved a peak signal-to-noise ratio of 26.874, a structural similarity index measure of 0.665, a learned perceptual image patch similarity of 0.212, and a Frechet inception distance of 117.652 for Roman Holiday. Moreover, it achieved a peak signal-to-noise ratio of 22.794, a structural similarity index measure of 0.585, a learned perceptual image patch similarity of 0.247, and a Frechet inception distance of 119.265 for Tracks in the Snowy Forest. It also achieved better results in comparison with existing image color restoration methods. The results demonstrate the usability of the improved cycle-consistent generative adversarial network algorithm in color restoration of old art films, which can be applied in practice.
Бесплатно
Research on robot motion control and trajectory tracking based on agricultural seeding
Статья научная
With the development of science and technology, agricultural production has been gradually industrialized, and the use of robots instead of humans for seeding is one of the agricultural industrializations. This paper studied the seeding path planning and path tracking algorithms of the seeding robot, carried out experiments, and compared the improved proportion, integral, differential (PID) algorithm with the traditional PID control algorithm. The results demonstrated that both the improved and non-improved control algorithms played a good role in tracking on the straight path, but the improved control algorithm had a better tracking effect on the turning path; the displacement deviation and angle deviation of the tracking trajectory of the improved PID algorithm were reduced faster and more stable than the traditional PID algorithm; the tracking trajectory was shorter and the operation time of the robot was less under the improved PID algorithm than the traditional one.
Бесплатно
Security detection of network intrusion: application of cluster analysis method
Статья научная
In order to resist network malicious attacks, this paper briefly introduced the network intrusion detection model and K-means clustering analysis algorithm, improved them, and made a simulation analysis on two clustering analysis algorithms on MATLAB software. The results showed that the improved K-means algorithm could achieve central convergence faster in training, and the mean square deviation of clustering center was smaller than the traditional one in convergence. In the detection of normal and abnormal data, the improved K-means algorithm had higher accuracy and lower false alarm rate and missing report rate. In summary, the improved K-means algorithm can be applied to network intrusion detection.
Бесплатно
Статья научная
The development and use of a service-oriented application for processing and analyzing meteorological data for solving environmental monitoring problems for the Baikal Natural Territory are considered. The application uses the Web Processing Service standard for creating services. This ensures the ability to work with spatio-temporal data which is typically used in environmental monitoring. As part of the research, algorithms for data normalization, detection of individual and contextual anomalies, and correction of missing and anomalous values, were developed. A distinctive feature of the developed algorithms is the use of a machine learning model based on decision trees to analyze data series when detecting missing values, and the analysis of temporal and seasonal patterns when identifying individual and contextual anomalies using specialized Python programming libraries. The application, utilizing web services, represents an effective tool for comprehensive work with meteorological data for environmental monitoring. The application's use in the study of an autonomous energy complex facilitated the selection of its rational structure and operating parameters to meet electricity demand while maintaining ecological sustainability and resource conservation.
Бесплатно
Study on the planning of rural land spatial utilization by improved particle swarm optimization
Статья научная
The planning of rural land space utilization is a very important problem. In this paper, the objective function of rural land use planning was analyzed firstly, and then the improved particle swarm optimization (IPSO) algorithm was obtained by improving the inertia weight for solution. The results showed that the land space use in the study area was more reasonable after the planning based on the IPSO algorithm, the forest land and construction land increased, the area of grassland, cultivated land and water area reduced appropriately, the aggregation degree of all types of land improved, and the space distribution was more planned, which was more conducive to production activities. The analysis results verify the effectiveness of the IPSO method in land space use planning, which can improve the efficiency and benefit of land space use, and it can be popularized in practical application.
Бесплатно
The basic assembly of skeletal models in the fall detection problem
Статья научная
The paper considers the appliance of the featureless approach to the human activity recognition problem, which exclude the direct anthropomorphic and visual characteristics of human figure from further analysis and thus increase the privacy of the monitoring system. A generalized pairwise comparison function of two human skeletal models, invariant to the sensor type, is used to project the object of interest to the secondary feature space, formed by the basic assembly of skeletons. A sequence of such projections in time forms an activity map, which allows an application of deep learning methods based on convolution neural networks for activity recognition. The proper ordering of skeletal models in a basic assembly plays an important role in secondary space design. The study of ordering of the basic assembly by the shortest unclosed path algorithm and correspondent activity maps for video streams from the TST Fall Detection v2 database are presented.
Бесплатно
The optimization of automated goods dynamic allocation and warehousing model
Статья научная
In the development of modern logistics, the role of automated cargo warehousing is gradually reflected, which is essential for the automatic distribution of goods. This paper briefly introduced the automatic location allocation model and the particle swarm optimization (PSO) algorithm used to optimize the model. At the same time, it introduced the concept of genetic operator and multi-group co-evolution to improve the algorithm, and then the simulation analysis of standard PSO and improved PSO was performed on MATLAB software. The results showed that the improved PSO iterated fewer times and get better solution sets; compared with the manual allocation scheme, the improved PSO calculation reduced more warehousing time, lowered more center of gravity height, and improved shelf stability. In summary, the improved PSO algorithm can effectively optimize the automated goods dynamic allocation and warehousing model.
Бесплатно
Towards monitored tomographic reconstruction: algorithm-dependence and convergence
Статья научная
The monitored tomographic reconstruction (MTR) with optimized photon flux technique is a pioneering method for X-ray computed tomography (XCT) that reduces the time for data acquisition and the radiation dose. The capturing of the projections in the MTR technique is guided by a scanning protocol built on similar experiments to reach the predetermined quality of the reconstruction. This method allows achieving a similar average reconstruction quality as in ordinary tomography while using lower mean numbers of projections. In this paper, we, for the first time, systematically study the MTR technique under several conditions: reconstruction algorithm (FBP, SIRT, SIRT-TV, and others), type of tomography setup (micro-XCT and nano-XCT), and objects with different morphology. It was shown that a mean dose reduction for reconstruction with a given quality only slightlyvaries with choice of reconstruction algorithm, and reach up to 12.5 % in case of micro-XCT and 8.5 % for nano-XCT. The obtained results allow to conclude that the monitored tomographic reconstruction approach can be universally combined with an algorithm of choice to perform a controlled trade-off between radiation dose and image quality. Validation of the protocol on independent common ground truth demonstrated a good convergence of all reconstruction algorithms within the MTR protocol.
Бесплатно
Статья научная
The three-dimensional perception applications have been growing since Light Detection and Ranging devices have become more affordable. On those applications, the navigation and collision avoidance systems stand out for their importance in autonomous vehicles, which are drawing an appreciable amount of attention these days. The on-road object classification task on three-dimensional information is a solid base for an autonomous vehicle perception system, where the analysis of the captured information has some factors that make this task challenging. On these applications, objects are represented only on one side, its shapes are highly variable and occlusions are commonly presented. But the highest challenge comes with the low resolution, which leads to a significant performance dropping on classification methods. While most of the classification architectures tend to get bigger to obtain deeper features, we explore the opposite side contributing to the implementation of low-cost mobile platforms that could use low-resolution detection and ranging devices. In this paper, we propose an approach for on-road objects classification on extremely low-resolution conditions. It uses directly three-dimensional point clouds as sequences on a transformer-convolutional architecture that could be useful on embedded devices. Our proposal shows an accuracy that reaches the 89.74 % tested on objects represented with only 16 points extracted from the Waymo, Lyft’s level 5 and Kitti datasets. It reaches a real time implementation (22 Hz) in a single core processor of 2.3 Ghz.
Бесплатно