Bitte benutzen Sie diese Kennung, um auf die Ressource zu verweisen: https://dspace.chmnu.edu.ua/jspui/handle/123456789/2916
Titel: Analysis of methods and algorithms for processing unstructured text data based on JSON technology
Autoren: Kucherenko, Y.
Kulakovska, I.
Stichwörter: automated system
crowding
CSV
ELT
ETL
intelligent system
JSON
unstructured data
validation
Erscheinungsdatum: 2024
Herausgeber: Technology Center
Zusammenfassung: The object of research is the process of automating systems for structuring data from several sources. The subject of the research is methods and algorithms for implementing a complete system for automated and parallel processing, validation and structuring of data. One of the most problematic areas is the merging of databases with different structures and several common fields into a generalized structure. The research was aimed at developing a system to increase the efficiency of automation of big data processing. As a result of the work, optimization methods were studied, the influence of their internal parameters on the operation of algorithms was analyzed, their main advantages and disadvantages were determined, and software was developed in which the corresponding methods were implemented. An algorithm for structuring data before processing has been obtained. Data structuring is achieved by performing the «mapping» operation. Mapping can take place by indexes of already cleaned data or using a defined dictionary with a given set of keys, which allows not to care about the sequence of storing values and their possible shift. The practical significance of the developed system lies in the improvement of methods of collecting and processing information for the purpose of its further validation, cleaning and accumulation in the following categories: geographic addresses and geo-coordinates, validation and automated addition of a mobile phone number to the international format, processing of car numbers (in modern and outdated format), VIN code of the engine and car brand, validation of urls of social networks, passport data and processing of personal data. Compared to similar methods for processing large volumes of data, the possibility of splitting the input file or stream into separate parts was used, the cleaned data from which is combined at the end of the system operation. Thanks to this, it is possible to process data whose size exceeds the available volume of the device’s RAM, and the method of working with loosely structured text files in CSV format has been improved.
Beschreibung: Kucherenko, Y., & Kulakovska, I. (2024). Analysis of methods and algorithms for processing unstructured text data based on JSON technology. Technology Audit and Production Reserves, 3 (2), 10–18. DOI: 10.15587/2706-5448.2024.306435
URI: https://www.scopus.com/pages/publications/105012466726
https://journals.uran.ua/tarp/article/view/306435
https://dspace.chmnu.edu.ua/jspui/handle/123456789/2916
ISSN: 26649969
Enthalten in den Sammlungen:Публікації науково-педагогічних працівників ЧНУ імені Петра Могили у БД Scopus

Dateien zu dieser Ressource:
Datei Beschreibung GrößeFormat 
ANALYSIS OF METHODS ANDALGORITHMS FOR PROCESSING.pdf564.91 kBAdobe PDFÖffnen/Anzeigen


Alle Ressourcen in diesem Repository sind urheberrechtlich geschützt, soweit nicht anderweitig angezeigt.