Формалізація та первинна експериментальна перевірка адаптивного підходу до вибору OCR послідовності для розпізнавання тексту на зображеннях
The paper addresses the problem of selecting an appropriate text recognitionpipeline for images by considering image preprocessing methods and the specificfeatures of modern optical c haracter recognition (OCR) models. The relevance of thestudy is determined by the fact that OCR quality depends not...
Gespeichert in:
| Datum: | 2026 |
|---|---|
| Hauptverfasser: | , , |
| Format: | Artikel |
| Sprache: | Ukrainisch |
| Veröffentlicht: |
Інститут прикладних проблем механіки і математики ім. Я. С. Підстригача НАН України
2026
|
| Schlagworte: | |
| Online Zugang: | https://www.fmmit.lviv.ua/index.php/fmmit/article/view/442 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| Назва журналу: | Physico-mathematical modeling and informational technologies |
| Завантажити файл: | |
Institution
Physico-mathematical modeling and informational technologies| Zusammenfassung: | The paper addresses the problem of selecting an appropriate text recognitionpipeline for images by considering image preprocessing methods and the specificfeatures of modern optical c haracter recognition (OCR) models. The relevance of thestudy is determined by the fact that OCR quality depends not only on the selectedrecognition model but also on the characteristics of the input image, including noise,contrast, illumination, resolut ion, text skew, and background complexity.The aim of the paper is to formalize an adaptive approach to OCR pipelineselection and to perform its initial experimental evaluation. The proposed approach isbased on generating several preprocessed versions of the same input image, applyingOCR models to each version, obtaining recognized text, text region coordinates,confidence scores, and processing time, and then evaluating the obtained results usinga multi criteria quality score. The study considers the following OCR tools: Tesseract,EasyOCR, PaddleOCR, RapidOCR, and Amazon Textract. The preprocessingconfigurations include the original image without preprocessing, grayscaleconversion, contrast enhancement, denoising with scaling, and Otsu binarization. Thequality assessment is based on Character Error Rate (CER), Word Error Rate (WER),processing time, model confidence score, and fuzzy matching score. The experimentalpart is considered as an initial experimental evaluation rather than a full scalesta tistical comparison of OCR models. Its purpose is to verify the logic of the proposedmethodology, identify the main parameters that should be fixed in further experiments,and prepare a basis for extended research on a larger dataset of images of differentquality. The obtained results demonstrate that the quality of OCR recognition may varydepending on the selected combination of preprocessing method and OCR model.However, the results should be interpreted as preliminary and cannot be considered afinal ranking of OCR models. The practical value of the proposed approach lies in itspotential use as a methodological basis for building OCR pipelines in automateddocument processing systems, digital archives, electronic document managementsystems, information retrieval systems, and applications for text recognition fromimages. |
|---|---|
| DOI: | 10.15407/fmmit2026.42.158 |