Error correcting HTR’ed Byzantine text

The automated correction of errors in the Handwritten Text Recognition (HTR) output can be challenging and is far from solved. To address this challenge, we set up a shared task on AIcrowd that received 271 submissions, of which very few succeed. This paper presents the datasets, the best methods, a...

Description complète

Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pavlopoulos, John (Auteur) , Kougia, Vasiliki (Auteur) , Platanou, Paraskevi (Auteur) , Shabalin, Stepan (Auteur) , Liagkou, Konstantina (Auteur) , Papadatos, Emmanouil (Auteur) , Essler, Holger (Auteur) , Camps, Jean-Baptiste (Auteur) , Fischer, Franz (Auteur)
Format: Article (Journal) Chapter/Article
Langue:anglais
Publié: 15 May 2023
Édition:Version 1
In: Research Square
Year: 2023, Pages: 1-15
ISSN:2693-5015
DOI:10.21203/rs.3.rs-2921088/v1
Accès en ligne:Verlag, kostenfrei, Volltext: https://doi.org/10.21203/rs.3.rs-2921088/v1
Verlag, kostenfrei, Volltext: https://www.researchsquare.com/article/rs-2921088/v1
Accéder au texte intégral
Notes sur l'auteur:John Pavlopoulos, Vasiliki Kougia, Paraskevi Platanou, Stepan Shabalin, Konstantina Liagkou, Emmanouil Papadatos, Holger Essler, Jean-Baptiste Camps, and Franz Fischer
Description
Résumé:The automated correction of errors in the Handwritten Text Recognition (HTR) output can be challenging and is far from solved. To address this challenge, we set up a shared task on AIcrowd that received 271 submissions, of which very few succeed. This paper presents the datasets, the best methods, and experimental analysis in error-correcting HTRed manuscripts and papyri in Byzantine Greek, the language that followed Classical and preceded Modern Greek. By using recognised and transcribed data from seven centuries, the two best-performing methods are compared, one based on a neural encoded-decoder architecture and the other based on linguistic knowledge. We show that the recognition error rate can be reduced by both, up to 2.5 points at the level of characters and up to 15 at the level of words, also highlighting the weak and strong points of each.
Description:Gesehen am 03.08.2026
Description matérielle:Online Resource
ISSN:2693-5015
DOI:10.21203/rs.3.rs-2921088/v1