Error correcting HTR’ed Byzantine text

The automated correction of errors in the Handwritten Text Recognition (HTR) output can be challenging and is far from solved. To address this challenge, we set up a shared task on AIcrowd that received 271 submissions, of which very few succeed. This paper presents the datasets, the best methods, a...

Descrizione completa

Salvato in:
Dettagli Bibliografici
Autori principali: Pavlopoulos, John (Autore) , Kougia, Vasiliki (Autore) , Platanou, Paraskevi (Autore) , Shabalin, Stepan (Autore) , Liagkou, Konstantina (Autore) , Papadatos, Emmanouil (Autore) , Essler, Holger (Autore) , Camps, Jean-Baptiste (Autore) , Fischer, Franz (Autore)
Natura: Article (Journal) Chapter/Article
Lingua:inglese
Pubblicazione: 15 May 2023
Edizione:Version 1
In: Research Square
Year: 2023, Pages: 1-15
ISSN:2693-5015
DOI:10.21203/rs.3.rs-2921088/v1
Accesso online:Verlag, kostenfrei, Volltext: https://doi.org/10.21203/rs.3.rs-2921088/v1
Verlag, kostenfrei, Volltext: https://www.researchsquare.com/article/rs-2921088/v1
Testo
Note sull'autore:John Pavlopoulos, Vasiliki Kougia, Paraskevi Platanou, Stepan Shabalin, Konstantina Liagkou, Emmanouil Papadatos, Holger Essler, Jean-Baptiste Camps, and Franz Fischer
Descrizione
Riassunto:The automated correction of errors in the Handwritten Text Recognition (HTR) output can be challenging and is far from solved. To address this challenge, we set up a shared task on AIcrowd that received 271 submissions, of which very few succeed. This paper presents the datasets, the best methods, and experimental analysis in error-correcting HTRed manuscripts and papyri in Byzantine Greek, the language that followed Classical and preceded Modern Greek. By using recognised and transcribed data from seven centuries, the two best-performing methods are compared, one based on a neural encoded-decoder architecture and the other based on linguistic knowledge. We show that the recognition error rate can be reduced by both, up to 2.5 points at the level of characters and up to 15 at the level of words, also highlighting the weak and strong points of each.
Descrizione del documento:Gesehen am 03.08.2026
Descrizione fisica:Online Resource
ISSN:2693-5015
DOI:10.21203/rs.3.rs-2921088/v1