How ‘Pagan’ Is My Text? Information Extraction from Untranscribed Data
| dc.creator | Griffiths, Rachael | |
| dc.creator | Meelen, Marieke | |
| dc.date | 2025-11-22T00:30:39Z | |
| dc.date | 2025-11-21 | |
| dc.date.accessioned | 2026-08-03T02:11:20Z | |
| dc.description | In this paper, we present our work-in-progress on Information Extraction and Text Classification from large manuscript collections that have not yet been transcribed. We propose a three-stage pipeline starting with digitisation using a collaborative Handwritten Text Recognition (HTR) workflow, followed by Normalisation and Segmentation of the texts to create searchable collections, and, finally, we discuss how Text Classification and Information Extraction can help us identify the texts with Tibetan ‘Pagan’ religious features that are hidden among texts that belong to the Buddhist and Bön religious traditions. | |
| dc.description | ERC Advanced Grant, Pagan Tibet, 101097364 | |
| dc.format | application/pdf | |
| dc.identifier | 3070-8931 | |
| dc.identifier | https://www.repository.cam.ac.uk/handle/1810/392812 | |
| dc.identifier.uri | https://repo.dare.co.zw/handle/123456789/164666 | |
| dc.language | eng | |
| dc.publisher | Association for Computers and the Humanities | |
| dc.publisher | Department of Theoretical and Applied Linguistics | |
| dc.publisher | https://doi.org/10.63744/ayiz0ulyis4f | |
| dc.rights | Attribution 4.0 International | |
| dc.rights | https://creativecommons.org/licenses/by/4.0/ | |
| dc.title | How ‘Pagan’ Is My Text? Information Extraction from Untranscribed Data | |
| dc.type | Conference Object |