How ‘Pagan’ Is My Text? Information Extraction from Untranscribed Data

dc.creatorGriffiths, Rachael
dc.creatorMeelen, Marieke
dc.date2025-11-22T00:30:39Z
dc.date2025-11-21
dc.date.accessioned2026-08-03T02:11:20Z
dc.descriptionIn this paper, we present our work-in-progress on Information Extraction and Text Classification from large manuscript collections that have not yet been transcribed. We propose a three-stage pipeline starting with digitisation using a collaborative Handwritten Text Recognition (HTR) workflow, followed by Normalisation and Segmentation of the texts to create searchable collections, and, finally, we discuss how Text Classification and Information Extraction can help us identify the texts with Tibetan ‘Pagan’ religious features that are hidden among texts that belong to the Buddhist and Bön religious traditions.
dc.descriptionERC Advanced Grant, Pagan Tibet, 101097364
dc.formatapplication/pdf
dc.identifier3070-8931
dc.identifierhttps://www.repository.cam.ac.uk/handle/1810/392812
dc.identifier.urihttps://repo.dare.co.zw/handle/123456789/164666
dc.languageeng
dc.publisherAssociation for Computers and the Humanities
dc.publisherDepartment of Theoretical and Applied Linguistics
dc.publisherhttps://doi.org/10.63744/ayiz0ulyis4f
dc.rightsAttribution 4.0 International
dc.rightshttps://creativecommons.org/licenses/by/4.0/
dc.titleHow ‘Pagan’ Is My Text? Information Extraction from Untranscribed Data
dc.typeConference Object

Files