- Version
- Download 1
- File Size 360.71 KB
- File Count 1
- Create Date September 3, 2026
- Last Updated September 3, 2026
AN EXTRACTIVE TEXT SUMMARIZATION MODEL FOR THE YORUBA LANGUAGE
ABSTRACT
Text summarization is a subfield of Natural Language Processing (NLP) focused on generating concise summaries from longer texts. Despite recent advances, research on African languages remains limited due to historical marginalization. Yoruba, spoken in South-West Nigeria and parts of the diaspora, has been classified by UNESCO as at risk of extinction, leading to scarce NLP resources compared to high-resource languages like English. This study proposes an unsupervised text summarization framework tailored for Yoruba, addressing the challenge of limited annotated data. The framework operates in multiple stages: it first constructs a graph-based representation of the document to capture relationships between words. Sentence embeddings are then generated using the Afriberta Transformer and enhanced with intrinsic vectors derived from the graph. These representations are clustered, and key sentences are selected using graph-based methods. Evaluation using Jensen–Shannon divergence achieved 72.03%, while ROUGE-based evaluation yielded an F1 score of 46.4%, demonstrating competitive performance.
Keywords: Text Summarization, Yoruba, African, Low Resource, Transfer Learning
