Afaan Oromo Text Retrieval System

dc.contributor.advisorMeshesha, Million(PhD)
dc.contributor.authorGutema, Gezehagn
dc.date.accessioned2018-11-16T08:37:56Z
dc.date.accessioned2023-11-18T12:43:58Z
dc.date.available2018-11-16T08:37:56Z
dc.date.available2023-11-18T12:43:58Z
dc.date.issued2012-06
dc.description.abstractThis study is mainly intended to make possible retrieval of Afan Oromo text documents by applying techniques of modern information retrieval system. Information retrieval is a mechanism that enables finding relevant information material of unstructured nature that satisfies information need of user from large collection. Afaan Oromo text retrieval developed in this study has indexing and searching parts. Vector Space Model of information retrieval system was used to guide searching for relevant document from Oromiffa text corpus. The model is selected since Vector space model is the widely used classic model of information retrieval system. The index file structure used is inverted index file structure. For this study text document corpus is prepared by the researcher encompassing different news article and experiment is made by using 9(nine) different user information need queries. Various techniques of text pre-processing including tokenization, normalization, stop word removal and stemming are used for both document indexing and query text. The experiment shows that the performance is on the average 0.575(57.5%) precision and 0.6264(62.64%) recall. The challenging tasks in the study are handling synonymy and polysemy, inability of the stemmer algorithm to all word variants, and ambiguity of words in the language. The performance the system can be increased if stemming algorithm is improved, standard test corpus is used, and thesaurus is used to handle polysemy and synonymy words in the language.en_US
dc.identifier.urihttp://etd.aau.edu.et/handle/12345678/14347
dc.language.isoenen_US
dc.publisherAddis Ababa Universityen_US
dc.subjectInformation Retrievalen_US
dc.titleAfaan Oromo Text Retrieval Systemen_US
dc.typeThesisen_US

Files

Original bundle
Now showing 1 - 1 of 1
No Thumbnail Available
Name:
Gezehagn Gutema.pdf
Size:
1.99 MB
Format:
Adobe Portable Document Format
License bundle
Now showing 1 - 1 of 1
No Thumbnail Available
Name:
license.txt
Size:
1.71 KB
Format:
Plain Text
Description: