Abstract
This paper introduces an open-access, userfriendly online thesaurus for the Welsh language, aimed at enriching digital resources for Welsh speakers and learners. Utilising advances in Natural Language Processing (NLP), our approach combines pre-existing word embeddings, a Welsh semantic tagger, and human evaluation to establish related terms. In this case, an initial list of 250 words was expanded
by adding 6,953 synonyms provided by linguists, creating a more extensive foundation for building the gold-standards. With this expanded list, when a user queries a particular word, the thesaurus presents all of its synonyms, allowing them to choose from a wider range of options. This is especially helpful when a
user is unsure of the exact word they want to use or wants to explore different ways to express a concept. The resulting thesaurus offers a comprehensive, reliable resource for Welsh language users, fostering enhanced communication and expression. Our work promotes Welsh NLP and showcases NLP’s potential to support under-resourced languages. The thesaurus will be accessible via a bilingual website, and the accompanying Python code will be available in a bilingual, public GitHub repository, and it will be available as a web service. Our approach presents a more efficient, cost-effective method for thesaurus creation, with potential applicability to other under-resourced languages.
by adding 6,953 synonyms provided by linguists, creating a more extensive foundation for building the gold-standards. With this expanded list, when a user queries a particular word, the thesaurus presents all of its synonyms, allowing them to choose from a wider range of options. This is especially helpful when a
user is unsure of the exact word they want to use or wants to explore different ways to express a concept. The resulting thesaurus offers a comprehensive, reliable resource for Welsh language users, fostering enhanced communication and expression. Our work promotes Welsh NLP and showcases NLP’s potential to support under-resourced languages. The thesaurus will be accessible via a bilingual website, and the accompanying Python code will be available in a bilingual, public GitHub repository, and it will be available as a web service. Our approach presents a more efficient, cost-effective method for thesaurus creation, with potential applicability to other under-resourced languages.
| Original language | English |
|---|---|
| Pages | 306–315 |
| Number of pages | 10 |
| Publication status | Published - Sept 2023 |
| Event | Proceedings of the 4th Conference on Language, Data and Knowledge - Vienna, Austria Duration: 12 Sept 2023 → 15 Sept 2023 |
Conference
| Conference | Proceedings of the 4th Conference on Language, Data and Knowledge |
|---|---|
| Country/Territory | Austria |
| City | Vienna |
| Period | 12/09/23 → 15/09/23 |
Fingerprint
Dive into the research topics of 'Open-source thesaurus development for under-resourced languages: a welsh case study'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver