The aim of this paper is to present guidelines to build the best and biggest possible silver corpus for word segmentation in Modern Standard Arabic, which should be open data, the most verified and correct possible, and built with a clear and reproducible methodology. Such a corpus can help train machine learning models for segmentation. This paper will present such a corpus and the methodology used to build it in order to satisfy these requisites. Our survey shows that Kalimat is the only morphologically tagged corpus that can be used to initiate this process. We have converted a small part of it to a gold corpus, used it for selecting the best segmenters available, and show how to combine them in order to diminish the number of misses, thus providing a silver corpus freely available for scientific purposes. © 2025 IEEE.
https://www.scopus.com/pages/publications/105041835823?origin=resultslist