TY - GEN
T1 - HarfoSokhan
T2 - 19th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2026
AU - Sarvestani, Hamid Jahad
AU - Ramezanian, Vida
AU - Saadat, Saee
AU - Serajeh, Neda Taghizadeh
AU - Razavi Taheri, Maryam S.
AU - Kasaei, Shohreh
AU - Fazli, Mohammad A.
AU - Asgari, Ehsaneddin
N1 - Publisher Copyright:
© 2026 Association for Computational Linguistics.
PY - 2026
Y1 - 2026
N2 - A wide array of NLP/NLU models have been developed for the Persian language and have shown promising results. However, the performance of such models drops significantly when applied to the colloquial form of Persian. This challenge arises from the substantial differences between colloquial and formal Persian and the lack of parallel data facilitating the robustness of the model to the colloquial data or to transform the data to formal Persian. In addressing this gap, our research is dedicated to the development of the HarfoSokhan dataset, a large-scale colloquial to formal Persian parallel dataset of 6M sentence pairs. Our proposed dataset is a critical resource for training models that can effectively bridge the linguistic variations between colloquial and formal Persian. To illustrate the utility of our dataset, we used it to train a GPT2 model, which exhibited remarkable proficiency in colloquial to formal text style transfer, outperforming both OpenAI’s GPT-3.5-turbo model and a leading rule-based system in this task. This conclusion is supported by our proposed ranking-based human evaluation. The results underscore the significance of the HarfoSokhan dataset in enhancing the performance of natural language processing models in the challenging task of colloquial to formal Persian conversion. Resources are available at huggingface.
AB - A wide array of NLP/NLU models have been developed for the Persian language and have shown promising results. However, the performance of such models drops significantly when applied to the colloquial form of Persian. This challenge arises from the substantial differences between colloquial and formal Persian and the lack of parallel data facilitating the robustness of the model to the colloquial data or to transform the data to formal Persian. In addressing this gap, our research is dedicated to the development of the HarfoSokhan dataset, a large-scale colloquial to formal Persian parallel dataset of 6M sentence pairs. Our proposed dataset is a critical resource for training models that can effectively bridge the linguistic variations between colloquial and formal Persian. To illustrate the utility of our dataset, we used it to train a GPT2 model, which exhibited remarkable proficiency in colloquial to formal text style transfer, outperforming both OpenAI’s GPT-3.5-turbo model and a leading rule-based system in this task. This conclusion is supported by our proposed ranking-based human evaluation. The results underscore the significance of the HarfoSokhan dataset in enhancing the performance of natural language processing models in the challenging task of colloquial to formal Persian conversion. Resources are available at huggingface.
UR - https://www.scopus.com/pages/publications/105040581219
U2 - 10.18653/v1/2026.eacl-long.346
DO - 10.18653/v1/2026.eacl-long.346
M3 - Conference contribution
AN - SCOPUS:105040581219
T3 - EACL 2026 - 19th Conference of the European Chapter of the Association for Computational Linguistics, Proceedings of the Conference, Vol. 1 - (Long Papers)
SP - 7380
EP - 7392
BT - Long Papers
A2 - Demberg, Vera
A2 - Inui, Kentaro
A2 - Marquez Villodre, Lluis
PB - Association for Computational Linguistics (ACL)
Y2 - 24 March 2026 through 29 March 2026
ER -