Skip to main navigation Skip to search Skip to main content

Gender Identification in Modern Greek Tweets

  • National and Kapodistrian University of Athens

Research output: Chapter in Book/Report/Conference proceedingChapterpeer-review

Abstract

The aim of this paper is to analyze tweets written in Modern Greek and develop a robust methodology for identifying the gender of their author. For this reason, we compare three different feature groups (most frequent function words, gender keywords, and Author Multilevel N-gram Profiles) using two different machine learning algorithms (Random Forests and Support Vector Machines) in various text sizes. The best result (0.883 accuracy) was obtained using SVMs trained with the AMNP feature group using 100-word tweet chunks. This methodology can lead to reliable and accurate gender identification results using tweet chunk sizes as small as 50 words each.

Original languageEnglish
Title of host publicationQuantitative Linguistics
EditorsArjuna Tuzzi, Martina Benesová, Ján Macutek
PublisherWalter de Gruyter GmbH
Pages75-88
Number of pages14
DOIs
Publication statusPublished - 2015
Externally publishedYes

Publication series

NameQuantitative Linguistics
Volume70
ISSN (Print)0179-3616

Keywords

  • author profiling
  • gender identification
  • Modern Greek
  • Multilevel Ngram Profiles
  • Random Forests
  • Support Vector Machines
  • twitter

Fingerprint

Dive into the research topics of 'Gender Identification in Modern Greek Tweets'. Together they form a unique fingerprint.

Cite this