Domain adaptation using neural network joint model

  • Shafiq Joty
  • , Nadir Durrani*
  • , Hassan Sajjad
  • , Ahmed Abdelali
  • *Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

2 Citations (Scopus)

Abstract

We explore neural joint models for the task of domain adaptation in machine translation in two ways: (i) we apply state-of-the-art domain adaptation techniques, such as mixture modelling and data selection using the recently proposed Neural Network Joint Model (NNJM) (Devlin et al., 2014); (ii) we propose two novel approaches to perform adaptation through instance weighting and weight readjustment in the NNJM framework. In our first approach, we propose a pair of models called Neural Domain Adaptation Models (NDAM) that minimizes the cross entropy by regularizing the loss function with respect to in-domain (and optionally to out-domain) model. In the second approach, we present a set of Neural Fusion Models (NFM) that combines the in- and the out-domain models by readjusting their parameters based on the in-domain data. We evaluated our models on the standard task of translating English-to-German and Arabic-to-English TED talks. The NDAM models achieved better perplexities and modest BLEU improvements compared to the baseline NNJM, trained either on in-domain or on a concatenation of in- and out-domain data. On the other hand, the NFM models obtained significant improvements of up to +0.9 and +0.7 BLEU points, respectively. We also demonstrate improvements over existing adaptation methods such as instance weighting, phrasetable fill-up, linear and log-linear interpolations.

Original languageEnglish
Pages (from-to)161-179
Number of pages19
JournalComputer Speech and Language
Volume45
DOIs
Publication statusPublished - Sept 2017

Keywords

  • Distributed representation of texts
  • Domain adaptation
  • Machine translation
  • Neural network joint model
  • Noise contrastive estimation

Fingerprint

Dive into the research topics of 'Domain adaptation using neural network joint model'. Together they form a unique fingerprint.

Cite this