Part of Speech and Universal Dependency effects on English Arabic Machine Translation

Ofek Rafaeli, Omri Abend, Leshem Choshen, Dmitry Nikolaev · arXiv · 2021

In one sentence

Manually aligning the content words of 1000 English-Arabic sentences in the Parallel Universal Dependencies corpus yields POS and dependency divergences for building translation challenge sets, but Google Translate scores higher on the extracted divergence sentences than on the corpus as a whole.

Abstract

In this research paper, I will elaborate on a method to evaluate machine translation models based on their performance on underlying syntactical phenomena between English and Arabic languages. This method is especially important as such"neural"and"machine learning"are hard to fine-tune and change. Thus, finding a way to evaluate them easily and diversely would greatly help the task of bettering them.

Questions this paper answers

if you pick sentences where English and Arabic grammar disagree, does a translation system actually do worse on them?
do POS and dependency-label divergences between English and Arabic select sentences with lower Google Translate BLEU than a random treebank sample?
how do I tell whether a divergence-based test set is genuinely harder than the corpus it was drawn from?
should I build my English-Arabic MT test set by filtering for syntactic divergences, or will the scores come out the same?
Google Translate's mean sentence-level BLEU on English-Arabic sentences selected for POS/dependency divergences is about 0.70, below its 0.7495 mean over all 1000 Parallel Universal Dependencies sentences. The divergence-based selection did not produce a harder-than-average challenge set.
Holds for: English-to-Arabic Google Translate output scored with NLTK smoothed sentence_bleu against the corpus's professional Arabic translation as the single reference; 1000 PUD sentences from news and Wikipedia.
The six divergence types tested give near-identical Google Translate BLEU means, from 0.6938 for amod -> nmod (113 sentences) to 0.7036 for xcomp -> obl (74 sentences). Verb -> Noun is the largest set, at 307 sentences.
Holds for: One reference translation per sentence, BLEU-4 with uniform weights; sentence counts are subsets of the 1000-sentence PUD corpus and overlap where a sentence contains several divergences.
which bits of grammar change shape when English sentences are translated into Arabic?
which English-to-Arabic dependency and POS divergences pass a frequency threshold in the Parallel Universal Dependencies treebank?
how do I find recurring grammatical mismatches between English and Arabic to write test-set extraction rules from?
which English-Arabic construction mismatches are frequent enough to be worth targeting in my evaluation set?
Six English-to-Arabic divergences were selected as candidate challenge-set rules from the manual alignment: obl -> nmod, amod -> nmod, Aux -> verb, obj -> nmod, Verb -> Noun and xcomp -> obl. The threshold was over 8% of English words with a given tag taking the divergent Arabic tag, and over 50 occurrences.
Holds for: Thresholds applied to the POS percentage matrix and UD count matrix built from one annotator's content-word alignment of the English-Arabic PUD corpus; not validated against a second annotator.
In the English-Arabic AUX -> VERB divergences, English inflections of "be" align with كان and its inflections, while English "can", "could" and "have" align with يمكن.
Holds for: Word pairs aligned more than 3 times in the manually tagged PUD corpus; a lexical regularity in this corpus, not a claim about Arabic generally.
how were the English and Arabic versions of the same sentences matched up word by word, and did any sentences defeat it?
how complete was the manual content-word alignment over the 1000-sentence English-Arabic PUD corpus, and what background does the annotation demand?
how do I hand-align an English-Arabic parallel treebank at the content-word level, and what language skill does it take?
do I need to be fluent in Arabic to redo this kind of cross-lingual word alignment myself?
Manual content-word matching between English and Arabic covered all but 2 of the 1000 Parallel Universal Dependencies sentences. Sentences 454 and 491 could not be fully tagged, because the Arabic rendering diverges in meaning from the English.
Holds for: Single annotator at ILR R-1+ reading proficiency in Modern Standard Arabic, matching content words rather than translating; annotation reliability is not independently measured.
Replicating this style of manual cross-lingual content-word alignment requires reading a newspaper or Wikipedia sentence in the target language without a dictionary for more than half the words. Formal grammatical study of both languages is also needed.
Holds for: The author's own recommendation based on the English-Arabic annotation experience; a suggested minimum, not an empirically validated annotator qualification.
are there Arabic grammatical forms that machine translation into English simply gets wrong because English has no equivalent?
do Arabic-specific constructions such as the passive participle, dual number and maf'ul mutlaq fail under Google Translate with no English structural parallel?
how do I probe an Arabic-English translation system on constructions that English cannot express directly?
can I expect Google Translate to handle Arabic dual forms and cognate accusatives in my documents?
Four Arabic constructions with no fixed English parallel, namely passive participle, dual form, maf'ul mutlaq and verb-preposition distance, each yield a hand-made Arabic sentence that Google Translate renders wrongly. One such example, ظفرت الجائزة بتصفيق, comes back as "The award won applause" instead of "The award was received with applause".
Holds for: Single hand-constructed example per phenomenon, Arabic-to-English direction, Google Translate at time of writing in 2021; no BLEU scores or automatic extraction for these cases.
why can't Arabic dual forms and cognate accusatives be pulled out of a treebank automatically?
what annotation gaps in the Arabic UD treebank block automatic extraction of dual number and maf'ul mutlaq, and how accurate are Arabic root extractors?
how do I automate extraction of Arabic dual number or maf'ul mutlaq examples from Universal Dependencies annotation?
can I rely on Arabic Universal Dependencies annotation to mine these constructions, or do I need manual work?
Two Arabic phenomena resist automatic extraction from Universal Dependencies: dual number is absent from the UD data, and maf'ul mutlaq needs word roots the Arabic UD parser leaves unfilled on the advmod. Off-the-shelf Arabic root extractors reach only around 75% accuracy.
Holds for: UD v2 annotation of the Arabic PUD treebank and publicly available Arabic root-extraction tools as of 2021; the 75% figure is quoted from those tools' reported accuracy, not measured in this work.
can you score a translation just by checking whether the expected grammatical form turned up in the output?
is a target-side divergent-tag presence check a sound substitute for BLEU when scoring English-Arabic challenge sentences?
how should I score an English-Arabic challenge set, by looking for the divergent dependency label or by an n-gram metric?
if I check my Arabic output for the expected divergent tag instead of computing BLEU, what goes wrong?
Checking whether a translated Arabic sentence contains the expected divergent tag is an unreliable substitute for BLEU. The divergences occur well below 100% of the time, so such a check would penalise correct translations that keep the English construction.
Holds for: Argument grounded in the divergence rates of the POS and UD correlation matrices for English-Arabic; the alternative check was reasoned about, not implemented or measured.
would using shorter or longer word sequences to score the translations change the comparison?
are the English-Arabic divergence-set BLEU results sensitive to n-gram order, comparing BLEU-3, BLEU-5 and BLEU-6 against BLEU-4 under uniform weights?
how do I check whether my BLEU comparison of English-Arabic divergence sentences survives a change of n-gram order?
do I need to report several n-gram orders when scoring Arabic translations, or is BLEU-4 enough?
Recomputing the English-Arabic divergence-set BLEU scores with n = 3, 5 and 6 instead of the default BLEU-4 did not greatly alter the results under uniform weights.
Holds for: Same 1000-sentence PUD material and single-reference smoothed NLTK sentence_bleu; no numeric table of the alternative-n scores is reported.
what should I read about using grammar differences between two languages to build translation test sentences?
which work applies syntactic-divergence challenge-set methodology to English-Arabic machine translation using Universal Dependencies?
where do I start if I want to build syntax-divergence challenge sets for a new language pair such as English-Arabic?
The work is an undergraduate-level case study in building syntax-divergence challenge sets for machine translation. It applies the challenge-set methodology of Choshen and Abend (2019) to English-Arabic, via manual annotation of the Parallel Universal Dependencies corpus.
Holds for: One language pair, one 1000-sentence parallel corpus and one MT system (Google Translate) as of 2021; arXiv preprint, not peer-reviewed.
A reusable rule for automatically extracting English-Arabic challenge sets from a parallel treebank is the target, rather than a fixed challenge set. Others could then generate their own diverse test sentences with little human labour.
Holds for: As stated by the author for the English-Arabic pair in 2021; the reported experiment tested the rule and did not confirm that it isolates hard sentences.
was the aim to hand out a finished set of tricky English-Arabic sentences, or a recipe other people can run?
is the English-Arabic contribution a fixed challenge set or a reusable extraction rule over a parallel treebank?
if I work on a different language pair, do I get a ready-made English-Arabic test set out of this or a procedure I can adapt?
A reusable rule for automatically extracting English-Arabic challenge sets from a parallel treebank is the target, rather than a fixed challenge set. Others could then generate their own diverse test sentences with little human labour.
Holds for: As stated by the author for the English-Arabic pair in 2021; the reported experiment tested the rule and did not confirm that it isolates hard sentences.

Claims and scope

Common misreadings

Terminology in this paper

Challenge set
A set of source-language sentences with gold target-language translations, chosen so that they are difficult enough that any competent machine translation engine must be able to translate them correctly.
PUD (Parallel Universal Dependencies)
A corpus of 1000 sentences parallel across languages, drawn from news and Wikipedia, with 750 originally English and the rest translated via English, annotated under Universal Dependencies v2 guidelines.
Syntactic divergence
A case where a word's part-of-speech or dependency label in the source sentence maps to a different label in the professional translation, such as an English adjectival modifier (amod) rendered as an Arabic noun modifier (nmod).
Maf'ul mutlaq (cognate accusative)
An Arabic verbal noun placed after the verb, resembling it in form or meaning, used for emphasis or to state a type or number, as in ترتبط ارتباطاً وثيقاً ('is linked a strong link').

How to cite

@article{DBLP:journals/corr/abs-2106-00745,author       = {Ofek Rafaeli and
                  Omri Abend and
                  Leshem Choshen and
                  Dmitry Nikolaev},
  title        = {Part of Speech and Universal Dependency effects on English Arabic
                  Machine Translation},
  journal      = {CoRR},
  volume       = {abs/2106.00745},
  year         = {2021},
  url          = {https://arxiv.org/abs/2106.00745},
  eprinttype    = {arXiv},
  eprint       = {2106.00745},
  timestamp    = {Mon, 25 Oct 2021 01:00:00 +0200},
  biburl       = {https://dblp.org/rec/journals/corr/abs-2106-00745.bib},
  bibsource    = {dblp computer science bibliography, https://dblp.org}
}

References

See the full reference list in the paper.