BLEU was one of the first metrics to report high correlation with human judgements of quality. The metric is currently one of the most popular in the field. The central idea behind the metric is that “the closer a machine translation is to a professional human translation, the better it is”.[1] The metric calculates scores for individual segments, generally sentences, and then averages these scores over the whole corpus in order to reach a final score. It has been shown to correlate highly with human judgements of quality at the corpus level.[2]
BLEU uses a modified form of precision to compare a candidate translation against multiple reference translations. The metric modifies simple precision since machine translation systems have been known to generate more words than appear in a reference text.
Notes
This guide is licensed under the GNU Free Documentation License. It uses material from the Wikipedia.
Need an webmaster? Click HERE
Discover more from MultiMedia
Subscribe to get the latest posts sent to your email.
Leave a Reply