EPIC QE model Version 3 is Now Available

EPIC QE model Version 3 is Now Available

Discover the enhanced features of EPIC QE model Version 3, offering improved multilingual support and greater efficiency in quality estimation for translations.

We are pleased to announce Version 3 of the off-the-shelf QE model available through the EPIC API. Building on previous releases, Version 3 introduces a redesigned multilingual model architecture that combines a substantially expanded training dataset and improved natural language processing (NLP) and linguistic techniques.

The new model delivers stronger Quality Estimation performance across a broad range of language pairs. TAUS evaluations indicate an average increase of more than 5% in Savings Percentage across the languages tested, allowing more translations to bypass unnecessary human post-editing. The model’s strongest and most extensively validated performance remains in translations from English into French, Italian, German, and Spanish (FIGS), while the expanded multilingual training data improves coverage across many additional high- and low-resource languages.

This improved performance can help reduce post-editing effort and costs while lowering the risk of poor-quality Machine Translation (MT) output reaching end users without human review.

5%+
average increase in Savings Percentage across languages tested, per TAUS evaluations.

At the heart of Version 3 is a redesigned multilingual model architecture engineered for greater performance and efficiency. Trained on longer sequences than Version 2, the model is better equipped to evaluate longer segments and provides a stronger technical foundation for supporting even longer source and target texts in future releases.

Version 3 also improves the identification of several important translation errors. Compared with previous versions, it handles missing source or target text more appropriately, penalizes untranslated content and identical source-and-target text, detects named-entity errors more reliably, and performs better when evaluating idiomatic expressions and overly literal translations.

The model also remains sensitive to numbers, names, additions, omissions, meaning-changing word choices, punctuation inconsistencies, and inappropriate shifts in formality, while accounting for natural differences in word order and language-specific conventions. Together, these capabilities help distinguish between translations that are merely fluent and those that accurately preserve the meaning, details, and intent of the source.

These advances represent an important step toward more scalable and reliable Quality Estimation and provide a stronger foundation for developing customized models tailored to specific customer data, domains, and use cases.

Language Coverage

Version 3 expands TAUS QE coverage across a broad range of high- and low-resource language pairs. English-to-FIGS remains the model’s most extensively trained and validated area, while the larger and more diverse training dataset delivers stronger performance across many additional English-to-target and non-English language pairs.

Evaluations show improvements across most tested languages, with particularly strong gains for Greek, Indonesian, Japanese, Korean, Portuguese, Vietnamese, and Chinese. TAUS experiments indicate an average increase of more than 5% in Savings Percentage across a range of languages. This broader coverage gives customers greater flexibility to apply Quality Estimation across multilingual workflows and explore new opportunities for automation.

Version 3 supports any-to-any language directions, including translations with non-English source languages. However, English-to-any remains the recommended direction because it has undergone the most extensive evaluation. The model has also been trained and tested on several language pairs with non-English source languages and can support bidirectional use cases. Customers working with these directions are encouraged to validate performance using representative content from their own workflows.

Migration and Thresholds

Version 3 has a slightly different score distribution from Version 2, so customers may need to adjust their decision thresholds. In most cases, a threshold reduction of between 0.02 and 0.05 will be appropriate.

As with previous versions, the threshold determines the balance between the Savings Percentage and False Positive Rate. Higher savings mean less human post-editing, while a lower False Positive Rate reduces the likelihood that poor-quality translations will be approved for automated use.

Customized Models

Version 3 will also serve as the foundation for the next generation of customized QE models trained on customer-specific data. Combined with TAUS’s ongoing improvements in training data, modeling, and evaluation techniques, it will support stronger customized models tailored to each customer’s languages, domains, terminology, style, and specific quality requirements.

Ready to try Version 3?
Connect through the EPIC API to start using the updated QE model today.
Get Started

 

Author
Amir Soleimani

Amir Soleimani is a Senior NLP Engineer at TAUS, where he focuses on enhancing the performance of NLP models, particularly in Quality Estimation (QE) and Automatic Post-Editing (APE). He earned his Ph.D. from the University of Amsterdam in 2024, where his research centered on natural language processing applications for information verification. His work combines machine learning and AI with a strong commitment to advancing language technologies that improve the quality of multilingual content and information.

Related Articles
Loading related articles…