taus-logo-normal
menu-icon
EPIC

OVERVIEW

What is EPIC?

The platform in a nutshell

Choosing a QE Model

Match model to content & risk

EPIC FEATURES

Quality Estimation

Score translations automatically

Automated Post-Editing

Clean up MT output at scale

Retrieval-Augmented Generation

Feed context with glossaries

Specialized Models

Language specific models

USE EPIC

Integrations

Connect with your TMS & Tools

API Docs

Technical docs & API references

EPIC APP

Interactive playground & sandbox
Resources

Blogs

Perspectives from our team

Reports

In-depth industry research

Webinars

Live and on-demand sessions

Case Studies

Real results from real teams
icon epic arrow
icon epic arrow
TAUS DeMT™ Evaluation Report 2023

Discover the latest insights in machine translation with TAUS' DeMT™ Evaluation Report. Explore the impact of large language models, domain-specific translation, and the evolution of customization in 30 language pairs and domains. Uncover how translation is becoming more precise, consistent, and essential in the ever-changing landscape of language technology.

Download the report to access these exclusive findings for free.

icon epic arrow

Key Takeaways

Between June 2022 and September 2023, there were slight improvements in the average BLEU scores for uncustomized translations by both the Microsoft Custom Translator and Amazon Active Custom Translation engines, with Microsoft's baseline score increasing by approximately 1.5 points to just above 48, and Amazon maintaining a score between 46 and 47. Furthermore, Microsoft demonstrated more consistent BLEU scores across languages in its latest tests, narrowing the gap with Amazon's consistently scored translations.

In contrast, the customized translation engines showed more significant improvements, averaging around 53-54 BLEU points. Amazon improved by 1.5 points, while Microsoft improved by nearly 3 points compared to the previous report. Despite variations among different domain and language combinations, these results indicate that customization consistently leads to improvements, with all customizations showing scores almost always more than 3 points higher than the baseline model, reflecting a more uniform and probable enhancement.