Quality Assurance in Translation: Moving Beyond Basic Proofreading
In our industry, "quality" is the most overused and least-defined word. Every LSP promises it. Few define what it actually means for a given project, and fewer still have systematic processes to measure and deliver it consistently.
After eight years as a linguist and quality lead, I've come to believe that translation quality isn't a single standard: it's a framework that must be calibrated to each project's purpose, audience and risk profile. Here's how we think about it at El Turco.
The Quality Spectrum
Not all translations need the same quality level, and pretending otherwise wastes resources or creates false expectations. We work with three quality tiers:
Tier 1: Premium (Human Translation + Senior Review)
Use case: Marketing, brand content, legal documents, life sciences, creative transcreation Error tolerance: Near zero for critical errors; minimal for minor issues Process: Translation by senior linguist → review by second senior linguist → PM quality check → client review Typical turnaround: 2,000-3,000 words/day per linguist
Tier 2: Professional (Human Translation + Light Review)
Use case: Product descriptions, support documentation, training materials, UI strings Error tolerance: Zero critical errors; minor issues acceptable within defined thresholds Process: Translation by qualified linguist → review by second linguist → automated QA Typical turnaround: 3,000-5,000 words/day per linguist
Tier 3: Rapid (MT + Professional Post-Editing)
Use case: High-volume content, internal communications, knowledge base articles, user-generated content Error tolerance: Accuracy must be maintained; style and fluency may be functional rather than polished Process: MT output → professional post-editing → automated QA Typical turnaround: 5,000-8,000 words/day per post-editor
The key insight: quality tiers aren't about caring more or less. They're about matching the process to the content's purpose and risk profile.
The MQM Framework
The Multidimensional Quality Metrics (MQM) framework, developed by QT21 and DFKI, is the industry's most comprehensive approach to translation quality evaluation. It categorises errors along multiple dimensions:
Accuracy
- Mistranslation: Target text conveys a different meaning than source
- Omission: Source information missing from target
- Addition: Target contains information not in source
- Untranslated: Source text left untranslated in target
Fluency
- Grammar: Morphological, syntactic or orthographic errors
- Punctuation: Incorrect or missing punctuation
- Spelling: Misspellings in target text
- Register: Inappropriate formality level or style
Terminology
- Inconsistent terminology: Same concept translated differently within a project
- Wrong terminology: Incorrect term used per approved glossary
Style
- Style guide violations: Deviations from documented style preferences
- Awkward phrasing: Grammatically correct but unnatural expressions
Design/Formatting
- Tag errors: Missing or misplaced formatting tags
- Truncation: Text cut off due to character limits
- Layout issues: Formatting inconsistencies
Each error is further classified by severity:
- Critical: Impacts meaning, creates legal/safety risk or is offensive
- Major: Noticeable to end users and affects quality perception
- Minor: Detectable but doesn't significantly impact user experience
Automated QA: What It Catches and What It Misses
Modern automated QA tools (Xbench, Verifika, QA Distiller, built-in TMS QA) are essential but not sufficient. Here's an honest assessment:
What Automated QA Does Well
- Tag validation: Catches missing, extra or misplaced formatting tags with near-perfect reliability
- Number checking: Verifies that numbers in source appear in target
- Terminology verification: Cross-references translations against approved glossaries
- Consistency checking: Flags when identical source segments have different translations
- Spelling: Standard spell-check functionality
- Length checks: Flags segments that exceed defined character limits
What Automated QA Cannot Do
- Semantic accuracy: An automated tool cannot determine if "bank" was correctly translated as a financial institution or a river bank
- Cultural appropriateness: No algorithm can evaluate whether a Turkish marketing translation resonates culturally
- Register evaluation: Detecting whether the formality level is appropriate for the audience
- Naturalness: Determining if a translation, while technically correct, sounds unnatural to a native speaker
- Contextual correctness: Evaluating whether a UI string makes sense in its actual on-screen position
This gap between automated and human QA is where quality is won or lost.
Building Effective QA Workflows
The Three-Layer Approach
We use a three-layer QA model:
Layer 1: Translator self-review (immediate) Before delivering any segment, translators run automated QA checks and perform a self-review against the project style guide. This catches the majority of mechanical errors.
Layer 2: Linguistic review (same day) A second linguist reviews for accuracy, fluency and style. The reviewer has access to the source text, reference materials and the project style guide. They use a standardized review form that maps to MQM categories.
Layer 3: Spot-check and metrics (delivery day) The quality lead performs a statistical spot-check on a random sample (typically 5-10% of delivered segments). Results are scored against project-specific quality thresholds and recorded for trend analysis.
Quality Feedback Loops
QA isn't just about finding errors: it's about preventing them. We maintain:
- Translator scorecards: individual quality metrics tracked over time
- Error pattern analysis: identifying recurring error types by linguist, language pair or domain
- Style guide evolution: updating guidelines based on common issues
- Calibration sessions: regular meetings where linguists and reviewers align on quality expectations
Quality for Turkish and Arabic: Language-Specific Considerations
Turkish QA Specifics
- Vowel harmony verification: Turkish suffixes must follow vowel harmony rules. Errors here immediately signal non-native quality.
- Consonant mutation checks: Certain consonant changes at morphological boundaries (e.g., kitap → kitabı) are frequent error sources.
- Formality consistency: Verify that sen/siz usage is consistent throughout the deliverable.
Arabic QA Specifics
- Dialect consistency: Ensure the target dialect is maintained throughout. Mixing MSA with dialectal forms is a common quality issue.
- Hamza placement: Correct hamza (ء) placement on alif carriers is a frequent error that automated tools often miss.
- Tashkeel (diacritics): When required (e.g., educational or religious content), verify diacritical mark accuracy.
Measuring Quality Over Time
We track three key metrics across all projects:
- Error rate per 1,000 words: total errors weighted by severity
- Critical error rate: tracked separately because a single critical error can invalidate an otherwise excellent deliverable
- Client satisfaction score: qualitative feedback that captures dimensions our internal metrics might miss
Over time, these metrics reveal patterns: which linguists excel in which domains, which content types are error-prone and where our processes need strengthening.
Melis is a Senior Linguist and Quality Lead at El Turco. She has reviewed over 5 million words of Turkish translation and developed QA frameworks for clients across gaming, e-commerce and technology.
