Inside Google PaLM: How Large Language Models Handle Translation
When Google introduced PaLM (Pathways Language Model) in 2022, it represented a significant architectural shift in how large language models process multilingual tasks. For language professionals, understanding how these models work isn't about academic curiosity: it's about making informed decisions about where AI fits (and doesn't fit) in professional translation workflows.
The Architecture Behind PaLM
PaLM is built on the Transformer architecture, the same foundation that powers virtually all modern language models. But PaLM introduced several innovations that specifically impact multilingual capabilities:
Pathways System
Traditionally, AI models are trained on one task or a narrow set of tasks. Google's Pathways framework trains a single model across many tasks simultaneously, allowing the model to leverage cross-task learning. For translation, this means the model doesn't just learn to translate: it simultaneously learns to summarise, classify and generate text, and these capabilities reinforce each other.
Scale and Parameter Count
PaLM 2 (the successor used in Gemini) operates at 340 billion parameters. To put this in perspective: every "parameter" is a numerical weight that the model has learned during training. More parameters generally mean more capacity to capture linguistic nuance, including the morphological complexity of languages like Turkish and the rich dialectal variation of Arabic.
Multilingual Training Data
PaLM's training corpus spans over 100 languages, but crucially, the distribution is not equal. English dominates, followed by other high-resource languages. Turkish, Arabic, Persian and Turkic languages are represented but at significantly lower volumes. This asymmetry directly impacts output quality for these language pairs.
How LLMs Actually "Translate"
It's important to understand what LLMs are doing when they translate, because they're not doing what you might assume.
They Don't Have Bilingual Dictionaries
Unlike traditional MT systems that maintain explicit source-target mappings, LLMs operate in a high-dimensional semantic space. The model has learned that certain concepts map to certain token sequences in different languages, but these mappings are probabilistic, not deterministic.
When you ask PaLM to translate "The company exceeded revenue expectations" into Turkish, the model isn't looking up "exceeded" → "aştı" in a dictionary. It's predicting, token by token, what the most probable Turkish sequence is given the English input and its training data. The prediction process considers:
- Semantic meaning of the source
- Syntactic patterns typical of Turkish
- Statistical co-occurrence of Turkish words in similar contexts
- The tokens already generated in the translation so far
The Cross-Lingual Transfer Effect
One of PaLM's most interesting properties is cross-lingual transfer: knowledge learned from data in one language can improve performance in another. This is particularly relevant for Turkic languages.
Turkish, Azerbaijani, Uzbek, Kazakh, Kyrgyz and Turkmen share structural features: agglutinative morphology, vowel harmony, SOV word order. When PaLM learns these patterns from Turkish (its best-represented Turkic language), some of that knowledge transfers to the lower-resource Turkic languages.
At El Turco, we've observed this effect in practice. Our Azerbaijani MT outputs benefit measurably from the model's Turkish training data, with approximately 8-12% better BLEU scores than expected based on Azerbaijani training data volume alone.
Implications for Turkish Translation
Turkish presents specific challenges for LLM-based translation that are worth understanding:
Agglutination
Turkish builds meaning by stacking suffixes. The word "evlerinizden" (from your houses) consists of:
- ev (house)
- -ler (plural)
- -iniz (your, formal)
- -den (from)
LLMs handle this better than previous statistical MT systems but still make errors with complex suffix chains, particularly when disambiguation depends on context beyond the current sentence.
Subject Dropping
Turkish is a pro-drop language: subjects are often omitted because the verb conjugation makes them clear. LLMs sometimes generate unnecessarily explicit subjects ("o yapıyor" instead of "yapıyor"), producing grammatically correct but unnatural output.
Word Order Flexibility
While Turkish's default order is SOV, the language allows considerable word order variation for emphasis and style. LLMs tend to default rigidly to SOV, missing the stylistic nuance that skilled human translators employ.
Implications for Arabic Translation
Arabic presents a different set of challenges:
Diglossia
The vast gulf between Modern Standard Arabic and the various dialects is a fundamental challenge. LLMs trained predominantly on written text (MSA, news, Wikipedia) produce output that sounds bookish and formal in contexts requiring dialectal Arabic.
Root-Based Morphology
Arabic derives words from three-consonant roots through pattern-based morphology. The root k-t-b generates kitab (book), kataba (he wrote), maktaba (library), kātib (writer). LLMs handle high-frequency patterns well but can stumble on rarer derivations.
Right-to-Left Processing
While this is primarily a rendering issue, mixed LTR/RTL content (Arabic text with embedded English brand names, numbers or technical terms) creates segmentation challenges for LLM tokenizers.
Practical Takeaways
-
LLMs are not replacing professional translation: they're becoming another tool in the workflow, alongside TMs, termbases and style guides.
-
Language pair selection matters more than model selection. The gap between English→Turkish MT quality and English→French MT quality is far larger than the gap between GPT-4 and Gemini for the same language pair.
-
Custom training data is the differentiator. A fine-tuned model trained on 100,000 verified Turkish gaming translations will outperform a general-purpose model with 340 billion parameters for Turkish gaming content.
-
Understanding how the model works helps you evaluate its output. Knowing that LLMs predict probable sequences (rather than "understand" meaning) helps you anticipate where they'll fail and design appropriate QA processes.
Cenk is the founder of El Turco and has been tracking LLM developments since GPT-2 in 2019. He regularly evaluates new models for practical applicability to the language services industry.
