LLM Translation Benchmark
An open comparison of large language models on translation quality. This measures the models, not translation products, and Fink is not one of the contestants.
Loading benchmark results
Models compared
Switch a model off to remove it from the breakdowns below.
Head-to-head
How this was measured
Every claim on this page comes from the run named above. The method, and its limits, in full:
Dataset
15 English source sentences from the FLORES-200 devtest split, drawn from Wikinews, Wikibooks and Wikivoyage, translated into 25 target languages. Difficulty is defined purely by sentence length: easy is under 15 words, medium 15 to 30, hard over 30. FLORES-200 is a public dataset, so we cannot rule out that these sentences appeared in the models' training data.
Judging
Each pair of translations is shown to the judge model twice, with the order swapped the second time. The judge also sees the professional FLORES-200 reference translation of that sentence, framed as one valid rendering rather than the only correct answer, so it can resolve what the source means instead of having to be the sole authority on 25 languages. A model only wins when both orders agree. When the two orders disagree, we record that as judge inconsistency rather than folding it into the ties, because the two mean different things: a tie is the judge seeing no difference, inconsistency is the judge contradicting itself. Numeric scores are averaged across both orders.
Automatic metrics
chrF++ and BLEU are computed against the professional FLORES-200 reference translations. Prefer chrF++ here: it works on character n-grams and stays meaningful for languages that do not separate words with spaces. BLEU is reported for languages where word tokenization is well defined, and suppressed where it is not.
Known limits
Fifteen sentences per language is a small sample, and all sources are English and Wikipedia-family, so nothing here says anything about translating into English or about the email, UI and document text most people actually translate. A single judge model rates every pair, with no human validation. Latency is measured end to end from Europe and reflects each model's reasoning configuration, not raw speed.
By difficulty
By writing system
Win rate over decisive comparisons, grouped by the target language's script. This is a pattern we found in the data rather than one we set out to test, so treat it as a lead worth following, not a settled result.
By language
Win rate per target language. Small samples per language, so read these as texture, not as a ranking.
Head-to-head matrix
Every comparison
Loading benchmark results
| Target | Source | Difficulty | Winner | chrF++ | BLEU |
|---|
Try the translator behind the benchmark
Fink translates, then explains what changed and why. No account needed to start.
Open App