Back to blog
Published

Which AI Model Translates Documents Best Right Now: GPT, Claude and Gemini Compared

The best AI model for document translation is rarely the whole story. What changes when GPT, Claude and Gemini run the same files, and how to test it.

Which AI Model Translates Documents Best Right Now: GPT, Claude and Gemini Compared

Every few weeks somebody asks us which model to use, and the honest answer is that the question has a short shelf life. Models ship, get quietly updated, and trade places on whatever list you were reading last month. A more useful question is what actually changes when the same document goes through GPT, Claude and Gemini, and how much of the difference comes from the model at all. In our experience the best ai model for document translation matters less than the shape of the request you send it. The differences that do exist are real and repeatable, though, and they are worth knowing before you standardize on one.

Why "best AI model for document translation" is the wrong first question

A benchmark sentence and a client document are different problems. In a WMT-style evaluation, a system receives a source sentence, sometimes a little surrounding context, and produces one target sentence that a human or a learned metric scores. Your job is not that. Your job is a 40-page DOCX, 900 segments, a glossary of sixty locked terms, tables whose figures have to come through untouched, and a client who will notice if the product name is translated in section 3 and left in English in section 11.

At the sentence level the gap between current frontier models is narrow enough that experienced translators often cannot tell which output came from which. Over 900 segments the gap widens, and it widens along dimensions that sentence-level benchmarks never look at. Does the model still respect the glossary at segment 700. Does it keep the register it started with. Does it stop writing when the source cell is empty instead of filling the silence.

We have watched two teams run the same model on similar files and land a full quality rung apart. One sent a glossary, a short domain instruction, and enough surrounding context for the model to know what kind of document it was in. The other pasted segments with "translate to German" and nothing else. Crediting or blaming the model in that comparison would be a category error.

So the order of operations matters. Fix the request, then compare models. If you compare models while your prompt is thin, you are measuring noise and you will pick a winner that stops winning as soon as anything about your workflow changes.

What actually differs between GPT, Claude and Gemini on long documents

We are not going to publish a scoreboard. Any ranking we print here is wrong within two months, and the sub-versions move faster than the brand names do. What holds up longer is the shape of the differences, because those come from how each family is trained and tuned rather than from a specific checkpoint.

Instruction adherence over hundreds of segments

This is where we see the most practical variation. Give a model a glossary with sixty entries and a document with 900 segments, and the interesting number is not whether it uses the right term in segment 5 but whether it still uses it in segment 750. Some models hold a constraint steadily across a long run. Others obey for the first few hundred segments and then relax back toward their default vocabulary, especially when the source phrasing drifts away from the glossary entry's exact form. That decay is worth measuring on your own files, because it decides how much post-editing you pay for.

How much the model rewrites by default

Models differ in how much they consider improving the source part of the job. One will keep a clumsy nominal German sentence clumsy in English because that is what the source said. Another will quietly restructure it into something a reader enjoys. Neither behavior is correct in the abstract. For marketing copy the rewriting instinct is a gift; for a specification, a contract, or anything a regulator reads, it is a liability. Test with the document type you actually deliver, not with a paragraph of general prose.

Refusals and over-caution on ordinary business content

This one surprises people who have not run high volumes. Every model family has content it handles cautiously, and the boundaries do not always match commercial reality. HR disciplinary letters, safety data sheets with hazard statements, medical device instructions, insurance claim narratives, occasionally a product catalog. When a model softens a hazard warning or declines a segment, the output is not wrong in an obvious way, which is what makes it expensive. We have found one refusal pattern per model family worth knowing before a deadline, and they are not the same patterns.

Long context windows are not the same as long attention

A large context window lets you send a whole document. It does not guarantee the model weights the middle of that document as carefully as the ends. If a run has to be consistent across 40 pages, feeding the whole file in one request helps, and it does not remove the need for a glossary and a terminology check on the output. Our earlier comparison of Claude and GPT-4 on translation work covers this in more detail for those two specifically.

How to test the three on your own documents in an afternoon

Public comparisons are built on public test sets. Your clients are not in them. Building a private eval set is a few hours of work once and then pays for itself every time a new model version lands.

Start with three files you have already delivered and QA'd, in the language pairs and subject areas you actually sell. From them, pull 50 to 60 segments that gave a human trouble: terms with a client-mandated rendering, sentences with a genuine ambiguity, numbers with units, a row where the source is a bare abbreviation with no context, a heading that could be a noun or an imperative. Skip the easy ones. Easy segments are where all three models look identical and your eval learns nothing.

Then run the identical prompt and identical glossary through each model, one at a time, and score the outputs blind. Strip the model names before scoring. The rating changes when you know which system you are looking at, and it changes in the direction of whatever you read last week.

Score by error category rather than by a single impression. The MQM framework gives you a serious typology if you want one, and a reduced version works fine: accuracy, terminology, fluency, and anything that broke the file. Count errors per thousand words so results stay comparable when the sample size changes. What comes out of this is not an absolute quality number, it is a ranking for your work, which is the only ranking that helps you.

This approach works best when you have enough volume in a stable domain to justify the setup. It does not apply well if every project is a different subject area and a different pair, in which case the more reliable investment is a stronger prompt and glossary routine that travels across models.

Two cases where the model choice changed the result

A Japanese-to-English equipment specification. The source was the usual dense, passive, heavily nominal style, with 前記 back-references all the way through. One model reproduced that structure faithfully into English and produced something technically defensible and close to unreadable. Another normalized it into clean engineering English, which the client preferred until we noticed that two hedges had disappeared in the process and one tolerance had become a plain statement rather than a conditional. The fix was not picking the nicer model. The fix was an instruction that said restructure for readability, never drop a modal or a conditional, and then the difference between the two models on that file became small.

An XLSX product catalog, roughly 4,000 short cells, a lot of them two or three words. Here the model choice mattered more than it did on the specification, and for an unglamorous reason: how each one treated cells it could not interpret. One model left an ambiguous abbreviation alone. One expanded it into a plausible full term that turned out to be the wrong component category, which propagated through 200 rows because the same abbreviation repeated. One returned an empty cell as an empty cell; another filled it with a translation of the column header. None of that shows up on a translation benchmark. All of it shows up in the client's review.

The pattern in both cases is the same. Short, context-poor segments amplify model differences. Long prose with plenty of surrounding context flattens them.

The failure modes all three still share

Terminology drift over long files without a glossary. Every model does this, and it is the single most common reason AI-translated documents come back from client review. The model has no memory of what it decided 300 segments ago unless you put that decision back in front of it, which is one of the reasons document-level context beats feeding a model one segment at a time when you have the choice.

Confident completion of ambiguous fragments. UI labels, column headers, isolated table cells. A human translator asks the project manager. A model produces the most statistically comfortable reading and moves on, with no flag to tell you it guessed.

Invented precision around numbers. The digits usually survive. What does not always survive is the decimal comma, the thousands separator, and the temptation to convert a unit that nobody asked to have converted. Check numeric cells and measurement-heavy paragraphs separately from the prose, because a QA pass tuned for language errors will read straight past a converted value.

Register defaults that nobody chose. Models have house styles. For formality-marked pairs, the default lands somewhere in the middle, which is wrong for both a legal notice and a consumer email. This one is cheap to fix in the prompt and expensive to leave alone across a client relationship.

What public benchmarks can and cannot tell you

Benchmarks are still worth reading, as long as you read them for what they measure. The WMT general translation task, run annually by the Conference on Machine Translation, is the most serious human evaluation in the field, and general-purpose LLMs have been competing in it since GPT-4 was submitted to the 2023 edition, where it ranked at or near the top in several out-of-English directions. That result told the industry something real. It did not tell anyone how the same system behaves on a 60-page manual with a client termbase.

Learned metrics like COMET correlate with human judgment far better than BLEU ever did, which is why they have largely replaced it in research. They are still segment-level and reference-based, so they answer "how close is this to one good translation of this sentence" rather than "is this document deliverable". Reference-free quality estimation is closer to what a production workflow needs and comes with its own caveats, which we went through in our piece on what a translation quality score actually measures.

Two further cautions. Vendor-published evaluations exist to sell the model, and the selection of pairs and domains in them is a marketing decision. Independent evaluations are more trustworthy and lag new releases by months, which means the public evidence for the model you are considering today is usually evidence about its predecessor.

What we would actually do

Pick a primary model by testing against your hardest recurring document type, not your average one. Average documents make every model look adequate. The difficult type is where the money leaks.

Keep a second model configured and working. Not for quality arbitrage on every job, but because refusal patterns and outages differ and a deadline does not care which.

Keep the glossary and the prompt outside the model, in files you version. When they live in a saved chat, switching models is a rewrite. When they live in a file, it is a parameter change, and your afternoon eval becomes something you can rerun whenever a new version ships.

Then set a recurring hour, once a quarter, to rerun those 60 segments across whatever is current. That hour is the whole investment. It replaces the question people keep asking us, because after two rounds of it you will have your own answer, specific to your pairs and your clients, and it will stay current in a way that no article about model rankings can.

One last thing, and it is the one we would most like to be believed. When output quality drops, check the prompt, the glossary, and the way the file was segmented before you change models. In the cases we have looked at closely, the request was the problem far more often than the model was.

Newsletter

Get the next article without checking back.

We send occasional product notes and workflow essays when there is something worth reading.

Need the product walkthrough instead? Read the docs.

We care about your data. Read our privacy policy.