Back to blog
Published

How AI Translation Quality Has Evolved from 2023 to 2026: What Actually Changed

What changed in ai translation quality 2026 versus 2023, what did not improve at all, and how to test the difference on your own translated documents.

How AI Translation Quality Has Evolved from 2023 to 2026: What Actually Changed

Three years is nothing in most professions and forever in machine translation. When we ask translators what they make of ai translation quality 2026 has delivered, the answers sort almost perfectly by one variable: when the person last ran a real test. Someone who benchmarked LLM output in early 2023 and formed a view still describes failures that have largely stopped happening. Someone who ran the same experiment last month describes different problems, quieter ones, harder to catch. Both are reporting honestly. What follows is our attempt to separate what actually changed from what people assume changed, and to say which complaints have gone stale.

Where things stood in 2023

The 2023 baseline needs restating, because a lot of published opinion still rests on it.

Context windows were small enough that a 40-page manual had to be chopped into pieces, and each piece was translated with no memory of the one before it. That constraint alone produced most of the era's signature errors. A term translated one way on page 3 came back different on page 27. A product name got translated in one chunk and transliterated in another. Section headings that repeated verbatim in the source came out in three variants.

Then there was formatting. Passing a DOCX through an LLM meant losing the DOCX. People pasted text into a chat window, got text back, and spent an hour rebuilding tables and heading levels by hand. We watched agencies price that rebuild as a separate line item, because it reliably took longer than the post-editing.

The third failure is the one that made experienced translators write off the whole category: fabrication under pressure. Feed a model a segment with an unfamiliar abbreviation or a truncated sentence and it produced something plausible instead of flagging the gap. A safety warning missing a clause came back as a complete, confident, wrong safety warning.

Something was already working in 2023, though, which is why adoption happened anyway. Fluency in high-resource European pairs was strong. Register control through instructions worked. For a translator who was going to rewrite the output regardless, a fluent draft beat a blank page. That trade was the entire value proposition, and it was enough.

What ai translation quality 2026 looks like on the same material

We periodically re-run archived jobs from 2023 and 2024 against current models. Same source, same pair, no glossary, no special prompt. The pattern repeats often enough that we now expect it.

Terminology drift inside a single document has mostly stopped. Not because models learned anyone's terminology, but because they can see the whole document, or a large chunk of it, while translating any given segment. A term established in section one stays established. That is the largest single quality change in the period, and it is architectural rather than linguistic.

Fabrication dropped sharply without disappearing. Current models are much more likely to leave an ambiguous fragment close to the source than to invent a resolution. The residual cases cluster somewhere predictable: source text that is itself broken. Truncated cells in an XLSX workbook, OCR noise from a scanned PDF, a table cell holding half a sentence. Give the model garbage and it still tries to help.

Structural handling improved through tooling, not through models. Nothing learned to preserve DOCX styling. What changed is that extraction and reassembly became a normal part of document translation workflows, so the model only ever sees text and the structure gets rebuilt from the original file. A translator running a DOCX import today expects heading levels and tables to survive. In 2023 that expectation was unreasonable.

The last shift caught us off guard. Long documents now come back more consistent than short batches of unrelated segments. Working document by document beats working segment by segment, which inverts the assumption most CAT tool workflows were built around.

One re-run makes the shape of this concrete. We kept a 2024 job for a machinery client: a maintenance procedure in DOCX, English into Russian, about 9,000 words, with a lot of repeated component references. The 2024 raw output rendered one recurring assembly name four different ways across the document and translated two model designations that should have stayed in Latin script. Re-run in 2026 with the identical source and no glossary, the assembly name was consistent throughout and the designations came back untouched. What did not change: the same three procedural sentences with ambiguous subject reference came back ambiguous both times, and both runs collapsed a two-column warning table into prose when the extraction step was skipped. The errors that vanished were consistency errors. The errors that stayed were comprehension errors and structural ones.

Context is the change that mattered most

If you want one explanation for the gap between 2023 output and 2026 output, it is scope. Almost everything else follows.

Take pronoun and gender resolution. In a segment-by-segment setup, "It must be replaced every 500 hours" is unresolvable when the target language marks gender, because the antecedent sits two segments back. The model guesses. In 2023 it guessed constantly and missed often enough to make post-editing a grind. With the surrounding paragraphs available, the guess turns into a lookup.

Register consistency improved by the same mechanism. A 60-page employee handbook translated in isolated chunks wanders between formal and informal address, because nothing anchors the choice. Translated with the document in view, the first choice propagates. On Russian and German handbooks we have seen formality corrections drop to a fraction of what they were.

There is a limit here that gets skipped over. Document-level context helps when the document is coherent. It does nothing for a spreadsheet of 4,000 unrelated product names, where no surrounding text exists to disambiguate anything. Agencies translating catalog data see far less improvement from newer models than agencies translating manuals, and some of them are quietly baffled that their experience does not match the general enthusiasm. Their content just does not benefit from the mechanism that drove the change. If your material is short disconnected strings, the 2023 advice about heavy terminology control still applies nearly unchanged.

Our longer piece on why document-level context beats segment-by-segment translation covers how this plays out inside a CAT tool workflow, including what it does to match rates and TM reuse.

Terminology moved from cleanup to setup

The second real shift is when terminology work happens.

In 2023, terminology was post-editing. You ran the translation, hunted for inconsistencies with a QA tool, then fixed them. Glossaries lived in CAT tools and had no route to the model.

By 2026 the working method is injection before the run. A glossary supplied as part of the translation instructions changes the output rather than the correction list. Compliance is not absolute, and anyone promising that is overselling. In our experience a well-formed glossary of 80 to 200 terms gets followed most of the time, with failures concentrated in inflected forms: the right term in the wrong grammatical case.

A case from an industrial client makes the shape of it clearer. A supplier manual into Russian, roughly 18,000 words, with a 140-term glossary covering component names and regulated safety vocabulary. Run without the glossary, the QA report flagged sixty-odd terminology deviations. Same source with the glossary attached: eleven, of which seven were declension errors on correctly chosen terms. Four genuine misses. That ratio has held across similar jobs for about a year now.

The planning consequence matters more than the numbers. Terminology preparation is now billable work that pays for itself, where in 2023 it was overhead you carried on top of unavoidable cleanup. Agencies that moved an hour of glossary preparation to the front of the project generally report shorter post-editing passes afterward. The effect is biggest on repeat clients, where the glossary gets built once and reused across every subsequent DOCX or XLSX workbook translation.

What did not improve

The non-changes are more useful than the wins, because this is where projects still fail.

Low-resource pairs did not close the gap. The gains of the last three years concentrate where training data is dense. If you work into Kazakh, Georgian, Amharic, or most African languages, 2026 output beats 2023 output and sits nowhere near the European baseline. MTPE effort estimates built on English-to-Spanish experience will be badly wrong.

Asian pairs improved unevenly. Japanese and Korean output reads better and still needs heavy intervention on register, honorifics, and sentence restructuring. Chinese fares better. The distance between "reads well" and "is right for this audience" stays wide enough that we would not price CJK post-editing at European rates.

Marketing copy did not become machine-translatable. Transcreation is a different job and models fail at it in a specific way: fluent, generic, on-brand-sounding text that says slightly less than the source and lands differently. Nothing in three years touched this.

The most consequential non-improvement is that fluency still outruns accuracy. A 2023 error announced itself. It was clumsy, or obviously broken, and your eye caught on it. A 2026 error reads perfectly and is simply wrong, so reviewer attention slides right past. Some people in the industry have started calling this output workslop, and the word fits. It is why every AI translation workflow we consider sound keeps a QA report and a human review step, and why we distrust any process that treats a quality score as a substitute for reading the text. Our earlier look at how accurate AI translation actually is in 2026 goes through the numbers behind that gap.

Nondeterminism persists too. Run the same document twice and you get two slightly different translations. For a one-off delivery, irrelevant. For a document updated monthly, unchanged source paragraphs come back reworded every cycle unless you reuse a TM or lock previously approved segments.

How to run your own 2023 comparison

Industry surveys from CSA Research, Nimdzi, and Slator are useful for direction. They cannot tell you what happened to your content. The comparison is cheap to run yourself and worth more than any published benchmark.

Pull one archived project from 2023 or 2024 where you still have three things: the source file, the raw machine output, and the delivered final version. Re-translate the same source with your current setup and no glossary, so you are measuring the model rather than your own preparation. Compare against the raw 2023 output, not the delivered file. Comparing against the delivered file measures your editing, which was always good.

Score by category, not by total error count, using whatever error typology your QA process already uses. Accuracy, terminology, fluency, style, and formatting all move independently. In every comparison we have run, terminology and formatting improved dramatically, accuracy improved moderately, and style barely budged. Your mix will differ by domain and pair. Knowing where yours sits tells you where to keep spending review time and where you are now over-reviewing out of habit.

One caution. Measure edit distance or edit time, not your impression. Fluent output feels finished. The 2023 output felt unfinished, so it got read carefully. Impressions systematically overstate the improvement, and we have caught ourselves doing it.

This test has limits worth knowing before you run it. Archived raw output from 2023 was usually produced with whatever prompt and chunking that tool used at the time, and you may no longer know what those were, so part of what you measure is tooling rather than model capability. If your archive only holds the delivered file, the comparison is not available to you at all and no reconstruction will fix that. And a single project tells you about a single domain. Two or three from different domains give you a far more usable picture, particularly if one of them is the disconnected-strings kind of content that benefits least.

The output of the exercise should be a short internal note, not a score. Something like: terminology checks on our German manuals can drop to spot-checking, formatting review can drop, numeric and date verification stays exactly where it is, and marketing content stays under full review. That note is what changes how your team spends its hours. A single quality number does not.

What to do with this

If you last evaluated AI translation before mid-2024, your objections probably describe solved problems and your quality gates probably point at the wrong targets. Run the archived-project comparison this month. It takes an afternoon.

Then change one thing. On your next repeat-client project, move terminology work from post-editing to preparation: build the glossary before the translation runs, attach it to the job, and compare the QA report against the last cycle. That single change captures most of the practical benefit of the last three years, and unlike the model improvements, it is entirely yours to make.

Newsletter

Get the next article without checking back.

We send occasional product notes and workflow essays when there is something worth reading.

Need the product walkthrough instead? Read the docs.

We care about your data. Read our privacy policy.