Back to blog
Published

How to choose which terms go in your AI translation glossary

Which AI translation glossary terms actually change the output? A practical selection method: tiers, sources, what to leave out, and how to test it.

How to choose which terms go in your AI translation glossary

Most glossaries we see attached to AI translation jobs are too big to do anything. Someone runs an extraction tool over a 60-page manual, gets 400 candidates, exports all of them, and hands the file over. The model then weighs 400 instructions against the sentence in front of it, and the four or five terms a reviewer would actually flag get the same attention as "pump" and "system". Picking AI translation glossary terms is a filtering problem, not a collection problem, and the filter is the part most teams skip.

Why bigger glossaries produce worse translations

We worked on a 60-page industrial pump manual last spring, English into Russian, with a 340-term glossary supplied by the client's engineering department. The delivered file had "impeller" rendered three different ways. The glossary contained "impeller". It also contained "pump", "pressure", "bolt", "temperature", and roughly two hundred other words that no current model gets wrong.

That file explains the whole problem. A glossary is an instruction set, and instructions compete with each other. When you inject 340 term pairs into a translation prompt, you are spending most of the model's terminology attention on words it would have handled correctly anyway, and the genuinely contested terms sit somewhere in the middle of a wall of text with nothing marking them as different.

The same thing happens to humans, which is the part worth sitting with. Hand a translator a 340-row spreadsheet and they will read the first screen, skim the rest, and consult it only when they are already unsure. Hand them fourteen rows of terms that are genuinely contested in this document and they will read all fourteen and remember most of them. Glossary length has an inverse relationship with glossary compliance, for people and for models alike.

There is a second cost. Long glossaries include general vocabulary, and general vocabulary gets forced into contexts where it does not belong. If your glossary says support = поддержка, the model will happily write "поддержка" when the sentence means a physical bracket. We have seen glossary entries actively cause the terminology errors they were added to prevent. A term pair with no context field is a blunt instrument, and the more of them you load, the more collateral damage you get.

How to pick AI translation glossary terms that change the output

We use three questions, in order, and a term has to pass all three.

First: would a competent translator plausibly choose differently? If two experienced linguists working on this document would produce the same target word without being told, the entry is noise. "Temperature" fails this test. "Seal" passes it, because a Russian translator can reasonably write "уплотнение" or "пломба" depending on whether they think it is a mechanical seal or a tamper seal.

Second: does the wrong choice cost something? Some inconsistencies are invisible to the reader and some end up in a complaint email. In a maintenance manual, a component name that changes between the parts list and the procedure means the reader cannot find the part. In a marketing deck, a slightly different word for the same feature is a style issue nobody notices. Rank by consequence, not by frequency.

Third: is the term stable across the document? Words that mean different things in different sections are bad glossary candidates, because a single-line entry cannot express "translate this way in chapter 4 and that way in chapter 9". Those belong in the prompt as an instruction with conditions, or in a context column if your format has one, or in the reviewer's notes. Forcing them into a flat term pair produces confident, consistent, wrong output.

Run those three questions over a 340-row extraction and you typically end up somewhere between 15 and 60 rows. That is not a failure of the extraction. Extraction tools are built to find every candidate; deciding which candidates matter is a judgment call about this client, this document, and what went wrong last time.

Three tiers that make a glossary easier to enforce

Flat glossaries treat every entry as equally binding, which is never true. We split terms into three tiers and label them, because the tier tells the model and the reviewer how much latitude exists.

Do-not-translate entries come first. Product names, feature names, model numbers, interface labels that appear in a monolingual UI, legal entity names, and anything the client has decided stays in English. These are the strictest entries and also the easiest to check afterward, since verification is a search for the source string in the target file. In a German handbook we handled for a software client, roughly 40 of the 55 total glossary rows were do-not-translate entries, and that list did more for the delivery than every other entry combined.

Fixed equivalents come second: terms with exactly one acceptable target rendering, usually because the client has a published translation, a regulatory requirement, or a term that appears in a UI the document describes. If the interface says "Абонемент", the manual cannot say "Подписка", regardless of which word is better Russian.

Preferred variants come third, and this tier is where most teams get confused. These are terms where more than one rendering is defensible and the client has a house preference. Marking them as preferences rather than rules matters, because it tells a post-editor not to fight the sentence when the preferred word does not fit grammatically. A glossary that treats a stylistic preference as a hard constraint produces stiff, obviously machine-shaped target text.

Not every project needs all three tiers. Short documents often need only the do-not-translate list. But when a glossary crosses about 30 entries, tiering stops being bureaucracy and starts being the thing that keeps the list usable.

Where the right terms actually come from

The source document is the obvious starting point and the weakest one on its own. Frequency counts tell you what is common, not what is contested, and the two rarely overlap. What works better is scanning the document for the vocabulary that made you hesitate on the first read. Those hesitations are unusually reliable signals; if you paused, a model will guess.

Previous approved deliverables are the strongest source we know. If this client has accepted translations before, the terms in those files are already ratified by the person who will review the new one. Pulling term pairs out of an old bilingual file takes twenty minutes and it captures decisions nobody wrote down. This is where a translation memory earns its keep, incidentally, even on projects that never touch a CAT tool: a TM is a record of what the client accepted.

The client's own published target-language material is the third source, and it is the one most teams forget. If the company already has a Russian website, a Russian datasheet, or a Russian safety notice, those documents contain their terminology whether or not anyone has written it into a glossary. Mining them is faster than asking the client for terminology, and it does not put the client in the position of having to decide things they have already decided.

Last, and most useful over time: the reviewer's corrections from the previous round. Every terminology change a reviewer made is a glossary entry you failed to supply. We keep a running list per client and add to it after each delivery. After three or four projects, that list is worth more than any extraction tool output, because it is composed entirely of terms that this particular reviewer cares about.

What to leave out

General vocabulary, first. If the word appears in a bilingual dictionary with one obvious sense, it does not belong in a project glossary. This includes most nouns in technical documents.

Whole phrases and sentences, second. Glossaries handle terms; sentence-level instructions belong in the translation prompt or the style guide. We see glossaries with entries like "This product is not intended for medical use = ...". That is a segment, and putting it in a glossary means it gets injected into every prompt regardless of whether the sentence appears in that chunk.

Style and register rules, third. "Use formal address", "avoid gerunds", "keep sentences under 25 words" are prompt instructions, not term pairs. Mixing them into a glossary makes both documents worse. There is a real division of labor between a glossary and a style guide, and blurring it costs you the ability to reuse either one.

Terms that are not in the document. Clients often send a corporate master termbase covering their entire product line. Ninety percent of it is irrelevant to the file in front of you, and the irrelevant ninety percent is exactly what dilutes the ten percent that matters. Filter the master list against the actual source text before using it. This takes a spreadsheet formula and about five minutes.

One caveat on all of this: it applies to documents with recurring terminology and a length that makes consistency a real risk, roughly 5,000 words and up. For a two-page marketing email, where nearly every sentence is a judgment call, a glossary is the wrong tool and a good brief is the right one. And if you are doing certified or regulated work where the client mandates a full termbase as a contractual deliverable, you deliver the full termbase, then build a working subset for the actual translation run.

How to tell whether your glossary did anything

Most teams never check, which is why bad glossaries survive for years. The check is not complicated.

After delivery, take each glossary entry and search the target file for the specified rendering. Count the hits and compare against the source-side occurrences. If "impeller" appears 34 times in the source and your target term appears 11 times, the glossary was not enforced and you now know it before the client does. A spreadsheet with source count, target count, and the difference takes ten minutes to build and it converts a vague worry into a number.

Then look at the misses individually. Some will be grammatical, where the term was inflected and your simple search missed it. Some will be legitimate, where the sentence genuinely required a different word. And some will be failures, which is the population you care about. That third group tells you whether the problem is the glossary, the prompt, or the model.

QA reports help here if yours breaks errors out by category. The MQM framework treats terminology as its own error type, separate from accuracy and fluency, and that separation is what makes the signal readable. A file with high fluency scores and repeated terminology flags has a glossary enforcement problem, not a translation quality problem, and the fixes for those two are completely different.

If you need to translate the document with AI and want the terminology step to be explicit rather than buried, SnapIntel runs glossary preparation as a separate stage: it generates a glossary from the source DOCX, XLSX, or PPTX, lets you edit or replace it entirely, and will not start translation until you have approved both the glossary and the prompt. The delivered results include a QA report and a quality rating, so the enforcement check above is something you read rather than something you assemble. There is also a free glossary generator if you only need the term list and are translating somewhere else.

A 30-minute selection pass

Here is the process we run on a new document, and it fits in half an hour.

Read the first 10 pages and write down every word that made you pause. Do not look at an extraction list yet. This usually produces 10 to 25 candidates and it captures the contested vocabulary better than any frequency count.

Pull the client's last approved delivery, if one exists, and extract the terms that overlap with your candidate list. Where the old file and your instinct disagree, the old file wins; the client already signed off on it.

Sort what remains into do-not-translate, fixed, and preferred. If the do-not-translate tier is empty, check again, because almost every business document has product names or interface strings that should stay in the source language.

Run your candidate list against the full source text and delete anything that appears fewer than three times, unless it appears in a heading or a warning. Low-frequency terms rarely cause consistency problems, and consistency is what a glossary is for.

Then stop. A finished glossary for a 60-page manual should be somewhere between 20 and 60 rows. If yours is 300, you have built a dictionary, and the translation will show it. If you want a structure to pour these into, our post on what columns a translation glossary actually needs covers the format side.

The single change that pays off fastest: after your next delivery, list every terminology correction the reviewer made and add those terms to the client's glossary. Do it three times and you will have a list that no extraction tool could have produced.

Newsletter

Get the next article without checking back.

We send occasional product notes and workflow essays when there is something worth reading.

Need the product walkthrough instead? Read the docs.

We care about your data. Read our privacy policy.