Back to blog
Published

Translation Glossary Template: What Columns You Actually Need and How to Fill Them

A translation glossary template needs fewer columns than you think. Here's what to include, what to skip, and how to fill it without stalling the project.

Translation Glossary Template: What Columns You Actually Need and How to Fill Them

Most translation glossary templates die the same death. Someone builds a spreadsheet with fourteen columns, fills in three of them, sends it to two translators, and nobody opens it again. We see this constantly in agency handovers and in the reference folders freelancers inherit from clients. The terms are usually fine. The template is the problem: it asks for information nobody has time to supply, and it leaves out the one field that would have caught the error a reviewer found two weeks later.

So this is a walk through the columns that earn their place in a bilingual glossary, the ones worth adding when a specific project justifies them, and the ones we would quietly delete. Along with a way to fill the thing that does not require a terminologist or a free afternoon.

What a translation glossary is for, and what it is not

A glossary is not a bilingual dictionary. A dictionary tells you what a word can mean. A glossary tells you what your client decided it means, in their documents, this year. That distinction changes what belongs in the file.

Think of each row as a recorded decision plus enough evidence for the next person to trust it. If a translator has to reopen the debate about whether "seal" is торцевое уплотнение or сальник every time a new pump manual arrives, the glossary has failed even if both options are listed in it.

We ran into this on an operations and maintenance manual for a pump manufacturer. The English source used "seal" for two physically different components. The first translator picked one Russian term and used it everywhere. The client's engineer rejected the file, correctly, because half the maintenance instructions now described the wrong part. The fix was not a better dictionary. It was a glossary row that said: "seal (mechanical, shaft assembly) → торцевое уплотнение", with a note pointing at figure 4.2, and a second row for the gasket-type seal. Two rows, one context field, problem gone for every future job from that client.

That is the standard to hold a template to. Does a column help the next translator act without asking a question? If yes, keep it. If it only helps someone audit the glossary later, it is optional at best.

The five columns that carry the weight

After enough projects, the same short set keeps proving itself. Source term. Target term. Context or definition. Part of speech. Status.

Source and target are obvious, with one rule that gets broken constantly: one term per row, in its base form. No "seal / sealing / seals" crammed into a single cell, no slashes offering the translator a choice. A slash in a target cell is an unresolved decision dressed up as terminology.

Context is where the value sits, and it is the field people skip. It can be a short definition, a source sentence, a figure reference, or a domain tag. Anything that tells the translator which of several possible senses this row covers. Without it, homonyms quietly corrupt the file.

Part of speech looks academic until you work into a language where it changes everything. English "monitor" is a noun and a verb. A glossary that does not say which one produces Russian rows where a translator confidently uses монитор in a sentence about supervising a process.

Status is the field that keeps a glossary alive across revisions. Approved, proposed, deprecated, do not translate. That last value does more work than the rest combined. Product names, model numbers, interface labels that ship in English, legal entity names: they all need an explicit instruction, because a translator who sees "SmartFlow 200" with no guidance will make a judgment call, and so will the next one, differently.

A minimal template looks like this:

Source termTarget termContextPOSStatus
mechanical sealторцевое уплотнениеshaft assembly, fig. 4.2nounapproved
gasket sealпрокладкаflange joints, section 3nounapproved
SmartFlow 200SmartFlow 200product nameproper noundo not translate
to monitorконтролироватьprocess supervision, not displayverbapproved

Five columns. A translator can work from it. A reviewer can check against it. Nobody needs training to fill a new row.

Columns worth adding when the project earns them

Everything past those five should be justified by a specific problem you have already hit.

Domain or subject field starts paying off the moment one client sends documents from more than one department. A manufacturer's HR handbook and its technical manuals share a company and share almost no terminology. Tagging rows by domain lets you maintain a single client glossary instead of five files that drift apart.

Forbidden variant is the column we recommend most often and see least. It records the wrong answer explicitly. On an employee handbook for a US company translating into Spanish, the client insisted on empleado and rejected asociado, which is the term their competitor uses. Writing "approved: empleado" is weaker than writing "approved: empleado / never: asociado", because the second version survives the moment a new translator reaches for the obvious alternative. It also gives your QA check something concrete to search for.

Owner and date matter once a glossary crosses the six-month mark. Not for bureaucracy. For the argument that happens when a client's new marketing manager wants to change a term that their predecessor approved. A date and a name settle it in thirty seconds.

Source of truth is useful in regulated work: a link or citation to where the approved term comes from, whether that is a client style guide, a published standard, or a previously delivered file. On pharmaceutical and legal jobs, "because a regulator's published glossary uses it" is a stronger answer than "because we always have".

Term ID becomes necessary only when you start moving terminology between systems. If you plan to export to TBX, the ISO 30042 format that termbases use for exchange, stable identifiers keep entries matched across versions. Before that point, an ID column is overhead.

Columns that look useful and mostly are not

Some fields survive in templates because they feel thorough.

Frequency counts are the clearest example. Knowing that a term appears 43 times in the source is interesting during extraction and useless afterward. It ages instantly and tells the translator nothing about what to write.

Confidence scores from an extraction tool have the same problem. They describe how sure the extractor was, not whether a human approved the term. Once someone reviews the row, the score is noise. Keep the human status field instead.

A free-text notes column tends to swallow everything the template failed to model. We have opened glossaries where notes contained context, forbidden variants, client emails, and a translator's complaint about the deadline. If a note is repeated across many rows, it belongs in its own column. If it appears once, it probably belongs in the style guide.

Full inflection tables are the trap in Slavic and Finno-Ugric target languages. It is tempting to record every case form. In practice translators do not need them, the columns are never complete, and maintaining them turns glossary upkeep into a second job. Base form plus grammatical gender, where gender actually affects agreement, covers almost every real need.

One more: separate columns for each target language in a multilingual file. It works up to three or four languages, then becomes unreadable and impossible to filter. Past that, one row per language pair with a shared concept ID is the version that survives.

How to fill the template without stalling the project

The failure mode here is not laziness. It is scope. Someone decides to build the glossary properly, extracts 900 candidate terms from a 40-page manual, and never finishes reviewing them, so the translation starts with no glossary at all.

A 40-page technical manual usually needs somewhere between 40 and 120 approved terms. Not 900. The rest are ordinary words that any competent translator handles without instruction, or one-off items that belong in the TM rather than the glossary.

The sequence that works for us starts with an automated first pass to surface candidates, then a fast human cut. We keep terms that meet at least one test: the client has a stated preference, the term is ambiguous in the source, an error would be expensive, or the term repeats across documents that different translators will handle. Everything else gets dropped without guilt.

Then comes the part most workflows skip: send the shortlist to the client before translation, not after. Twenty terms in a table, with your proposed target and a one-line context. Clients who will not review a 900-row spreadsheet will answer a 20-row email, and the answers you get are the ones that would otherwise have surfaced as rejected deliveries.

For recurring clients, treat the glossary as something the project feeds rather than something you build once. After each job, add the terms that caused a question, a correction, or a reviewer comment. That habit builds a better client glossary in four months than any upfront extraction session, because every row exists for a documented reason.

One limitation worth naming: this approach assumes reasonably consistent source documents. If a client sends material written by twelve different authors with no internal standard, terminology work on the source side has to happen first, and no glossary template will paper over it.

Keeping the template usable in CAT tools and AI prompts

A glossary that only exists as a beautifully formatted spreadsheet is a glossary that will not be used under deadline pressure.

Practical constraints, all learned the boring way. Save as CSV or XLSX in UTF-8. No merged cells, no color coding as the only carrier of meaning, no header rows above the header row. One term per row. Keep the column order stable across versions, because import mappings break silently when it changes.

For CAT tool import, most tools want a flat source/target table or a TBX file, and most will let you map extra columns to custom fields. If you need to move between formats, the conversion path between CSV, TBX and Excel is its own small discipline, and we wrote up the format conversion steps separately.

AI translation changes what a glossary has to do. When terminology is injected into a translation prompt rather than surfaced as an editor suggestion, the model reads the context and status fields as instructions. This is why the do-not-translate status and the forbidden variant column pay off more now than they did five years ago: a human translator infers that a product name stays in English, while a model needs to be told. Long documents add a second problem, since terminology adherence tends to drift as context accumulates, which makes a compact high-priority glossary more effective than an exhaustive one.

If you need a first-pass glossary quickly, our free Glossary Generator extracts a bilingual candidate list from a document without an account. Inside SnapIntel itself, the glossary is not optional decoration: a project cannot start translation until the glossary and prompt exist and are approved, and both artifacts stay attached to the project alongside the translated DOCX, XLSX or PPTX output and the QA report. The point is that the terminology decisions and the delivered file stay in the same place, so a reviewer can check one against the other.

Where a glossary template stops helping, and what to do this week

A glossary controls terms. It does not control tone, register, sentence length, formality of address, or how a brand refers to its own users. Those belong in a style guide, and confusing the two produces glossaries full of instructions that no import will ever parse.

It also struggles with terms whose correct translation depends on syntax rather than meaning. Adjective agreement, verb aspect in Slavic languages, honorifics in Japanese and Korean: a two-column mapping cannot express those rules, and trying makes the file worse. Post-editing and MTPE workflows are where that gap gets closed, which is an argument for keeping the glossary tight rather than exhaustive.

And a glossary is only as current as its last review. A file that has not been touched in two years is a liability, because translators trust it and it may be wrong.

The concrete step: open the last glossary you delivered to a client and check whether it has a status column and a context column. If either is missing, add them and backfill only the rows that have ever caused a question or a correction. That is usually fifteen to thirty rows, it takes under an hour, and it converts a passive reference list into something the next translator can act on without asking you anything.

Newsletter

Get the next article without checking back.

We send occasional product notes and workflow essays when there is something worth reading.

Need the product walkthrough instead? Read the docs.

We care about your data. Read our privacy policy.