Back to blog
Published

Why SnapIntel requires glossary and prompt approval before translation starts

The SnapIntel glossary prompt approval gate blocks translation until both fields are filled. Here is why we built it and what to put in them.

Why SnapIntel requires glossary and prompt approval before translation starts

A translator wrote to us in the spring, annoyed. She had uploaded a 40-page maintenance manual, picked her language pair, and then hit the SnapIntel glossary prompt approval gate: the Approve control was greyed out and the Start button under it did nothing. Her question was reasonable. Why would a translation tool refuse to translate?

The honest answer is that the gate exists because we watched the same project failure happen over and over, always in the same way, and a disabled button turned out to be a better fix than a support article nobody reads.

What the SnapIntel glossary and prompt approval gate actually blocks

The rule is narrow. Before a translation job can start on a project you prepared by hand, two fields have to contain something: the glossary and the translation prompt. Not "something good". Just something, after whitespace is trimmed off both ends.

That check runs in two places. In the browser, the Approve and Start controls stay disabled until both fields hold text, with an inline note saying which one is missing. On the server, the endpoint that starts a translation job runs the same check again and rejects the request with a 400 if either field is empty. The second layer matters more than it sounds. Without it the gate would be a UI suggestion that anyone with a session cookie could route around, and admin accounts get no exemption either.

Domain analysis sits outside the gate on purpose. We recommend running it, and the interface says so next to the action, but it blocks nothing. You can skip it, write a glossary from memory, paste a prompt you keep in a text file, and start.

One more behaviour surprises people the first time: editing the glossary or the prompt after you have approved them resets the approval, and you have to approve again. That is deliberate. An approval that survives the thing it approved being rewritten is not an approval.

Projects created before we introduced the requirement still work as they always did. If a project already has non-empty values stored, it starts translation without asking anyone to redo anything.

Why an empty glossary produces output that reads fine and fails review

Here is the pattern we kept seeing. A spare parts price list arrives as an XLSX workbook, five sheets, a few thousand rows of short cell text. Someone runs it with no glossary. The output comes back fluent. Every cell is grammatical. Every cell is plausible.

Then the client's engineer opens it and finds that "housing" became three different target terms across the five sheets, because in isolation a two-word cell gives a model almost nothing to anchor on, and the model made a locally sensible choice five separate times. Nothing in that file reads like an error. It reads like five people worked on it, which is exactly what it is.

Short-segment content is where this bites hardest. Cell text, slide labels, table headers, warning lines in a manual. A model translating a 200-word paragraph has context to work with. A model translating the string "Cover" has the surrounding document and whatever you told it, and if you told it nothing, the surrounding document is all it gets.

The second version of this problem is quieter. A term is translated consistently and consistently wrong, because the client's internal name for a part does not match the industry-standard name and nobody wrote that down. The QA report will flag inconsistency, and the quality rating gives you a signal that something needs a look, but neither can tell you that a client calls a component something unusual. Only the glossary can.

A PPTX deck we saw last year had the opposite failure. The company name appeared on eleven slides, and on four of them it had been translated, because on those slides it sat inside a sentence rather than alone in a title placeholder and read like an ordinary noun. The fix was one glossary line marking it as untranslatable. Without that line there is nothing in the file telling anyone, human or model, that this particular word is a name.

This is why the gate asks for a glossary rather than trusting the model to work it out. Terminology is not a language problem the model can solve on its own. It is a project fact somebody has to supply, and the only person who reliably has it is whoever took the job.

The prompt is where the project's rules live

The glossary handles words. The prompt handles everything else, and "everything else" turns out to be most of what makes a translation acceptable or unacceptable to the person who ordered it.

Take two DOCX files that arrive in the same week. One is a supply contract heading for internal legal review. One is a sales deck for a trade show. Both are English to German. The correct output for those two files differs in register, in address form, in how literally to handle idiom, in whether a sentence can be split for readability, and in what happens when the source is ambiguous. A model with no instructions will pick reasonable defaults for a generic document, and a generic document is not what either of these is.

The things worth writing down are boring and specific. Who reads this. Formal or informal address. Whether numbers and dates get localised or left alone. Which strings never get translated at all, such as product names, part codes, and placeholders. What to do when the source text is genuinely unclear, because the model will do something, and you would rather choose what.

We removed the pre-filled default prompt from the product for exactly this reason. Earlier versions shipped with sensible-looking default text sitting in the field. People accepted it. The results were generically fine and specifically wrong, and worse, nobody could tell afterwards whether the prompt had been a decision or a default nobody read. The field now starts empty with placeholder guidance, so there is nothing to rubber-stamp.

Something we did not expect: the prompt turns out to be the most reusable artefact in the workflow. A glossary is usually specific to one document. A prompt is usually specific to one client, and once you have written a good one for a client's technical documentation, you paste it into every subsequent project for that client and edit two lines. Several agencies we talk to now keep a folder of client prompts the way they keep style guides, which is roughly what a prompt is: a style guide written for a reader that follows instructions literally and has no memory of the last job.

What changes when SnapIntel prepares everything for you

Not every project deserves five clicks. If you are translating a supplier manual once a month and have no linguists on staff, being asked to approve a glossary you have no basis to judge is a ritual, not a control.

So auto mode runs the chain unattended: domain analysis, glossary, prompt, translation, all from one action. There is no Approve button in that path, because there is nothing waiting for your confirmation.

The requirement did not disappear, though. It moved. In auto mode the same non-empty assertion runs in the worker, immediately before translation executes. If glossary generation comes back with nothing usable, the job fails there and stops. It does not fall back to a default, because we deleted the defaults, and it does not translate anyway. An auto project either has document-specific preparation or it has no translation.

We should be straight about the trade, because it is real. Provenance is not approval. A pipeline that generates a glossary and immediately consumes it satisfies "non-empty" while dropping "a human looked at this". Our answer is that the artefacts are written back to the project as each stage finishes and delivered alongside the results, so you can open a finished job and read the exact glossary and prompt that produced the translation. Review moves after delivery instead of before it. That is a weaker guarantee than manual approval, and it is the right default for a document nobody was going to review beforehand anyway.

Manual mode stays first-class. Every project that existed before auto mode shipped is a manual project, and anything without an explicit mode resolves to manual.

What to actually put in the two fields

Here is the advice we give when someone asks what "enough" looks like.

For the glossary: aim for the terms that would cost you something if they came out wrong, not a dump of every noun in the file. Twenty to sixty entries covers most single documents. Product and part names. Terms the client uses differently from the rest of the industry. Anything with a legal or safety consequence. Anything a reviewer has already corrected once on a previous job for this client, which is the highest-value entry there is and the one people forget to record. We went deeper into the selection question in how to choose which terms go in your AI translation glossary, and if you are starting cold, the free glossary generator gives you a first bilingual draft to edit instead of a blank page.

For the prompt: five or six sentences beats a page. Say who reads the output and what they do with it. Say formal or informal. List the do-not-translate strings explicitly, because "obviously don't translate the product name" is not obvious to a model reading a cell that says "Vector". Say what happens with units and dates. If the document has an unusual structure, say that too, as in "the left column is a control name in the software interface and must match the target-language interface exactly".

Then read both fields once before you approve. That read is the whole point of the gate.

A note on generated content, since most people generate rather than write from scratch. The glossary that comes out of generation is a draft, and drafts have a predictable weakness: they over-collect. You get forty entries where fifteen matter, and the twenty-five extras are ordinary words that did not need pinning down. Delete them. A bloated glossary is not neutral, because every entry you keep is an instruction the model tries to follow, and instructions that were never needed compete with the ones that were. Two minutes of deleting is usually worth more than two minutes of adding.

Where the gate does not help

It checks presence, not quality. A glossary with two useless entries passes. A prompt that says "translate well" passes. If you are determined to defeat the check it takes about four seconds, and we are fine with that, because the person who types "translate well" has at least been shown that a prompt is a thing that exists and affects the result.

It also does nothing about a glossary that is confidently wrong. If a term pair is incorrect, the model applies it consistently across the document, and a wrong term applied consistently is harder to spot in review than a wrong term applied once. Terminology errors that survive the gate come back in the QA report as consistency findings, which will not tell you the term is wrong, only that it is everywhere.

And it does not fit every job. A one-page internal note that somebody wants the gist of does not need a prepared glossary, and asking for one is friction with no payoff. That is what auto mode is for.

The takeaway

Before you start your next job in SnapIntel, spend the four minutes: pull the 20 terms that matter from the source, write five sentences of instruction naming the audience and the register, read both once, approve. That is the difference between a translation you review and a translation you rewrite, and it costs the same four minutes whether the file is a DOCX manual, an XLSX workbook, or a PPTX deck.

If you want to test that claim on your own files, the SnapIntel trial gives you 5,000 words over seven days with no card required. Run the same document twice, once with a real glossary and prompt and once with the thinnest thing that clears the gate, then compare the two QA reports. The gap between them is the argument for the gate, and it is more convincing than anything we can write here.

Newsletter

Get the next article without checking back.

We send occasional product notes and workflow essays when there is something worth reading.

Need the product walkthrough instead? Read the docs.

We care about your data. Read our privacy policy.