Back to blog
Published

Pre-translation in a CAT tool: when it speeds you up and when it backfires

CAT tool pre-translation saves a review pass on some files and costs you one on others. How to read the analysis report and set the threshold per project.

Pre-translation in a CAT tool: when it speeds you up and when it backfires

A project manager we work with pre-translates every file before assigning it, as policy. A freelancer we work with refuses to open a pre-translated file at all and asks for clean source instead. Both are good at their jobs. Neither can tell you the rule they are actually following, because CAT tool pre-translation is one of those batch tasks that lives in a setup dialog and gets switched on out of habit. Usually the habit is harmless. Sometimes it costs a full review pass. We have watched a team cut two weeks off a manual revision with it, and watched another team pay for the same 40 segments twice because of it. What follows is the decision as we would make it, file by file, and the signals we read before touching the setting.

What cat tool pre-translation actually does to a file

Pre-translation is a batch operation that runs after import and segmentation, before anyone opens the editor. The tool walks the segment list in order, queries the attached translation memories for each segment, and writes a target into any segment whose match percentage clears the threshold you set. Depending on configuration it also stamps a status on that segment: pre-translated, confirmed, or locked. Where the TM has nothing, machine translation or an AI engine can fill the gap, if you let it.

Two things follow from that, and the second is the one teams miss.

The first is mechanical. Nothing about the source changes. Segmentation, tags, and structure were fixed at import; pre-translation only populates target fields. If the segmentation is wrong, pre-translation will not reveal it and will not fix it.

The second is behavioural. The person who opens that file is no longer translating. They are judging. Those are different tasks, with different speeds and different failure modes, and the person doing the work has usually not been told which one they were hired for. A translator quoting on new words and receiving a file that is 60% filled has been handed a review job at a translation rate, or the reverse, and either way somebody is unhappy by Friday.

In Smartcat's terminology, pre-translation means running AI translation automatically before any human reviews the file, and its pipeline runs TM lookup first: exact matches apply automatically at no cost, fuzzy matches are suggested, and AI fills what is left before QA checks run. Most tools follow roughly that order. The order is the point. Running an engine across content you already own in TM is work you are paying for a second time.

The settings that decide whether pre-translation helps or hurts

Three settings do almost all of the work, and only one of them gets any attention.

The fuzzy threshold is the one everybody adjusts. Set it at 70 and nearly every segment comes back filled. Set it at 100 and only identical segments do. Industry habit sits around 75, the bottom of the usual fuzzy band, and we think that habit is wrong for most files. A 76% match is not a draft. It is a sentence that is mostly about something else, and dealing with it means reading the suggestion, reading the source, working out which fragments survived, and rewriting the rest. Writing the sentence from a blank field is frequently faster, and it is almost always better.

The second setting is status, and it decides what happens downstream. Whether pre-translated segments arrive confirmed, unconfirmed, or locked determines who looks at them again. Confirmed segments get skipped by most QA profiles and by most reviewers, who read the status as a signal that somebody has already been through. Locked segments cannot be touched without an unlock, which is right for approved regulatory wording and wrong for nearly everything else.

The third is whether an engine fills the non-matching remainder, and at what status it lands. This is where we see real damage. Machine translation arriving confirmed is a quiet instruction to the reviewer that it has been checked. Nobody checked it.

If you want a default to start from: TM fill at 85 and above, nothing auto-confirmed, engine fill only when the person opening the file has been told it is there. Then adjust per file, which is the rest of this article.

Where pre-translation earns its place

The clearest case is versioned documents. Not similar documents, versioned ones: revision three of the same equipment manual, this year's edition of a policy, a contract template with four clauses changed and the rest untouched.

We handled a 180-page machine manual last year that was the second revision of a file translated eighteen months earlier. The analysis report came back at 71% exact matches and 9% fuzzy. Pre-translating the exact matches only, locked, left the editor about 20% of the file to work through, and the pass took four days rather than the three weeks a cold translation would have needed. The locking was deliberate. The client had already approved that wording, had paid a reviewer to approve it, and did not want it re-litigated by a new translator with different taste in verbs.

The second good case is internal repetition inside a single file with no TM history at all. Product catalogues do this constantly. So do specification tables in an XLSX workbook translation, where the same dozen phrases turn up in 400 cells with nothing but a part number between them. Pre-translating against the memory you build as you go, or using a repetition-propagation setting if your tool has one, means you translate each phrase once and spend the rest of the time on the content that differs.

The third is content where reuse is the requirement rather than the saving. Safety warnings, regulated labelling, approved disclaimers: these need to come back word for word, and pre-translation from a locked TM is a more reliable mechanism for that than asking a translator to remember what was agreed in March.

This works best when the memory is yours, you know who confirmed the entries, and the new file is genuinely the same document rather than the same subject. It does not apply when the TM is a mixed bag inherited from a client with no provenance, which, in our experience, is the more common situation.

Where pre-translation quietly costs you money

The expensive failure is not a wrong translation. It is anchoring.

Give an editor a target field that is 80% right and they will edit toward the suggestion rather than toward the source. The shape of the suggested sentence survives even when the source sentence was shaped differently. The register of the previous client's TM survives into the new client's document. This is not carelessness, and it is not a sign of a weak linguist. It is how reading works, and it happens to people who would have written something better from an empty field.

A concrete case: a marketing DOCX for a new client, pre-translated at 75 against a TM built on a different product line in the same industry. 160 segments came back filled. The edited file read like a repair job, because that is what it was, and the client's reviewer flagged tone across the entire document rather than segment by segment, which is the worst kind of feedback to receive. We retranslated roughly 40 segments from source with the TM detached. The pre-translation had saved maybe two hours and cost a day and a half.

The second failure mode is status laundering, and it is worse because it leaves no trace. If engine output arrives confirmed, the QA profile skips it and the reviewer trusts it. A number error or an inverted negation inside a confirmed segment walks straight into delivery. We have seen 15 mg come through as 1.5 mg in a confirmed segment in a file that passed QA cleanly, because QA had been configured to check unconfirmed segments only. The file was technically compliant with the process. The process was the problem.

The third is pricing, and it is the one that damages relationships. Pre-translate and then pay a full new-word rate, and you have paid for the engine and for a human to ignore it. Discount against match percentages nobody sampled, and you have priced a repair job as a review. Translators notice the second one immediately, they rarely argue about it the first time, and they quote higher or stop answering the next time.

A five-minute check before you run the batch task

Read the analysis report before you decide anything, and read the distribution rather than the weighted total. A weighted figure of 62% can mean 60% exact matches and almost nothing in between, which is a gift, or a thick band sitting at 75 to 84, which is a trap. The single number cannot tell those apart, and it is the number most quoting templates are built on.

Then sample. Pull thirty segments out of that mid band, open them beside the source, and read them the way the editor will. Ask one question of each: would I rather fix this or write it? If the answer is write it more than a third of the time, the threshold is too low for this file. The whole exercise takes about ten minutes, and we have yet to see it fail to change somebody's mind about at least one project in a batch.

Check where the memory came from while you are in there. Who confirmed these entries, for which client, and when? An entry confirmed by a reviewer you trust and an entry imported from an alignment nobody has looked at are the same colour in the editor, and they should not be treated to the same threshold.

Last, check segmentation. If the TM was built in a tool with different segmentation rules, matches come back low on content that is word-for-word identical, and you will conclude the memory is useless when the rules are the problem. Twenty segments is enough to tell. Our longer piece on how translation memory reuse compounds across projects covers the measurement side of this in more detail.

Then set the threshold for this file, rather than for your account.

When there is no translation memory to pre-translate from

Most of the jobs we see do not have one. First project from a client, a one-off DOCX, a PPTX deck somebody needs by Thursday. TM pre-translation returns nothing on those, and the only pre-fill available is an engine.

That is a legitimate use, with conditions attached. Segment correspondence has to survive the round trip, so the reviewer can work line by line instead of comparing two documents in two windows. Terminology has to be settled before generation rather than corrected afterwards, because an engine that chose the wrong term chose it consistently and you now have 300 instances of it. And the output should not arrive confirmed, for all the reasons above.

If you need the AI step to produce something a CAT tool can actually ingest, that is the gap SnapIntel was built for. You upload a DOCX, XLSX or PPTX, review and approve the glossary and the translation prompt before translation starts, and get back the translated file with a quality rating and a QA report, plus a neutral source/target XLSX export: source in one column, target in the next, one segment per row, which is the shape most CAT tools will take as TM content without a conversion step. The approval gate is the part that matters for this argument: terminology gets decided before the engine runs, not patched after. The import mechanics are in how the neutral XLSX export moves into any CAT tool.

What none of this does is replace a memory for a repeat client. The first project is where the memory comes from. Confirm the segments, export them, attach them next time, and the second project becomes the versioned-document case from earlier, which is where the real saving has always lived. MTPE on an engine-filled file is a reasonable way to get through project one. It is a poor substitute for owning the memory by project four.

What to change on your next project

Stop treating pre-translation as a setting on your account and start treating it as a decision per file. Three rules cover most of it.

Lock 100% matches only when you know who approved them, and name that person in the project notes so the next PM does not have to guess. Everything else arrives unconfirmed, with no exceptions for engine output. And raise the fuzzy floor to 85 for a month, then compare how long files take to review against the month before. Most of the teams we have talked into that have not gone back.

Then write the decision down somewhere the quoting side of the business can see it. The reason pre-translation keeps going wrong is not that project managers choose badly; it is that the choice gets made in a dialog box, by one person, thirty seconds before assignment, and never reaches the person who priced the job or the person who has to review it. Two lines in the project record fix that: what was pre-translated, at what threshold, and at what status. A translator who can read those two lines will tell you when the setting is wrong, which is cheaper than finding out from the client's reviewer.

We will also say the obvious thing, since nobody in this industry says it often enough: on a genuinely new document for a genuinely new client, the right threshold is sometimes nothing at all. Leaving the target fields empty is a legitimate configuration, not a failure to use your tooling.

If you do only one thing, do the sampling check: thirty segments from the mid band, read against source, before the batch task runs. Ten minutes, once per project, and it is the difference between pre-translation that pays for itself and pre-translation that quietly bills you for the privilege.

Newsletter

Get the next article without checking back.

We send occasional product notes and workflow essays when there is something worth reading.

Need the product walkthrough instead? Read the docs.

We care about your data. Read our privacy policy.