How to Use Claude for Document Translation: A Practical Guide for Translators
How to use Claude for translation in real projects: prompt structure, glossary control, chunking long documents, and the review checks to run before delivery.

A translator emailed us last month asking whether she could drop her paid document translation tool and just paste files into Claude instead. Fair question. For some work, yes. The limits show up faster than most people expect, though. This guide covers how to use Claude for translation in real projects: what belongs in the prompt, how to get a DOCX in and back out without wrecking it, how to hold terminology steady across sixty pages, and what to check before a file reaches a client. Across the translators and agencies we work with, one pattern repeats. People who get usable output treat Claude like a capable junior translator who needs a proper brief. People who get mush type "translate this to German" and press enter.
Why translators reach for Claude in the first place
The appeal is instruction-following. A neural MT engine gives you a target sentence and no way to argue with it. You can tell Claude that this is a pump maintenance manual for field technicians in Kazakhstan, that "housing" means the pump casing and not a residential building, that the client uses formal address throughout, and that anything inside curly braces is a placeholder. It will mostly do those things, and when it gets one wrong you can ask why and get an answer that is sometimes useful.
The second reason is that it reads whole passages rather than isolated segments. Segment-level MT is blind to the sentence three paragraphs up that defined the abbreviation. An LLM working on a full section can carry that across. We have watched this matter most in legal and technical text, where a term is defined once and then referenced twenty times in shorthand. Slator has been tracking the shift from segment-based MT toward LLM-based document translation since 2023, and this context effect is the substance behind it.
The third reason is more mundane: register. Marketing copy translated by a generic engine reads like a specification sheet. Claude will adjust tone if you describe the audience, which saves real post-editing time on anything persuasive.
None of this makes it a translation environment. It has no translation memory, no segment status, no QA rules engine, no way to prove what changed between two versions. If you want a model-level comparison before you commit, we wrote up Claude versus GPT-4 for translation separately. This guide is about the workflow around whichever model you pick.
How to use Claude for translation: the prompt structure that works
The biggest quality difference we see comes from the brief, not the model. "Translate naturally" is not a brief. It tells the model nothing it did not already assume.
A prompt that produces reviewable output has six parts, and they are worth writing once per client and reusing.
Start with the document type and its purpose. "This is a supplier quality manual that will be read by production line supervisors" constrains vocabulary more than any style adjective.
Name the language pair with the variant. "Spanish" is not a target language. "Spanish for Mexico, formal usted throughout" is.
State the register explicitly, including the one thing translators forget: whether to keep the source's sentence structure or rewrite for readability. Technical documentation usually wants the former. A press release wants the latter.
Paste the glossary inline as a plain source-to-target list. Not an attachment, not a description of where the glossary lives. Inline.
Add a do-not-translate list. Product names, internal system names, code identifiers, anything in a fixed format like part numbers. This is the cheapest single intervention available and it prevents the most annoying class of error.
Finally, specify the output shape. Ask for a two-column table with source on the left and target on the right, or for the translation only with no commentary. If you do not say, you will get preambles like "Here is the translation" and occasional explanatory notes woven into the target text, which then have to be stripped out by hand.
One more thing that helps: tell the model what to do when it is unsure. "If a term is ambiguous, translate it and add [CHECK] after it" gives you a searchable flag instead of a silent guess.
Getting the document in and out without losing the formatting
This is where the chat-window approach breaks first, and it has nothing to do with translation quality.
Claude reads text. A DOCX is not text; it is a zip archive of XML with styles, numbering definitions, tracked revisions, comments, headers, footers, footnotes, text boxes, and table structure. Copy-paste from Word into a chat window flattens most of that. Tables arrive as ragged lines. Numbered lists lose their numbering scheme. Footnotes either vanish or get dumped at the end with no anchors. Content in headers, footers, and text boxes usually does not come along at all, which means you can deliver a file that looks complete and still has the company address untranslated on every page.
Two routes work in practice.
The first is the bilingual table. Before translating, restructure the source into a two-column table: segment on the left, empty on the right. Translate the right column only. Then rebuild the document from the filled table. This is laborious and it is what people actually do for short files.
The second is a spreadsheet round trip. Extract the translatable strings into an XLSX with a source column, translate into a target column, and write the results back into the original file. This is how most XLSX workbook translation is handled anyway, and it has the pleasant side effect of giving you a bilingual file you can reuse as a translation memory import later. If you work in a CAT tool, this is also the format your reviewer can open without installing anything.
Whichever route you choose, build a coverage check into it. Search the delivered file for source-language characters. In Cyrillic-to-English work that is trivial; in Spanish-to-English it means spot-checking headers, footers, footnotes, chart labels, and anything inside a shape.
Chunking a long document so the context survives
A large context window is not the same thing as sustained attention across it. We have seen terminology drift inside a single long conversation often enough to treat it as the default expectation rather than an edge case.
The worst example we ran into was a 62-page pump maintenance manual, Russian into English, translated in one long chat in roughly eight passes. By pass six, "корпус" had become "housing", then "casing", then "body", all three within the same chapter. Nothing was wrong at the sentence level. The document was unusable as a reference because a technician searching for "casing" would find two thirds of the mentions.
Three habits fix most of it.
Split on structural boundaries, not word counts. A chunk that ends mid-procedure produces a target that reads as though it started mid-thought, because it did. Chapter and section breaks give the model a complete unit of meaning.
Restate the brief with every chunk. Not a reference to it, the whole thing: glossary, do-not-translate list, register instruction. It costs you a paste and it is the difference between chunk seven behaving like chunk one and chunk seven improvising.
Keep a running glossary and grow it. When the model coins a term you approve of, add it to the list before the next chunk. By the end of a long document your glossary is longer than it was at the start, and that is the point. This is also the fastest way to build a termbase for a client who never gave you one.
For documents that get reissued, save that final glossary. The next revision of the same manual should start from it rather than from nothing.
Keeping terminology consistent with a glossary Claude will actually follow
Glossary adherence is the complaint we hear most about LLM translation, and it is usually a glossary problem rather than a model problem.
Three things go wrong. The list is too long, so it gets diluted; past a few hundred entries, adherence degrades noticeably. The list contains terms that do not need controlling, so the ones that matter lose priority. And nobody verifies afterwards, so violations ship.
Tier the list. Put hard constraints first, marked as non-negotiable: regulated terminology, client-mandated renderings, product names. Put preferences second. Put the do-not-translate items in their own block, because they are a different instruction: leave these alone, do not render them at all. We wrote a longer piece on how to choose which terms go in an AI translation glossary that covers the selection criteria.
Then verify. Take your ten most important terms and search the delivered target for each approved rendering and for the obvious wrong ones. Ten searches, two minutes, and it catches the errors that damage client trust most. Terminology errors are the category clients notice and remember, well out of proportion to their frequency.
One habit worth adopting: when a term genuinely has no settled equivalent, decide it yourself rather than letting the model decide it eight times. Write your decision into the glossary with a one-line rationale. You will need that rationale when the client's reviewer challenges it in four months.
Reviewing the output before anyone else sees it
Never deliver a first pass. This is not a statement about AI; the same holds for human first drafts. What changes is the order of the checks, because LLM output fails differently than human output.
Run completeness first. Skipped segments are the failure that hurts most and the easiest to miss, because the surrounding text reads fine. Compare source and target segment counts. Look specifically at anything that was hard to extract: tables, footnotes, headers, speaker notes in a PPTX, hidden or filtered rows in a workbook.
Run numbers, dates, units, and currency second. These are mechanical and objectively verifiable, and LLMs get them wrong in ways that MT engines do not, particularly in digit-grouping and decimal separators. A figure rendered as 1,500 instead of 1.500 in a German target is a real error with a real cost.
Run terminology third, using the searches described above.
Then check tags and placeholders. Anything in braces, percent signs, angle brackets, or double-brace template syntax must survive byte-identical. This is why the do-not-translate block exists.
Only then read for fluency and accuracy. By this point you are reading for meaning rather than hunting mechanical defects, which is a faster and more accurate way to read.
The failure mode that worries us most is invention. In a contract translation we reviewed, the target contained a cross-reference to a clause number that does not exist in the source. It was fluent, plausible, and formatted correctly. A reviewer skimming for readability would pass it. This is why accuracy review against the source is not optional, and it is why ISO 18587, the post-editing standard, requires comparison with the source text rather than monolingual review. If you want a framework for recording what you find, MQM error categories are the usual starting point.
Where a chat window stops being enough
Everything above works for a single document, a few pages to maybe ten thousand words, one or two language pairs, translated by someone who will personally review the result. That is a real and common job, and doing it in a chat window is a reasonable choice.
It stops working on three axes. Volume, because pasting and reassembling forty files by hand is a worse use of a translator than translating. Repeatability, because a workflow that lives in your memory produces different results next month. And evidence, because when a client asks what glossary was used and what the quality check found, "I checked it myself" is not an answer that survives procurement.
If you need the structured version of this workflow rather than the manual one, that is roughly what we built SnapIntel for. You upload a DOCX, XLSX, or PPTX, run domain analysis, generate or paste the glossary and the translation prompt, approve both before anything starts, and get back the translated file in its original format plus a neutral source/target XLSX for TM import and a QA report with a quality rating. The preparation steps are the same ones described above; the difference is that they are recorded rather than retyped. There is a trial if you want to test it on a real file: snapintel.io.
What to do next
Take the last document you translated and write its brief down properly: document type, audience, target variant with formality, register instruction, glossary as a source-to-target list, do-not-translate block, and output format. Save it as a text file named after the client.
That file is the thing that improves your output, and it is reusable. Next time, paste it, translate one chapter, and run the five checks in order: completeness, numbers, terminology, placeholders, then meaning. If terminology drifted, your glossary was too short. If numbers broke, add an explicit instruction about separators. Fix the brief rather than the translation, and the second document costs you a fraction of the first.