Back to blog
Published

How to translate legal documents with AI without losing accuracy

Legal document translation AI works well on some files and fails badly on others. How to set terminology, confidentiality and review rules that hold up.

How to translate legal documents with AI without losing accuracy

Legal translation is where the distance between "this reads fine" and "this is correct" gets expensive. Most of the legal document translation AI conversation skips over that distance. It asks whether a model can produce fluent target-language prose, which current models can, and skips the part where one mistranslated limitation of liability clause changes what a party owes. We talk to agencies and freelancers who handle contracts, filings and corporate records every week. The ones getting usable results from AI are not the ones with the cleverest prompt. They are the ones who decided in advance which parts of the job a model is allowed anywhere near.

Where legal document translation AI holds up, and where it doesn't

The honest split is not by document type. It is by how much of the meaning depends on the legal system the text came from.

AI handles repetitive commercial paperwork well. NDAs, standard supplier agreements, purchase terms, board resolutions, powers of attorney, corporate registry extracts. These documents repeat themselves across clients and across years, the sentence structures are conventional, and the terminology is stable. One agency we work with runs a quarterly batch of roughly forty supplier contracts for the same manufacturing client. The contracts differ in parties, amounts and a schedule or two. AI drafting plus a focused review pass cut the turnaround from two weeks to four days, and the reviewer's notes shrank with each cycle because the glossary kept absorbing the corrections.

Then there is the other category. Litigation documents where the formal requirements of the receiving court matter. Anything drafted around a concept that exists in one legal system and not in the other. Instruments where the translator is making an interpretive choice that a lawyer will later rely on.

The failure here is not fluency. A model will produce a confident, grammatical English sentence for a German Geschäftsführer and call that person a "director", which quietly imports a set of duties and liabilities from company law that do not apply. Common-law "consideration" has no clean equivalent in most civil-law systems, and a model will usually supply a word for "payment" or "compensation" and move on. The output looks finished. That is the problem: nothing in the file signals that a decision was made.

So the rule we would give is narrow. Use AI for volume and for first drafts on conventional commercial text. Do not use it to resolve a concept that does not exist in the target system. Those stay with a human who can explain the choice.

Build the terminology layer before you translate anything

Legal translation terminology is not decoration on top of a translation. In a contract set it is most of the work.

A contract defines its own vocabulary. "Agreement", "Services", "Affiliate", "Confidential Information" are not ordinary words in that document, they are defined terms, and once defined they must render identically on every one of the next sixty pages. Models do not reliably do this on their own. They vary their word choice, because varying word choice is what fluent writing usually requires. Here it is a defect.

Before you translate, pull the defined terms out of the definitions clause and fix their target equivalents. Add the party names exactly as they should appear, including the legal form suffix. Add any statute or regulation titles cited in the document, with the official target-language title where one exists and the source title preserved where it does not. Add the client's house preferences: whether they want "shall" rendered as an obligation or a present tense, whether a company name stays in Latin script. If you want a longer treatment of what belongs on that list and what does not, we wrote about choosing which terms go in your AI translation glossary.

Two practical notes from doing this badly first. Keep the glossary per client and per document family, not one global legal glossary. A term that is right for a Swiss client's employment contracts will be wrong for a Spanish one. And record the reason next to the entry, even four words of it. Six months later nobody remembers why that term was chosen, and without the reason the next translator overrides it.

This works best on document sets. For a single one-off contract from a client you will never see again, building the terminology layer costs more than it returns, and a careful read-through serves you better.

Jurisdiction, not language, decides the right term

The question a legal translator answers is not "what is this word in French". It is "what does this instrument do, and what is the French instrument that does that".

Sometimes the answer is a functional equivalent. Often there isn't one, and then you have a choice to make, and the choice has to be consistent across the document and defensible to the client. The usual options: keep the source term and add a short explanatory gloss on first use, use a descriptive rendering, or borrow the term and italicise it. All three are legitimate. Mixing them inside one document is not.

A concrete case. A corporate client sent an agency a set of English-language shareholder agreements for translation into Russian for a group restructuring. The agreements used "equitable relief" throughout. There is no equity jurisdiction in Russian law to map onto. The AI draft produced a phrase meaning roughly "fair compensation", which is wrong in a way that changes the remedy from an injunction into money. The fix was a descriptive rendering, agreed with the client's counsel, applied once and then locked into the glossary so every later file in that restructuring matched.

Second case, smaller and more common. A Kazakh notarial document translated into English listed a date as 03.09.2026 and an amount as 1 500 000,00. A model will often carry the source punctuation straight through, and an English-language reader will read that amount as fifteen hundred thousand with a comma decimal or stall on it entirely. Dates, amounts, registration numbers, identity document numbers: check all of them by hand, every time. They are the cheapest errors to catch and among the most damaging to miss.

Confidentiality is a contract question before it is a technical one

Most of the confidentiality discussion around AI translation gets stuck on encryption and data centres. The question that bites first is simpler: what did you already promise this client in writing.

A lot of translation NDAs and framework agreements were signed before AI tooling was normal, and they say something like "the Supplier shall not disclose the Confidential Information to any third party without prior written consent". Putting a litigation bundle through a third-party AI endpoint is a disclosure to a third party. Whether the vendor trains on it does not change the analysis of that clause. We have seen agencies discover this during a client security review rather than before it, which is a bad moment to find out.

So the order of operations runs: read the client contract, then decide the tooling. Where the contract is silent or permissive, get the technical facts in writing from the vendor anyway. What is retained, for how long, where it is processed, whether the content is used for model training, who can access it internally, and what happens on deletion request. For EU personal data you also need to know whether you are a processor or a sub-processor in that chain and whether your own agreement reflects it.

Where the contract blocks it, say so and quote the clause. Several agencies we know now go further and add an AI clause to new client agreements, declaring that AI-assisted drafting with human review may be used unless the client opts out in writing. That conversation is easier before a project than during one.

Free consumer-grade endpoints deserve a separate warning, and we covered the reasoning in more detail in is Google Translate safe for confidential documents. The short version for legal work: paste nothing privileged into anything whose terms you have not read.

Certified and sworn work follows its own rules

There is a persistent confusion worth clearing up, because clients cause it and then translators inherit it.

A certified translation, in most English-speaking practice, means the translator or agency attaches a signed statement that the translation is complete and accurate. A sworn translation, in much of continental Europe and in a number of other jurisdictions, means a translator formally authorised by a court or ministry signs and stamps it, and that person's authorisation is what the receiving body recognises. Notarisation is a third thing again: the notary verifies the identity of the person signing, not the quality of the translation. Clients routinely ask for one while needing another.

The relevant point for AI: in a sworn or certified workflow, the signature carries personal liability for the content. Nothing in that arrangement cares how the draft was produced. A sworn translator can use an AI draft and sign the result, and the responsibility sits entirely with them, which in practice means they have to read every line as if they had written it. The time saved is in typing, not in verification.

The exception to watch is procedural. Some receiving authorities, courts and bar associations have started issuing rules on machine-translated submissions, and those rules differ by country and change. If a document is going to a court, an immigration authority or a registry, confirm the requirements for that specific body before you decide how the file gets produced. We would not make a general claim here in either direction, because the picture genuinely varies and a confident generalisation would be wrong somewhere expensive.

A review process built around what AI actually gets wrong

Generic proofreading is the wrong instrument. AI output in legal text fails in a specific, predictable set of ways, and a checklist aimed at those ways catches far more per hour than a careful read.

What to check, in roughly this order. Numbers, dates, amounts, percentages, and all reference numbers, against the source. Negations, because a dropped "not" produces a perfectly fluent sentence with the opposite meaning. Modal verbs: shall, may, must, is entitled to, is required to. The distinction between an obligation and a permission is the whole content of a clause and models flatten it. Defined terms, checked for consistency across the full document rather than within a paragraph. Internal cross-references, since "as set out in Section 7.2" has to still point at the right section after translation. Party names and their roles, which get swapped more often than you would expect in documents where Seller and Buyer appear in alternating clauses. And omissions, which are the quiet one: a segment that vanished leaves no trace in the target file.

A QA report that flags number mismatches, glossary violations and untranslated segments mechanically will clear most of that first group before a human opens the file, which leaves the reviewer's attention for the parts that need judgement.

For the clauses where real money sits, liability, indemnity, termination, governing law, dispute resolution, a back-translation of those clauses alone is worth the time. Not the whole document. Back-translating a forty-page agreement is a way to spend two days and learn little. Back-translating five clauses tells you quickly whether the obligation survived the trip.

Decide the post-editing level per document and tell the client which one they are buying. Full post-editing on an agreement that will be signed and relied on. Light post-editing on discovery material being read for relevance, where the client needs to know what a document says, not to publish it. Charging the same for both, or silently doing the lighter one, is where the trust goes.

What to set up before your next legal file

Start with the two documents you already have and probably have not read together: the client's confidentiality agreement and the vendor's data terms. If those two conflict, nothing downstream matters. Fix that first, in writing, before the file arrives.

Then take your most repetitive client, the one whose contracts look the same every quarter, and build the terminology layer for that family: defined terms, party names, statute titles, the client's standing preferences, with a short reason beside each entry. Run the next batch as AI draft plus full post-editing, and keep every correction your reviewer makes. Feed them back into the glossary before the following batch. The second cycle is noticeably cheaper than the first, and the fourth is where the economics actually change.

Keep the review checklist short enough that it gets used. Numbers, negations, modals, defined terms, cross-references, omissions. Six items, in that order, every file.

And hold one line without negotiating it: anything that turns on a concept the target legal system does not have goes to a human who can explain the decision and stand behind it. That is not a limitation of current models that will be patched next year. It is a question about law, and the document needs someone who can be asked why.

Newsletter

Get the next article without checking back.

We send occasional product notes and workflow essays when there is something worth reading.

Need the product walkthrough instead? Read the docs.

We care about your data. Read our privacy policy.