Translating Safety Data Sheets and Compliance Documents: What You Can and Cannot Automate
Learn what you can safely automate when you translate a safety data sheet, where regulated phrases must be reused verbatim, and what still needs a human.

Someone in procurement forwards a twelve-page document and asks how quickly it can be in Polish. It is a safety data sheet for an industrial solvent, and the shipment leaves Thursday. If you have ever had to translate a safety data sheet under that kind of deadline, you know the uncomfortable part already. Most of the file is ordinary technical prose that any competent AI handles well. A small fraction of it is regulated text where a fluent, sensible, well-written rendering can still be a compliance failure. Telling those two parts apart is most of the work, and it is the part nobody documents.
Why a safety data sheet behaves differently from other documents
An SDS is not a description of a product. It is an instrument that travels with the product and that other people act on. A warehouse supervisor reads section 7 to decide where the drums go. A nurse reads section 4 while someone is holding their eye open under a tap. A downstream formulator reads sections 2, 3, and 8 to build their own risk assessment. The document has legal weight in every market it enters.
That is why the format is fixed rather than suggested. The sixteen-section structure comes from the UN Globally Harmonized System and is written into law in most jurisdictions that matter for trade. In the EU it sits in Annex II to the REACH Regulation, revised by Commission Regulation (EU) 2020/878. In the United States it sits in OSHA's Hazard Communication Standard at 29 CFR 1910.1200. The section numbers, their order, and their headings are prescribed. You cannot merge section 12 into section 13 because the target language makes them read awkwardly, and you cannot drop a section because the supplier left it blank.
For anyone coming from general business translation, this takes some adjusting. In a marketing brochure, a translator who restructures a clumsy paragraph is doing good work. In an SDS, the same instinct produces a document that reads better and passes inspection worse. The structure is part of the content.
There is a second consequence that matters for automation. Because the format is rigid and repeated across a supplier's whole catalogue, an SDS set is unusually well suited to a controlled, template-aware process. The same rigidity that makes free translation risky makes machine handling reliable, as long as the process knows which parts are frozen.
What the rules actually require in each language
The language requirement is easy to state and easy to get wrong. Under REACH Article 31(3), an SDS must be supplied in an official language of each Member State where the substance or mixture is placed on the market, unless that Member State decides otherwise. Placing a product on the market in five countries can therefore mean five language versions, and in some countries more than one. Belgium alone can require Dutch and French, and German depending on the region you are supplying.
The United States works differently. OSHA requires the SDS in English. A Spanish version is welcome and often sensible for a workforce, but it is an addition rather than a substitute, which means a US-bound file and an EU-bound file are not the same translation problem at all.
Two practical points follow. First, the source document you are translating from may itself be a translation. We have seen a chain where a Japanese manufacturer's SDS was rendered into English by a distributor, then into three European languages from that English version, with a classification error introduced at step one and faithfully preserved through every downstream file. If the English version is an intermediate, treat it as suspect rather than authoritative.
Second, the requirement is jurisdictional, not linguistic. Supplying a correct Polish translation of a document written for the German market does not make the document correct for Poland. Section 15 in particular describes which regulations apply to that substance in that territory, and translating it produces a Polish-language account of German obligations. That is not a translation defect. It is a scoping mistake that translation cannot fix.
If your regulatory position is at all unclear, this is the point to ask a compliance advisor rather than a language provider. The rules change, and they differ by product category.
The phrases you must never let AI translate freely
This is the single most useful thing to know about SDS work, and it surprises people who come from other document types.
Hazard statements and precautionary statements are not free text. Under the CLP Regulation (EC) No 1272/2008, hazard statements sit in Annex III and precautionary statements in Annex IV, and both annexes are published in every official EU language. The Polish wording of H315 already exists. Your job is to retrieve it, not to produce it. The same applies to signal words, to hazard class and category names, and to precautionary combinations such as P305+P351+P338.
Here is the failure mode in practice. On a German-to-Polish set we reviewed, an AI produced a Polish rendering of P305+P351+P338 that was accurate, natural, and clearly understandable. A worker reading it would have done exactly the right thing. It was also not the Annex IV wording, and an inspector comparing the file against the official list would have marked it. Nothing about the output looked wrong. That is what makes this category dangerous: the error is invisible to every quality check that asks whether the translation is good.
Identifiers behave the same way. CAS numbers, EC numbers, UN numbers, index numbers, and the Unique Formula Identifier introduced by Annex VIII to CLP are codes. They are copied, never converted, never reformatted, never "cleaned up". An AI that helpfully normalises a UFI's grouping or strips a leading zero from a CAS number has broken the document in a way that no linguistic review will catch.
The practical response is to move these strings out of the translation problem entirely. Build them once as a lookup, keep them per language, and treat any deviation as an error rather than a stylistic choice.
What automation handles well when you translate a safety data sheet
Now the good news, because the pessimistic half of this article is not the whole picture.
Sections 9 through 14 are largely descriptive prose. Physical and chemical properties, stability and reactivity, toxicological and ecological information, disposal considerations, transport information. This material is technical, repetitive, heavily templated, and almost identical across products in the same family. A supplier with two hundred SDS files in a catalogue will usually find that most of the wording in these sections repeats.
That profile suits AI translation well, and it suits it better than a rotating pool of human translators does. Consistency across two hundred files is a machine strength and a human weakness. A team of five freelancers working a catalogue over eighteen months will produce five defensible ways of saying the same sentence about thermal decomposition. A single controlled process with a fixed glossary will produce one.
Sections 7 and 8 sit in between. Handling and storage advice is ordinary prose and translates cleanly. Exposure control values in section 8 are numbers with national scope, and they are not translation material at all. More on that below.
Two conditions make this work. The source needs to arrive as DOCX or XLSX rather than a flattened PDF, because extracting text from a PDF adds a step where table structure and superscripts get lost before translation even starts. And the process needs a glossary that is attached to the job rather than living in someone's head, so that "flash point" and "auto-ignition temperature" resolve the same way in file 199 as they did in file 1.
This does not apply cleanly if your SDS files are scans of paper documents, or if the supplier has embedded the tables as images. Fix the source before automating anything downstream of it. Our walkthrough on translating technical manuals with their tables and warning labels intact covers the same extraction problem in more detail.
Where automated translation of compliance documents goes wrong
Number handling is the most common failure and the least dramatic-looking. Decimal separators differ between locales, and a concentration written as 0,5 % in a German source can become 0.5 % or, worse, 5 % after a careless pass. We hit this on an XLSX substance register where the workbook's own locale settings and the translated cell values disagreed, and the resulting file listed a component at ten times its actual concentration. It took an hour to find and it would have taken a very bad afternoon to explain.
Unit conversion is the second. An AI asked to make a document natural in the target market may convert °C to °F, or mg/m³ to ppm, without being asked. In an SDS this is a content change wearing the costume of a translation choice. The prompt has to forbid it explicitly, because the default behaviour of a helpful model is to help.
Substituted values are the third and worst. Occupational exposure limits are national. Germany's AGW values, an EU indicative limit, and a Polish NDS value for the same substance are three different numbers that exist for three different legal reasons. A model that understands it is producing a Polish document may reach for a Polish-looking number. Anything that changes a value in section 8 is regulatory work performed by an unqualified party.
Then there is the empty field. SDS files are full of gaps, because for many substances the data genuinely does not exist. "No data available" is a legitimate, meaningful entry. Language models dislike blank space, and a model given a sparse section 11 will sometimes produce fluent, plausible, entirely invented toxicological prose. This is the failure that worries us most, because unlike a wrong number it looks exactly like a good translation.
None of these are arguments against automation. They are arguments for automation with the boundaries written down, which is a different thing.
A workflow that puts the machine and the human in the right places
Start by locking the regulated strings. Pull the H and P statements your product range actually uses from the official annex texts in every target language, build them into a glossary, and reuse it across every file forever. This is a one-time job that removes the highest-risk category from the translation entirely.
Freeze the identifiers next. CAS, EC, UN, index numbers, UFI, internal product codes, and the emergency telephone number in section 1.4 go on a do-not-translate list. The phone number deserves separate thought: it has to be reachable from the market the document is going to, which is a business decision rather than a linguistic one.
Then translate the descriptive sections with the glossary attached and an instruction set that says what not to do. Do not convert units. Do not reformat numbers. Do not substitute national limit values. Do not fill empty fields. Written prompts sound like a formality until the first time one of them catches something.
Check structure before you check language. Count the sections, count the subsections, count the table rows, and count the numeric values in the source and the target. A structural comparison catches dropped rows and merged sections in seconds, and those are exactly the errors a fluent reader skims past.
Finally, put a human on sections 1 through 4, section 8, and section 15. That is a small enough slice to be affordable even on a rush job, and it covers identification, hazards, composition, first aid, exposure limits, and regulatory scope. If nobody on your team reads the target language, our guide on checking translation quality without a speaker of the language sets out what you can verify anyway.
This is roughly the shape we built SnapIntel around. You upload the DOCX or XLSX, review and approve a glossary and a translation prompt before anything runs, and get back a translated file with a QA report and a quality rating rather than just text. The approval step exists precisely because documents like these need their terminology settled before translation starts, not argued about afterwards.
What to verify before the file leaves your hands
Run four checks, in this order, on every language version.
Confirm that all sixteen sections are present, numbered correctly, and in order. Compare every H and P statement against the official annex text for that language, character for character. Compare every number and unit in sections 3, 8, and 9 against the source, including decimal separators. Confirm that section 15 describes the destination market rather than the source market, and that the emergency number in section 1.4 can actually be dialled from there.
If any of those four fails, the problem is upstream of translation and no amount of linguistic review will resolve it. If all four pass, you have a document that a machine did most of the work on and that a person is willing to sign.
The one thing not to do is treat an SDS as a document translation job that happens to be technical. It is a regulated form with some prose in it, and the prose is the easy part.