The Last Word

Can you trust AI translation with regulated content? Only if you can prove it.

Written by Mathias Saint-Jean | Sep 2, 2026, 1:47:32 PM

Start with what is true. AI translation in 2026 is fast, cheap and, for most content, genuinely good. One of the market's largest platforms reports that enterprise usage of AI translation grew 218% in a single year, and the contracts have followed: quality guarantees that would have sounded reckless five years ago are now standard line items. The language industry itself, worth an estimated USD 72.6 billion in 2025 according to the 2026 Nimdzi 100, is being repriced around this reality, with buyers demanding cost reductions of up to 75% through AI and vendors re-engineering their pipelines to deliver them.

For a support article, an internal memo or a product listing, this is simply progress. The error rate is low, the cost of a miss is lower, and human review of every sentence was never a good use of anyone's budget. If that were the whole story, this article would end here. It is not the whole story, because the economics that make AI translation cheap are the same economics that quietly remove the one thing regulated content cannot lose: a person who is accountable for the sentence.

The remaining failures of AI translation share one property: they read beautifully. The 2026 Nimdzi 100 names the problem in three words, warning that AI output is prone to 'hallucinations, deceptive fluency, and a lack of the cultural nuance required for critical texts'. Deceptive fluency is the operative phrase. A clumsy error announces itself. A fluent error looks like everything around it. The American Translators Association made the same point in its May 2025 statement on AI: the greatest danger of AI-generated translation is that it 'may appear accurate to the general observer'.

This is not a theoretical risk, and the people it fools are not careless. In May 2025, in Lacey v. State Farm, a federal court in California sanctioned two law firms USD 31,100 after AI-fabricated citations survived the review of multiple attorneys and reached the court. The special master's words describe the trap better than any study could: 'Plaintiff's use of AI affirmatively misled me. I read their brief, was persuaded (or at least intrigued) by the authorities that they cited, and looked up the decisions to learn more about them, only to find that they didn't exist.' A trained expert read fluent, confident, wrong text and believed it. Swap the legal brief for a translated dossier reviewed under deadline and the mechanism is unchanged: what fooled a federal special master is exactly what a post-editor faces on every page, in a language the final approver often cannot read. As of August 2026, a public database maintained by legal researcher Damien Charlotin tracks 1,933 court decisions involving AI-fabricated content, and it grows daily.

Healthcare has its own version of the same story. An Associated Press investigation found that Whisper, a speech-to-text model used in tools deployed across roughly 7 million medical visits, invented text that nobody said, including nonexistent medications, and a peer-reviewed study rated 38% of the hallucinations it identified as potentially harmful. One deployment detail matters more than every statistic: the original audio was erased after transcription. Fluent output, no way back to the source.

You might expect the industry's automated scoring tools to catch what tired human eyes miss. The benchmark evidence says otherwise. When the field's main evaluation campaign, WMT25, put these tools to the test, it found that 'accurate error detection' remains a persistent challenge, that robustness across the diversity of languages is still a major weakness, and that at the level of the individual sentence, which is exactly where liability lives, LLM-based evaluators are still outperformed by older baseline metrics.

Do not expect the market to correct this on its own, because every incentive points the other way. In a flat market where buyers demand cuts of up to 75%, the only cost a vendor can meaningfully compress is human attention per word, which means the industry is structurally pushed to remove exactly the control that deceptive fluency defeats. Even the research community rations rigorous evaluation: WMT25 applied full MQM scoring to just two of the fifteen language pairs it assessed with human judges, because doing it properly is expensive. A vendor promising MQM at industrial scale is promising something the field that invented the method cannot afford to do itself. Which brings us to the numbers the vendors publish about themselves.

Here is the part most AI translation buyers have not priced in: U.S. regulators have already written down what they expect, and while good quality is assumed everywhere in their rules, what the rules actually demand is proof. Federal health regulation names machine translation explicitly. Under Section 1557 of the ACA, 45 CFR 92.201(c)(3), if a covered entity uses machine translation when the text is 'critical to the rights, benefits, or meaningful access' of a limited-English-proficiency individual, when 'accuracy is essential', or when the source material is complex or technical, 'the translation must be reviewed by a qualified human translator'. That provision is in force today. It survived the 2026 court vacatur that struck other parts of the rule, and it survived the 2025 executive order designating English as the official language, which explicitly did not require agencies to stop providing multilingual services and left Title VI obligations intact.

Finance reads the same way. FINRA's 2026 Annual Regulatory Oversight Report states that its rules 'continue to apply when firms use GenAI or similar technologies in the course of their businesses, just as they apply when firms use any other technology or tool', and then describes what sound practice looks like: 'storing prompt and output logs for accountability and troubleshooting; tracking which model version was used and when'. Read that list again, because prompt logs, output logs, model versions and timestamps amount to a regulator publishing the specification of an audit trail in the same year the market is proposing to remove the human from the loop. In pharma, the standard has existed for decades: 21 CFR 11.10(e) requires 'secure, computer-generated, time-stamped audit trails' recording who created, modified or deleted every electronic record, and the FDA's January 2025 draft guidance on AI extends the same logic through a risk-based credibility assessment for a defined context of use, documented and reviewable.

Now put the vendor contract next to those requirements. CSA Research, the industry's own analyst firm, states it plainly: the small print for all LLMs 'carries disclaimers about accuracy and shifts responsibility for any errors to the user'. A high score makes errors rare, and a defensible record makes them survivable; you need both, but only one of the two appears in your contract. The gap between those documents is yours to carry, because none of the standards the industry certifies against will close it. ISO 5060, the 2024 standard for evaluating translation output, scores error types and explicitly excludes quality assurance processes from its scope. ISO 17100, the flagship translation services standard, dates from 2015, excludes machine translation post-editing entirely, and has been flagged for revision since 2025. MQM, the framework behind most quality dashboards, classifies errors and says nothing about provenance. Every translation standard measures the output, every regulated-records rule traces the process, and nobody's certificate covers the distance between the two.

So, can you trust AI translation with regulated content? Yes, if you stop treating your content as one bucket. CSA Research's Alison Toon puts the rule in one sentence: 'You can't just generalize and put all of your content through the same process, assuming the same level of accuracy, and therefore risk, for every type of content and every language pair.' Route by consequence, not by word count, and in practice the line is easier to draw than it sounds. Support macros, interface strings, product listings and internal documentation belong in the fast lane, where you take the discount with a clear conscience. Package inserts and labeling, informed consent forms, prospectuses and fund documents, benefit letters and claim decisions belong in the accountable lane, where a qualified human reviews the content and the review leaves a trace: reviewer name, timestamp, the approved source the translation is grounded in, the terminology version applied.

That last clause is the one to negotiate, because a human in the loop who leaves no record is a hope, not a control. Section 1557 does not ask whether a person was somewhere in your workflow; it asks whether this translation was reviewed by a qualified human translator, and the only way to answer is to be able to show it.

This is, in full transparency, the problem OTTO was built around. Rather than adding a score on top of a generic pipeline, OTTO grounds every translated segment in your approved source content, applies your terminology as a versioned asset rather than a suggestion, routes high-stakes content to qualified human review, and captures the whole chain, reviewer, timestamp, model and provenance, as a native part of the deliverable. Not because quality scores do not matter, they are measured there too, but because when the question comes, and in regulated industries it eventually comes, the answer has to be a record.