What AI translation actually does well in 2026
Start with what is true. AI translation in 2026 is fast, cheap and, for most content, genuinely good. One of the market's largest platforms reports that enterprise usage of AI translation grew 218% in a single year, and the contracts have followed: quality guarantees that would have sounded reckless five years ago are now standard line items. The language industry itself, worth an estimated USD 72.6 billion in 2025 according to the 2026 Nimdzi 100, is being repriced around this reality, with buyers demanding cost reductions of up to 75% through AI and vendors re-engineering their pipelines to deliver them.
For a support article, an internal memo or a product listing, this is simply progress. The error rate is low, the cost of a miss is lower, and human review of every sentence was never a good use of anyone's budget. If that were the whole story, this article would end here. It is not the whole story, because the economics that make AI translation cheap are the same economics that quietly remove the one thing regulated content cannot lose: a person who is accountable for the sentence.
The question has widened rather than changed. Whether the quality is good enough still matters, and the honest answer still depends on the language, the domain and the stakes. But a second question now sits on top of the first, and it is the one your regulator will ask: when a single translated sentence is challenged, what exactly can you produce?
Where it still fails, and why you will not see it
The remaining failures of AI translation share one property: they read beautifully. The 2026 Nimdzi 100 names the problem in three words, warning that AI output is prone to 'hallucinations, deceptive fluency, and a lack of the cultural nuance required for critical texts'. Deceptive fluency is the operative phrase. A clumsy error announces itself. A fluent error looks like everything around it. The American Translators Association made the same point in its May 2025 statement on AI: the greatest danger of AI-generated translation is that it 'may appear accurate to the general observer'.
This is not a theoretical risk, and the people it fools are not careless. In May 2025, in Lacey v. State Farm, a federal court in California sanctioned two law firms USD 31,100 after AI-fabricated citations survived the review of multiple attorneys and reached the court. The special master's words describe the trap better than any study could: 'Plaintiff's use of AI affirmatively misled me. I read their brief, was persuaded (or at least intrigued) by the authorities that they cited, and looked up the decisions to learn more about them, only to find that they didn't exist.' A trained expert read fluent, confident, wrong text and believed it. Swap the legal brief for a translated dossier reviewed under deadline and the mechanism is unchanged: what fooled a federal special master is exactly what a post-editor faces on every page, in a language the final approver often cannot read. As of August 2026, a public database maintained by legal researcher Damien Charlotin tracks 1,933 court decisions involving AI-fabricated content, and it grows daily.
AI output is prone to hallucinations, deceptive fluency, and a lack of the cultural nuance required for critical texts. The 2026 Nimdzi 100
Healthcare has its own version of the same story. An Associated Press investigation found that Whisper, a speech-to-text model used in tools deployed across roughly 7 million medical visits, invented text that nobody said, including nonexistent medications, and a peer-reviewed study rated 38% of the hallucinations it identified as potentially harmful. One deployment detail matters more than every statistic: the original audio was erased after transcription. Fluent output, no way back to the source.
You might expect the industry's automated scoring tools to catch what tired human eyes miss. The benchmark evidence says otherwise. When the field's main evaluation campaign, WMT25, put these tools to the test, it found that 'accurate error detection' remains a persistent challenge, that robustness across the diversity of languages is still a major weakness, and that at the level of the individual sentence, which is exactly where liability lives, LLM-based evaluators are still outperformed by older baseline metrics.
Do not expect the market to correct this on its own, because every incentive points the other way. In a flat market where buyers demand cuts of up to 75%, the only cost a vendor can meaningfully compress is human attention per word, which means the industry is structurally pushed to remove exactly the control that deceptive fluency defeats. Even the research community rations rigorous evaluation: WMT25 applied full MQM scoring to just two of the fifteen language pairs it assessed with human judges, because doing it properly is expensive. A vendor promising MQM at industrial scale is promising something the field that invented the method cannot afford to do itself. Which brings us to the numbers the vendors publish about themselves.
What a quality score answers
How good the output is on average, across a sample, as measured by a model or a reviewer. That is a genuinely useful engineering signal, and it deserves a place in every pipeline. What it describes, though, is a population of sentences.
What a regulator asks
Who reviewed this particular sentence, and when, and against which approved source document, under which version of the terminology, and whether you can reproduce the result. These are questions about one sentence, and a population average answers none of them.
Quality is a property of the output; evidence is a property of the process. You can sample the first after the fact, but you can only design the second in advance. That is why a pipeline retrofitted with a score never becomes defensible, while a pipeline built around a record is defensible even on its bad days.
What a regulator will ask you if it goes wrong
Here is the part most AI translation buyers have not priced in: U.S. regulators have already written down what they expect, and while good quality is assumed everywhere in their rules, what the rules actually demand is proof. Federal health regulation names machine translation explicitly. Under Section 1557 of the ACA, 45 CFR 92.201(c)(3), if a covered entity uses machine translation when the text is 'critical to the rights, benefits, or meaningful access' of a limited-English-proficiency individual, when 'accuracy is essential', or when the source material is complex or technical, 'the translation must be reviewed by a qualified human translator'. That provision is in force today. It survived the 2026 court vacatur that struck other parts of the rule, and it survived the 2025 executive order designating English as the official language, which explicitly did not require agencies to stop providing multilingual services and left Title VI obligations intact.
Finance reads the same way. FINRA's 2026 Annual Regulatory Oversight Report states that its rules 'continue to apply when firms use GenAI or similar technologies in the course of their businesses, just as they apply when firms use any other technology or tool', and then describes what sound practice looks like: 'storing prompt and output logs for accountability and troubleshooting; tracking which model version was used and when'. Read that list again, because prompt logs, output logs, model versions and timestamps amount to a regulator publishing the specification of an audit trail in the same year the market is proposing to remove the human from the loop. In pharma, the standard has existed for decades: 21 CFR 11.10(e) requires 'secure, computer-generated, time-stamped audit trails' recording who created, modified or deleted every electronic record, and the FDA's January 2025 draft guidance on AI extends the same logic through a risk-based credibility assessment for a defined context of use, documented and reviewable.
Now put the vendor contract next to those requirements. CSA Research, the industry's own analyst firm, states it plainly: the small print for all LLMs 'carries disclaimers about accuracy and shifts responsibility for any errors to the user'. A high score makes errors rare, and a defensible record makes them survivable; you need both, but only one of the two appears in your contract. The gap between those documents is yours to carry, because none of the standards the industry certifies against will close it. ISO 5060, the 2024 standard for evaluating translation output, scores error types and explicitly excludes quality assurance processes from its scope. ISO 17100, the flagship translation services standard, dates from 2015, excludes machine translation post-editing entirely, and has been flagged for revision since 2025. MQM, the framework behind most quality dashboards, classifies errors and says nothing about provenance. Every translation standard measures the output, every regulated-records rule traces the process, and nobody's certificate covers the distance between the two.
The FDA has required time-stamped, attributable audit trails on electronic records since 21 CFR Part 11, and its 2025 draft guidance on AI asks sponsors to establish the credibility of any AI output for a defined context of use, with documentation to match. A translated consent form or labeling change lives inside that regime whether or not your language vendor has heard of it.
FINRA treats AI-generated communications exactly like any other communication: fair, balanced, not misleading, supervised and retained. Its 2026 oversight report goes as far as describing prompt and output logs and model-version tracking as sound practice. If your multilingual investor material cannot produce that trail, the gap is yours, not your vendor's.
Section 1557 names machine translation directly: when the text touches rights, benefits or meaningful access, or when accuracy is essential, a qualified human translator must review it. The provision survived the 2026 litigation that struck other parts of the rule, and the English-official-language executive order left these obligations standing.
Stanford Medicine researchers argued in a February 2026 paper that the right metric for AI translation in healthcare is 'patient comprehension and safety, not merely linguistic accuracy'. Substitute your own industry: the right metric is the outcome you are accountable for, not the fluency of the sentence.
The decision rule, and what to ask your vendor
So, can you trust AI translation with regulated content? Yes, if you stop treating your content as one bucket. CSA Research's Alison Toon puts the rule in one sentence: 'You can't just generalize and put all of your content through the same process, assuming the same level of accuracy, and therefore risk, for every type of content and every language pair.' Route by consequence, not by word count, and in practice the line is easier to draw than it sounds. Support macros, interface strings, product listings and internal documentation belong in the fast lane, where you take the discount with a clear conscience. Package inserts and labeling, informed consent forms, prospectuses and fund documents, benefit letters and claim decisions belong in the accountable lane, where a qualified human reviews the content and the review leaves a trace: reviewer name, timestamp, the approved source the translation is grounded in, the terminology version applied.
That last clause is the one to negotiate, because a human in the loop who leaves no record is a hope, not a control. Section 1557 does not ask whether a person was somewhere in your workflow; it asks whether this translation was reviewed by a qualified human translator, and the only way to answer is to be able to show it.
This is, in full transparency, the problem OTTO was built around. Rather than adding a score on top of a generic pipeline, OTTO grounds every translated segment in your approved source content, applies your terminology as a versioned asset rather than a suggestion, routes high-stakes content to qualified human review, and captures the whole chain, reviewer, timestamp, model and provenance, as a native part of the deliverable. Not because quality scores do not matter, they are measured there too, but because when the question comes, and in regulated industries it eventually comes, the answer has to be a record.
Do
- Tier your content by consequence before you tier it by volume. The routing decision is the compliance decision.
- Demand segment-level traceability: source sentence, approved reference, terminology version, reviewer identity, timestamp.
- Keep the prompt and output logs FINRA describes. If your vendor cannot export them, that is your answer.
- Take the AI savings on low-stakes content. The discount is real and refusing it helps nobody.
Don't
- Do not accept a quality score as a compliance artifact. A percentage is a claim about a population, not proof about a sentence.
- Do not assume certifications cover you. ISO 17100 excludes MT post-editing; ISO 5060 excludes process; MQM ignores provenance.
- Do not let review happen off the record. An undocumented human in the loop dissolves the moment a regulator asks.
- Do not sign indemnification clauses unread. The LLM small print shifts error liability to you.
Six questions for your next RFP
- The sentence test For any published sentence in any language, can you show me the approved source it derives from?
- The reviewer test Can you name the qualified human who reviewed this segment, with a timestamp, as 45 CFR 92.201 requires?
- The log test Can you export prompt logs, output logs and model versions, per FINRA's 2026 description of sound practice?
- The reproducibility test Same input, same configuration: can you reproduce the output, or explain why not?
- The audit-log test Do your audit logs cover translation events, or only logins and access? These are not the same product.
- The liability test Show me the indemnification clause. Who owns the error when one severe mistake in a hundred ships?
Our vendor guarantees 98+ MQM. Is that not enough?
It is a real achievement and worth demanding, so keep it in the contract. Just do not ask it to do a job it was never designed for: MQM classifies errors, it does not record who reviewed what against which source, and a regulator investigating one sentence gets no answer from a population average.
We already have a human in the loop.
In the loop where, and where is it written down? Under time pressure, fluent AI output defeats expert readers; the attorneys in Lacey v. State Farm were humans in the loop. A review that leaves no record cannot be distinguished, later, from no review at all.
Our regulator has never asked us for any of this. Why prepare now?
Because the first request arrives after the incident, not before it. Section 1557's human-review requirement is already in force, FINRA's 2026 report already describes prompt logs and model-version tracking as sound practice, and an audit trail is the one control you cannot build retroactively.
Is this not just the language industry defending its margins?
The pressure is real, and so is the discount. But the requirements quoted here come from HHS, FINRA and the FDA, not from translators. The vendors who should worry are the ones selling scores to buyers who need records.
- Nimdzi, The 2026 Nimdzi 100, May 7, 2026 (USD 72.6B market, 75% cost-cut demands, 'deceptive fluency'). nimdzi.com/nimdzi-100-2026
- Vendor platform announcements, May and June 2026 (218% usage growth; automated review figures of 90%, 99% and 85% cited in the chart). smartling.com/company-news ; smartcat.com
- Damien Charlotin, AI Hallucination Cases database, consulted August 19, 2026 (1,933 decisions). damiencharlotin.com/hallucinations
- Lacey v. State Farm General Insurance Co., U.S. District Court, C.D. Cal., sanctions order, May 6, 2025 (USD 31,100).
- Associated Press, October 2024, Whisper in medical transcription; Koenecke et al., 'Careless Whisper', FAccT 2024 (38% of hallucinations potentially harmful). apnews.com
- American Translators Association, Statement on Artificial Intelligence, May 20, 2025. atanet.org/advocacy-outreach/ata-statement-on-artificial-intelligence
- 45 CFR 92.201(c)(3), Section 1557 final rule, current text. ecfr.gov/current/title-45/section-92.201
- FINRA, 2026 Annual Regulatory Oversight Report, GenAI section. finra.org
- 21 CFR 11.10(e), current text. ecfr.gov/current/title-21/chapter-I/subchapter-A/part-11
- FDA, draft guidance, Considerations for the Use of AI to Support Regulatory Decision-Making for Drug and Biological Products, January 2025. fda.gov
- WMT25, Findings of the Shared Task on Automated Translation Evaluation Systems, 2025. aclanthology.org/2025.wmt-1.24
- CSA Research, Alison Toon, 'Reliable Automation Needs Measured Language Risk', May 8, 2026. csa-research.com/l/blog/article/reliable-automation-needs-measured-language-risk
- ISO 5060:2024; ISO 17100:2015 (flagged for revision, 2025); ISO 18587:2017; MQM framework (themqm.org).
- Slator, 'Research: Why Medical AI Translation Validation Should be Different', March 10, 2026 (Stanford Medicine paper). slator.com/medical-ai-translation-validation
- Executive Order 14224, March 1, 2025, Federal Register. federalregister.gov/documents/2025/03/06/2025-03694