Your AI isn't the risk. What's missing around your multilingual content is.
Ninety five percent of companies that piloted generative AI for their content saw no return on it, according to MIT NANDA's State of AI in Business 2025. Most people read that number and assume the AI itself wasn't good enough. It usually was. It could draft a product description or translate a paragraph into German just fine. What was missing wasn't a smarter AI. It was everything around it: deciding which content needed extra care, catching a wrong translation before it reached a customer, and proving, after the fact, that someone actually checked it.
That's easy to miss because the AI is the part everyone notices, the tool that writes or translates the text. What's missing is invisible, call it governance, and it's exactly what most AI rollouts skip. An AI with nothing built around it isn't a shortcut to multilingual content. It's usually how a company ends up publishing something nobody actually verified.
It comes down to five specific things an AI model, on its own, was never going to do.
An AI that only produces text is doing half a job. Something, or someone, still has to do the rest.
What a model handles
Writes, translates, and reformats your content quickly, at volume, in whatever language you need.
What still needs to happen without a system
Deciding which content needs a closer look, catching mistakes in translation, and keeping proof of what was checked.
Five things your AI can't do alone with your multilingual content
1. Decide what needs it before it runs.
Without a routing step, everything takes the same path through the same AI model. A low risk product blurb and a high risk regulatory filing get the exact same amount of scrutiny, which in practice means the filing gets far too little. Nothing upstream of the model was ever built to tell the two apart.
That's the first gap, and it makes the second one worse.
2. Ground itself in what the organization already knows.
A model that hasn't been pointed at your translation memory, your glossary, or your style guide isn't drawing on your knowledge, in any of the languages your multilingual content ships in. It's guessing at it, based on whatever it happened to see during training. The result usually reads fine, right up until someone who actually knows the terminology looks closely, and by then the same drift has probably crept into other markets nobody checked as carefully.
With the first two gaps open, everything now depends on a third one holding.
3. Catch its own mistakes.
It doesn't. One model, one pass, no second opinion. Whatever comes out becomes the final answer by default, not because anyone evaluated it. For low stakes content that's a minor risk. For regulated content, it's the gap between a mistake your team catches internally and one a regulator catches for you.
Even that might be survivable, if the model at least knew its limits. It doesn't do that either.
4. Recognize when it doesn't actually know something.
Models are built to produce a confident sounding answer even when the evidence underneath is thin. OpenAI has said as much about its own systems. Nothing in that design tells a model to pause and flag uncertainty instead of guessing, so left alone it simply answers, and the answer looks exactly as confident as a correct one would.
Which leaves one last problem, the one that shows up after everything else has already gone wrong.
5. Leave behind proof of what happened and why.
A model hands back text. It doesn't hand back a record of what it was allowed to use, what score its output got, or who reviewed it before it went out the door. When an auditor eventually asks how a piece of content came to exist, there's nothing to show them except the output itself and a guess at how it got there. That absence is exactly what governance is supposed to prevent.
Five gaps, and not one of them gets smaller with a better model.
Check what's true today, not what's planned. Then see exactly what covers the rest.
A model that can't say no doesn't fail loudly. It fails quietly, and the bill shows up later.
What actually catches what your AI misses.
None of the five gaps above get solved by using a different or newer AI tool. They get fixed by adding a governance system around the one you already use, one that works the same way no matter which language your content is in.
OTTO isn't positioned as a replacement for the model your team uses today. Keep the model, keep the vendor relationship. OTTO sits around both, as a model agnostic orchestration and governance layer for multilingual content, with full visibility into cost and latency, and the ability to swap the underlying model in under a second. Nothing already in place has to be torn out.
OTTO RCC (Risk, Compliance & Category)
Looks at each document before any model touches it, then routes it by risk, complexity, and regulatory sensitivity. This closes the first gap.
OTTO ALL (Augmented Linguistic Layer)
Pulls in your own translation memory, glossary, and style guide, document by document and language by language. This closes the second.
OTTO QEE (Quality Estimation & Evaluation)
Scores every output with a stated reason, routes anything weak or uncertain to a qualified reviewer, and logs the decision. This closes the third, fourth, and fifth gaps at once, and it's what turns the whole workflow into something you can actually govern.
Do
- Ask to see one real document move through routing, grounding, and scoring, start to finish.
- Ask for a logged, reasoned score on every output, not just a polished final draft.
- Test the system on the regulated content that actually carries risk, not on safe marketing copy.
Don't
- Treat "the model tested well" as proof the workflow around it is sound.
- Let the same model that wrote the answer be the only thing grading it.
- Assume the fix is a bigger model. It's usually a missing layer, not a missing parameter.
The one question worth asking before you trust any AI tool with your content.
Skip the accuracy slide in the next vendor call and ask something more specific instead: walk me through exactly what happens the first time this system is wrong. A team that's actually built for that moment points to a routing rule, a confidence score, a named reviewer, and a log entry. A team that hasn't will talk about the model instead, how big it is, how it was trained, how well it scored on a benchmark that has nothing to do with your multilingual content.
Won't a more advanced model eventually close this gap on its own?
Probably not. These are gaps in decision making and governance, not accuracy. A better model will guess better, but guessing better isn't the same as deciding what to check, grading itself, or keeping a record. That work happens outside the model no matter how good it gets.
If we already use retrieval, doesn't that solve the grounding problem?
It helps, but on its own it doesn't score its own confidence or know when to hand something to a person. A well grounded looking answer can still be wrong, and without a scoring step it can still ship as if it were right.
Doesn't adding a human reviewer just bring back the old bottleneck?
Only if every document goes to a person by default. The whole point of classification and scoring is making sure a reviewer's time goes to the handful of documents that genuinely need it, not to everything that comes through the door.
The point isn't that the model can't be trusted. It's that trust was never the model's job to begin with, and something still has to do it.
- MIT NANDA, State of AI in Business 2025. Enterprise GenAI pilot return statistic.