Patent Translations Inc. · Notes
Identifying and Avoiding AI Translation Risks

Law firms are switching from conventional machine translation to AI translation for first-pass understanding, and the data clearly support this move. In Japanese-to-English patent translation, frontier large language models produce one or two minor errors per sentence and one or two major errors per paragraph, similar to first drafts produced by well-trained human translators. This is a major advancement over MT output, in which we would typically see major errors in essentially every sentence. In even better news, LLMs sometimes allow attorneys who do not speak Japanese to dig deeper into ambiguities in important sentences. For example, when they spot something that looks like it might constitute anticipatory disclosure, they can drill down on that particular passage and check for other potential translations of the same phrase, such as by asking for the broadest and narrowest reasonable readings. They can even do the same thing with an examiner's machine translation, or a translation produced by the other side in litigation, probing it for possible grounds for impeachment.
The Japanese-to-English pair has always been difficult for MT. There are differences, not only in the structures of the languages, but also in the content of the communication itself. For example, Japanese does not ordinarily use the plural or singular and does not have definite and indefinite articles. From an English language point of view, this information is simply missing. Because old machine translation worked sentence-by-sentence and could not consider the context of the entire publication, it simply defaulted to whichever was statistically more common. LLMs can consider context and so make better (or at least better informed) guesses.
This improvement can be a double-edged sword, as it can be more difficult to tell when an LLM is making the wrong guess. Conventional MT tended to fail loudly by producing the occasional nonsensical sentence and outputting untranslated words in the original language, making it obvious that human attention was necessary. LLMs are models that guess what a human would most likely say next, on the basis of everything that has been said so far. This propensity to output exactly what you would expect masks failures and hallucinations, so that the overall sentence sounds reasonable even if the translation is wrong.
Context awareness, for example, is an upgrade, but it carries its own risk. AI tends to give considerable weight to whatever is in its "context window," including the output that it has generated so far. That is to say, it treats its own writing as a source of truth simply because it is there. As an example of how this can create risk, consider the Japanese word "軸," which can be translated as "shaft," "axle," or "axis." If the word appears six times in your document and the LLM has translated it as "shaft" in the first five instances, it is very likely to translate it as "shaft" in the sixth, even if "axis" is actually the correct choice in that sentence. It may actually go so far as to anchor its translation of that sentence around the word "shaft," and remold the rest of the sentence to fit. What is more, the LLM is likely to defend its choice when queried simply because the previous context predisposes it to do so (autoregressive bias). Because confirmation bias also makes human readers more likely to accept what we have seen in the text so far, this sort of failure is rarely noticed.
AI translation has also inherited some of the problems of earlier neural network machine translation, such as Google Translate, which burst onto the scene ten years ago. At the time, I wrote about the specific risks to patent attorneys posed by assumption, smoothing, replacement, omission, false inclusion, and scrambling. LLMs keep all of these and add drift, inconsistency, false consistency, gullibility, and tokenization failure.
Today, the conventional response to unreliable machine output carries a new risk of its own. When a first-pass computer translation makes it clear that a foreign reference is going to be important, attorneys often turn to translation agencies, where the document is either translated afresh by hand or the machine output is edited by a human translator. Although the first approach leaves productivity gains (cost reductions) on the table, the second approach, which has become standard, can involve significant risk. Agencies typically charge much less for the service known as "postediting" because the heavy lifting of looking up words in dictionaries and typing things out has already been done. To justify the lower rates, the human in the loop must work much faster than they would on a manual translation, and often for a lower hourly wage. This situation encourages the translator to trust the machine output and make changes only where there are obvious errors or the translation reads badly. Not only do genuine errors slip by, but perfunctory editing of the English target text without sufficient analysis of the Japanese source can amplify the problem of smoothing, making it even more difficult for monolingual attorneys to recognize errors in the finished product.
If it sounds as if LLMs are uniquely problematic, it is important to remember that human translators actually fail in similar ways. Omission, replacement, false inclusion, false consistency, misunderstanding, and even autoregressive bias caused problems long before computers existed. This is because, at the level of information theory, translation is inherently "lossy." The encoding and decoding involved when a person reads a sentence written in one language and then writes it in another language is subject to mathematically inherent loss, known as equivocation, as well as unavoidable noise and distortion. In professional human translation, these problems are identified and corrected through understanding the technology described, training, supervision, rule sets, and multiple-linguist review.
Now, for the first time, attorneys can use AI to bring some of these same techniques to bear on machine output. For example, you can ask another AI model, or another instance of the same model with a different prompt, to review an LLM translation. Keep in mind that, if you are not fluent in the source text language, this technique will only be useful for surfacing errors and will not let you arrive at a translation that will stand up to scrutiny by an expert translator. It can tell you that something might be wrong, but it cannot tell you that something is definitely right. We have tested these techniques and integrated some of them into our own workflow for expert-certified translations produced by a human team. The next posts in this series will look at specific techniques and results.
One thing that we learned early is that it is best to avoid extended conversations about a translation with the same AI. In a long enough exchange, you can get an AI to translate almost any way that you want, and to confidently give hallucinated reasons as to why the translation is correct. These creative endeavors will not withstand scrutiny by a human expert.
The propensity toward agreeableness is perhaps the greatest weakness of LLMs when compared to non-interactive MT. An examiner will sometimes accept a machine translation of a foreign reference, particularly if it is on a public site such as the EPO Espacenet service. But they generally know that an LLM translation provided by an attorney or an applicant will depend entirely on undiscoverable prompts, parameter settings, and context. Unverifiable and unduplicatable, LLM translation is neither a matter of record nor an opinion. In fact, sending text to a datacenter for translation is not even an option for some documents.
None of this reduces the real advantages of a tool that can provide instantaneous insights into what Japanese documents disclose, as long as we remember that it is a tool and not an authority. The question is not whether to use AI translation but when to rely on the output and when to call in a human eye. In the next installments in this series, we look at specific techniques for catching the failures described above, and dig into results from testing current frontier models against a translation that was ruled on by the Southern District of New York.