Who Protects a Book’s Voice? Literary Translation and the Limits of AI
- Esra OBUT
- 5 days ago
- 11 min read

Monika Kim’s novel The Eyes Are the Best Part opens with the sentence, “Umma tells me that the eyes are the best part.” When translating the novel into Turkish, I rendered this as, “Umma, bana gözlerinin en güzel yeri olduğunu söyler.” At first glance, translating “Umma” as annem, or “my mother,” might have seemed like the most natural choice. The Korean word means mother and its Turkish equivalent was clear. Yet on the same page, only a few sentences apart, the same woman was referred to by a different term. Watching her mother clean a fish, the narrator said, “The flesh is still steaming hot, but my mother doesn’t seem to feel anything at all.” Immediately afterward, the text returned to “Umma.”
The parts of a translation that require the most thought are not always the words without dictionary equivalents. Sometimes the equivalent is perfectly clear. The real difficulty begins when we notice why an author uses one word for a person in one place and a different word elsewhere.
Why were two different terms used for the same woman on the same page? For me, that was the real translation question. “Umma” was not simply the Korean equivalent of “mother.” It carried the language spoken inside Ji-won’s home, her family’s immigrant experience, her cultural identity and the intimacy of her relationship with her mother. “My mother,” by contrast, appeared when the narrator was observing her mother, describing her behavior and looking at her from a certain distance. I therefore established a principle on the very first page: when the author wrote “Umma,” I would preserve it; when she wrote “my mother,” I would use annem. Collapsing both into a single Turkish word would not have altered the basic meaning of the sentences. The novel would still have been comprehensible. Yet the small but important movement that allowed Ji-won to approach her mother from within the language of the family at times and to look at her from the outside at others would have disappeared.
The decision I made on the first page returned some 270 pages later in the novel’s final judgment. Ji-won first reflected on the men in her mother’s life: “Umma allowed the men in her life to control her…” I kept “Umma” here. The next sentence read, “My mother may be too weak to protect herself, and my sister too young. But I’m neither of those things.” Here, I used annem. The choice of word was especially meaningful in this final instance. Ji-won was no longer simply a daughter addressing her mother. She was assessing her, naming her weakness and distinguishing herself from her. The psychological distance between them was embedded in the word itself. A decision I had made on the novel’s first page found its meaning on the last.
Translating a book does not mean translating hundreds of independent sentences one after another. A small decision made on the first page can become a promise that must be kept until the last. A translator holds in mind not only the sentence in front of her, but also what the book has already said and what it has yet to say.
I share this example to show why the “correct equivalent” is not always enough in translation. “Umma” and “my mother” both refer to the same person. If a translation system renders both as annem, it has not necessarily made a dictionary-level error. It has nevertheless missed what the text is doing. In a literary work, words tell us not only whom or what they refer to, but also where the character is looking from.
This is where AI’s most surprising limitation in translation becomes visible. A system does not have to choose the wrong word. It can translate every word correctly and still miss the essential movement within the sentence. The result may be grammatically flawless, easy to read and apparently free of problems. What has been lost is simply harder to point to than an incorrect word.
For a long time, the debate around machine translation centered on the claim that “machines translate incorrectly.” In his well-known 2018 essay for The Atlantic, Douglas Hofstadter argued that beneath Google Translate’s impressive fluency there was no understanding, and that the system did not even know that words stood for things. Systems have advanced so dramatically since then that this critique no longer applies in its original form. Today’s large language models recognize idioms, follow context and imitate style. A stylometric study by researchers at Hong Kong and Swinburne universities even shows that, at the level of word choice and syntax, GPT-4 translations can statistically imitate human translations quite closely. It is becoming increasingly difficult to distinguish between the two by looking only at the surface.
What does a machine miss when it translates correctly?
What makes this question important is not that AI produces bad texts. A bad translation is easy to recognize: an incorrect word, a broken sentence or a meaningless phrase immediately stands out. The real difficulty lies in texts that do not reveal their problems. A fluent sentence appears correct. A natural expression seems like an appropriate choice. Yet fluency tells us only that a text is easy to read. It does not prove that the voice, cultural context or web of meaning constructed by the author has been preserved.
The most striking answer to this question so far came from a reader study published in June 2026. Researchers from Simon Fraser University, UQAM and Microsoft asked readers to compare human translations of 15 contemporary novels translated from French, Polish and Japanese into English with translations produced by an advanced AI pipeline. At first glance, the results appeared favorable to the machine. Readers found the machine translations “fine” and could not reliably distinguish human work from machine output: only 17 of 30 guesses were correct, barely above chance. When the same readers compared the two translations closely, however, they clearly preferred the human version in 522 of 772 comparisons. Perhaps the study’s most important finding was that, within a single book, the quality of machine translation fluctuated far more than the quality of human translation.
Readers may not always be able to name the problem, but they can sense that something in the text has changed. A character’s voice may feel strong on some pages and fade on others. A word translated in a particular way earlier may receive a different equivalent later. Sentences that seem acceptable in isolation may sound, when placed side by side, as if they do not belong to the same person, narrator or book.
This finding deserves attention. The reader may not detect the difference at sentence level, but senses it in the whole. The loss, then, does not reside in individual sentences. It accumulates in the consistency, or inconsistency, of the decisions. My choice regarding “Umma” was not the decision of a single sentence. It was a decision made again across 270 pages, one that found its meaning at the end of the novel. The equivalent I selected when a word first appeared also determined the meaning it would carry hundreds of pages later. A machine, by contrast, solves each sentence anew and on its own. Every solution may be reasonable, but a series of reasonable solutions does not necessarily create a coherent voice.
For this reason, loss in literary translation is not always located within a single sentence. Sometimes it accumulates as a voice gradually changes, a repetition disappears or a meaning introduced on the first page fails to reach the last. These small shifts, invisible until the translation is read as a whole, either create or weaken the book’s memory.
Research confirms this picture from other angles. A comprehensive study published at ACL 2026 tested 23 large language models on their ability to understand source texts and make creative translation decisions. The models performed well in understanding the text but struggled to turn that understanding into creative choices such as preserving a metaphor, recreating wordplay or sustaining the narrator’s voice. Human translators received a creativity score of 0.246, while the great majority of the models scored below 0.1. This suggests that the central problem is not simply whether a machine understands a text. It is whether it can decide what within that text is worth preserving. The machine’s first translation can also influence human choices. According to research conducted at Ghent University, translators who edit machine translations produce more literal and less creative solutions. In other words, AI does more than offer a suggestion. It can quietly narrow the range of possibilities the translator is able to see.
We can think about this more simply. The same English sentence can be rendered correctly in Turkish in several different ways. Every option may communicate the basic meaning. Only some of them can carry the character’s voice, the emotion of the scene and the overall atmosphere of the book at the same time. A translator’s real work begins where dictionary meaning ends and a decision must be made among these possibilities.
The machine’s first suggestion is not as neutral as we may think. When a polished, ready-made sentence appears before us, we may become less inclined to look for other possibilities. AI therefore does not merely offer its own choice. It can also narrow the field of options available to the human. Using AI in translation does not reduce the translator’s responsibility to evaluate. On the contrary, it creates a need for a more attentive reader who can think again without being carried along by the first suggestion.
This brings me to what I find most thought-provoking: we cannot measure this loss. According to a study published at NAACL 2025, the automated metrics used to assess translation quality, as well as the increasingly common practice of asking AI to evaluate translations, can score the work of experienced literary translators lower than machine translations. The same pattern appeared in the reader study: automated metrics failed to capture reader preferences and favored machine translations. This is not a flaw so much as a design choice. Metrics measure the sentence, and at sentence level machine translation really is smooth. Yet if meaning lives not in the sentence but in the whole, no tool that measures sentences can see what has been lost. In my work evaluating AI systems, I encounter this tension every day: what we can measure and what matters are not the same thing. When measurement belongs to the part and meaning belongs to the whole, only human judgment can bridge the gap.
This is not a problem unique to translation. With all AI outputs, we tend to place too much importance on the qualities we can measure easily. Is the grammar correct? Does the text contain the necessary information? Does it follow the instruction? These are important questions. Yet it is not as easy to measure what the text is actually saying, whether it suits its context, whom it leaves out or where it has made a premature assumption. The distance between an output that looks good and one that truly is good opens up precisely here.
The value of expertise becomes more visible in the age of AI for the same reason. An expert is not only someone who knows the correct answer, but someone who can detect what is missing from an answer that appears correct. A translator’s knowledge often lies not in the final sentence, but in the alternatives rejected on the way to it. The reader sees only the selected equivalent. The translator also sees the other words that could have been used, the associations that would disappear if the choice changed and the decision that might create a problem several chapters later.
The economic consequences of this gap are now becoming visible. Last January, Harlequin France ended its translators’ contracts and moved to a model based on machine translation. A similar controversy arose in Türkiye when it was revealed that a publisher had translated several books with DeepL and released them under translators’ pseudonyms, immediately turning the issue into a debate about transparency. These arguments generally produce two sides: those who view the automation of translation as an inevitable matter of efficiency and those who oppose it on principle. In light of the reader study, however, I believe the real issue lies elsewhere. A reader does not audit a book sentence by sentence; a reader trusts the book. The difference between a translation that is merely “fine” and one that carries a voice across 300 pages may not be captured in any single sentence, yet it permeates the entire reading experience. That difference is difficult to price and its absence nearly impossible to prove.
It is therefore not enough to reduce the debate to the question, “Can AI replace the human translator?” The more important question is what kind of text we are willing to accept as readers. Will we consider a translation sufficient if it communicates the basic meaning and reads smoothly? Or will we continue to regard an author’s voice, the distance between characters and the memory a book accumulates across its pages as part of translation itself?
I faced a similar question of coherence while working on the novel’s title. In English, “the best part” can mean both the finest part and, in the context of food, the tastiest piece. Introduced in the opening scene as the mother eats a fish’s eye, the phrase goes on to connect images of beauty, desire, the body and consumption. The title does not refer only to a part that can be eaten. It also carries the disturbing boundary between finding eyes beautiful and wanting to possess and consume them. The Turkish title, Gözleri En Güzel Yeri, sought to preserve both meanings as far as possible. I was not translating a single sentence. I was tying the first knot in a web of imagery the book would construct within itself.
Making a sentence more fluent does not always make it better. Sometimes the harshness of the source text must be preserved. Removing repetition does not always strengthen a text; repetition may be the rhythm of a thought growing in a character’s mind. Explaining an ambiguous expression can make the reader’s task easier. If the author intends to leave the reader inside that ambiguity, clarity takes something away from the text. At this point, the translator is not simply someone working between two languages. The translator is the person who reads what the text is trying to protect.
Thinking about translation is therefore one of the most concrete ways of thinking about AI. The essential question in translation is not only, “Is this sentence correct?” It is also, “Is this sentence faithful to the memory of the book it belongs to?” That question extends far beyond translation. Nearly all the tools we use to evaluate AI systems today ask the first question: Is this output accurate, fluent and acceptable? We still do not know how to ask, much less measure, the second: Is this output faithful to the memory of the whole to which it belongs? A decision made when a word first appears may find its meaning 270 pages later. What machines still cannot do, and what we still do not know how to measure, lies not in any individual sentence but exactly there.
This does not mean that AI has no place in translation. It can quickly suggest different equivalents, recast a sentence in several tones and reveal a possibility the translator may have overlooked. Its value lies not in taking over the entire decision, but in its ability to expand the space for thought. The human still decides which equivalent belongs to the book, what must be preserved and when the sentence is truly complete.
Machines are writing better and better sentences. This development does not remove human responsibility toward language. It moves that responsibility somewhere else. It is no longer enough merely to produce the sentence. We must look at it carefully, notice invisible losses and decide what is good enough. In some sentences, the essential meaning lives not in the words themselves, but in the memory those words carry throughout the book.
References
Hofstadter, D. (2018). “The Shallowness of Google Translate”. The Atlantic.
Yao, X., Kang, Y.-B. & McCosker, A. (2025). “Missing the human touch? A computational stylometry analysis of GPT-4 translations of online Chinese literature”. arXiv.
“AI translation of literary texts is ‘fine’, but readers still prefer human translations” (2026). Simon Fraser University, UQAM & Microsoft. arXiv.
Zhang, R., Eger, S., Tezcan, A., Zhao, W., Ponzetto, S. P. & Macken, L. (2026). “Beyond Reproduction: A Paired-Task Framework for Assessing LLM Comprehension and Creativity in Literary Translation”. Findings of ACL 2026.
Macken, L., Ruffo, P. & Daems, J. (2025). “The Role of Translation Workflows in Overcoming Translation Difficulties”. Proceedings of the Workshop on Creative-text Translation and Technology.
Budimir, B. (2025). “The Challenge of Translating Culture-Specific Items: Evaluating MT and LLMs Compared to Human Translators”. MT Summit 2025.
Zhang, R., Zhao, W. & Eger, S. (2025). “How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs”. NAACL 2025.
Wang, Y. (2023). “The Role of Emotion in the Translation Process from the Perspective of Embodied Cognition”. Psychology.
“Embodiment in Translation Studies: Different Perspectives”. inTRAlinea, “Embodied Translating” special issue.
“Harlequin is firing its human translators…” (2026). Literary Hub.
“‘Yapay zekâyla çeviri’ tartışması”. Milliyet.



Comments