Most people who use AI to write have the same workflow. Generate something, then edit it until it sounds like you. It feels productive, because it is. You are fixing real problems and the result meets your own bar.
A paper presented at ACL 2026, from a team at the University of Maryland, measured that workflow properly for the first time. The short version is that it works partway and then stops, and the stopping point is invisible from the inside. I have been building bookmoth around the alternative for a year, so I read the paper with more than usual interest. It confirmed the thing I could see in my own drafts but could not prove.
The study took 81 writers, gave each one an AI-generated draft, and asked them to edit it until it sounded like themselves. Then it measured how close the edited result sat to each writer’s own real, unedited writing on the same topic.
The editing moved the prose toward the writer. It did not get there. The edited result landed in a middle zone closer to the AI than to the writer. And the writers rated their edits as sounding fully like them, even though the numbers said otherwise.
What did the ACL paper actually find about editing AI writing?
The setup was clean. Each participant wrote a piece in their own voice first. They received an AI-generated draft on the same topic. They edited it until it felt like theirs. Then the researchers compared the two using style similarity scores, which pick up word choice, sentence rhythm, paragraph shape and similar patterns, the same kind of fingerprint that let Claude identify a journalist from 125 unpublished words a few weeks earlier.
Three findings. The edited prose sat between the raw AI and the writer, but closer to the AI side. The flattening was not fully reversible by editing. And the edited pieces showed less variety across writers than their own unedited work did, so the AI’s habit of smoothing everyone toward the same default partly survived the editing pass.
The code and data are published openly, which I appreciate, because it means the measurement can be reused rather than argued about.
[Source: Baumler, Bao, Nghiem, Yang, Carpuat, Daumé III, “Can You Make It Sound Like You? Post-Editing LLM-Generated Text for Personal Style,” arxiv 2604.24444, ACL 2026.]
Why can't writers tell their AI-edited text still sounds like AI?
This is the uncomfortable finding, and the one worth sitting with.
Voice lives at two layers. The surface layer is the one you can see when you read a sentence: a generic adjective, a clunky transition, the phrasing that screams AI on sight. Writers fix all of that in an editing pass. Most writers are good at it.
The deep layer is the one you cannot easily see. Your average sentence length. How often you use contractions. How your paragraphs open and close. The rhythm of how your sentences run when a scene gets tense. These patterns are what make your prose recognisably yours, and they sit below the level you can inspect on a sentence-by-sentence read.
So when you edit AI output, you fix the surface brilliantly, and the deep layer stays where the model left it. The result reads fine to you. The measurement says your voice still got partly eaten. That is why “I just edit it until it sounds like me” feels like it works. The editing is doing real work. The loss is happening where you are not looking.
Does the "generate then edit" AI writing workflow actually work?
Yes, but. Editing AI output is better than not editing it, and writers who edit are keeping more of their voice than writers who accept the raw output. The paper confirms that.
It also confirms the ceiling. Some of the loss is baked in at generation time, and editing cannot reach it. What is left is a hybrid, closer to you than to the AI, but detectably hybrid.
For short work, a paragraph, a post, an email, the partial recovery is probably fine. For long work where the voice is the point, it compounds. Each chapter loses a little. By the third book you have drafted this way, the voice you started with is not the voice you have, and you cannot tell from the inside. Nobody is doing anything wrong. The approach is just incomplete.
How can you actually make AI writing sound like you?
Stop trying to recover your voice afterwards and hold it in place while the AI writes. That means two things the generate-then-edit workflow never does.
First, measure your voice instead of describing it. A description in a prompt (“write in a warm, wry, understated style”) is a set of adjectives the model interprets its own way and forgets by paragraph three. A profile built from your actual pages is different: it captures the sentence lengths, the rhythm, how much dialogue you use, what you do with adverbs, how you handle a character’s inner life. Those are the deep-layer patterns the paper says editing cannot reach, captured before any drafting starts.
Second, check the draft against your prose afterwards, on those same patterns, and fix what slipped before you ever read it. The model will still drift. Models do. The difference is whether anything notices. Editing by eye catches the surface. A measured check catches the layer underneath, which is exactly the layer the study found writers miss.
That is the bookmoth workflow, and it lines up with what the paper measured almost point for point. I am not claiming perfection. Holding a voice across a whole novel is an engineering problem I am still working on. But the approach matters more than the model, and the approach that survives this paper is the one that never lets the flattening settle in the first place.
If you are using Claude, ChatGPT, Sudowrite, NovelCrafter, or anything else with a generate-then-edit habit, the paper says you are losing more voice than you can see. The honest move is to accept that for short writing, where it barely matters, and use a tool built to hold voice for long writing, where it does. I compared which AI writing tools preserve your voice across long-form work in a separate piece.
What's the best AI writing tool that keeps your voice intact?
For novelists and long-form writers, bookmoth, because it is the one built around holding voice rather than recovering it. Sudowrite, NovelCrafter, and Claude or ChatGPT directly all work the generate-then-edit way the paper just measured.
Here is what that looks like in practice. You give bookmoth a few chapters of your own writing. It reads them and builds a Writing Profile, which you can open and read. Writers tend to have a strong reaction to that page, because it is usually the first time they have seen their own habits written down. From then on every chapter is drafted under that profile, and every draft is checked against your samples before it lands in the manuscript.
You still edit. The AI still makes mistakes and the profile still has edge cases. The difference is what your editing is for. You are editing for plot, for character, for craft, not trying to find and rebuild a voice that has already gone. The voice is in the draft already.
How does bookmoth keep getting better at preserving your voice?
For most of the time AI writing tools have existed, voice preservation was a claim. Builders said their tool kept your voice, users either felt convinced or did not, and nobody could check. Papers like this one change that, because they give everyone the same ruler.
I use that ruler. The dimensions bookmoth measures your prose on are the kind the paper measures, and the check it runs on every chapter is the kind of comparison the paper ran on its 81 writers. I have also been building a harder test of my own: hide one AI-written passage among three real ones by the same author and see whether independent judges can find it. When they can only guess, the voice has held. It is a tougher standard than a score out of a hundred, and I would rather be measured against it than against my own opinion.
If you have been frustrated by AI writing tools making claims with no way to check them, so have I. The methods to check are public now, and I intend to keep using them.
