Have you ever eaten a dish where the flavor was all there but the texture was just… off? Maybe it’s perfectly seasoned, but the mouthfeel is gritty or somehow wrong. Texture matters for good food.
That’s often my feeling when reading completely AI-generated content, even if the underlying content is decent. I wrote a prior piece criticizing a trend where writers would cite ChatGPT or Claude to implicitly offset blame if their research was wrong. This isn’t that.
An experienced and prolific author recently described to me the process of composing words as similar to composing music. There’s a rhythm and flow to it. I agree. This “rhythm” is hard to quantitatively measure. But we tend to know it when we see it.
Humans are very good at spotting “unnatural” things, especially those that imitate humans. This is the entire origin of the uncanny valley, where hyperrealistic robots or animations mimicking humans feel off-putting to us, and we can intuitively catch the tiniest mismatches. In music, producers often need to microshift their timings to make it deliberately less perfectly on-beat—like there would be in human performance—to prevent it from feeling robotic.
What is it about AI writing exactly?
In line with societal angst about AI, especially in writing and publishing circles, people talk about how much they hate AI writing. Pangram’s rollout on Substack has created mini-witch hunts, and influencers like Hank Green who are accused of using AI face being “cancelled.”
None of this actually describes what is wrong with AI writing, though. Being AI is sufficient to hate it.
As longtime readers know, that’s not my stance. Although I didn’t use AI at all in my book (publisher requirement, plus it was too early in 2022 and 2023 to have AI be useful even as a research assistant), I’ve been transparent about AI use before it was cool and tried to see how far I can push it in my Substack.
The answer to its usefulness is quite mixed. It never gets anywhere close to passing the bar on its own and basically gets totally rewritten—enough so that even when I use AI, usually Claude, to do a first draft, it never comes up as AI writing on Pangram. It’s not bad as a carefully monitored research assistant (especially with a lot of adversarial reviews of its work!).
As a side note: I pretty much entirely stopped having AI try its hand at a first draft. I always ended up changing it so much that, for now, it is far easier just to do it myself.
Pangram, love it or hate it, is interesting in part because while we often know AI writing when we see it, it isn’t really that clear—quantitatively speaking—what makes writing AI.
What it is not
The typical hallmarks of AI, like the oft-complained-about em dash—which I will never give up—tend to be poor indicators. That should make sense. If something is so obvious, the next model version will inevitably remove it. For example, recent Claude models have made “load-bearing” a punchline given how much they seem to love it. That’s inevitably going to disappear, if not naturally, then definitely through post-training.
Additionally, you get a lot of variation just by using different models. They don’t all write the same. While they all do feel like AI to me, there’s no static characteristic that’s simple to measure. That is, of course, Pangram’s entire claim to fame. Whether or not you believe them, I do expect that they’ll need to continuously train their classifiers because models will keep changing their habits.
Music is a little ahead of us here, mostly because audio is easier to measure than prose (no offense to music, but notes don’t have explicit connotation and denotation when stuck together, even if it is a “universal” language in a way!).
Researchers tested a commercial AI music detector across 30,000 tracks. The detectors worked almost perfectly on the generators they were trained on. Then they tried the detectors on a generator they hadn’t seen before. The result: 3 of 50 AI tracks were caught. Oh, but it gets worse!
Funny enough, a bit of mangling also threw off the detectors. They simply lowered the sample rate on one of the generators’ tracks (Suno’s), which the detector had been trained on… and the detector caught none of the tracks. The detector was basing its verdict on encoding artifacts rather than anything musical.
Although I don’t know what’s under Pangram’s hood, I’d expect they need a pretty robust retraining loop just to keep up, given how much models change from release to release.

What it is—at least for now
Again, it’s really, really hard to precisely say what creates the AI writing uncanny valley. I expect that whatever metrics one comes up with will expire in a year or two—either simply by model progress or by deliberate action from the AI labs.
Nevertheless, we do have some ability to measure why AI writing seems to lack the unique rhythmic quality of human writing.
While I did say that normal “AI tells” are unreliable, certain models make themselves quite obvious. If I said “load-bearing” everywhere in this piece, even if I wrote it myself, a daily user of Claude would likely have PTSD flashbacks. As it stands, even with a fairly long style guide, I found that Claude consistently had stylistic tics that I would immediately rip out of its drafts.
Usually, the reason is this kind of short sentence—“mic drop” (as I call it)—is meant to emphasize the climax of an argument. Claude, in particular, loves to put mic drops everywhere. And, of course, if everything is emphasized, nothing is emphasized.
This has pretty obvious parallels in both storytelling and music. You need to have moments where you pull back, create tension, and vary things up. Constantly hitting the crescendo climax non-stop just means that you’re sonic noise and an ear-sore.
Nevertheless, a lot of us experience this particular tic through Claude mostly because Claude is what we use. It isn’t exclusive to it, though—models do vary here. Every model has its own habits.
But broadly speaking, all of them tend to have similar “rhythmic” issues stemming from the fact that they are token generators. For the same reason there’s context rot—where LLMs get dumber the longer you chat with them—LLMs inevitably have limitations in being able to perfectly track long writing in the way a human would.
We’ve seen some studies that basically show this. LLMs tend to converge on similar arguments. After all, their training set is similar because they all just ate the internet. Additionally, their sentences, as one would expect, are templates which also all resemble one another more than human sentences do. LLMs, as it says in their very name, are language models and you… well, model… sentences in language.
Again, this is a “point in time” metric. I have zero doubt that if this ends up being the main hallmark of AI writing, the labs will find a way to “jitter” sentences to “break” these measurements. It’s like benchmaxxing—once someone creates a benchmark, inevitably the labs will have releases that “game” those metrics.
As an aside, this is why certain models (say… like Gemini) seem to top benchmarks… but are terrible to use in actual practice.
Of course, breaking up this weird templating will likely help make the text feel more human, benchmaxxing or not. Going back to music, it’s possible to have musical notes hit exactly at the right timing. This is called quantization. Many music producers, even in electronic music, add in “swing” or simply hand-play certain parts to add the right “imperfection” and make their pieces feel more human. Similarly, less “quantized” AI text will probably read as more human. But, again, like music, there’s still a gap between just random jitter and true human performance.
It comes down to semantics. There’s a reason why certain rhythms are adopted for certain emphases and emotional impact. Just as true AI or algorithmically generated music often feels “weird” to listeners because the meaning or intent doesn’t seem to be there, I do expect that ineffable quality will be a gap for a long, long time in AI writing.
In fact, without more fundamental changes in how the models work, I don’t think that gap will be closed at all. While “embedding” (translating tokens into numerical vectors… which also creates numerical associations between concepts) kind of gets at semantics, it’s not really the same thing as composing an overall argument using language and the flow of a writing piece in mind.
Why Does it Matter?
To some degree, I’m in the camp that I don’t care if you use AI if I can’t notice it. I say “to some degree,” perhaps because, as a writer, I’m more sensitive to the “musical” rhythm of good writing.
If you use AI to do research, so long as you weren’t lazy and did the equivalent of “Let me Google That For You” and added something I can trust to my understanding, I’m happy enough. It doesn’t matter if you used an encyclopedia, Google, or ChatGPT. You put your stamp on it, which means you stake your reputation on it, which, in turn, allows me to trust what you put out. It literally is the original purpose and meaning behind “brands” in adding a mark of trustworthiness.
If you draft the piece with AI… but do such extensive edits—similar to my own process with Substack for a while—that it’s unrecognizable as AI, I also don’t really care. You basically inserted enough “human performance” that everything I care about, even “musically,” is present anyway.
Admittedly, given my experience, I’m not sure a first AI draft actually saves time with models as they are today… but that’s irrelevant to the end result. Whether you waste time doing it that way or not is beside the point.
No matter what, if I’m reading you as an author, I want it to be you as the author. You need to add what makes you you. Otherwise, I could just go to Claude or ChatGPT and just go prompt it myself.
Who is the Author, Exactly?
The concept of authorship is something I wrangled with in my book. My conclusion, between Death of an Author (a totally “ChatGPT written” book by @Stephen Marche), pop artists like Andy Warhol often using assistants to make the “actual art,” and the very art form of photography itself… authorship or creative ownership is more about the idea than the tool used to execute it—including AI.
Death of an Author is an example I especially love, since the author kept regenerating it until he was satisfied with what he got (and, besides that, wrote the plot himself). In an extreme case, if you “type” your book by hitting regenerate millions of times until it looks like what you would write—or at least what passes your quality bar—is that really that much different than just opening Microsoft Word and typing it? Wasteful/inefficient, sure, but it’s the same end result.
There are less extreme thought experiments, but I think most of them intuitively come out on the side of “it’s equivalent to typing” if you think hard enough about it. While there is a spectrum, if the human is the arbiter of quality and taste and lets it go out into the world having satisfied that bar… well, again, that’s why I follow specific authors anyway.
AI Watermarks
This particular point has some recent news salience thanks to Anthropic. Recently, in response to EU regulations (which, as is typical, created many more unintended bad consequences than things it solved…), Anthropic declared that they’d be using watermarks for AI writing going forward. The relevant AI Act provisions took effect on August 2nd, and per Anthropic’s own documentation, models launched on or after that date carry marks at launch. While they only needed to do it within the EU, they rolled it out worldwide.
Now, what’s bad about this? Well, two things.
One is that watermarks, in theory, will degrade the quality of writing. The way you can have “invisible” watermarks is by deliberately introducing certain non-randomness in word choice (well, token generation, but you get what I mean) that can be picked up. Google’s own research on the technique concedes it causes “some reduction to inter-response diversity.”
I just spent the entire piece talking about this kind of textual or rhythm quality causing AI generated writing to feel “off.” The proof will be in the pudding, but it’s hard to imagine this watermark not moving Claude in the wrong direction in “good writing.” But my reaction to that is somewhat of a shoulder shrug. I expect if it makes Claude’s writing feel really bad, Anthropic will change it. It’s not like I don’t already have to massively edit for word choice and rhythm anyway, so making it slightly worse is somewhat irrelevant to my use case.
The second is worse and has to do with “authorship.” Using their models, even if you’re doing a copywriting pass purely, you can have Claude “claim ownership” using watermarks. I think it’s silly in general, but certain other authors like Ben Thompson and John Gruber have been fairly pissed off by the entire thing.
After all, as per Anthropic itself:
Claude may not be the original author. People often use Claude to proofread, translate, summarize, or convert files. The output can carry a Claude mark even if the underlying ideas, text, or data originated from another source
In Dithering, Ben Thompson asserts that this is conveniently aligned with Anthropic’s view that Claude is basically its own entity and deserves credit for everything it outputs. After all, once the entire world is using Claude, all of it should partly belong to it—and Anthropic.
This is somewhat typical AI lab arrogance. It is certainly convenient for Anthropic to use the EU’s regulation to try to put its mark on everything.
In the Longer Term
Perhaps the watermark only matters insofar as the public still has extreme reactions to AI writing. After all, if every single word processor in the world used an “AI copyeditor” like Claude that put in small watermarks, its presence would be entirely irrelevant.
I do think we’ll eventually go toward a world in the future where AI use is as incidental as computer use. Basically, you’d find a new employee who comes up to you and says, “I don’t use AI,” just as weird as one today who comes up to you and says, “Well, I don’t really do ‘computers’” and hands you a stack of handwritten notes… which you’d likely then need to transcribe into a computer.
It’s part of the reason I’m not really that fussed about the watermark, even if I do think it’s misguided.
But putting aside practical considerations, there’s a deeper philosophical reason I end up where I am. And that is that AI doesn’t replace humans, at least not with the AI that we have today. It has different strengths and weaknesses… and especially for matters of taste, I still care about the human behind the artistic piece, whether it’s visual, audio, or written. That really shouldn’t be surprising. After all, there’s no necessarily “objective” reason why a human-performed music piece is better than a jittered or quantized piece.
If AI can do a lot of the practical stuff better and improve productivity in the world, I’m all for it. That’s been the way progress works, from shovels to steam engines to nuclear-powered ships. None of these obsoleted humans; they simply changed the nature of work.
If you think, from the creative side, that AI will simply wipe out humans… well, perhaps you have less faith in humanity than I do.
But even if you do have less faith, I think you should logically expect that humans—by matters of pure taste—will prefer things with a human touch… because we are human and AI is not.
Thanks for reading!
I hope you enjoyed this interview. If you’d like to learn more about AI’s past, present, and future in an easy-to-understand way, I’ve published a book titled What You Need to Know About AI.
You can order the book on Amazon, Barnes & Noble, Bookshop, or pick up a copy in-person at a local bookstore.











I am using Gemini in novel writing. I give it a short draft with all the ideas, and it spits it back with some more staging, repetition of motifs, and dialogue that sounds like my characters. I consider the two, together, as a first draft. When I rewrite, I think there's never a Gemini sentence that remains untouched. I have never been satisfied with a first draft of anything, so I have no expectations for this first draft. As an old professor said, "The best writing is re-writing." I asked Gemini to compare itself to Claude, and it says that Claude writes with more power and concision, but it says it gives me more to work with as a re-writer, and I agree. As it offered in conclusion, after pointing out that I came up with all the good ideas and all it did was string it out "We are a great team!" I don't expect it to give me finished work, but it does prompt me to add a lot of points that I want to add, in my own way. I don't mind people criticizing me for using the boost, but I do mine people who say I'm not a legitimate writer, and that the product is not mine.
Great post on *why* AI writing is subpar!
I’ve realised that when I read AI generated content out loud and then read human generated content, there’s a difference in cadence/flow and doesn’t seem to flow as well as it would if it was human written (assuming said person can write well).