The rise of AI-generated content is an intriguing phenomenon that has many writers and linguists spooked. It's like a ghost writer has invaded our linguistic realm, capable of crafting everything from poetry to corporate jargon with ease. The versatility and speed of these AI models are astonishing, and they are leaving their mark on various platforms, from inboxes to LinkedIn feeds, and even in academic papers.
One of the challenges we face is identifying AI-written content. It's not as simple as spotting a few peculiar words or dashes; it requires a deeper analysis of writing style and language use. Linguist Karolina Rudnicka highlights that, just as there is no single style of human writing, there is no uniform style of AI writing either.
To tackle this, detection algorithms have been developed, claiming near-perfect accuracy. However, these algorithms are like black boxes, providing no insight into their decision-making process. Researchers have also attempted to identify AI-generated texts by searching for suspicious words or comparing papers written before and after the introduction of LLMs. Yet, these methods have limitations, especially when it comes to distinguishing AI quirks from natural language trends.
Unveiling the AI Writing Style
The Economist conducted a study to uncover the unique hallmarks of AI writing. They compared AI-generated articles with human-written pieces from their own publication, as well as from other reputable sources like CNN, the New York Times, and the Washington Post. Additionally, excerpts from popular novels were included in the analysis.
The findings were eye-opening. While AI prose is distinguishable by word choice, punctuation, and sentence structure, the quirks are not as expected. This is because AI writing styles have evolved with software updates. It's a dynamic and ever-changing landscape.
One of the key differences lies in vocabulary. AI models, particularly Gemini and Claude, overuse polysyllabic words like "significant" and "consequences", and they have a penchant for rare words and scientific terminology. They also love nominalizations, turning verbs into nouns. This style of writing, as George Orwell would describe it, is "pretentious diction", a way for writers to sound clever by using complicated words and jargon.
Punctuation is another giveaway. Contrary to popular belief, the latest AI models do not overuse em-dashes. In fact, only Claude uses more em-dashes than human writers. Instead, AI writing is characterized by a lack of punctuation, with fewer commas and semicolons, and almost no parentheses. This is partly due to their preference for longer sentences and their tendency to avoid quoting experts.
When it comes to sentence structure, AI writing tends to be long and uninterrupted. They often employ rhetorical devices like "not X but Y" and "not only but also" to add variety. ChatGPT and Claude, in particular, overuse these constructions.
The Future of AI Writing Detection
The study by The Economist shows that AI writing is becoming increasingly similar to human prose with each software update. While detection algorithms like Pangram may have been successful in the past, they might struggle to keep up with these rapid advancements. As Tommie Juzek from Florida State University notes, AI models learn from human feedback and adapt their writing to what is considered impressive by humans.
The ability of AI to learn and adapt quickly is remarkable. Take ChatGPT, for example. It has drastically reduced its use of em-dashes, demonstrating a rapid shift in writing style.
In conclusion, the world of AI-generated content is an ever-evolving landscape. As AI writing becomes more sophisticated, the challenge of distinguishing it from human writing will only intensify. It's a fascinating development that raises questions about the future of writing and the role of AI in our linguistic landscape.