Brief IA

ChatGPT Outperforms Human Authors, But the Illusion Crumbles

🤖 Models & LLM·Tom Levy·

ChatGPT Outperforms Human Authors, But the Illusion Crumbles

ChatGPT Outperforms Human Authors, But the Illusion Crumbles
Key Takeaways
1A study reveals that readers cannot distinguish texts from ChatGPT from those written by humans.
2Over 2,500 participants rated AI-generated stories more favorably.
3The scores for the texts drop when readers discover their artificial origin.
💡Why it mattersThis raises questions about the perception of creativity and trust in artificial intelligence in literature.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

ChatGPT Surpasses Human Authors, But the Illusion Crumbles

Readers struggle to distinguish between stories generated by ChatGPT and those written by humans. They even rate the AI-generated texts higher, but only as long as they are unaware that a machine produced them.

In three experiments involving over 2,500 participants, subjects failed to better differentiate between human-written stories and those generated by ChatGPT than by pure chance.

In the first experiment, each of the 1,682 participants read one of six stories, each about 1,000 words long. Three came from well-known literary magazines and short story collections. The other three were generated using ChatGPT 4.0, with instructions based on the theme, style, and narrative perspective of the human originals.

Half of the participants were informed that the story was written by a human. The other half were told it came from ChatGPT. This information was accurate for only half of the participants in each group, according to researchers Sydney Sears and Deena Skolnick Weisberg in their study published in the journal Judgment and Decision Making.

The stories from ChatGPT were rated significantly higher than the human-written texts in terms of perceived quality and immersion. For quality, the average score for AI-generated stories was 1.54 compared to 0.97 for human stories on a scale of -3 to +3. For immersion, the gap was 1.42 versus 1.00.

Participants' attitudes toward AI also influenced their ratings. Regardless of who actually wrote the story, participants assigned higher scores when told a human was the author.

Participants with a positive attitude toward AI generally gave higher ratings overall. When they were also informed that the story came from ChatGPT, their scores increased further. Among participants skeptical of AI, this effect reversed. A previous study on AI-generated poems revealed a similar bias.

In two other experiments with 905 participants, researchers made the task more challenging. Each person read both a human-written story and an AI-generated story, then had to determine which was which. Even with a direct comparison, participants did not perform better than by pure chance.

Self-reported experience with AI systems was positively correlated with the ability to correctly identify the origin of the stories. In contrast, self-reported experience with fiction did not help participants distinguish them.

Higher Ratings Do Not Necessarily Mean Better Writing

AI-generated texts tend to be more fluid, easier to read, and more emotionally optimistic than human-written texts. According to the authors, these characteristics could explain the higher ratings without the AI stories being genuinely better in literary terms. People tend to prefer materials that are easier to process.

High-quality literary fiction, on the other hand, is often intentionally difficult to access and pushes readers to work to extract meaning. A story can be of high quality but not engaging, and vice versa, the researchers suggest.

The short story format likely also plays in favor of AI. Telling a coherent story in 1,000 words is a very different challenge than doing so over hundreds of pages. Nevertheless, the researchers' conclusion is clear: AI can generate creative works that people perceive as at least on par with those of humans, but people do not believe that AI is capable of this.

All data and materials from the study are freely available on the Open Science Framework.

A study from last year conducted by Stony Brook University and Columbia Law School shows that reader expertise only matters up to a certain point. With simple instructions, professional readers clearly preferred texts written by humans. However, when models were specifically trained on the styles of individual authors, experts preferred AI-generated texts eight times more often for style imitation and twice as often for writing quality.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.