Start with two numbers: 1.54 against 0.97. On a scale that runs from minus 3 to plus 3, that is how readers scored the quality of short stories produced by ChatGPT 4.0 versus short stories written by people. The machine came out ahead.
The immersion scores tilted the same direction, 1.42 for the AI-written pieces against 1.00 for the human ones. And in three experiments involving more than 2,500 participants, not one group beat chance at working out which was which.
The setup left little room to bluff
The opening experiment gave each of its 1,682 participants a single short story of roughly 1,000 words, drawn from a pool of six. Three had been lifted from well-known literary magazines and short story collections. The other three were generated by ChatGPT 4.0, prompted to mirror the theme, style and narrative perspective of those human originals.
Then came the trick. One half of the participants was told the story had a human author. The other half was told ChatGPT had produced it. Within each of those groups the label was only correct for half the people, according to researchers Sydney Sears and Deena Skolnick Weisberg, whose work appeared in the journal Judgment and Decision Making.
The label moved the score more than the author did
Whoever had actually written the text, ratings climbed when readers believed a person was responsible. The byline was carrying weight the prose itself was not.
How people felt about AI going in bent things further still. Participants who arrived with a positive view of the technology scored stories higher across the board, and higher again when told ChatGPT was the author. Among those who distrusted AI, the effect reversed. A previous study using AI-generated poems surfaced exactly the same bias.

Reading both side by side didn’t help
Two follow-up experiments, 905 participants between them, made the job look far simpler on paper. Every person read one human story and one AI story and then named which was which. Straight comparison, nothing left to memory, a coin-flip baseline.
They still performed at chance.
Exactly one factor predicted who got it right, and it is not the obvious one. Self-reported experience with AI systems tracked positively with identifying the true origin. Self-reported experience with fiction did nothing whatsoever. Getting through a lot of novels apparently teaches you to enjoy prose rather than inspect it.
Easy to read isn’t the same as good
The authors offer a straightforward reading of those scores, and it is not that the model is the better writer. Text produced by AI tends to run smoother, easier on the reader and more emotionally upbeat than human prose. Since people gravitate toward material that is simple to process, those qualities can pull ratings up without anything literary sitting underneath them.
Serious literary fiction frequently operates on the opposite principle. It is difficult to enter by design and demands effort from you. As the researchers put it, a story can be high in quality while barely engaging, and the reverse holds too.

Length is a factor as well. Sustaining 1,000 words is a wholly different task from sustaining hundreds of pages, and the short story is a format that suits what the model does well. Even so, the conclusion holds: AI can turn out creative work that people rate as at least the equal of human work, while those same people do not believe AI is capable of it.
Train the model on one writer and the experts flip
Expertise helps, though not as much as you might assume. Research published last October by Stony Brook University and Columbia Law School found that when the prompts were simple, professional readers came down firmly on the side of the human-written texts.
Then the models were trained on the styles of individual authors. At that point the experts favoured the AI-generated texts eight times more often on style imitation, and twice as often on writing quality.

Every piece of data and material behind the Sears and Weisberg study sits openly on the Open Science Framework, so the comparison is one you can run on yourself instead of taking anybody’s word for the result. Try it before concluding you would have caught the machine. That is what everybody assumes.


















STAY ALWAYS UP TO DATE