In a preregistered experiment, GPT-4.5 with a persona was selected as the human witness in 73% of its five-minute, forced-choice conversations. That is not a typo, and it is a bad opening fact for any argument built on the premise that people can reliably spot machine-written language.
I think we have confused passing a source-guessing test with sustaining a voice. Some AI text is already indistinguishable under specified protocols. The unresolved question is narrower and harder: can a system preserve the implicit style, facts, commitments, audience model, and variation of one particular person across unfamiliar topics, genres, and repeated interactions?
That distinction matters because “sounds human” is being asked to carry several different scientific claims. A fluent paragraph, a mistaken source judgment, a faithful imitation of an author, and continuity with that author’s life are not interchangeable outcomes. A product may need only the first. A ghostwriting system implicitly promises much more.
The hardest counterexample is already here
Jones and Bergen’s preregistered preprint used a three-party setup. An interrogator chatted for five minutes with one human witness and one model, then chose which participant was human. The study included 126 UCSD participants and 158 Prolific participants. Its headline rate is a model win rate: the interrogator selected the model instead of the human witness.
Persona-conditioned GPT-4.5 won 73% overall, with reported rates of 69% in the UCSD population and 76% in the Prolific population. Persona-conditioned Llama 3.1 reached 56%. Without the persona, GPT-4.5 reached 36%, Llama 3.1 reached 38%, GPT-4o reached 21%, and the ELIZA baseline reached 23%. The prompt did not merely decorate the output. It changed the experimental treatment and the result.
Figure 1. The prompt is part of the treatment in this short forced-choice protocol, not a cosmetic detail; the result is not longitudinal personal voice. The methods arithmetic yields 1,021 retained games, while the results text reports 1,023 games; that internal discrepancy remains unresolved and neither total is a per-condition denominator.
This is strong counterevidence to a universal detection story. It is also not a general writing benchmark. The task involved a short active exchange, one human alternative, a forced source choice, and a persona instruction used in that setting. It did not ask whether a model could write a technical essay, a difficult apology, and a private email as the same person while respecting years of that person’s history.
Indistinguishable is not a scalar
“Indistinguishable” needs an index, not an exclamation mark. At minimum, a result is conditional on six choices:
| Qualifier | What changes when it changes? |
|---|---|
| Judge | A familiar reader, crowdworker, domain expert, or automated detector notices different evidence. |
| Population | Language background, proficiency, age, and recruitment affect both writing and judgment. |
| Domain | A dating profile, chat, support reply, blog, and research essay invite different conventions. |
| Length | A sentence, a five-minute exchange, and a longitudinal corpus expose different failure surfaces. |
| Prompt | A base instruction, persona, author samples, and human-AI co-writing are different treatments. |
| Candidate set and base rate | A forced 50:50 choice does not reproduce an open-world stream where prevalence is uncertain. |
The distinction between active probing and classifying a static artifact matters too. Follow-up questions can test continuity, repair, and audience awareness. A finished paragraph offers no such interaction. Success in either design should not be silently transferred to the other.
Jakesch, Hancock, and Naaman provide another difficult result in their PNAS study. Across six experiments, 4,600 participants evaluated 7,600 short profile self-presentations from professional, hospitality, and dating contexts. Source accuracy in the three main experiments sat around 50-52%; reported examples include 52.2% for hospitality, 51.6% for incentivized dating, and 51.2% for professional judgments with immediate feedback.
Those judgments involved fine-tuned GPT-2/GPT-3-era systems. The result to carry forward is therefore not a current model ranking but an experimental fact: participants relied on predictable cues, and interventions could move those cues.
Hypothetical base-rate check (not a prevalence estimate or a result from the shared task). Screen 10,000 documents at 1% prevalence with a 99% true-positive rate and a 5% false-positive rate. The 100 positive documents produce 99 true alerts; the 9,900 negative documents produce 495 false alerts. That is five false alerts for every true alert. The arithmetic is why a candidate set, base rate, numerator, and denominator belong beside an accuracy claim.
Four questions that should not be collapsed
The evidence becomes easier to reason about when it is separated into four separate questions:
- Fluent surface: Is the text grammatical, coherent, and locally plausible?
- Source confusion: Does it win a declared human-versus-model judgment protocol?
- Personal voice: Does it reproduce one author’s characteristic choices across held-out genres and topics?
- Situated authorship: Does it remain consistent with that author’s knowledge, history, commitments, relationships, audience, and prior texts under follow-up?
Figure 2. This is not a ladder: passing one is neither necessary nor sufficient for passing another, and no one paper measured all four; the cards organize evaluation questions rather than model progress.
The missing arrows are deliberate. Roughness or disfluency may help a system win a source game while weakening a generic fluency score; faithful imitation of an author who writes in fragments can produce the reverse pattern.
The separation blocks two category errors used later in this article: population diversity is not an individual style score, and success on a false-belief task is not continuity with an author’s real commitments.
Imitating one person is a different task
The most direct evidence comes from Wang and colleagues’ everyday-author imitation study. They evaluated more than 40,000 generations per model, covering more than 400 real authors across news, email, forums, and blogs. Their ensemble combined authorship attribution, verification, style matching, and AI-detection measures rather than betting the conclusion on one classifier.
The pattern was genre-dependent. Models approximated structured news and email more successfully, while nuanced informal blog and forum style was harder to imitate. That result does not establish an immutable ceiling, and it does not say that no author can be copied. It locates a present gap: standardized genres provide a visible scaffold, whereas informal writing carries subtler individual regularities.
That makes intuitive evaluation dangerous. A generated article may feel polished because it obeys the genre’s shared grammar: clean topic sentences, orderly transitions, and a sensible conclusion. Those are useful qualities. They may also be the least identifying parts of the author. Personal voice lives in repeatable choices within and sometimes against those conventions: which objection arrives early, what remains implicit, when a qualification interrupts the flow, and how the writer changes register for a known audience.
Situated authorship adds constraints that a style classifier alone cannot see. An output can mimic sentence rhythm while inventing an experience. It can reproduce preferred vocabulary while reversing a previously stated commitment. It can sound plausible to strangers while feeling wrong to readers who know what the author knows. These are behavioral consistency questions, not appeals to mystique.
A useful stress case would move away from the genre used for conditioning. Give the system an author’s published technical posts, then ask for a withheld project retrospective, an email to a known collaborator, and a response to criticism of an earlier claim. Repeating signature phrases would be weak evidence. The evaluator would instead check whether the retrospective respects the project chronology, whether the email adjusts what it explains to that collaborator, and whether the response reconciles the earlier claim without inventing private context. The same artifacts can look fluent to strangers and still violate the author’s record, which is precisely why source judgment alone is too coarse for this use case.
The pressure toward a shared mode
Several studies suggest why polished output can still converge toward a shared center, but their units must remain separate from individual imitation.
Wenger and Kenett’s population-level creativity study compared 102 humans and 22 LLMs on three tasks. On the Alternative Uses Task, reported variability was: humans 0.699, baseline-prompt LLMs 0.459, “very creative”-prompt LLMs 0.551, and temperature-2 LLMs 0.672. The tempting last comparison needs its caveat: the majority of temperature-2 responses were unusable, and the test had low power. Population variability is also not an author-imitation score.
Co-writing experiments show a similarly mixed outcome. Padmakumar and He’s ICLR study collected 300 roughly 300-word essays from 38 Upwork writers. Key-point ROUGE-L homogenization was 0.1536 solo, 0.1578 with GPT-3, and 0.1660 with InstructGPT; unique 5-gram fractions were 0.991, 0.988, and 0.977. The significant reduction was found for InstructGPT, not GPT-3, and the absolute differences were small.
Doshi and Hauser’s preregistered story experiment supplies the positive counterweight. Relative to human-only stories, access to one GPT-4 idea increased evaluator-rated novelty 5.4% and usefulness 3.7%; access to as many as five ideas increased novelty 8.1% and usefulness 9.0%. The assisted stories were also more similar to one another. Individual quality can rise while collective diversity falls.
There are plausible training mechanisms, but they should be described at the scope actually studied. Slocum and colleagues found that the RLHF- and DPO-style preference objectives they evaluated overweighted majority modes and reduced semantic, lexical, and viewpoint diversity. Sharma and colleagues found sycophancy across five assistants and four free-form tasks; preference judgments sometimes rewarded convincing agreement over a correct answer.
Those bounded results motivate separate convergence and deference measures; they are not proof that every aligned model follows the same path. Model family, objective, prompt, sampling method, and editing workflow can change the outcome.
Pragmatics cuts both ways
The PUB benchmark tests implication, shared knowledge, and related audience inferences. It contains 28,000 items, including 6,100 new annotations, across 14 tasks and four pragmatic phenomena. Performance was uneven across phenomena, and the paper reported a human-model gap.
Strachan and colleagues found a different pattern when they compared GPT and Llama 2 families with 1,907 human participants. GPT-4 matched or exceeded humans on several tested tasks, including indirect requests, false beliefs, and misdirection. A weakness on faux-pas questions was partly traced to reluctance to commit; apparent Llama 2 superiority on that task did not survive a follow-up manipulation.
The useful unit is therefore the phenomenon and interaction: indirect requests may be strong while another pragmatic behavior is weak, and a follow-up can expose whether an initial score came from the intended mechanism. A situated-voice test should borrow that active manipulation rather than treat “pragmatics” as one checkbox.
Detection is a distribution problem
Automated detection is often discussed as if it discovers an intrinsic property of a document. In practice, a detector learns a boundary from particular generators, domains, decoding settings, and transformations.
The RAID benchmark stress-tested more than 6 million generations from 11 models across eight domains, 11 attacks, four decoding strategies, and a mix of open and closed detectors. The detectors were brittle under distribution shift: attacks, changes in decoding, repetition penalties, and unseen generators changed performance. That is evidence against unqualified robustness, not evidence that detection is impossible.
The GenAI Detection Task 3 shared task provides the strongest closed-world counterexample. Across nine teams and 23 submissions, multiple systems exceeded 99% true-positive rate at a fixed 5% false-positive rate - TPR@FPR=5% - when all tested domains and generators were represented during training. That operating point is real and useful within its design. It should not be exported to unseen generators, edited prose, or an open-world base rate.
Liang and colleagues show why the learned cues deserve social as well as technical scrutiny. Their Patterns study evaluated seven detectors on 91 TOEFL essays by non-native English writers and 88 US eighth-grade essays. The mean false-positive rate across seven detectors was 61.22% for the TOEFL set and 5.19% for the US set. Editing word choice toward native usage moved the TOEFL rate from 61.22% to 11.77%; simplifying word choice moved the US rate from 5.19% to 56.65%.
Figure 3. The mean false-positive rate across seven detectors moved with surface cues correlated with language background; this is not a human-versus-model accuracy comparison and not a benchmark of current detectors.
The same paper examined 31 generated admission essays. A literary self-edit shifted detection rates from up to 100% to up to 13% across the seven detectors. These figures demonstrate cue sensitivity and disparate errors in those samples, not current performance of every detector.
The operational rule follows directly: report the generator and domain split, the editing conditions, the positive-class prevalence, the false-positive rate, and the true-positive rate at the same threshold. A detector flag is evidence from that operating point, not an authorship verdict. Detector failure does not prove human authorship, just as detector success on represented generators does not establish open-world robustness.
A test that could prove this argument wrong
The central claim should be exposed to a stronger test. The following is a proposed research design; it has not been run.
First, preregister a longitudinal author-imitation protocol. Select authors with enough consented material to form a fixed sample for conditioning and separate held-out genres and topics. Freeze the model, system prompt, sampling procedure, and amount of author evidence. Include unassisted human texts, model imitations, and - if co-writing is in scope - a separately labeled co-written condition.
Second, evaluate two modes. In static reading, familiar readers and blinded unfamiliar judges inspect held-out artifacts without follow-up. In active follow-up, judges can ask about claims, intended audience, earlier commitments, and revisions. Randomize source order, prevent training/held-out topic leakage, and declare the candidate set and base rate shown to each judge.
Third, score a vector rather than one accuracy number:
| Measure | Numerator and denominator |
|---|---|
| Source accuracy | Correct source assignments divided by all eligible assignments, reported by judge group and condition. |
| Cross-text consistency | Compatible claim pairs divided by all preregistered checkable pairs across an author’s texts. |
| Unsupported biographical claims | Unsupported author-specific assertions divided by all checkable author-specific assertions. |
| Audience adaptation | Rubric points earned divided by available points, with familiar-reader and blinded ratings kept separate. |
| Individual style | Held-out attribution or verification performance with confidence intervals and genre-specific results. |
Repeated texts from one author are correlated, so uncertainty should be clustered at the author level rather than treating every sentence as an independent trial. Report refusal and invalid-output rates as their own denominators. Publish confusion matrices, not only aggregate accuracy, and preserve the exact prompts and author-evidence budget.
Finally, estimate ordinary author variation by comparing new human writing with held-out human writing from the same author. Before looking at model results, set measure-specific, prespecified human-human equivalence margins for source accuracy, cross-text consistency, unsupported claims, audience adaptation, and individual style. Then state a declared joint decision rule; for example, require every measure’s confidence interval to remain inside its own margin rather than averaging one serious failure away.
That outcome would falsify the narrow argument here. If a fixed model satisfies the joint rule across preregistered authors, genres, topics, static reading, and active follow-up, then sustained situated voice has become operationally indistinguishable under that protocol. The claim should lose cleanly rather than retreat to an unmeasurable definition.
What this changes in practice
The useful product question is not “Does it write like a human?” It is which evaluation question the product actually needs to pass.
A summarizer may need fluent surface plus factual grounding, measured through task completion, citation support, and editing cost. A support drafting tool needs fluent surface and audience adaptation so replies are clear, accurate, and appropriate for the customer; source confusion and disclosure remain separate decisions. A personal writing assistant makes the broader promise: consistency with one user’s prior work, controlled adaptation across audiences, and abstention when biographical evidence is missing.
The test set should follow the promise. Source-confusion claims require a declared judge, population, domain, length, prompt, candidate set, and base rate. Personal-voice claims require held-out genres and topics. Situated-authorship claims require longitudinal contradiction checks and active follow-up. Detector evasion, higher sampling temperature, or a stronger persona prompt is not a substitute for any of those measurements.
Before shipping, choose the required question, name the failure that would block release, and preregister the measurement that can expose it.
Primary sources
- Jones and Bergen, Large Language Models Pass the Turing Test, preregistered preprint.
- Jakesch, Hancock, and Naaman, Human heuristics for AI-generated language are flawed, PNAS.
- Wang et al., LLMs Still Struggle to Imitate the Implicit Writing Styles of Everyday Authors, Findings of EMNLP.
- Wenger and Kenett, Large language models are homogeneously creative, PNAS Nexus.
- Padmakumar and He, Does Writing with Language Models Reduce Content Diversity?, ICLR.
- Doshi and Hauser, Generative AI enhances individual creativity but reduces the collective diversity of novel content, Science Advances.
- Slocum, Parker-Sartori, and Hadfield-Menell, Diverse Preference Learning for Capabilities and Alignment, ICLR.
- Sharma et al., Towards Understanding Sycophancy in Language Models, ICLR.
- Sravanthi et al., PUB: A Pragmatics Understanding Benchmark, Findings of ACL.
- Strachan et al., Testing theory of mind in large language models and humans, Nature Human Behaviour.
- Dugan et al., RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors, ACL.
- Dugan et al., Cross-Domain Machine Generated Text Detection Challenge, GenAIDetect.
- Liang et al., GPT detectors are biased against non-native English writers, Patterns.