Graphite finds nearly 13,000 linguistic tells in AI writing
“This matters” was 116 times more common in AI text, while the signatures of Claude, GPT and Gemini changed in different directions.
Graphite has identified 12,877 words, phrases and sentence frames that appeared at least twice as often in AI-generated articles as in comparable human writing. The most pronounced example was “this matters,” which appeared 116 times as often in AI text as in the human baseline, according to coverage of the study.
Graphite calls such features “AI tells.” Its study used 10,000 human-written Common Crawl articles published before ChatGPT’s Nov. 30, 2022 launch, then had nine models produce 90,000 articles from summaries of the same source material. That design was intended to match topics between the human and machine-written sets without copying the original articles.
The company examined single words, two- and three-word sequences, recurring sentence frames, sentence-length variation and other textual features. Each of the nine models generated 10,000 articles, and each model was found to have between 2,355 and 3,746 tells.
The findings suggest that conspicuous AI habits can be removed without eliminating model-specific patterns. Graphite found that 65% of all tells were unique to a single model family, and that 55% to 72% of tells were unique to one version when it compared the two latest tested versions in each family.
Claude Opus 5.5, for example, used the construction “why X matters” 92 times as often as the human samples. It used “dependable” 23 times as often, according to TechCrunch’s account of the research. Claude Opus 5 also made “less X, more Y” comparisons 82 times as often as human writers in Graphite’s analysis.
GPT-6 Astra displayed a different style. Graphite found that it used “may” three times as often as GPT-4.1, while reporting on the study identified hedges such as “may provide” and “can provide” as common Astra phrasing. Astra also used “does not establish” 275 times as often as Claude output, according to BigGo Finance.
Astra frequently used what Graphite termed corrective framing, such as describing something as “not simply X” or choosing one approach “rather than relying on X.” The company’s report put Astra’s overall use of this framing at 12 times the human rate; TechCrunch reported that particular corrective constructions exceeded the human rate by more than 100 times. Astra also commonly invoked “another dimension” of the subject under discussion.
The study found divergent trajectories among model families. Under Graphite’s word-distribution measure, Claude was the only family to become more human-like over time, while the divergences for GPT and Gemini rose overall. Between the earliest and latest versions tested, the number of tells fell 29% for Claude and 32% for Gemini, but rose 48% for GPT.
That does not mean the broadest and most familiar signals are becoming more prevalent. Across nine well-known features, their combined frequency fell by 41% to 86% in every model family from the earliest to the latest version examined. Claude Opus 5.5 used em dashes 99% less often than Opus 5, while Astra’s em-dash rate was 88% below the pre-ChatGPT human rate. Gemini 3.1 Pro had almost eliminated them.
GPT-6 Astra had 3,687 identified tells, slightly more than GPT-5.6 Sol, but the two shared only 45% of their tells. Sol used “in addition” 14 times as often as Astra, whereas Astra used “need not” 17 times as often as Sol. Graphite said GPT-6 Astra, Claude Opus 5 and Gemini 3.1 Pro together accounted for 7,043 distinct tells.
The study also identified broader differences in prose. AI articles used wider vocabulary within individual pieces under Graphite’s MTLD measure, and their words averaged 5.2 to 5.7 characters, compared with 4.9 in human writing. Their sentence lengths varied less, while human writers used exclamation marks more than 100 times as often as each of the three current models in one comparison.
Claude Opus 5 and Gemini 3.1 Pro recorded mannered-prose scores about 2.5 times the human score, while Astra was much closer to the human result. Every model’s word- and phrase-use distribution was closer to that of at least one other model than to the human baseline, although Graphite cautioned that shared prompts may account for part of the similarity.
The research is from a marketing firm with a commercial interest in AI detection and content analytics, rather than an independent academic institution. Separate peer-reviewed research on Japanese texts nevertheless found that stylometric features could distinguish its 350 AI-generated samples from 100 human public comments with 99.8% random-forest accuracy on those materials, even as human participants had limited success identifying AI writing themselves.
Graphite’s Greg Druck said the results showed that removing recognizable habits can simply replace them with new ones. “These are giant models with billions of parameters,” he told TechCrunch. “They have some finite number of tests they can run, and things slip through.”