How AI Detection Scores Vary Across Content Formats: Blogs, Essays, Emails, Social Posts, and More Compared
By Blog AI 18-09-2026 1
Run the same detector across a 6,000-word thesis and a four-line LinkedIn post, and there is a real chance the LinkedIn post scores higher. Phrasly’s 2026 study of 50 real content samples found LinkedIn posts averaging a 79.3% AI-generated score, almost double what marketing web pages returned in the same test (44.2%), with the median LinkedIn score sitting at a full 100% (Phrasly, 2026). Length did not rescue the theses either. What the data actually shows is that detection scores rise and fall on pattern density inside a piece, not word count or “long-form” status. Here is where those numbers land across formats, and why one format is so much easier to flag than another.
Shorter is not safer, and longer is not smarter
The instinct most writers bring to a detector is wrong. Long-form gives you room to sound human, the thinking goes, and short posts are too tight for a detector to grip. Phrasly’s data reverses both. Across 20 blogs, 10 LinkedIn posts, 10 theses, and 10 web pages, length showed almost no relationship to score. Some of the shortest LinkedIn posts in the sample returned a full 100% AI-generated (Phrasly, 2026). A 6,000-word thesis can score as high as a 200-word LinkedIn post because detectors read pattern density, not paragraph count.
The broader context helps here. An Ahrefs corpus of about 900,000 pages found that 74% of new web content in 2025 carried AI-generated text (Ahrefs, 2025), and HubSpot’s State of AI report put marketer AI use at 80% (HubSpot, 2024). The volume is enormous. The fingerprint just shows up unevenly.
The detection numbers, format by format
Here is what the same detector scored across the four sampled formats, straight from Phrasly’s Average AI Detection Score by Content Format study, which ran all 50 samples through the Phrasly AI Detector in Q2 2026.
LinkedIn posts, 79.3% average, 100% median
LinkedIn was the highest-scoring format in the study, and the only category where the median score reached a full 100%. At least half of the sampled posts were fully flagged (Phrasly, 2026). Two habits drive the number. The “Whether you…” opener appears in nearly every AI-drafted LinkedIn post, and the format tends to run three neat paragraphs of similar length, which reads to a detector as machine-tidy rhythm. Em dash use peaks here, averaging 3.6 per post.
Academic theses, 69.4% average, 72% median
Theses came in second. The words “foster” and “ultimately” appear so frequently in AI-drafted academic writing that both landed inside the study’s top 15 flagged phrase patterns. A 2024 arxiv paper analyzing 26,700 words of academic text recorded “delves” appearing 28 times more often after 2023 than before, “underscores” nearly 14 times more often, and “showcasing” more than 10 times more often (Kobak et al., arxiv 2406.07016). Formal register is where models train hardest, and the vocabulary shift shows up on the page.
Blog posts, 58.3% average, 52% median
Blogs landed near the study average. The dominant giveaway was the “not just X, but also Y” construction, which appeared in more than a third of the blog sample. Blogs also carried the second-highest em dash count in the study (3.15 per piece), a punctuation habit that spiked after ChatGPT hit mainstream use.
Web and product pages, 44.2% average, 38% median
Marketing web pages scored the lowest. That surprised the research team, since web copy is written to sell and often reuses persuasive templates. What likely helps web pages is heavy human editing, brand-voice constraints, and short, punchy sentences that break the smooth rhythm most detectors weight. “Seamless” was still the most common flagged word inside this format (Phrasly, 2026).
What about emails and social posts beyond LinkedIn?
The Phrasly study did not include email samples, but the surrounding research points a clear way. AI’s contraction rate flips almost entirely with the register of the prompt: three chatbots asked for a formal task wrote zero contractions across nine test answers, while a conversational-prompt task from ChatGPT averaged 48.9 contractions per 1,000 words (Phrasly, 2026). Formal emails, dense with nominalizations and abstract nouns, would likely score close to theses. Casual emails and short social posts on X or Instagram, being contraction-heavy and structurally messier, tend to score closer to web pages or lower.
Why the same detector reads formats so differently
Every mainstream AI detector runs on the same underlying idea: measure how predictable the writing is at the sentence and token level, then score how far it drifts from average human variance. Formats that push a writer toward smooth, uniform rhythm register as more AI-like even when a human wrote them. That is why LinkedIn wins the detection race. The platform culture rewards clean parallelism, one-idea-per-paragraph structure, and hook openers, which happens to be the exact shape a large language model produces by default. Long-form theses trail LinkedIn only because they carry more variation across sections like methods, results, and discussion, which dilutes pattern density.
A useful frame comes from Phrasly’s AI Writing Pattern research on contractions: contraction rate flips almost entirely with register, and register flips with format. Detectors do not see intent, they see rhythm and vocabulary consistency. Format is what enforces both.
How to bring your detection score down, whatever you write
Four moves work across every format tested.
Break the rhythm
Write one three-word sentence in a paragraph of long ones. Fragment a clause. Detectors weight sentence-length variance heavily, and this is the fastest lever you have.
Kill the giveaway phrases
“Not just X, but also Y” in blogs. “Whether you…” in LinkedIn. “Foster” and “ultimately” in theses. Search your draft for the top 15 flagged patterns and replace them with sharper, native phrasing.
Trade abstraction for specifics
Instead of “significantly improves engagement,” write “moves engagement from 2% to 3.4%.” Detectors read generic detail as low burstiness, and low burstiness reads as AI.
Use a sentence-level detector, not just a total score
An overall percentage tells you almost nothing about where the problem lives. Phrasly’s AI Detector highlights individual flagged sentences, which is where the actual fix happens.
Format decides where the gauge starts
Detection is a function of format, not effort. A LinkedIn post is fighting the platform’s own conventions the moment it exists, and a marketing web page has structural room to look human that a thesis does not. That does not mean LinkedIn writers are cheating or web copywriters are more original, it means the pattern density baked into each format decides where the detector’s gauge starts. If you write across multiple formats, treat your scores as format-relative benchmarks rather than moral verdicts on your writing. Break the rhythm, cut the tell-tale phrases, and run the passage through a sentence-level check before you publish. The patterns are predictable. Once you know which one is coming for you, they stop being a real problem.
Tags : .....