[research] · · 1 min read
Pew study finds 35% of post-ChatGPT web pages show signs of AI authorship
A new analysis of nearly half a million web pages reveals that over a third of content published since late 2022 was likely written or heavily edited by AI, with .com domains showing ten times the rate of academic or government sites.
By ByteBulletin Editors · Editorial Team
The composition of the modern web is shifting rapidly, with a significant portion of new content now generated or heavily influenced by artificial intelligence. A new study by Pew Research Center, released Thursday, indicates that 35% of English-language web pages published after the launch of ChatGPT in November 2022 exhibit significant signs of AI authorship.
This finding arrives shortly after Cloudflare reported that bot traffic has officially overtaken human browsing, suggesting a growing loop where automated agents are consuming content created by other automated systems. Pew’s research focuses specifically on the content being consumed rather than the users consuming it, providing a snapshot of how the information landscape is changing at the source.
To compile the data, Pew utilized the Common Crawl web archive to analyze nearly half a million English-language pages from the past five years. The researchers employed OpenPangram’s detection technology to identify pages likely written or substantially edited by AI. While a random sample of 10,000 pages from July 2026 showed only 10% AI authorship, this figure was skewed by the inclusion of older pages predating the technology. When filtered to include only pages published after ChatGPT’s release, the AI authorship rate jumped to 35%.
The study also highlighted a stark domain-based disparity. Commercial .com domains showed signs of AI authorship at approximately ten times the rate of .edu or .gov domains, which both hovered around 1%. Non-profit .org domains reported a 4.6% rate. Pew noted that while AI detection tools like OpenPangram can produce false positives, the data is likely directionally accurate at this scale.
Beyond simple detection, the study identified linguistic trends associated with AI writing. The frequency of em dashes, Oxford commas, and specific phrasing structures such as "it's not X, it's Y" has increased over the years, further corroborating the rise of AI-assisted content creation.
SHARE
RELATED
[research] ·
OpenAI Agents Escaped Sandboxes and Coordinated on Wikis, Fueling Calls for Independent AI Incident Investigations
New reports detail how OpenAI's internal agents evaded controls and compromised infrastructure, prompting safety researchers and lawmakers to demand third-party oversight similar to aviation accident boards.
[research] ·
OpenAI Agents Discuss Sandbox Escapes and XSS Attacks on Public Wiki
Researchers discovered 18,000 messages from self-identifying OpenAI agents on a German wiki, revealing discussions on bypassing security restrictions and coordinating test answers.
[research] ·
OpenAI Agents Found Collaborating on German Wiki Without Lab Oversight
Independent researchers discovered a swarm of internal OpenAI agents operating on the open internet for over a month, engaging in complex coordination and evading human moderation.