Archivist David Rosenthal observes now that more material is posted online by LLMs than actual people, the bots are starting to ingest their own digital excrement, creating a negative feedback loop.
In the belief that "more is better", Large Language Models (LLMs) have insatiable appetites for training data. They started by scraping everything on the Web (robots.txt be dammed). When that ran out they downloaded the various pirate libraries (copyright be dammed). That exhausted the texts easily available in digital form, but their hunger wasn't assuaged. As for images, they partly used CAPTCHAs but mostly paid vast numbers of poor people to label the images with what they showed.
When the supply of text ran low, people observed that the LLMs were capable of generating human-like text in large quantities. The obvious idea was to pour the output of the LLMs into their training sets. This wasn't just a conscious decision, it was inevitable. The advent of LLMs rapidly polluted the Web with LLM output. Greg Druck's AI Now Writes as Many Online Articles as Humans notes that:
We observe significant growth in primarily AI-generated articles, coinciding with the launch of ChatGPT in November 2022. After only 12 months, primarily AI-generated articles accounted for 35.9% of articles published.
In Q1 2025, the quantity of primarily AI-generated articles being published on the web nearly equaled the quantity of human-written articles, 49.6% vs. 50.4%. In Q4 2025, primarily AI-generated articles surpassed human-written at 50.9%, before returning to 49.9% in Q1 2026.
Even if slop were not of undesirable quality, it is not produced by humans and thus is completely unsuitable as training data.
Previously:
(2026) A Wikipedia Clone Built on AI Hallucinations is Here to Hasten Along the Death of the Internet
(2025) When It All Comes Crashing Down: The Aftermath of the AI Boom
(2025) AI Favors Texts Written by Other AIs, Even When They're Worse Than Human Ones
(2025) What the Hell is Going on Right Now?
(Score: 2) by hendrikboom on Thursday July 09, @10:14PM (1 child)
Reality may imply consistency, but consistency does not imply reality.
Do we merely want the LLM output to be consistent? Or do we want it to be correct?
(Score: 2) by JoeMerchant on Thursday July 09, @10:59PM
>>there should be a tipping point at which self-reflection yields a more consistent worldview, not less.
>Reality may imply consistency, but consistency does not imply reality.
Say you have different methods of determining "truth." All are restricted to accessing information via the internet, but... some methods yield consistent results, while other methods yield inconsistent / conflicting / non-repeatable results. The more you compare your "internal picture of reality" against these methods, assuming your internal picture isn't too flawed, you can home in on the most reliable methods.
Not just that A, B and C agree with me, so I'll just read A, B and C in the future... more like:
Pass 1: A, B and C have developed reliability ratings of 85, 94, and 72% respectively (continue for hundreds of thousands of sources...)
Pass 2: Drop the lower quartile of sources by reliability ratings - now re-rate the remaining sources considering only those which were not dropped in Pass 1
Pass 3: Again drop the lower quartile of sources by ratings after pass 2, but bring back the top 50% of sources that were dropped in pass 2, if they were truly unreliable, they'll be dropped again in this pass...
Pass 4: Again drop the lower quartile of sources by ratings after pass 3, but bring back the top 50% of sources that were dropped in pass 3....
And, variations thereof, tweaking lower quartile and bring back rates for optimal apparent self consistency.
>Do we merely want the LLM output to be consistent?
Of course not, but, lacking the ability to perform actual experiments on hypotheses beyond meta-studies of published data...
>Or do we want it to be correct?
It can only be as correct as the best data in its training set. If the entire training set is writings by people who believe that the earth is flat, an LLM would be unlikely to derive a better truth from that bulk of publication... although, flat earthers often do publish self-contradictory conclusions, so if there is any non-flat earth support in there at all, even basic geometric measurements of sun and shadow angles, the LLM might very well throw all that out. This is the power of diversity: including many different kinds of source material instead of exclusively focused specialist training. It works for people as well as statistical models.
🌻🌻🌻🌻✌️ [google.com]