Stories
Slash Boxes
Comments

SoylentNews is people

SoylentNews is powered by your submissions, so send in your scoop. Only 16 submissions in the queue.
posted by janrinok on Friday May 15, @07:24PM   Printer-friendly

A Wikipedia Clone Built on AI Hallucinations Is Here to Hasten Along the Death of the Internet:

There's a theory that a rising tide of LLM-generated nonsense will eventually drown both LLMs themselves and the internet as a whole. The idea goes like this: The first generation of LLMs is trained entirely on "real" material: the Gutenberg project, 4chan, that one article from Thought Catalog a decade ago, and everything in between. But as the output of those LLMs spreads across the internet, it also becomes part of the training data of future LLMs—and much of it is bullshit .

As a result, the quality of newer LLMs' training data is inferior to that of their predecessors—and by extension, so is their output. And as that output accumulates on the internet, it becomes part of future training data, and the cycle continues. With each passing day, the proportion of the internet that's low-quality LLM-generated bullshit increases, until eventually all that's left to train LLMs is the gibberish created by their predecessors.

The end result is a sort of RAM-hoovering, water-guzzling, bullshit-munching ouroboros, an unholy circular undulant with Jensen Huang's face at one end and Sam Altman's at the other, slowly human-centipeding both itself and the internet into oblivion. If humanity hasn't set fire to the planet by that point, then we start a new internet, hopefully with lessons learned along the way.

And even if the doomsday scenario of the internet drowning in a sea of em dashes and it's-not-just-x-it's-y constructions never comes to pass, people are starting to take the idea of using LLMs to poison LLM training data and run with it.

Take, for example, Halupedia , an absurdist Wikipedia-esque site whose pages are entirely populated by content that an LLM has made up—sorry, hallucinated— on demand. If you search for a topic that someone has previously entered, you'll get the existing nonsense. If your search is the first of its kind, the LLM will carefully assemble your very own small mound of nonsense from a list of possible topics.

According to the site's tips-for-tokens page , Halupedia appears to be the work of one Bartłomiej Strama. The page also provides a little more insight into the purpose of the project, which isn't 100% clear at face value—Strama tells one contributor, "Your contribution towards polluting LLM training data will surely benefit society!"

Of course, quibblers might argue that there's more than enough LLM-generated rubbish on the internet already without sites deliberately adding to the pile. Google pretty much anything these days and you'll find umpteen long-winded articles that purport to explain the topic in question, but really just waffle for paragraph after paragraph without saying anything at all. This is certainly true, but there's some virtue in the fact that Halupedia's output is openly and exuberantly absurd as opposed to content that is superficially credible and doesn't reveal its true nature without closer inspection.

Although... you may also find yourself wondering which topics other users have been entering into Halupedia. After all, you can basically enter any subject into the site's "search" bar and have it write an article for you. The answer lies in the site's list of trending topics, and... sigh.

Yep, it's the usual mix of shitposts, nonsense, and unabashed racism—or, in other words, it's basically the internet's id in microcosm. In fairness, some of these pages have been deleted—click on "niggabutt" and you get this:

But since the page title still shows up in the sidebar, it's not like it's been entirely banished. On the tip page, Strama also comments on the challenges of moderation: "The moderation sometimes is too restrict, but at least it's not griefed now." That's as it may be, but it's hard to see this ending well once 4chan gets a hold of it. This is why we can't have nice things, etc.


Original Submission

Related Stories

David Rosenthal on the LLM Negative Feedback Loop 29 comments

Archivist David Rosenthal observes now that more material is posted online by LLMs than actual people, the bots are starting to ingest their own digital excrement, creating a negative feedback loop.

In the belief that "more is better", Large Language Models (LLMs) have insatiable appetites for training data. They started by scraping everything on the Web (robots.txt be dammed). When that ran out they downloaded the various pirate libraries (copyright be dammed). That exhausted the texts easily available in digital form, but their hunger wasn't assuaged. As for images, they partly used CAPTCHAs but mostly paid vast numbers of poor people to label the images with what they showed.

When the supply of text ran low, people observed that the LLMs were capable of generating human-like text in large quantities. The obvious idea was to pour the output of the LLMs into their training sets. This wasn't just a conscious decision, it was inevitable. The advent of LLMs rapidly polluted the Web with LLM output. Greg Druck's AI Now Writes as Many Online Articles as Humans notes that:

We observe significant growth in primarily AI-generated articles, coinciding with the launch of ChatGPT in November 2022. After only 12 months, primarily AI-generated articles accounted for 35.9% of articles published.

In Q1 2025, the quantity of primarily AI-generated articles being published on the web nearly equaled the quantity of human-written articles, 49.6% vs. 50.4%. In Q4 2025, primarily AI-generated articles surpassed human-written at 50.9%, before returning to 49.9% in Q1 2026.

Even if slop were not of undesirable quality, it is not produced by humans and thus is completely unsuitable as training data.

Previously:
(2026) A Wikipedia Clone Built on AI Hallucinations is Here to Hasten Along the Death of the Internet
(2025) When It All Comes Crashing Down: The Aftermath of the AI Boom
(2025) AI Favors Texts Written by Other AIs, Even When They're Worse Than Human Ones
(2025) What the Hell is Going on Right Now?


Original Submission

This discussion was created by janrinok (52) for logged-in users only, but now has been archived. No new comments can be posted.
Display Options Threshold/Breakthrough Mark All as Read Mark All as Unread
The Fine Print: The following comments are owned by whoever posted them. We are not responsible for them in any way.
(1)
  • (Score: 4, Funny) by BsAtHome on Friday May 15, @08:13PM (3 children)

    by BsAtHome (889) on Friday May 15, @08:13PM (#1442560)

    This is translate a phrase from English to German to Chinese to Spanish to Hindi to Portuguese to Hungarian to French to Russian to Turkish to Bengal to Thai to Swahili to Dutch to Greek to Persian to Finnish to Arabic to Malay to Vietnamese to Kalanga to Latin to Nepali to English.

    Try to see if anything sensible remains... Then try the reverse order of translations and see if you get back to the original.

    • (Score: 2) by chucky on Friday May 15, @08:20PM (1 child)

      by chucky (3309) on Friday May 15, @08:20PM (#1442561)

      And the phrase is “table the idea”.

      • (Score: 4, Funny) by BsAtHome on Friday May 15, @09:21PM

        by BsAtHome (889) on Friday May 15, @09:21PM (#1442573)

        And "table the idea" is apparently equal to: "lights are suffering from chains using sand correlated to the waterways swimming violin with friendly" when put forward and reverse through the translations.

        Ah well, could've been worse if a micro-wormhole had transmitted the words to a neighboring universe.

    • (Score: 2) by mcgrew on Sunday May 17, @08:37PM

      by mcgrew (701) <publish@mcgrewbooks.com> on Sunday May 17, @08:37PM (#1442733) Homepage Journal

      A story from the 1950s holds that they were developing a Russian/English translator with the infant tube computers. When written, they fed "the spirit is willing, but the flesh is weak" and then fed the Russian translation back to English, where it read "The wine is good but the meat is spoiled."

      --
      The Epstein Memorial Golden Presidential Ballroom: Let them eat cake.
  • (Score: 5, Informative) by turgid on Friday May 15, @08:21PM

    by turgid (4318) Subscriber Badge on Friday May 15, @08:21PM (#1442562) Journal
  • (Score: 3, Funny) by turgid on Friday May 15, @08:36PM

    by turgid (4318) Subscriber Badge on Friday May 15, @08:36PM (#1442564) Journal

    Burn, bay, burn [github.com], token inferno. Complete with glitter balls and everything.

  • (Score: 4, Interesting) by VLM on Friday May 15, @08:46PM (4 children)

    by VLM (445) on Friday May 15, @08:46PM (#1442567)

    With each passing day, the proportion of the internet that's low-quality LLM-generated bullshit increases

    Don't forget the other side of the ratio Clankers repel humans. I don't go to facebook or reddit anymore because I have zero interest in talking to clankers. Hackernews aka "orange reddit" has a clanker infestation but its not as bad. 4chan is mostly paid glowies and mentally ill "no its totally not a mental illness" people at this point, but it used to have lots of good folks and still has quite a few good folks, just not many. /out/ /diy/ most of /g/ and of course /x/ are good boys. /b/ /s/ /gif/ are mostly mental illnesses as a biography kind of people. /pol/ is like 90% glowies trying to entrap some moron for a promotion, I've seen some wild stuff there.

    Anyway, this means the proportion of bot traffic increases, speeding to the point the advertising fraud is uncovered at which point they'll collapse. You can't make money by paying to send ads to clankers.

    Long before Reddit is 100% clanker, it'll be defunded by advertisers. Clankers don't buy stuff and humans don't go to Reddit. So there will be no training data for LLM of the year 2030 because they'll be no scrapeable human posts on the internet once Reddit is dead. Its already an issue. LLMs can debug pre-2025 technical problems using shitty stackexchange, but stackexchange died, so LLMs will never be able to troubleshoot any post 2025 problem, there's no StackExchange anymore to plagiarize. You'll always be able to ask a LLM how to subnet an IPv4 address because there's fifty bazillion stackexchange posts telling morons that a /24 is a class C or whatever for the hundredth time; someday they'll be a IPv7 and LLMs will be clueless about it without anything to plagiarize.

    I suspect the bullshite way something like UBI will be rolled out alongside Big Brother is we'll get suspiciously just enough money to barely survive if we let Big Brother record our every social interaction. Then "they" can train their LLMs to better control us off that data. Go ahead upload shite to the internet all day, doesn't matter if humans read or watch it as long as the clankers get verified training data. Fits in with de-anonymization moves; they can't train clankers unless they know we are human. The Big Brother is watching chilling effects are a free bonus.

    I would not be surprised to see a social revolution along the lines of "neo-amish" sorta like the fictional (so far) Butlerian Jihad perhaps pre-dating the internet. I'd certainly think about moving there. They'd need EEs and computer programmers it would just be illegal to deploy and operate packet forwarder software or LLMs. There's a lot of fun to be had aside from those two attractive nuisances. Maybe they'd allow something like Fidonet or UUCP because its not realtime. I'd be just the kind of idiot getting shunned from the Neo-Amish Co-op because I set up UUCP between two emulated VAX retrocomputers for the sheer hell of it and got caught. Yeah, I'd move there; I'd probably even (mostly) behave myself. Possibly a good dividing line would be no interconnected systems; we could run Soylent News as a dial up modem BBS. That would be pretty awesome.

    • (Score: 3, Touché) by PiMuNu on Saturday May 16, @12:16AM (3 children)

      by PiMuNu (3823) on Saturday May 16, @12:16AM (#1442583)

      > I'd certainly think about moving there.

      Until the guys next door with AI powered drones see some nice things and take over. Might makes right.

      • (Score: 0) by Anonymous Coward on Saturday May 16, @01:53AM (2 children)

        by Anonymous Coward on Saturday May 16, @01:53AM (#1442593)

        Maybe. Or maybe the "neo-amish" would allow shotguns, as long as they were shooting at drones...

        • (Score: 2) by PiMuNu on Saturday May 16, @11:36AM

          by PiMuNu (3823) on Saturday May 16, @11:36AM (#1442601)

          Try telling Ukraine (or Iran, for that matter) that all they need is a couple of shotguns.

        • (Score: 2) by VLM on Saturday May 16, @12:03PM

          by VLM (445) on Saturday May 16, @12:03PM (#1442605)

          By my definition, everything up to a Phalanx CIWS would be fine as long as you don't network it and don't connect it to a LLM.

          Something like "Space 1999" would be fine on a moonbase; again no internet no LLMs.

  • (Score: 1, Flamebait) by Nobuddy on Wednesday May 20, @04:24PM

    by Nobuddy (1626) on Wednesday May 20, @04:24PM (#1443006)

    I thought they had found Conservapedia. Not AI, human generated hallucinated nonsense.

(1)