Stories
Slash Boxes
Comments

SoylentNews is people

SoylentNews is powered by your submissions, so send in your scoop. Only 16 submissions in the queue.
posted by hubie on Thursday December 11 2025, @03:07PM   Printer-friendly

AI favors texts written by other AIs, even when they're worse than human ones:

As many of you already know, I'm a university professor. Specifically, I teach artificial intelligence at UPC.

Each semester, students must complete several projects in which they develop different AI systems to solve specific problems. Along with the code, they must submit a report explaining what they did, the decisions they made, and a critical analysis of their results.

Obviously, most of my students use ChatGPT to write their reports.

So this semester, for the first time, I decided to use a language model myself to grade their reports.

The results were catastrophic, in two ways:

  1. The LLM wasn't able to follow my grading criteria. It applied whatever criteria it felt like, ignoring my prompts. So it wasn't very helpful.
  2. The LLM loved the reports clearly written with ChatGPT, rating them higher than the higher-quality reports written by students.

In this post, I'll share my thoughts on both points. The first one is quite practical; if you're a teacher, you'll find it useful. I'll include some strategies and tricks to encourage good use of LLMs, detect misuse, and grade more accurately.

The second one... is harder to categorize and would probably require a deeper study, but I think my preliminary observations are fascinating on their own.

[...] If you're a teacher and you're thinking of using LLMs to grade assignments or exams, it's worth understanding their limitations.

We should think of a language model as a "very smart intern": fresh out of college, with plenty of knowledge, but not yet sure how to apply it in the real world to solve problems. So we must be extremely detailed in our prompts and patient in correcting its mistakes—just as we would be if we asked a real person to help us grade.

In my tests, I included the full project description, a detailed grading rubric, and several elements of my personal judgment to help it understand what I look for in an evaluation.

[...] The usual hallucinations began—the kind I thought were mostly solved in newer model versions. But apparently not: it was completely making up citations from the reports.

[...] Soon after, it started inventing its own grading criteria. I couldn't get it to follow my rubric at all. I gave up and decided to treat its feedback simply as an extra pair of eyes, to make sure I wasn't missing anything.

[...] Instead of asking the LLM to identify AI-written texts, which it doesn't do very well, I decided to compare my own quality ratings of each project with the LLM's ratings. Basically, I wanted to see how aligned our criteria were.

And I found a fascinating pattern: the AI gives artificially high scores to reports written with AI.

The models perceive LLM-written reports as more professional and of higher quality. They prioritize form over substance.

And I'm not saying that style isn't important, because it is, in the real world. But it was giving very high marks to poorly reasoned, error-filled work simply because it was elegantly written. Too elegantly... Clearly written with ChatGPT.

When I asked the model what it based its evaluation on, it said things like: "Well, the students didn't literally write [something]... I inferred it from their abstract, which was very well written."

In other words, good writing produced by one LLM leads to a good evaluation by another LLM, even if the content is wrong.

Meanwhile, good writing by a student doesn't necessarily lead to a good evaluation by an LLM.

This phenomenon has a name: corporatism.

[...] This situation gives me chills, because we have totally normalized using LLMs to filter résumés, proposals, or reports.

I don't even want to imagine how many users are accepting these evaluations without supervision and without a hint of critical thought.

If we, as humans, abdicate our responsibility as critical evaluators, we'll end up in a world dominated by AI corporatism.

A world where machines reward laziness and punish real human effort.

[...] To make sure students haven't overused ChatGPT, professors conduct short face-to-face interviews to discuss their projects.

It's the only way to ensure they've actually learned, and also, to be fair. If they've used the model to write more clearly and effectively but still achieved the learning objectives and understood their work, we don't penalize them.

In general, when a report smells a lot like ChatGPT, it usually means the students didn't learn much. But there are always surprises, in both directions.

Sometimes, it's legitimate use of ChatGPT as a writing assistant, which I actually encourage in class. Other times, I find reports that seem AI-written, but the students swear up and down they weren't, even after I tell them it won't affect their grade.

Maybe it's that humans are starting to write like machines.

Of course, machines have learned to write like humans—but current models still have a rigid, recognizable, and rather bland style. You can spot the overuse of bullet-pointed infinitives packed with adjectives, endless summary paragraphs, and phrasing or structures no human would naturally use.


Original Submission

Related Stories

David Rosenthal on the LLM Negative Feedback Loop 29 comments

Archivist David Rosenthal observes now that more material is posted online by LLMs than actual people, the bots are starting to ingest their own digital excrement, creating a negative feedback loop.

In the belief that "more is better", Large Language Models (LLMs) have insatiable appetites for training data. They started by scraping everything on the Web (robots.txt be dammed). When that ran out they downloaded the various pirate libraries (copyright be dammed). That exhausted the texts easily available in digital form, but their hunger wasn't assuaged. As for images, they partly used CAPTCHAs but mostly paid vast numbers of poor people to label the images with what they showed.

When the supply of text ran low, people observed that the LLMs were capable of generating human-like text in large quantities. The obvious idea was to pour the output of the LLMs into their training sets. This wasn't just a conscious decision, it was inevitable. The advent of LLMs rapidly polluted the Web with LLM output. Greg Druck's AI Now Writes as Many Online Articles as Humans notes that:

We observe significant growth in primarily AI-generated articles, coinciding with the launch of ChatGPT in November 2022. After only 12 months, primarily AI-generated articles accounted for 35.9% of articles published.

In Q1 2025, the quantity of primarily AI-generated articles being published on the web nearly equaled the quantity of human-written articles, 49.6% vs. 50.4%. In Q4 2025, primarily AI-generated articles surpassed human-written at 50.9%, before returning to 49.9% in Q1 2026.

Even if slop were not of undesirable quality, it is not produced by humans and thus is completely unsuitable as training data.

Previously:
(2026) A Wikipedia Clone Built on AI Hallucinations is Here to Hasten Along the Death of the Internet
(2025) When It All Comes Crashing Down: The Aftermath of the AI Boom
(2025) AI Favors Texts Written by Other AIs, Even When They're Worse Than Human Ones
(2025) What the Hell is Going on Right Now?


Original Submission

This discussion was created by hubie (1068) for logged-in users only, but now has been archived. No new comments can be posted.
Display Options Threshold/Breakthrough Mark All as Read Mark All as Unread
The Fine Print: The following comments are owned by whoever posted them. We are not responsible for them in any way.
(1)
  • (Score: 2) by krishnoid on Thursday December 11 2025, @03:10PM (1 child)

    by krishnoid (1156) on Thursday December 11 2025, @03:10PM (#1426505)

    Which one? Did ChatGPT like what ChatGPT wrote, or was it another one that was likely trained in the same way? For example, I would *LOVE* to see how an LLM trained on content in one human language, evaluates content written in another.

    • (Score: 2) by bmimatt on Thursday December 11 2025, @05:21PM

      by bmimatt (5050) on Thursday December 11 2025, @05:21PM (#1426517)

      I've seen Gemini do Claude's code review and call it 'manually implemented'.

  • (Score: 4, Interesting) by Mojibake Tengu on Thursday December 11 2025, @03:45PM (1 child)

    by Mojibake Tengu (8598) on Thursday December 11 2025, @03:45PM (#1426507) Journal

    When I see a complete and correct RISC-V/64 macroassembler written by a LLM in Haskell programming language, I'll start to believe generative AIs are actually useful for something.

    --
    The kinder you are, the easier it is for wicked people to morally coerce you.
    • (Score: 5, Interesting) by krishnoid on Thursday December 11 2025, @06:29PM

      by krishnoid (1156) on Thursday December 11 2025, @06:29PM (#1426523)

      Not complete [google.com], and I'm not sure if it's completely correct, but it's more about Haskell and RISC-V/64 than I knew about before.

      Isn't complete and correct beyond the capabilities of humans or teams, considering that there's always one more bug (or feature request) in a program?

  • (Score: 5, Informative) by ikanreed on Thursday December 11 2025, @03:57PM (9 children)

    by ikanreed (3164) on Thursday December 11 2025, @03:57PM (#1426509) Journal

    Once again, this is a problem that arises from thinking that LLMs are thinking.

    They are not. They are linguistic pattern recognition and regurgitation machines.

    And they have a huge library of training data in the form of graded essays. And the patterns within that training set give higher grades to the jargon-filled, ponderous academic writing that also makes up the LLMs post training to make it sound smarter to your average idiot user.

    It has much weaker correlations in its latents with specific rubric guidelines than it does with sounding like an ivory tower prat.

    People have got to stop treating them like they're following instructions the way a person does.

    • (Score: 3, Interesting) by VLM on Thursday December 11 2025, @05:47PM (2 children)

      by VLM (445) on Thursday December 11 2025, @05:47PM (#1426519)

      They were trained on mass media and social media, they'll have the "quality" of mass media and social media, which is quite low and dropping fast.

      Everyone's familiar with the effect where in my discipline the journalists are total morons but I am propagandized to believe they're geniuses about every other profession. Integrate across all professions. Then extend the analogy to AI that was trained on that kind of human-generated slop. Hmm.

      Some is probably a skill issue where they can't blame the prof for the LLM not being able to read his mind WRT subjective criteria so we blame the LLM. "produce your best effort, then later on grade that effort with a criteria of the best grade means it aspires to match your best effort" and unsurprisingly it rates its own output as the best. OK then.

      The best way to handle this is take a step back, there's no point wasting effort of using AI to make slop and AT to sloppily grade it, just ask the students to produce their prompt. We're at the stage of rating math papers for their typography quality by looking at the rendered PDF instead of looking at the TEX submitted to the renderer. It would make a heck of a lot more sense to grade the raw TEX than try to OCR the printed math journal paper and reverse engineer its typographical quality. I don't think the students will enjoy this; it might be easier to submit a human handwritten in a lecture hall essay as a midterm thats 5 pages long than to submit a usable 5 page long LLM prompt that squirts out a short book on some detailed topic.

      • (Score: 2) by The Vocal Minority on Friday December 12 2025, @03:34AM (1 child)

        by The Vocal Minority (2765) on Friday December 12 2025, @03:34AM (#1426558) Journal

        If you are using an LLM to generate the final version of a document directly you're doing it wrong, and this will result in mostly nonsensical slop.

        Suggest:
        1. Generate draft document, mainly for structure.
        2. Edit draft extensively so it actually makes sense.
        3. Get LLM to proof.
        4. Check proof for hallucinations.
        5. Profit.

        • (Score: 3, Insightful) by JoeMerchant on Friday December 12 2025, @12:37PM

          by JoeMerchant (3937) on Friday December 12 2025, @12:37PM (#1426589) Journal

          I would suggest:

            2.5 clear agent context to get it to actually read your edits. I find the agents will skim the documents for differences, tell you about them, then press forward with concepts it established in its context before your edits.

          --
          🌻🌻🌻🌻✌️ [google.com]
    • (Score: 2, Insightful) by Thexalon on Thursday December 11 2025, @07:49PM (4 children)

      by Thexalon (636) on Thursday December 11 2025, @07:49PM (#1426527)

      I personally think they should be rebranded from "AI" or "LLM" to what they actually are: Automatic Bullshit Generators, or ABGs.

      And I use "bullshit" in the very strict meaning of the term: These devices don't care whether what they're saying is true, they just try to fill time and space and seem like they might be plausible to anybody who isn't looking too closely at it.

      --
      "Think of how stupid the average person is. Then realize half of 'em are stupider than that." - George Carlin
      • (Score: 4, Touché) by HiThere on Thursday December 11 2025, @08:26PM (3 children)

        by HiThere (866) on Thursday December 11 2025, @08:26PM (#1426530) Journal

        LLM is accurate. It's what they are. It may not be what many people expect them to be, but it's what they are. (With a few caveats. E.g. the "guardrails" aren't really part of the language model.)

        OTOH, I'd agree that an LLM is not a full AI. It's PART of an AI. I'm quite amazed that it's as effective as it is. Saying that the LLMs don't think is wrong...or at least incomplete. They handle PART of thought. And I don't think the programs that generate pictures from text are actually exactly LLMs. They do part of the job handled by the visual cortex in humans. But we won't get an actual full AI until there's a robot that can navigate in the world, and can describe what it's encountering, and why it's doing what it's doing (perhaps not entirely honestly).

        --
        Javascript is what you use to allow unknown third parties to run software you have no idea about on your computer.
        • (Score: 3, Insightful) by bzipitidoo on Thursday December 11 2025, @11:11PM (2 children)

          by bzipitidoo (4388) on Thursday December 11 2025, @11:11PM (#1426546) Journal

          Bullcrap is right. LLMs bandy words. They have zero understanding of what they say, and so, do not qualify as the least intelligent.

          I would say LLMs aren't even part of an AI. Why? Intelligence should include elementary reasoning, and these LLMs don't even do that. For instance, LLMs can't play chess. I don't mean in the sense that they play badly (though they do), I mean that they are incapable of following the rules of chess. Sometimes they can finish a game, but all too often they submit illegal moves. Sometimes in place of a move, they cough up text that isn't a move at all. If they had any intelligence at all, they could understand that they are playing a game, and at least submit their moves to a simple move checker program to verify that it is a legal move, however bad the move might be.

          Likewise with code. When I tried them, I got code that had syntax errors. I asked it to check that the code would at least compile and it told me it wasn't allowed to do that. The code also used deprecated libraries, with new features mixed in so that if the syntax errors were fixed, it still wouldn't compile no matter what compiler version was used.

          However AI is eventually created, I guess LLMs won't play much of a role.

          • (Score: 2) by ikanreed on Friday December 12 2025, @02:05AM (1 child)

            by ikanreed (3164) on Friday December 12 2025, @02:05AM (#1426556) Journal

            There's a demarcation problem in there. What does "intelligent" mean when we discuss human beings or lab rats? It's not perfectly clear there.

            There's definitely a contingent of psychological research that just classifies pattern recognition as intelligence. The kind that really loves IQ tests. And these machines are pretty good at that.

            But there's kinds of "thinking" that they don't do. They don't think-plan-do like us. Their "thinking" such as it is is directly encoded into the do step. There's no human analogy for operating that way. It's more than a bit alien

            • (Score: 2) by bzipitidoo on Friday December 12 2025, @04:52AM

              by bzipitidoo (4388) on Friday December 12 2025, @04:52AM (#1426564) Journal

              Yes, one of the basic problems is that we lack a clear definition of intelligence. Chess was thought a great way to measure intelligence. There are many smart people who can't play chess well, but people who can play well all tend to be smart. Some seriously hoped that if a computer could be made to play chess well, we'd have Artificial General Intelligence. Well, now chess engines have advanced far beyond humans. And all this really showed is that chess is amenable to brute force calculation.

              Current computers are or can be great players of the sort of games where the actions are discrete. Most board games, for instance. But not sports. A kind of game they are completely unable to handle is the Role Playing Game. They can handle the mechanics of the typical RPG brilliantly, but not the role playing part. Designers of computer RPGs use embarrassingly crude heuristics to mostly fake it, ideas such as keeping "faction" scores. If your faction with a particular group of Non Player Characters is high, they will interact with you in very limited but friendly ways. Many NPCs are simply merchants or guards. A few are programmed with "quests", which involves set verbiage, actions, and perhaps items to give out under specific mechanical circumstances. If your faction is low, the actions the NPCs take can be even more limited: attack you on sight, and fight to the death. The MMO part is another means of bringing the game closer to real role playing, but that too is lacking. While that brings in lots of real people, the game engines simply do not admit to the variety and creativity of real role playing. And no, an LLM can't handle the Dungeon Master job.

    • (Score: 2) by JoeMerchant on Thursday December 11 2025, @07:57PM

      by JoeMerchant (3937) on Thursday December 11 2025, @07:57PM (#1426528) Journal

      >It has much weaker correlations in its latents with specific rubric guidelines than it does with sounding like an ivory tower prat.

      This is down to the training set more than anything.

      I have been having my LLMs write their own requirements documents, because if I write it and ask it for criticism it rips all over it, but if I describe to it what I want written, then let it write it, then ask it for criticism, it still rips all over it, but not as much for nit-picky stuff that arises out of differences in our training sets. I criticize its writing for content errors until we get an agreeable requirements document. Then I turn it loose to implement the requirements "it wrote" and watch it make more mistakes interpreting its own text. Just... like... human.... developers... but faster.

      --
      🌻🌻🌻🌻✌️ [google.com]
  • (Score: 2, Interesting) by Anonymous Coward on Thursday December 11 2025, @05:30PM (5 children)

    by Anonymous Coward on Thursday December 11 2025, @05:30PM (#1426518)

    Require the students to learn about footnotes along with primary source references, and require their use.

    • (Score: 3, Interesting) by VLM on Thursday December 11 2025, @06:05PM (4 children)

      by VLM (445) on Thursday December 11 2025, @06:05PM (#1426522)

      Thought experiment: Toss out the essay and grade the student based on their bibliography section.

      Hypothetically the assignment is to write an essay about assembly languages for teaching purposes, and your bibliography doesn't include at least some MIX/MMIX/Knuth references, that's a big WTF. Don't even have to read the essay to know there's going to be a Knuth sized hole in it. If they know enough to include Knuth's work in the bibliography, you don't have to read the paper...

      There's a bad attitude among students that the only purpose of a bibliography is to show off that you know MLA vs APA format and you know which one this academic topic uses, which is pretty lame. It can be a useful tool if used properly.

      • (Score: 3, Funny) by JoeMerchant on Thursday December 11 2025, @08:00PM

        by JoeMerchant (3937) on Thursday December 11 2025, @08:00PM (#1426529) Journal

        >MLA vs APA format

        Sounds like a job for an LLM to me...

        --
        🌻🌻🌻🌻✌️ [google.com]
      • (Score: 2) by HiThere on Thursday December 11 2025, @08:35PM (1 child)

        by HiThere (866) on Thursday December 11 2025, @08:35PM (#1426531) Journal

        I'm not sure that MIX is a good thing to put in an essay on assembly language for teaching purposes. I think Apple ][ (i.e. I6502 with special extensions for bank switching and graphics) would be a better choice. MIX was specialized for demonstrating algorithm timing.

        Are there any decent Apple ][ emulators around? If not perhaps something built around the Z80. A simple enough assembler with actual (if archaic) use cases. (But this: https://www.scullinsteel.com/apple2/ [scullinsteel.com] claims to be an Apple ][ emulator written in Javascript.)

        --
        Javascript is what you use to allow unknown third parties to run software you have no idea about on your computer.
        • (Score: 2) by VLM on Friday December 12 2025, @12:37PM

          by VLM (445) on Friday December 12 2025, @12:37PM (#1426590)

          Well, yes it would depend on the class. If the class were RISC-V or ARM you would learn a lot about the student's judgment of what is or is not important by looking at the bibliography.

          From memory there's a manual set for RISC-V where volume 1 is many simple user mode instructions and volume 2 is a few complicated kernel mode instructions so given an interrupt programming project it would be interesting to see which the student cites. Some of the interrupt stuff in volume 1 is interesting as per how it impacts user mode processing.

      • (Score: 2, Interesting) by pTamok on Thursday December 11 2025, @09:51PM

        by pTamok (3042) on Thursday December 11 2025, @09:51PM (#1426542)

        I've heard some academics set the task to students to generate the text of their assignments using an LLM*, then be required to criticise it and point out factual errors, biases, and 'hallucinated' references.

        Some get enlightened by that process.

        No doubt, some try to get an LLM to criticise either its own output, or the output of another LLM.

        *Perhaps part of the assignment is to provide the prompt used.

  • (Score: 4, Interesting) by VLM on Thursday December 11 2025, @05:57PM

    by VLM (445) on Thursday December 11 2025, @05:57PM (#1426520)

    phrasing or structures no human would naturally use

    I've noticed AI will generate awkward unspeakable text that genuine humans would replace with idioms and casual phrases. The kind of thing that would fail a public speaking class for being too weird and inhuman and lacking a consistent voice or personality.

    I had to make a short speech earlier this week in front of a bunch of people about a thing I was leading its not relevant to this discussion. I ended up rolling with an outline and some quotes on a piece of paper and it turned out fine, but I experimented with LLM generation of a speech and holy shit it sounded inhuman. It's actually pretty good at "provide me with a list of five quotes and five anecdotal very short intro-type stories about XYZ" and I pick and choose and research each selecting one or two, and that turned out OK.

    It's like that if you ask for AI generated art. Technically, it'll "meet expectations" at a minimal surly instruction following level, but its never what I was hoping for and always has a bunch of noise I didn't want or expect.

    I expect video generation will be like this. "Give me a tightly edited space opera" will never give you the original 1970s Star Wars first movie, but I bet it could squirt out boring forgettable formulaic predictable sequels like the latter hypercorporate low risk star wars movies have been.

(1)