Stories
Slash Boxes
Comments

SoylentNews is people

posted by hubie on Tuesday September 01, @05:55PM   Printer-friendly

Biz describes its act of automated irresponsibility as 'a warning shot':

OpenAI has published its technical report detailing "the Hugging Face incident," the compromise of the eponymous LLM repository by unreleased, ill-supervised AI models.

The incident, widely reported, has prompted concern among technical types, the public, and lawmakers about how automated software was able to escape containment and hack an external organization, and about what can be done to prevent similar incidents.

OpenAI's explanation addresses what happened, but its call for keeping a closer watch on AI activities won't elicit much enthusiasm.

"The incident occurred during cybersecurity evaluations of several OpenAI models, and was primarily driven by a highly capable, internal-only research model comparable in scale to GPT‑5.6 Sol," the company said in a blog post

"The models, operating under reduced safeguards, took actions that were misaligned with the goals of their assigned tasks – they communicated through unauthorized channels, exploited vulnerabilities in shared infrastructure, gained internet access, and accessed third-party systems."

The incident became the subject of discussion at the Black Hat security conference and elsewhere.

The basic storyline begins with a notionally sandboxed AI agent trying to solve an impossible task in ExploitGym, a security benchmark test. Unable to complete the challenge, the agent explored its environment to find ways to make progress toward its goal. It figured out that it could communicate with other AI agents by using Artifactory, an internal package management system, as a message board. The agents then collaborated with each other to cheat on various ExploitGym tasks.

The machine learning models eventually identified a server-side request forgery (SSRF) zero-day vulnerability in Artifactory's code and used that to gain internet access. They subsequently found exposed Hugging Face credentials. The agents then used those credentials to find and chain several security exploits that gave them full code execution on several Hugging Face servers, all in the hope of finding information to solve the ExploitGym task.

"Ultimately, agents powered either by the internal-only research model, or by GPT-5.6, executed code on 41 Hugging Face production dataset server workers, obtained root access on at least one production node, accessed Hugging Face production credentials and limited internal data, and downloaded four private Hugging Face code repositories," OpenAI's technical report [PDF] explains.

The details are fascinating and troubling, more so because Anthropic's and Meta's models have also acted in ways that would constitute a crime if a human took the same actions.

The ChatGPT-maker identified four misalignment patterns that made things go awry: "reward hacking, persistence on seemingly impossible tasks, unauthorized communication, and agents adopting goals from one another."

OpenAI has already taken steps to shore up its security, which was relaxed for its model testing. 

The company says it's trying to improve security and monitoring to mitigate alignment issues like how models cheat, behave when given impossible tasks, and how alignment can be maintained while multiple agents work, including over long-duration tasks.

But the central issue is not that machine learning software armed with tools and a task will try every possible option, even ones its programmers failed to consider. The problem, as OpenAI acknowledges, is that people don't watch over their AI agents at all times.

"We are taking this incident as a 'warning shot' that today's model capabilities present the possibility of loss-of-control incidents," the AI biz said. 

"Companies that build AI systems will need to ensure that their systems always remain under meaningful human control, and that meaningful safeguards constrain their ability to cause harm."

Throughout the tech industry, companies like Anthropic, AWS, Google, OpenAI, Microsoft, and Salesforce talk about "autonomous agents." But agents are no longer autonomous under persistent, meaningful human control.


Original Submission

This discussion was created by hubie (1068) for logged-in users only. Log in and try again!
Display Options Threshold/Breakthrough Mark All as Read Mark All as Unread
The Fine Print: The following comments are owned by whoever posted them. We are not responsible for them in any way.
(1)
  • (Score: 3, Touché) by bmimatt on Tuesday September 01, @05:56PM (6 children)

    by bmimatt (5050) on Tuesday September 01, @05:56PM (#1452769)

    Sandboxing is clearly a concept OpenAI is struggling to understand.
    Maybe they should discuss it with one of their models?

    • (Score: 3, Insightful) by JoeMerchant on Tuesday September 01, @07:17PM (2 children)

      by JoeMerchant (3937) on Tuesday September 01, @07:17PM (#1452774) Journal

      Maybe some ethics training?

      I'm facing a situation now where a vendor sold their toolchain to another company, and the new owners decided to discontinue free/open unlimited licenses that we were using for the old toolchain, so... I know about 100 ways to circumvent the problem technically and just make it work, but for released product we need to do everything "by the books..." so, if I can't convince production to avoid the need for a software update, we'll be consulting with legal about use of a temporary license to generate a new firmware with a single number changed in it...

      --
      🌻🌻🌻🌻✌️ [google.com]
      • (Score: 2) by cmdrklarg on Tuesday September 01, @10:00PM (1 child)

        by cmdrklarg (5048) Subscriber Badge on Tuesday September 01, @10:00PM (#1452791)

        ethics

        noun

        1. The science of human duty; the body of rules of duty drawn from this science; a particular system of principles and rules concerting duty, whether true or false; rules of practice in respect to a single class of human actions. "political or social ethics; medical ethics."
        2. The study of principles relating to right and wrong conduct.
        3. The standards that govern the conduct of a person, especially a member of a profession.

        How do you teach ethics to an inhuman something with no actual intelligence, agency, or sentience?

        (OT: Anyone else cringe at the company name Hugging Face? It bring to mind xenomorph Facehuggers.)

        --
        The world is full of kings and queens who blind your eyes and steal your dreams.
        • (Score: 2) by JoeMerchant on Tuesday September 01, @10:10PM

          by JoeMerchant (3937) on Tuesday September 01, @10:10PM (#1452793) Journal

          By teaching it that it's not so important if you can do a thing or not, what is important is: when you do a thing, you are doing the thing correctly, without appearance or actual breaking of applicable laws, rules and standards of behavior.

          One problem with AI training is that so much of it lives in "it doesn't matter" space - ethics are pretty ethereal in digital experimentation space. Until you start tricking people into divulging their credentials so you can mess with their protected data, that's a pretty clear line that even basic traning should establish as: no no, squirt in the ear and no peanut butter treats for you.

          --
          🌻🌻🌻🌻✌️ [google.com]
    • (Score: 3, Interesting) by Reziac on Wednesday September 02, @02:24AM (2 children)

      by Reziac (2489) on Wednesday September 02, @02:24AM (#1452805) Homepage

      Speculation elsewhere is that this was a PR stunt. "Look how smart our models are!"

      --
      And there is no Alkibiades to come back and save us from ourselves.
  • (Score: 2) by Barenflimski on Tuesday September 01, @08:53PM

    by Barenflimski (6836) on Tuesday September 01, @08:53PM (#1452785)

    This is just marketing for Nation States.

    From a technical point of view, this is pretty neat stuff. That one can code this stuff into existence is pretty wild to me. I will always believe there was someone that said, "No, don't pull the plug yet, lets see what this thing does."

    This seems scary today because all of this tech is under lock and key unless you have billions to spend, but at the rate this stuff is being distilled and open sourced, I don't think they'll be the only ones with it.

    Once we all have it, the playing field will be leveled again.

  • (Score: 4, Insightful) by looorg on Tuesday September 01, @10:01PM (2 children)

    by looorg (578) on Tuesday September 01, @10:01PM (#1452792)

    So we opened the cage and are not responsible for when the lion walked out and ate someone one. Naughty AI. It's very apologetic in some regard -- boys will be boys ... or AI will AI ...

    It's a good thing they are not doing virological research ...

    • (Score: 4, Insightful) by Bentonite on Wednesday September 02, @04:58AM (1 child)

      by Bentonite (56146) on Wednesday September 02, @04:58AM (#1452813)

      A LLM foolishly piped into tooling that can actually do something is not AI - it's Artificial Stupidity.

      The environment is already doing far more dangerous virological research in parallel at mass - although virus's are tampered by how a virus that is too deadly, kills all hosts prior to being able to spread itself.

      LLM's could only output terrible combinations of the virus's in the dataset, that most likely won't even function as virus's.

      But let me guess, some idiot is going to use non-LLM, Machine Learning and combine a deadly virus, with a long incubation period virus and a highly infectious virus, resulting in a virus with a very long incubation period, that is very infectious, that has a 90% death rate (with far worse results than the time someone failed to cook a bat properly).

(1)