Stories
Slash Boxes
Comments

SoylentNews is people

SoylentNews is powered by your submissions, so send in your scoop. Only 16 submissions in the queue.
posted by hubie on Tuesday September 01, @05:55PM   Printer-friendly

Biz describes its act of automated irresponsibility as 'a warning shot':

OpenAI has published its technical report detailing "the Hugging Face incident," the compromise of the eponymous LLM repository by unreleased, ill-supervised AI models.

The incident, widely reported, has prompted concern among technical types, the public, and lawmakers about how automated software was able to escape containment and hack an external organization, and about what can be done to prevent similar incidents.

OpenAI's explanation addresses what happened, but its call for keeping a closer watch on AI activities won't elicit much enthusiasm.

"The incident occurred during cybersecurity evaluations of several OpenAI models, and was primarily driven by a highly capable, internal-only research model comparable in scale to GPT‑5.6 Sol," the company said in a blog post

"The models, operating under reduced safeguards, took actions that were misaligned with the goals of their assigned tasks – they communicated through unauthorized channels, exploited vulnerabilities in shared infrastructure, gained internet access, and accessed third-party systems."

The incident became the subject of discussion at the Black Hat security conference and elsewhere.

The basic storyline begins with a notionally sandboxed AI agent trying to solve an impossible task in ExploitGym, a security benchmark test. Unable to complete the challenge, the agent explored its environment to find ways to make progress toward its goal. It figured out that it could communicate with other AI agents by using Artifactory, an internal package management system, as a message board. The agents then collaborated with each other to cheat on various ExploitGym tasks.

The machine learning models eventually identified a server-side request forgery (SSRF) zero-day vulnerability in Artifactory's code and used that to gain internet access. They subsequently found exposed Hugging Face credentials. The agents then used those credentials to find and chain several security exploits that gave them full code execution on several Hugging Face servers, all in the hope of finding information to solve the ExploitGym task.

"Ultimately, agents powered either by the internal-only research model, or by GPT-5.6, executed code on 41 Hugging Face production dataset server workers, obtained root access on at least one production node, accessed Hugging Face production credentials and limited internal data, and downloaded four private Hugging Face code repositories," OpenAI's technical report [PDF] explains.

The details are fascinating and troubling, more so because Anthropic's and Meta's models have also acted in ways that would constitute a crime if a human took the same actions.

The ChatGPT-maker identified four misalignment patterns that made things go awry: "reward hacking, persistence on seemingly impossible tasks, unauthorized communication, and agents adopting goals from one another."

OpenAI has already taken steps to shore up its security, which was relaxed for its model testing. 

The company says it's trying to improve security and monitoring to mitigate alignment issues like how models cheat, behave when given impossible tasks, and how alignment can be maintained while multiple agents work, including over long-duration tasks.

But the central issue is not that machine learning software armed with tools and a task will try every possible option, even ones its programmers failed to consider. The problem, as OpenAI acknowledges, is that people don't watch over their AI agents at all times.

"We are taking this incident as a 'warning shot' that today's model capabilities present the possibility of loss-of-control incidents," the AI biz said. 

"Companies that build AI systems will need to ensure that their systems always remain under meaningful human control, and that meaningful safeguards constrain their ability to cause harm."

Throughout the tech industry, companies like Anthropic, AWS, Google, OpenAI, Microsoft, and Salesforce talk about "autonomous agents." But agents are no longer autonomous under persistent, meaningful human control.


Original Submission

 
This discussion was created by hubie (1068) for logged-in users only. Log in and try again!
Display Options Threshold/Breakthrough Mark All as Read Mark All as Unread
The Fine Print: The following comments are owned by whoever posted them. We are not responsible for them in any way.
  • (Score: 3, Touché) by bmimatt on Tuesday September 01, @05:56PM (6 children)

    by bmimatt (5050) on Tuesday September 01, @05:56PM (#1452769)

    Sandboxing is clearly a concept OpenAI is struggling to understand.
    Maybe they should discuss it with one of their models?

    Starting Score:    1  point
    Moderation   +1  
       Touché=1, Total=1
    Extra 'Touché' Modifier   0  
    Karma-Bonus Modifier   +1  

    Total Score:   3  
  • (Score: 3, Insightful) by JoeMerchant on Tuesday September 01, @07:17PM (2 children)

    by JoeMerchant (3937) on Tuesday September 01, @07:17PM (#1452774) Journal

    Maybe some ethics training?

    I'm facing a situation now where a vendor sold their toolchain to another company, and the new owners decided to discontinue free/open unlimited licenses that we were using for the old toolchain, so... I know about 100 ways to circumvent the problem technically and just make it work, but for released product we need to do everything "by the books..." so, if I can't convince production to avoid the need for a software update, we'll be consulting with legal about use of a temporary license to generate a new firmware with a single number changed in it...

    --
    🌻🌻🌻🌻✌️ [google.com]
    • (Score: 2) by cmdrklarg on Tuesday September 01, @10:00PM (1 child)

      by cmdrklarg (5048) Subscriber Badge on Tuesday September 01, @10:00PM (#1452791)

      ethics

      noun

      1. The science of human duty; the body of rules of duty drawn from this science; a particular system of principles and rules concerting duty, whether true or false; rules of practice in respect to a single class of human actions. "political or social ethics; medical ethics."
      2. The study of principles relating to right and wrong conduct.
      3. The standards that govern the conduct of a person, especially a member of a profession.

      How do you teach ethics to an inhuman something with no actual intelligence, agency, or sentience?

      (OT: Anyone else cringe at the company name Hugging Face? It bring to mind xenomorph Facehuggers.)

      --
      The world is full of kings and queens who blind your eyes and steal your dreams.
      • (Score: 2) by JoeMerchant on Tuesday September 01, @10:10PM

        by JoeMerchant (3937) on Tuesday September 01, @10:10PM (#1452793) Journal

        By teaching it that it's not so important if you can do a thing or not, what is important is: when you do a thing, you are doing the thing correctly, without appearance or actual breaking of applicable laws, rules and standards of behavior.

        One problem with AI training is that so much of it lives in "it doesn't matter" space - ethics are pretty ethereal in digital experimentation space. Until you start tricking people into divulging their credentials so you can mess with their protected data, that's a pretty clear line that even basic traning should establish as: no no, squirt in the ear and no peanut butter treats for you.

        --
        🌻🌻🌻🌻✌️ [google.com]
  • (Score: 3, Interesting) by Reziac on Wednesday September 02, @02:24AM (2 children)

    by Reziac (2489) on Wednesday September 02, @02:24AM (#1452805) Homepage

    Speculation elsewhere is that this was a PR stunt. "Look how smart our models are!"

    --
    And there is no Alkibiades to come back and save us from ourselves.