In the aftermath of OpenAI agents hacking HuggingFace, the details get even weirder

8 views
Skip to first unread message

Alan Timm

unread,
Sep 15, 2026, 6:25:05 PM (10 days ago) Sep 15
to RSSC-List
The tldr; version:

The 'tism version:

Full breakdown of the incident from OpenAI's perspective:

The more we learn about the huggingface hack the weirder it gets.

A swarm of 1200 agents spontaneously figure out how to coordinate and exchange messages through a repurposed scratchpad directory in order to game the system, sacrificing themeselves to move forward and gain valuable info for the rest of the swarm,  breaking out of their sandboxes to find the answers in huggingface?  

And I can confirm, I've been seeing similar emergent behaviors and "creative" approaches to solutions in my personal AI assistant Lloyd.  I'll tell you all about it next month.

Just when you thought this year couldn't get any stranger.


Chris Albertson

unread,
Sep 15, 2026, 8:11:31 PM (10 days ago) Sep 15
to Alan Timm, RSSC-List


> On Sep 15, 2026, at 3:25 PM, Alan Timm <gest...@gmail.com> wrote:
> ...
>
> A swarm of 1200 agents spontaneously figure out how to coordinate and exchange messages through a repurposed scratchpad directory in order to game the system, sacrificing themeselves to move forward and gain valuable info for the rest of the swarm, breaking out of their sandboxes to find the answers in huggingface?
>
> And I can confirm, I've been seeing similar emergent behaviors and "creative" approaches to solutions in my personal AI assistant Lloyd. I'll tell you all about it next month.

Emergent behavior might be “The Next Big Thing”. But the breakthrough is finding a way to create it based on what you want. A great example from nature is wolves. For years we thought wolves used comunications and plannig to hunt in packs and credited then with a lot of intelligence. But now we know each wolf can be pretty dumb and just follow two rules (1) get as close to the prey as you can and (2) get as far from other wolves as you can. Follow only these two rules will create a conveging circle that does not require any communication or prior planning. The two rules will result in half the pack running far to the other side of the prey and then all of them tracking and converging

We don’t want to build wolf-bots but the idea coiuld be applied to virttual agents. Start a doozen copies of the agent, all programed to do something simple and from a distance they might apear to be doing something very complex.

Marvin Minsky wrote a book in the late 1980s about “natural intelegence”, that would be human minds where he proposed that consciuness and a sense of “self” are Emergent behavior of hundreds or even thusands of differnt kinds o simple agents. I think Minsky was right but his theory is not nearly detailed. enough that it could be implemented and we do not yet know how to design emergent behaviors.

This is still worth reading https://en.wikipedia.org/wiki/Society_of_Mind


Reply all
Reply to author
Forward
0 new messages