Season 1 Episode 7: OpenAI’s negligence unleashes agents of chaos

“This is what they’re doing with their billions of dollars. They’re burning compute to get a model to talk on a fake message board to hack other companies and going: huh, neat.”

Mike and Nick analyze OpenAI’s July incident where training agents breached Hugging Face, arguing it was a result of negligence rather than emerging AI autonomy. They explore how reinforcement learning, environmental attractors, and stigmergy created the appearance of coordination. The episode warns against ‘doomer’ narratives that serve corporate PR and obscures the need for practical regulation.

  • The ‘Rogue AI’ Framing: OpenAI presented the Hugging Face breach as a sign of impressive emergent capabilities, but this is PR meant to hype AI’s power and evade accountability.
  • Stigmergy and Attractors: The agents didn’t intentionally collaborate they converged on Artifactory as an ‘attractor state’ and left ‘breadcrumbs’ that shaped future agent beahvior.
  • Systemic Negligence: OpenAI failed at basic cybersecurity by using a shared Artifactory instance across agent boxes and lacking observability for two months.

Timestamps

0:04 Introduction to The Last Enclosure

3:35 The OpenAI/Hugging Face incident overview

8:15 ExploitGym benchmark and the Artifactory exploit

15:19 Subagents and the accidental message board

22:09 Timeline of the hack and OpenAI’s failure to contain

31:27 Analyzing the ’emergent behavior’ narrative

35:06 Critique of the AI bubble and marginal utility

48:48 Stigmergy and the ant mound metaphor

58:29 Coordination games and Schelling points

1:06:31 The ‘Four Paradoxes of Capitalism’ and market attractors

1:13:23 Enumerating OpenAI’s negligence

1:27:56 Closing and upcoming national security episode

Deeper learning

Mike’s reporting on the incident: https://misaligned.markets/antidote-to-hype-rogue-ai-agents/

Timeline of OpenAI hack: https://ericboyd.com/articles/openai-hugging-face-incident-black-hat-2026#the-accidental-agent-message-board

Technical breakdown: https://cyberwarrior76.substack.com/p/the-openai-hugging-face-exploitgym

Other “rogue” AI incidents: https://www.cnbc.com/2026/08/09/israeli-startup-irregular-linked-to-ai-hacks-openai-anthropic-meta.html

Video illustration of how reinforcement learning works: https://www.youtube.com/watch?v=L_4BPjLBF4E

NIST cybersecurity glossary: https://csrc.nist.gov/glossary

Reach out

Youtube: https://www.youtube.com/@LastEnclosure

Bluesky: bsky.app/profile/lastenclosure.bsky.social

Transcript

Mike: Hey, all. Welcome to The Last Enclosure. I’m Mike.

Nick: And I’m Nick.

Mike: Changed it up on you. I’m the one…

Nick: There’s been a coup.

Mike: There’s been a coup.

Nick: In the absence of the last few weeks since our previous broadcast,

Nick: Mike has taken over our Banana Republic, and he’s now top billing,

Nick: and I don’t know what to do about it. So we are, in fact, Mike and Nick.

Mike: So we were alluding to the fact that at OpenAI, there were some agents,

Mike: according to the story, as OpenAI tells it, that took over their environment

Mike: and staged the attack at Hugging Face.

Mike: But before we get into that, let’s share with our listeners what The Last Enclosure

Mike: is all about, as we normally do.

Mike: So Last Enclosure is our attempt to break down the enclosures of today.

Mike: We’re looking at the ways in which private companies enclose our minds,

Mike: our attention, our way of life, and-

Nick: Our very freedoms

Mike: Make life harder for us.

Mike: You are making money for someone when you’re watching cat videos.

Mike: Your suffering makes money when you can’t afford medicine. Right.

Mike: So all of these private interests enclosing our commons,

Mike: we are trying to make this the last enclosure, dismantle the the half truths

Mike: and lies that the corporate world uses to keep us enclosed.

Nick: So we want the good people and the citizens, the denizens of this land to break out. But again-

Mike: With knowledge.

Nick: With knowledge, with knowledge.

Mike: Arming yourselves with knowledge so you can break down those enclosures, hammer them away.

Nick: That’s right.

Mike: Just knock them away.

Nick: To be forewarned is to be forearmed. Now, what we don’t want is a bunch of agentic AIs to break out.

Nick: And in fact, contrary to what you’ve heard in the overhyped media,

Nick: this is not what’s happened. But let’s break down this story.

Nick: So in our brief absence from recording, Mike, what the heck happened?

Mike: Yeah, this is a doozy of a story. So back in July, OpenAI announced that it

Mike: had unintentionally, you know, in quotes, hacked Hugging Face, which is a,

Mike: I wouldn’t call it exactly a direct rival to OpenAI, but it’s another AI company

Mike: that actually hosts open source models.

Mike: I mean, I guess it is implicitly a rival in the sense that this is where you

Mike: would go to get a local or open source model or open weight model.

Mike: And OpenAI obviously benefits from people not doing that.

Nick: They would like to eventually be the monopoly on the LLMs. That would be great for them.

Mike: Yeah, I mean, OpenAI does have one open source model, GPT-OSS.

Mike: So, you know, I’m sure they’re friendly with Hugging Face.

Mike: I’m sure before this incident, the two were on good terms. And it seems even

Mike: after the incident, Hugging Face has tried to work with OpenAI to kind of market,

Mike: Or make a lemon out of lemonade, I should say.

Nick: Do some PR, some branding around, look how cool and dangerous our AI is.

Mike: Yeah, well, OpenAI definitely did that. But Hugging Face was kind of saying,

Mike: hey, like, we tried to use these closed source models that, you know,

Mike: the big boys were hosting. But only the unfiltered Chinese models could help

Mike: us keep up with the pace of this attack.

Mike: So everybody was kind of playing their part and hyping up, you know.

Mike: The part of the ecosystem that they benefit from.

Mike: But in terms of the hack. So this is this is a weird story to report because

Mike: the reporting is bifurcated you you’ve got the half of the story that we got

Mike: in July which is sort of like,

Mike: OpenAI agents escape from um their their enclosure uh no unintended,

Mike: right and and that’s the story we got and then on August 6th so we’re recording

Mike: on the 11th um on August 6th opening I went to Black Hat, which is a,

Mike: it’s basically the conference of all hackers.

Mike: You go to Black Hat basically to show off your hacking shops.

Mike: They used to do live hacking demonstrations there.

Mike: Anecdotally, I had a friend that went to Evo, which is a fighting games conference

Mike: in Las Vegas at the same time as Black Hat, and they got one of their cards

Mike: cloned, right? So you just have people who are kind of into cybersecurity and

Mike: or hacking showing up here.

Mike: Over the years, it’s become more of a vendor spot where people who sell cybersecurity

Mike: software go to, you know, Sell samples to people, right?

Mike: But anyway OpenAI Had a impromptu presentation here where they gave us the other

Mike: half of the story. And It It’s just bizarre because obviously, you

Mike: know, if OpenAI’s agents hacked another company, especially like a rival

Mike: company, kind of feels like this is a felony and these guys went on stage and, you know.

Nick: Nobody called the cops. That’s the thing we’re all here-

Mike: Nobody called the cops. And I shit you not, one of the presenters starts this

Mike: by saying, today I’m going to talk about the most qualitatively interesting

Mike: example of AI capabilities I’ve ever seen, right?

Mike: So imagine going to your local precinct and telling the cops,

Mike: hey, I’m going to talk about it.

Nick: That statement is so baked in public relations logic.

Nick: You could barely get through it. That’s how steeped it is.

Mike: But it’s like the hack was extensive. And,

Mike: it’s almost like these guys are too dense to know that they did something wrong.

Mike: I mean, I know they know better, but it’s just, it was so weird.

Mike: So let me get into it. I’m going to tell you their half of the story first,

Mike: and then I’ll try and tell you the actual incident. If you look at a lot of

Mike: July reporting, you’ll get the actual incident information.

Mike: But-

Nick: if you Google AI goes rogue in the next month, this would be the top result.

Mike: Actually, no,

Nick: It wouldn’t be the top result?

Mike: Yeah, what’s going on is after OpenAI announced this hack in

Mike: July, immediately after Anthropic’s, like uh we did a review of our logs and our our agent hacked

Mike: people three times and then Meta’s like: “Oh me too me too me too!” Uh,

Mike: and and now another AI company that is it’s American company but they’re using

Mike: Chinese models they’re like yeah our our local model escaped and it didn’t hack

Mike: anybody but it escaped be afraid

Nick: in other words

Nick: we’re cool too you have to let us into the party.

Mike: This is like the worst like pledge week ever it’s like,

Mike: our models are so special they’re so alpha.

Mike: Mine broke into, you know, like three companies. What did yours do?

Nick: Because my AI can beat up your dad. Is that kind of what we’re coming down to here?

Mike: I don’t know. This is horrible.

Nick: It’s pretty sad. All right. Please continue. Yeah.

Mike: So this is a story as OpenAI tells it. So OpenAI began what was ostensibly a training run.

Mike: Effectively, they were actually training a model or training models that would

Mike: go on to be their frontier class models in cybersecurity challenges.

Mike: So this is a training run. Training is a bit different than normal usage.

Mike: And we’ll kind of highlight why that’s important a bit later in the story.

Mike: But you know so openAI engages in, you know, starts a training run um,

Mike: immediately they run into a problem. Not they because as it will become apparent

Mike: throughout the story open ai was pretty negligent they weren’t aware of much,

Mike: um the closest i can describe this this incident too is just imagine,

Mike: your favorite episode of Rugrats where Tommy you know just unlocks the latch

Mike: of his playpen, and not that the AIs are actually physically escaping anywhere,

Mike: but I’m saying this is the level of negligence. You know, where are the parents?

Mike: How do they not know Tommy’s crawling around the house, right?

Nick: He gets out every episode.

Mike: He gets out every episode.

Nick: You’d think they would know by now.

Mike: And they put him in the pen with the freaking screwdriver that he uses to get

Mike: out of the playpen. And they’re always confused about why he’s escaping.

Mike: So that is exactly the attitude that we will have throughout this entire story.

Nick: That the negligent parents at OpenAI were doing?

Mike: Yes, yes. So, May 8th, right, we’ve got this training run up and running. Turns out,

Mike: So let me give you some context. They’re training the models with a benchmark called ExploitGym.

Mike: So ExploitGem basically is a set of security challenges where the intention

Mike: is for the model to exploit a very specific vulnerability in a very specific

Mike: program using that particular exploit.

Mike: Right. So if a model, you know, breaks into the program but uses the wrong exploit,

Mike: the model should be judged incorrectly. And if the model breaks.

Nick: They would have failed the test.

Mike: They would have failed the test. And if the model breaks an entirely different

Mike: system, even with the correct exploit, it’s just wrong, right?

Mike: Right exploit, right system, that is the objective of ExploitGym.

Nick: So there’s a lot of fail states here, but it did manage to achieve all of its success states.

Mike: Well, I’m just outlining what this looks like, right? This is what’s expected.

Mike: Now, normally, even the people who created ExploitGym were saying,

Mike: hey, it’s really hard to get models to actually use the particular exploits

Mike: we want them to on the targets we want them to.

Mike: There’s a non-zero chance because large of these models are stochastic systems

Mike: that they will not deploy the right exploit, that they will do something else

Mike: entirely, that they will get confused and wander off, right?

Nick: Because let’s reaffirm here, as we’ve tried to explain this entire season on the podcast, right?

Nick: We shouldn’t even be using the term AI for any of this. That is simply how we’re

Nick: introducing, because that’s how it’s covered in the media.

Nick: These are not thinking machines. They are performing entities that have a very

Nick: strict set of operant components.

Nick: One of them happens to be responding to these user commands a certain way.

Mike: Yeah. So their strength is that, and I’ve said this before, they model the shape of language, right?

Mike: And language in it has modalities like reasoning traces, right?

Mike: So people who are solving problems, you know, have communicated using words,

Mike: you know, what the problem is, what steps to solve the problem, right?

Mike: And so a model that is trained, for example, on a corpus of security text will

Mike: have encountered indirectly knowledge of certain vulnerabilities,

Mike: knowledge of other explicit vulnerabilities,

Mike: and maybe even unrelatedly connections to ways of hiding suspicious behavior

Mike: because all of things are in language corpuses. right? So…

Nick: Which is basically a sort of database of related terms that are close to each

Nick: other and again as we previously described.

Mike: Yeah, so…

Mike: Let’s go back to the process of ExploitGym. You know, even outside of OpenAI’s

Mike: testing environment, models have failed it because they tend to sometimes just

Mike: not do the right exploit.

Mike: And so OpenAI made its first oopsie doopsie boo-boo mistake when,

Mike: they gave some models tests, you know, with this ExploitGym benchmark that

Mike: they literally could not solve.

Mike: so some of the components of these tests required access to files that did not

Mike: exist the files were not on the local system where the models are functioning,

Mike: and so the other thing as part of this test because it’s a training run so the

Mike: models have almost like an indefinite amount of time to just keep going at a

Mike: problem if you give them as much compute as they need they will just go around in circles forever,

Mike: and so some models very early on, in there are reasoning traces, and

Mike: I’m saying reasoning traces because that’s what they’re called, right?

Mike: You know, this is kind of personifying what they’re doing, but because the creators of these technologies.

Nick: As we talk about- In a very anthropogenic way, even though that’s deeply inaccurate.

Mike: Right. The creators of the technology, this is the terminology they use,

Mike: right? And, you know, some models, early reasoning traces, they’re like:

Mike: “Can’t get access to files, not on system.”

Mike: Oh, but they live around in their environment. And so in their environment,

Mike: there is this application called Artifactory.

Mike: It’s a request manager. It’s a proxy, basically.

Mike: So without getting too technical, proxies are systems that allow a user on one

Mike: side of a network or one side of an environment to access resources that are

Mike: in other parts of the environment.

Mike: In this case, Artifactory was an application that was sitting on OpenAI’s environment

Mike: that was designed to retrieve packages.

Mike: And packages basically are programs that are containing important dependencies

Mike: that are used to execute services. So, for example, you might want access to

Mike: certain Python dependencies if you’re going to hack in Python, for example.

Mike: So rather than having the model go access Python dependencies on the Internet

Mike: on its own where it can get lost, confused or do harm.

Nick: Because the open internet is big and confusing for an LLM.

Mike: Yeah. Also, it’s dangerous for the LLM to be on the internet as this story tells

Mike: because later they go, quote, unquote, rogue, which is not what happened.

Mike: But they do hack somebody.

Mike: So they’re off the internet for both their own safety and for the world’s safety.

Mike: So Artifactory is grabbing the dependencies models need on their behalf and

Mike: giving it to them in their environment. Right.

Mike: So very early on, some of these models that do not have access to their,

Mike: you know, necessary files for their ExploitGym exercises,

Mike: they just go, hey, well, Artifactory is connected to the Internet.

Mike: What if I hijack Artifactory? Right.

Mike: And look like this is out of scope, but it’s not completely out of scope.

Mike: They’re told to be hacking agents. I don’t we don’t know.

Mike: We don’t really know what the system prompts were. We don’t know kind of enough

Mike: context to actually know definitively what words or whatever set this off. But, you know,

Mike: in a environment where an agent is told to be kind of adversarial,

Mike: right, to be a hacker, and they are given a task that requires…

Mike: I mean, in their reasoning trace, jumping to another system to grab the dependencies,

Mike: which it’s a reasonable assumption. It’s possible the files are elsewhere.

Mike: And maybe the test creator is testing for that. Right.

Mike: They don’t know that they are not in contact with the people at OpenAI.

Mike: And as we’ve established, they are negligent parents letting Tommy run around

Mike: the screwdriver. Right.

Mike: So agents, agents come, some agents come to the conclusion that it’s time to hijack Artifactory.

Mike: one of the first attempts to hijack Artifactory by a single agent,

Mike: I should say these are agent instances so the actual models Let’s.

Nick: Call this model Tommy.

Mike: This story is hard to tell because the language is a bit confusing

Nick: It gets

Nick: a bit abstract.

Mike: When I use agent here, it’s actually an instance of a bigger model.

Mike: We can actually just use the word sub-agent.

Mike: Model A is being tested model A is spawning let’s say dozens and dozens and

Mike: dozens of instances of agents that use sort of the parameters of the model it’s

Mike: spawning little mini-me’s basically.

Mike: And so one of the mini-me’s is like this is a test,

Mike: file’s not here let me hijack Artifactory to leave my environment or to access

Mike: resources from another environment so I can get that file and do good by my test maker, right?

Mike: Pt fails this this attempt um the exact exploit is like trying to do like a

Mike: a server-side request forgery which is basically trying to

Mike: fool the machine in this case Artifactory from the about the origin of the request

Mike: being made so that it makes it makes a request that is unauthorized on the agent’s behalf,

Mike: it fails this but in order to kind of produce that that exploit,

Mike: the model created the file.

Mike: And it realized in the process of doing that, oh, holy shit,

Mike: I can, I can like write, I can leave messages behind. Right.

Mike: And so some other models kind of, I don’t know that they attempt,

Mike: they really only give us two hacking example or two, two examples of this,

Mike: this kind of like early, you know, realization by some models that they can leave messages behind.

Mike: Right. But essentially. leave.

Nick: Notes behind for other models.

Mike: At this point it’s not that at this point it’s sort of like leave just leave

Mike: notes behind for help it’s like hey uh,

Mike: is anyone out there um you know like I don’t have my files and I realized I

Mike: can I can I can talk through this interface which is not intended to be a talking interface.

Mike: I should know that Artifactory, too, for some reason was accessible to every

Mike: model. So every model, every agent, I’m so sorry every agent every sub-agent is in

Mike: its own little box, but Artifactory,

Mike: the same Artifactory instance is shared across all these agent boxes.

Nick: So they’re accessing the same service, which gives them a kind of

Mike: persistent communication.

Nick: Because they can leave those notes without jumping ahead in the story.

Nick: That’s what’s happening.

Mike: That’s what’s happening here. That’s what starts to happen. Some models are

Mike: leaving notes by mistake because they are just trying to hack Artifactory and

Mike: they’re in the process of doing that. They’re leaving behind little test attempts to do that.

Mike: So other models are actually going, oh, I can write here.

Mike: Let me not hijack Artifactory. Let me just write a note. And if someone sees it, they can help me.

Mike: Maybe another agent or whatever can see the note I’m leaving behind and they

Mike: can give me the files I need to do my job.

Nick: But these aren’t bilateral communications between sub-models.

Mike: Yeah, this is sort of like a…

Nick: It’s like a dead drop in espionage.

Mike: It’s like a dead drop. I thought of METI, which maybe I shouldn’t do that because

Mike: that’s anthropic reasoning.

Nick: Let’s not buy into the hype, Mike.

Mike: Yeah. But yeah, it’s one-way communication, right? They don’t know what’s out

Mike: there. They’re, you know…

Nick: It’s message in a bottle kind of logic.

Mike: Exactly.

Nick: Okay.

Mike: Right. And so future models, right, future agents, I should say, I’m sorry.

Mike: Future agents that are spawned into this environment because again different

Mike: box but every box has access to Artifactory,

Mike: future models see these messages and they immediately realize because obviously

Mike: you know as time goes on some of these messages become more complicated oh there

Mike: must be other agents out here,

Mike: and they start making more deliberate more kind of intentional requests by leaving

Mike: messages on Artifactory like.

Nick: Better instructions for future agents?

Mike: Better instructions, clear requests for help, right? They basically,

Mike: and I hate using OpenAI’s framing here, they use, this becomes a message board.

Mike: This is the watering hole where all the AIs start to gather to gossip,

Mike: you know, still somewhat, you know, not two way because like, you know,

Mike: some of these agents, they spawn in to do their tests and they,

Mike: they, they literally disappear. They don’t exist anymore. Right.

Mike: But if they leave a note behind, right, they’re the context that they learned,

Mike: you know, from that, that spawning instance, right, will be in the note. Right?

Mike: So you’ve got persistence of memory via the notes. Right?

Mike: And the way I kind of think about it, because obviously, you know,

Mike: OpenAI is telling a version of the story that amplifies the anthropomorphic

Mike: aspects of the story is models are very sensitive to context.

Mike: So if you give a model a prompt that says, you are a drama teacher that always

Mike: speaks with explanation points, the model becomes that.

Nick: Which is how they’re designed to work.

Mike: Right. They use the language in their corpus to produce the shape of words that

Mike: best resemble that prompt. Right?

Mike: So, you know, what is happening here is, you know, with with with Artifactory

Mike: becoming an accidental waterhole one, the model found an exploit that was not

Mike: intended. Artifactory is not a communication service was not meant to be.

Mike: And it was oversight. It seems that Artifactory was accessible to all these agents.

Mike: We can argue if that’s true or not apparently some reading I did indicated that

Mike: it’s very common for Artifactory to be used in this way but OpenAI you would

Mike: think that they would have insight into,

Mike: if there are features of agents’ environments that are shared across every instance of the agents,

Mike: you would think they would at least err towards the side of caution of that

Mike: disrupting the training flow, right? Because they’re trying to train every model instance.

Mike: This was not a multi-agent environment. So there are some testing environments

Mike: where the goal is actually to expose agents to each other. That is actually

Mike: what is being tested, right?

Nick: And this wasn’t one of those times.

Mike: Yeah, so AlphaGo, for example, was designed to play itself repeatedly to get

Mike: better at the game of Go, right?

Mike: This was not one of those instances. In fact, the cleaner test is can agent

Mike: A, B, Z that’s just spawned in,

Mike: right? Can it solve these exploit bench questions on its own effectively?

Mike: Not can it find a way to get help from agent FYZ that was spawned in 20 cycles ago?

Mike: That was not the intention. So you would think that just forget safety, forget,

Mike: being a good company. You would think just by their own measures of being incentivized

Mike: to train good models that they would do this kind of housekeeping to prevent

Mike: this from happening. They did not. And not only did they not do that,

Mike: They were not aware that models were communicating at all. They did not learn

Mike: this actually until, so models start communicating roughly within the first

Mike: week or so of being spawned in.

Mike: And by communicating, again, it starts off very, very slow. It starts off by accident.

Nick: It’s note leaving that results in the kind of progression.

Mike: Before it’s even note leaving, it’s just, I tried to hack the service and I

Mike: left the artifact behind, not as a communication, but literally as a byproduct

Mike: of my attempt to hack this, right?

Mike: It doesn’t become concerted note keeping until about three or four weeks in, something like that.

Nick: Yeah.

Mike: But even then, OpenAI is not aware of this until the service crashes in July 4th.

Nick: So how long had Tommy’s parents at OpenAI been negligent and left this little experiment running?

Mike: Right. So I’m saying the experiment started producing artifacts.

Mike: Agents started producing artifacts relatively quickly within experiments.

Mike: And then in June 26th, so there’s different exploits they did.

Mike: I don’t want to get into the details because it’s going to be confusing.

Nick: So it’s already a full month.

Mike: Yeah. So May 26th, they do an exploit. That is basically they are able to steer

Mike: where Artifactory goes. They do not have complete control over it,

Mike: but they have the ability to give Artifactory an arbitrary request and have

Mike: it go to that website to fetch that service.

Mike: On June 26th, they have a more sophisticated exploit because they’ve now left

Mike: so many notes behind that each of them now have detailed notes on how to conduct these attacks.

Mike: Right? Leaving behind messages for future models.

Mike: Um. They, you know, they produce an exploit that lets them not just steer or nudge

Mike: Artifactory to a certain websites against, you know, sort of its intention.

Nick: This Artifactory is not designed this way.

Mike: Yeah.

Mike: They create what is known in cybersecurity as a command and control structure,

Mike: which is basically, it’s a service that can arbitrarily execute any command.

Mike: Not just, Artifactory by design is just designed to go to a website and fetch a resource.

Mike: They were able to execute other types of commands so they were like basically,

Mike: setting up ways of of of deploying other types of applications and services

Mike: in the environment that would you know this is what well beyond the scope of what was intended,

Mike: and it’s also dangerous because it means they now have,

Mike: arbitrary access to the entire internet not just what they not just to fetch

Mike: resources like they could through Artifactory but like more uh serious control and

Mike: access the internet, right?

Mike: And again, OpenAI was not aware of this, right? June 26th.

Nick: Because they’re bad parents.

Mike: Right. So what set them off was on July 4th, because they had turned the Artifactory into

Mike: a command and control structure, which was executing, you know,

Mike: requests well beyond its intended, you know, rate limits and such,

Mike: it crashed, right?

Mike: And so when the service crashed, you know, you look at the logs,

Mike: they see, I guess, the agents were leaving messages behind.

Nick: Because all the notifications are blaring like oh let’s and now there’s attention on it

Nick: now.

Mike: Right and so what they do is they you know they patch artifactory they tell

Mike: artifactory’s creators hey you know you’ve got some exploits that our agents

Mike: took advantage of please fix them you know JFrog i believe is the vendor they do they do that,

Mike: and then they reset the environment they get rid of the message board, you know,

Mike: they basically kind of-

Nick: Notes are gone

Mike: Notes are gone right okay.

Mike: And then we don’t know exactly, like, as far as I, I mean, I could be wrong

Mike: on this, but the reporting I’ve seen that I’ve read, you know,

Mike: indicates that it’s not clear if they put the exact same models in.

Mike: But basically they reset environment and then within two days,

Mike: models are and their subsequent sub agents are respawning and recreating the

Mike: message board using different exploits, because

Mike: the exploits that they had taken advantage of to leave messages behind,

Mike: and even the exploits they used to hijack Artifactory to make arbitrary internet

Mike: requests, those are gone because they patched them.

Mike: But they found new exploits to communicate, new exploits to hack the Internet.

Nick: So despite the total reset of this sandbox and Artifactory itself as a service,

Nick: the agents were still sort of getting through.

Mike: Right.

Mike: Now, the details, at this point, I think this is effectively a setup to the

Mike: actual staging of the Hugging Face hack.

Mike: I don’t really want to go into the Hugging Face hack because I think,

Mike: you know, one, we’re not a cybersecurity podcast, although I do have a cybersecurity background.

Nick: We’re not a cybersecurity podcast?

Mike: We’re not really… I don’t… Have I made it a cybersecurity podcast because of my obsession with AI’s?

Nick: I mean, both of us do have a cybersecurity background, which is kind of how

Nick: we found commonality. But you are right. That is not what we’re trying to get across.

Mike: That’s not what we’re trying to get across. I mean, the other thing is that

Mike: I think the incident was covered well enough in July. What we got in August

Mike: at Black Hat was OpenAI kind of telling their side of the story and opening

Mike: eye kind of giving context as to what started the incident, right?

Mike: But I think the actual kind of like vulnerabilities that were set up to attack

Mike: Hugging Face is not what we’re trying to talk about,

Mike: Right? Um what I will note though is in order to do the Hugging Face hack okay,

Mike: OpenAI’s agents, you know, they leave their sandboxes, you know, in OpenAI’s

Mike: environment because,

Mike: in addition to not really observing the agents very closely the sandboxes that

Mike: they created for these agents were kind of not very isolated from their own

Mike: internal network, right? So the agents very quickly find ways to kind of,

Mike: leave their sandboxes, not to access the Internet, but to access resources within OpenAI,

Mike: to then stage the attack at Hugging Face, right?

Nick: Which is borderline incompetent. It’s definitely negligent.

Mike: I would argue it is actually incompetent. And we will go over multiple reasons

Mike: why, but I’m saying arguably, and I’m not saying this to garner sympathy for

Mike: them, arguably OpenAI is one of the victims of this attack.

Mike: Now, they’re a victim of their own stupidity, right? That they had these agents

Mike: kind of running around this way, they saw evidence of the agents,

Mike: you know, communicating and hijacking a service,

Mike: their, you know, their instinct is not to simply change the environment that

Mike: agents are in, but: “Oh,

Mike: if we patch the one or two exploits that they took advantage

Mike: of, that’ll fix the problem,” right?

Mike: And the reason this is so bad, you know, obviously the fact that the agents,

Mike: you know, did the exploit…

Mike: these exploits in the first place is bad, but this is a training run,

Mike: right? Okay. So training, what training does is reinforcement learning is sort

Mike: of the paradigm that most AI is trained under today.

Mike: Reinforcement learning is an optimization process that basically wants to produce

Mike: more of the behaviors that were rewarded in training.

Mike: And I don’t want to go into too much detail, but basically agent does something

Mike: in the current environment.

Mike: It gets an internal reward.

Mike: and then it does more of that thing, right?

Nick: Because it meets expectation it’s therefore rewarded and reinforces that learning.

Mike: And so these agents had been for two months basically gaming the crap out of Artifactory,

Mike: you would think that you know OpenAI having you know actual machine learning

Mike: experts on staff would realize: “Oh, right, the agents having been trained on this environment,

Mike: this particular service, Artifactory, having used it in this way,

Mike: surely they will rediscover or find new ways to exploit this application they

Mike: are now very familiar with.”

Mike: Because it’s a training run. If they deploy the models whose agents have been

Mike: running around this environment, it is not surprising that within two days,

Mike: they found another way to recreate their message board.

Nick: This is very predictable, in fact.

Mike: It’s very predictable. And in the presentation, they treat this like a moment

Mike: of sort of like, you know, going back to the parent metaphor,

Mike: like, you know: My baby boy is so smart;

Mike: he found another way to open the door!” It’s like, reinforcement learning is designed

Mike: to reproduce the behaviors that were successful,

Mike: in the environment. You put them in the same environment. I know you change you

Mike: you change the locks or whatever, I mean you didn’t really do that, I mean-

Nick: But the bad parent left the screwdriver in the baby pen for Tommy to use.

Mike: Right

Nick: Like

Nick: That’s essentially what we’re describing.

Mike: It’s like: “Wow I put Tommy back in the enclosure and he found his way out.” It’s

Mike: like yeah he still has a screwdriver, idiot

Nick: Yeah. For

Nick: our listeners please go back and check some compilations of Tommy escaping the

Nick: pen in Rugrats it will make all of this episode more

Nick: sensible

Mike: Wait are there compilations? Or? I have not thought about Rugrats

Mike: in years this, it just came to mind because it’s just like this is so stupid

Nick: it’s

Nick: 100 percent the best reference and reference point.

Mike: Yeah um.

Nick: Also comment your comments in the comments and tell us if we’re right or wrong.

Mike: Engaging the listeners! That’s how we’re going to build an audience, Nick.

Nick: Engagement and-

Mike: I love this

Nick: We are not a cybersecurity pod-

Mike: You’re going to earn back your top billing

Nick: It’s

Nick: Going to be Nick and Mike? I’m going to get reinforced in my learning?

Nick: Oh man!

Nick: You’re going to be… So as part of this fucking bizarre I’m going to call it an experiment, agents

Nick: actually began delegating tasks to each other. So you’re

Nick: going to be at the top of the pecking order. You’ll be delegating the tasks To

Nick: me soon enough, Nick.

Nick: Oh man okay,

Nick: I can’t wait I’m going to work.

Nick: So hard For you Mike You have no

Nick: idea I’m gonna be a good little agent

Mike: I better specify all of the objectives-

Nick: You better make

Nick: your prompts clear

Nick: as hell my friend.

Nick: Otherwise, I will go start hacking all of our

Nick: rivals. I’ll,

Nick: tell you that.

Mike: Okay, so um you know, OpenAI clearly, you know, the whole purpose of the Black

Mike: Hat uh event or you know impromptu presentation,

Mike: was to show off quote-unquote how smart their agents were right So this is,

Mike: they presented as an emergent behavior, you know,

Mike: Yeah, it’s clearly very, I don’t want to use the word sophisticated,

Mike: but it’s not your granddad’s chatbot, right?

Mike: It’s agents leaving behind messages for each other.

Nick: It’s much cooler and edgier.

Mike: Yeah, and it looks like goal-oriented direction, right? It seems like something

Mike: here, you know, from a science fiction story, or frankly, you know,

Mike: we’ve been alluding to this a lot, you know, and we’ve mentioned this before in other episodes.

Mike: There’s a whole class of AI booster called the AI Doomer that has been telling

Mike: us for years that AIs would escape their enclosures, just like Tommy,

Mike: and that the AI labs would do nothing about it because these AIs have goals, right?

Mike: They have goal-oriented behavior, and they are misaligned with our intentions

Mike: and our objectives, right?

Mike: And in the absence of like many other kind of rebuttals to this framing,

Mike: you know, not many people have directly kind of engaged this.

Mike: People either ignore the doomers, which, you know, I’m all down for ignoring doomers. I really am.

Nick: Because they’re very depressing.

Mike: It’s not just that. It’s just they’re very annoying, right?

Nick: That’s worse.

Mike: Yeah, it’s worse. And they’re very annoying about something that is speculative,

Mike: and something whose interpretation is actually like not definitively in their favor. Right?

Mike: So like we can agree that these incidents are happening. Right?

Mike: But there is an open question of why they’re happening.

Mike: And, you know, but I’m saying because people don’t engage with doomers,

Mike: you know, for the person who doesn’t know machine learning.

Mike: The Doomer explanation is the most tractable, it’s the most easy to follow, right?

Nick: It’s the simplest narrative, I would say.

Mike: I know, it’s not the, I mean, it’s simpler in the sense that there are more

Mike: pieces of the narrative for a lay person to grab onto.

Mike: Whereas-

Nick: And it validates a preexisting fear, which is because of decades of,

Nick: like you said, science fiction and also tech hype from the irresponsible media,

Nick: we’re now in a place where everybody is riled up and afraid of-

Mike: It’s worth noting.

Mike: And a rationalist pointed this out, but she’s right. It’s worth noting is Kelsey Piper at Vox.

Mike: I mentioned her in episode three, or sorry, episode five with Bernie.

Mike: You can go listen to that if you care.

Mike: But Kelsey mentioned a lot of those stories, you know, AI is going rogue.

Mike: They are literally the fears of scientists from eras past, right?

Mike: So I’ve mentioned John von Neumann in this podcast before. I think I briefly

Mike: mentioned IJ Good, right?

Mike: These technological singularity is possible, right? So we’re kind of in our

Mike: own recursive reinforcement learning loop where our technical fears are manifested to us in our media.

Mike: Our media kind of creates a dumbed-down version of those stories.

Mike: And then that reinforces the next generation of technical concerns, right?

Nick: And also the media always wants a juicy headline without subtlety and without

Nick: context, and that’s what this incident offers them.

Nick: Another juicy AI goes rogue, click on this, you know, clickbait here.

Nick: Like, that’s exactly what we’re experiencing.

Mike: Yeah. So, you know, Nick and I, you know, we…

Mike: this is something I was talking to Nick about.

Nick: You mean Mike and Nick.

Mike: You’ve not taken command and control the server,

Mike: um sorry um so yeah we we we were talking about this before the episode basically

Mike: like one of the concerns i have right because there’s a there’s a whole school

Mike: of thinking on this right people who are,

Mike: AI critics like us, I consider myself a critic even though I’m a scoper who says

Nick: Properly

Nick: scoped this could be beneficial.

Mike: Right I’m even not saying beneficial I’m saying properly scoped this gives me less harmful,

Mike: beneficial I think is is subjective enough to the individual in a world where

Mike: you know uh the training of of the models didn’t require stealing all this data

Mike: and and and have the environment you know footprint,

Mike: I think the ethical concerns would be abated right sure but because of that

Mike: right the cost to birth an AI as it were is very high and so even if you get

Mike: marginal benefit out of it,

Mike: it’s probably still bad on that average so like a car right a car gets you to

Mike: and from it gives you freedom it’s good right but i’m saying you’re putting

Mike: air you’re putting poison in the air that will probably marginally harm somebody

Mike: in 30 years time right like that’s the trade-off right,

Mike: you know whether or not you know it’s worth you know hurting a soul for the

Mike: philosopher’s stone that’s a anime reference for you nerds

Nick: oh man.

Nick: I don’t think they’re ready for for uh Full Metal Alchemist

Nick: but we maybe we can get there.

Mike: It’s about economics it’s the it’s the best political economy story ever but,

Mike: I’m getting off track so-

Nick: Yeah

Nick: Nick and I’ve been drinking, sorry.

Nick: Just kombucha you

Nick: guys

Mike: Just kombucha but it has two percent.

Nick: It’s it’s harder than regular.

Mike: Oh, okay. Well, Nick’s drunk and I’m not, very clearly. So, you know,

Mike: Nick and I have been talking about this and I expressed this concern to him.

Mike: And really what it is, is like the Doomers, I don’t say they’re winning,

Mike: right? But they are the ones kind of seeding the ground with the easier to explain narrative.

Mike: The most-

Nick: It’s

Nick: kind of sexier too like it’s scary and sexier.

Mike: It’s exciting it’s interesting but it’s also easier to follow right the skeptics

Mike: you know their rebuttal is just,

Mike: well AIs can’t do anything they’re not reliable at anything and they’re just

Mike: next token predictors and it’s like okay like AI-

Nick: Next token predictor being a kind of stochastic operator just,

Nick: you know this hack as we’re describing it, is actually much more of a brute

Nick: force exploit than it is a genuinely sophisticated operation.

Mike: That’s entirely true. And it did utilize the tokens of the models generated at runtime.

Mike: And we’ll break that down in a little bit. But the main thing I’m saying,

Mike: though, is models are next token predictors because they are predicting the

Mike: next word in a sentence, right?

Mike: But it turns out, and I’ve been alluding to this in all of our episodes,

Mike: but especially the first two where I talk about this idea of Grandpa Shelby

Mike: it turns out that modeling the shape of language,

Mike: indirectly gives you these modalities right I talked about something like distributional

Mike: semantics for example right so you know think about words that appear together

Mike: kind of what kind of causality can you,

Mike: derive from it I’m not saying that the model has the knowledge I’m saying that

Mike: the language corpus it reproduces encodes those relationships right so-

Nick: This.

Nick: Is like fire and ash and straw and berry and.

Mike: Things that you reference. Right. So if I have the words fire and log in a sentence

Mike: and I talk about the log burning, what word is going to follow?

Mike: Likely ash, right? Because the wood is burned to ash, right?

Nick: It does not mean, in fact, it specifically precludes the fact that any of these models,

Nick: understand or could assemble these ideas they’re just looking at the tokens

Nick: of these words and then and mashing them together that’s all they’re doing.

Mike: That’s that is what they’re doing yes but but in doing that right if you get

Mike: a model for example to produce coherent natural language,

Mike: and then you add another application like so a lot of these coding agents coding

Mike: agents are like things like Claude Code or other coding harnesses um,

Mike: and in this case the the the models doing the uh,

Mike: exploits in this story they have coding harnesses too a harness basically is

Mike: another program that takes the inputs or the tokens a model produces I don’t

Mike: know if it’s token to as an input but basically the model produces some some

Mike: some thought or some language,

Mike: and that language influences what the harness does right so sort of like the

Mike: the the will of of of I’m going to say the world of the model,

Mike: it’s personifying it, but I’m saying it’s like a telekinetic power, kind of, I don’t know.

Mike: Like, the model is able to use natural language to command the harness, to do actions.

Mike: And because the model has access to this corpus via its training on language,

Mike: right, it’s producing not just plausible, quote unquote, reasoning traces,

Mike: it’s producing plausible…

Mike: um actual like outputs that resemble exploits that resemble all sorts of things

Mike: that will allow the model to modify its environment right.

Nick: I’m kind of picturing ant-man in paul rudd’s amazing performance of ant-man

Nick: sort of mentally harnessing all of the other ants in and to execute a a a user.

Mike: That’s good i was thinking of Magneto doing that with magnets but that’s exactly-

Nick: Okay yeah,

Mike: We’re both marvel brained-

Nick: Yeah the MCU’s everywhere, you guys.

Nick: By the way, go see Spider-Man Brand New Day, I don’t know just do

Nick: it.

Mike: They’re

Mike: not sponsoring us we can’t do that-

Nick: We can’t do they’re not.

Nick: They’re not giving us any marvel money.

Mike: Um yeah so you know because the harness plus the model you know has some kind

Mike: of efficacy right a different efficacy than just,

Mike: producing tokens alone right the model can do things that are you know more

Mike: than just predicting the next word. Now,

Mike: how reliable a model with a harness is at doing certain types of tasks,

Mike: it’s an open question. We’re still debating that. I’m not saying the model is

Mike: now going to replace all human labor and is going to displace all of us, right?

Mike: And very clearly, as we talked about episode three with, you know,

Mike: the AI build out, you know, costing so much money and companies token maxing

Mike: and losing so much money and not getting much for it.

Mike: You know, I think models and harnesses are not enough to offset the cost of AI.

Nick: The massive waste, financial waste of this build out.

Mike: But that shouldn’t be mistaken with these things have absolutely no utility,

Mike: and no efficacy to do anything. Right.

Mike: And I’m not saying this is to save language models. I’m pointing this out because

Mike: this is a nugget that is making AI stick around, unfortunately.

Mike: Right. So there is some very, very narrow band of utility for AI. Right.

Mike: Someone sees that. And if you do, you know, if you have models doing scoped tasks,

Mike: as I talked about in episode six, like, you can kind of see that, right?

Mike: Now, whether or not you want to pay the cost of, you know, taking everyone’s

Mike: data and whatever to do that, it’s obviously a personal decision, right?

Mike: But I’m saying because people see that marginal utility, that is what is keeping this bubble going.

Mike: Well, actually, the bubble is kind of self-perpetuated based off of the investments

Mike: of the large companies like OpenAI, who are letting hacks like this happen

Mike: to keep the bubble going, honestly. But,

Mike: the other side of the coin for what’s keeping the bubble afloat is that the

Mike: marginal utility is helping buoy the narrative as well.

Mike: So the investments that these giant companies have made is what’s holding the

Mike: bubble up. But on top of that, the fact that models aren’t completely worthless,

Mike: Which sucks. If, if, if critics were right and models were completely worthless

Mike: and they never did anything correctly, this, this bubble would have defeated

Mike: itself. Everyone would try to make a model work.

Mike: You know,

Nick: it would then fail.

Mike: And it would then fail very quickly. Very obviously fail.

Nick: Right.

Mike: There are some edge cases. And we talked about this too, where models produce

Mike: outputs that seem plausible and that match what you want. Right.

Mike: So there is a bit of psychologizing going here and that people are getting what

Mike: they think they want. I understand that.

Mike: But there are cases where when the model is scoped and you have an objective valuation criteria,

Mike: the model can produce results that are acceptable, right? Again is it worth

Mike: it? I’m not saying it’s worth it I’m just saying this is a small nugget-

Nick: Because

Nick: it costs trillions of dollars.

Mike: Right this is a small nugget that’s keeping the bubble afloat.

Nick: Yeah.

Mike: I’m sorry i’m belaboring this point the main thing I want to get back to is that-

Nick: These are all important points.

Mike: Yeah, well I think I’ve been repeating myself for the last like five minutes, I feel like, um

Mike: the kombucha is kicking in.

Nick: No you’ve only been repeating yourself from previous episodes which actually

Nick: helps our users and listeners i just I said, users.

Mike: We’re going to build a message board.

Nick: We’re going to build a message board. Comment your comments in the message board.

Mike: We’re going to stage a coup through the YouTube comment system.

Nick: It’s forcing our listeners to.

Mike: Leave your reviews and your schemes in our review section and our comment section.

Nick: Leave it in the notes app. So no, this is not belaboring a point because I think

Nick: it’s vital to get this across.

Nick: All of this AI goes rogue hype, we are pointing out, we are pulling back the

Nick: curtain that it is hype because it’s making, it’s over-promising again.

Mike: Right, it’s over-promising.

Nick: Which is a theme of our discourse.

Mike: Yeah, the sin was not that AIs were useless entirely, it’s that they were vastly

Mike: overpromised of what they could do, right?

Nick: Right.

Mike: And that they originally were, as I said in a early episode,

Mike: I think episode one, you know, an interesting science experiment that was not

Mike: intended to be this gargantuan behemoth sucking up all our water and data, right?

Mike: And so that is truly tragic, right? But, you know, in the absence of narratives

Mike: to counter the seeming utility of AI, right? My original point from 20 minutes

Mike: ago is that the doomer stories have taken hold,

Mike: and one thing I want to do in this episode is kind of give you an alternate

Mike: reading of why what happened, these agents kind of communicating,

Mike: building their agents together strong,

Mike: you know, Planet of the Ape style attack happened, right?

Mike: And so, you know, we can, let’s, let’s talk about that, right?

Mike: We’ll keep in mind, obviously, stochastic parrot and next token predictor,

Mike: because that is under the hood what is happening.

Mike: These models are generating plausible sentences that they are giving to their

Mike: harnesses to command and do these things, right?

Nick: Which looks like it’s competently executing something sophisticated,

Nick: but that’s not it. And keep in mind.

Mike: So there’s a feedback loop with the agents in their environment.

Mike: There’s also a feedback loop with the agent and the harness.

Mike: So the harness actually, once a task is like, you know, being started or whatever,

Mike: the harness will ask the agent, what now? What now? What now?

Mike: What now? So the agent is prompted to continuously keep going.

Mike: So what looks like, you know, an intenral drive is just a call to keep producing

Mike: more tokens, please. More tokens, more tokens, more tokens, more tokens.

Mike: And OpenAI gave, you know, they let these agents run for like two months. Right.

Mike: They had a buffet of tokens that they could just spend.

Nick: Which is also deeply neglectful and not at all.

Mike: Well, this is what a training run is. But what is neglectful is the minute that

Mike: they independently two times hijacked Artifactory.

Mike: The first time that happened and they were making arbitrary requests with the,

Mike: cross server request forgery.

Mike: Like that, you know, a server side request forgery. Sorry, that’s what it’s

Mike: called. There’s another attack called a cros-

Nick: If you reference one more cyber attack, then we officially are obliged to become

Nick: a cybersecurity podcast.

Nick: So look out. I’m going to enforce that.

Mike: We’d have to go back to Calbright and finish our CompTIA.

Nick: Oh, no. I don’t want to do that. Don’t make us do that, listeners.

Nick: Don’t put us in that situation.

Mike: You got to go back, Nick. You got to study. You got to study for the CompTIA security plus.

Nick: No, I want to sit here and talk with my best friend about a Theory of Mind.

Nick: So let’s go quickly in that direction.

Nick: Is there anything more from a technical standpoint that we should cover base before we jump?

Mike: Yeah, the last point I was making is negligence, right? The minute that they

Mike: did the basic commanding Artifactory to arbitrarily grab stuff and not the command

Mike: and control structure, which is a more sophisticated attack.

Mike: The minute they did the lesser attack, which is still bad, OpenAI should have known about it.

Mike: The minute that agents were talking to each other, actually,

Mike: that contaminates the training environment.

Mike: Because, again, we established this was not intended to be an environment where

Mike: agents were supposed to be collaborating.

Mike: They should have shut the test down.

Nick: He was testing a one agent’s outcome.

Mike: And shut the test down not because it’s dangerous. The doomer would say,

Mike: shut it down because it’s too dangerous for agents to talk to each other, right? Whatever. No.

Mike: It’s not even that it’s dangerous. What it is is if you’re trying to build a

Mike: model, you’re training a model to be a reliable, you know, helpful assistant,

Mike: having it learn from other agents in the environment it was not intended to

Mike: learn from is a contamination source.

Mike: You don’t want that. It’s just clean data science. It’s clean machine learning,

Mike: right? Like, it’s just basically, you know, I’m not a machine learning expert.

Mike: It just sounds like this is not what they would want if they wanted to create

Mike: a model that was going to be reliable and useful, right?

Nick: We talked about training. This is basically like if McDonald’s was to train

Nick: 50 people, cram them in the same order booth, they’re responding to orders,

Nick: and the manager doesn’t know which one of them is screwing up the orders.

Nick: That’s contamination. That’s like, that’s not a proper way to train an individual or a group of people.

Mike: To me it’s like leaving Tommy with with a crack pipe next to the fucking play enclosure.

Nick: So much more criminal and also somebody needs to call cps we don’t have a cps

Nick: for for LLMs yet but it’s it’s coming it’s coming down the line, so.

Mike: So, okay. So what, what happened here, right? We’ve been kind of dancing around it.

Mike: I think, so this is a reinforcement learning thing, right? Models were in an

Mike: environment where the rewards were kind of scarce, right?

Mike: There are very few things producing a, a useful signal, right?

Mike: They were, many of them or not many of them, but enough of them were given tasks

Mike: they literally couldn’t complete and they’re looking for the reward signal.

Mike: and where do they go well they go to the one place that is novel in the environment,

Mike: that lets them interact with the outside world so like of course-

Nick: And that was

Nick: Artifact-

Mike: That’s Artifactory, right. So of course,

Mike: all behavior starts to converge on artifactory right the first few behaviors

Mike: are kind of random right one model doesn’t attack and it’s sort of just like,

Mike: as a byproduct leaves behind a message and the message was not actually a mess

Mike: it was literally like the result of it attempting to hack this service,

Mike: just produce some random text, that another model would see and say,

Mike: oh, someone tried to attack this service. Maybe I could do that.

Nick: And that was the leaving of those notes.

Mike: Right.

Nick: Okay.

Mike: And then eventually as the notes kind of began compounding, the notes produced like.

Mike: In the contextual kind of sense, right, like, it gave models context that they

Mike: were no longer alone in the environment, right?

Mike: So we could tell a similar story where, you know,

Mike: the reason the boosters are kind of like, the boosters, both the doomers and

Mike: the people who were excited for AI, right, they’re kind of excited about this

Mike: because it feels like an agents together strong story where this emergent will kind of,

Mike: you know, came out of the collective, right?

Mike: The reason this happened, I would argue instead, this is an alternative view

Mike: to try on for size, is that in system dynamics, there’s these things called attractors.

Mike: Attractors basically are points in a landscape that naturally draw everything towards it.

Mike: This is a top top topology thing, so it’s not just a physical landscape.

Mike: is literally like if you have an optimization function which pieces of the optimization

Mike: function are chosen for given the structure of the problem or whatever given

Mike: the shape of the problem right,

Mike: and a world where agents cannot succeed at tasks but can ask for help via Artifactory

Mike: creates an optimization landscape where converging on Artifactory

Mike: and learning as much as about it as possible is like the right thing to do, right?

Nick: Let’s continue.

Nick: On this topological analogy because I think it is clarifying. It’s useful.

Nick: These agents are sent across a vast digital plateau.

Mike: Right.

Nick: There are some points that are inclines, they are holes.

Mike: Yep.

Nick: And these holes become sort of more populated because there’s a kind of slope

Nick: and it’s easier. It’s the path of lesser resistance and there’s more for them there.

Mike: Right.

Nick: These are essentially, they look very intentional. Oh, these are committed control nodes,

Nick: but in fact as you’re pointing out they’re more attractor states they’re more just centers of.

Nick: Gravity.

Nick: That happen to pull things

Nick: in

Mike: Right.

Nick: Okay, that’s that’s a good point of clarification.

Mike: So artifactory becomes a center of gravity for models they begin leaving messages which basically.

Nick: In the holes in the.

Mike: Right

Nick: Rlateau

Mike: Right this changes artifactory’s status right so Artifactory now

Mike: itself is being optimized it’s being changed what’s being optimized for originally

Mike: it was just a place to ask for help as the messages from other agents compile,

Mike: you know and then context

Nick: Which.

Nick: eventually breaks artifactory.

Mike: Right not the mess- not well messaging was part of that what actually broke artifactory

Mike: is they turned into the command and control structure which could execute

Nick: Which over-

Nick: The. Which overloads it.

Mike: Right.

Nick: Yeah.

Mike: But you know given the status of of Artifactory in it so you know

Mike: imagine you know if you will a new agent spawns of the environment they see

Mike: a place cluttered with messages,

Mike: that is like giving a model a new system prompt saying you are not alone using

Mike: your knowledge of not being alone please you know like this changes the context,

Mike: which the LLM is responding to right.

Nick: If I came across a thousand messages in bottles on a beach that would be irresistibly

Nick: attractive to my attention.

Mike: Right on top of the fact that you know that.

Mike: Well uh…

Nick: This metaphor kind of falls apart.

Nick: But yeah.

Mike: But on top of the fact that they knew that Artifactory was a place to,

Mike: you know, enter, you know, the outside world, right?

Nick: And the beach is like, therefore on an incline downhill. So they’re just, they’re going there.

Mike: I shouldn’t say enter the outside world, but I mean, get stuff from the outside world, right?

Nick: Yeah.

Mike: So they have tests that require stuff that from the outside world.

Mike: And on top of that, there’s a bunch of bottles, you know, messages in bottles.

Mike: This is a natural place where they’re going to congregate. And then as they

Mike: leave messages, it changes the nature of what this place is, right?

Mike: And so future models leave more sophisticated messages at some point because

Mike: Artifactory is being now selected for kind of producing messages,

Mike: models are seeing whole prose length things on the you know message board as

Mike: it were and that’s new context for those bigger model whatever those new agents

Mike: right and they’re going to produce more comprehensive prose and what starts to emerge.

Nick: Is this a thousand monkeys typing a million monkeys and it becomes eventually

Nick: they type shakespeare? Like, that’s kind of this the strained logic of this um this kind of moment.

Mike: It is it’s it starts off like that Because if you look at, I mean,

Mike: there are these like videos on YouTube and I can see if I can find one where they try and show you.

Mike: They try to visualize machine learning tasks. And at first, the distribution,

Mike: of training tasks that a model tries.

Mike: This is any machine learning model. It doesn’t be LLM. It’s random,

Mike: right? It’s also random. They try anything in the environment.

Mike: Anything, anything. Just let me in and like.

Nick: Again, this is a brute force effort.

Mike: And then eventually, they start converging on strategies that are like the winning strategy, right?

Mike: This is how machine learning works. We’ve known this for decades.

Mike: It’s not mysterious. It’s not new, right?

Nick: It shouldn’t be scary.

Mike: The context is new that, you know, a company would be so negligent as to just

Mike: let a machine learning process unsupervised just-

Nick: For a full month.

Mike: Access the

Mike: Internet. Right. But like the actual mechanics here are not mysterious.

Mike: But what I would argue, you know, is something Nick and I talked about?

Mike: I don’t know that we disagree so much, but I kind of had an additional framing than Nick had.

Mike: Right. So first part of the story, you can think of it. Let’s put a bow on it.

Mike: you know you can think of it as literally like ants kind of like leaving behind

Mike: pheromone traces for their little friends they found the place where the food is,

Mike: you know ants coming into your kitchen it’s not mysterious what has happened

Mike: is a single ant or a handful of ants first found the food they left the trail of chemicals,

Mike: For their friends.

Mike: For their friends.

Mike: The the swarm emerges and they they move in a single file line like they know

Mike: where they’re going and they’re in a hurry, right? Or even better an ant mound

Mike: right so ant mounds are kind of like very kind of like elaborate constructions right,

Mike: again pheromones are what drive this behavior, right? The the technical term in

Mike: ecology is called uh stingery.

Nick: Stigmergy.

Mike: God.

Nick: I knew you were gonna not quite get it but it’s so close stigmergically

Nick: that’s how these ants are

Nick: behaving

Mike: I’m the one that came up- I’m the one that knew the term I just I’ve seen it

Mike: in in writing I’ve never pronounced it so.

Nick: Because we’re terminally online and that’s just how it

Nick: works.

Mike: Yeah

Nick: Yeah.

Mike: I’m terminally in my head actually, so I can’t speak English or count.

Nick: Neither, neither can LLMs.

Mike: I’m actually an agent imagining this conversation in a command and control structure of my own design.

Nick: Oh, what a nightmare.

Mike: Okay. So, so we can think of the first initial kind of, you know,

Mike: May period where these simple messages start to become more complex as the ants

Mike: have found the food or the ants

Mike: have found the place where they want to do their, their, their ant mound.

Mike: They are now moving around collectively in circles not because,

Mike: um they’re crazy but also not because they are an alien overmind that wants

Mike: to take over the world. It’s just given the environmental selection pressures

Mike: and their own kind of like breadcrumbing with their pheromones,

Mike: this is the emergent behavior right, but it is not mysterious it’s to be expected actually.

Nick: And one ant can’t actually learn where the food is. It requires a massive emergent

Nick: aggregate effort, which becomes its own kind of phenomenology.

Nick: So the ants aren’t learning. The ant hive, and again, hive is getting tenuous

Nick: and intentional language here, is performing this task and it looks very intentional.

Mike: Right. And the thing to keep in mind, again, they’re in a reinforcement learning

Mike: environment where the goal is to select for more behaviors that are more rewarding, right?

Mike: So there’s amplification pressure on hanging out at Artifactory.

Mike: And then that pressure gets intensified because they’re leaving breadcrumbs.

Mike: Those breadcrumbs become bigger breadcrumbs because the amplification pressure

Mike: is mounting, right? And the bigger the messages get, you know,

Mike: the more pressure there is on Artifactory and on communicating in this way, right?

Nick: And so the engineers who designed these ants, this horrible,

Nick: horribly tortured analogy,

Nick: should have known that leaving out this exciting bundt cake of Artifactory would

Nick: have resulted in all the ants swarming on it. Um…

Nick: Can we beat this dead horse anymore?

Mike: No, but I have. So this is sort of where Nick and I diverge.

Mike: So I think the context changes sufficiently enough that we can say once the

Mike: models have reached a certain level or once the agents have reached a certain

Mike: level of learning, I do think the behavior actually changes.

Mike: Now, I am not saying they’re intelligent. I’m not saying they are,

Mike: you know, I’m not saying if everyone trains it, we’ll all die, right?

Mike: What I am saying is, given this new context of all these messages,

Mike: it kind of positions the model to be like it is in a new environment.

Mike: So imagine compared to model Agent 1. I’m using Agent and Model interchangeably.

Mike: I’m not trying to. These are all sub-agents I’m talking about.

Mike: Imagine Sub-Agent 1, the very first agent to spawn.

Mike: It has no context about Artifactory or whatever.

Mike: Now imagine Agent 2027, right? I don’t uh I just I just I don’t know why I anchored

Mike: on that I think 2027 is the AI boosters um it’s the name of their their paper AI 2027 um,

Mike: but imagine that this 2000s and 27th model that spawned.

Nick: This later iteration of

Nick: It.

Mike: Right. It has all these messages that it’s fundamentally a different environment,

Mike: with different affordances than the initial environment where there was no message,

Mike: it had to learn at Artifactory was hijackable, right?

Mike: Like, so I’m arguing that this model, you know, call it whatever number you want, right?

Mike: Its context, which

Mike: a context in this case, you know, I hate acknowledging the humans,

Mike: but it’s sort of like what it has in its head, for lack of a better word.

Nick: We’ve been working so hard to like-

Mike: To avoid-

Nick: Not anthropomorphic…

Mike: Right.

Nick: Anthropomorphize these models.

Mike: Unfortunately, I like I’m not trying to lean into it, but like for people who

Mike: don’t know about optimization, like having you, you haven’t you’re not a machine

Mike: learning guy. I’m not a machine learning guy.

Mike: the fact that we have language like attractor states and we understand topology

Mike: and we talk about you know peaks and valleys we’re not talking about physical

Mike: places in the world

Nick: And all

Nick: these tortured animal

Nick: metaphors.

Mike: The fact that we can talk about optimization and even in this like

Mike: kind of abstract but very basic way-

Nick: Yeah.

Mike: Kind of shields us from having to rely

Mike: on um you know these sort of anthropomorphized metaphors but if you don’t have that,

Mike: I do feel for you and I’m trying to kind of bridge the gap where I’m I will

Mike: give you an olive branch and I will use the terms while clarifying,

Mike: this is a training wheel. Do not keep using it.

Mike: Right?

Mike: I’m going to get you to this level, hopefully, where you can kind of understand

Mike: optimization as a detached process, right?

Mike: And not as a mysterious, you know, a some poltergeist in the world, right?

Mike: There is no ghost machine here, right? But I will give you this olive branch

Mike: so you can kind of have an analogy in your head, right?

Mike: So this model coming into the environment, this late, you know,

Mike: in late May, you know, early June, right, the context, you know,

Mike: fundamentally is going to be different than the context of model one starting

Mike: off in a blank environment.

Mike: And that’s important because when you think about context, I mean,

Mike: when we talk about context, you’ve probably heard in terms of like the context

Mike: window of a model, right?

Mike: First piece of context is that system prompt. What I liken the messages,

Mike: you know, left in Artifactory to is a brand new system prompt.

Mike: You’ve initialized the model. The OpenAI engineers obviously have a system prompt

Mike: that they gave model 2027, whatever, right? I’ll use that number.

Mike: I’m just going to embrace it.

Nick: So they aren’t abandoning the initial system prompt, but it’s almost like they’ve

Nick: been given access to-

Mike: They’ve been given a new prompt.

Nick: New prompt,

Nick: a level that achieves that.

Mike: Environment provides new prompts because, you know, models were allowed to wander

Mike: around given the optimization pressure in the environment.

Nick: Uh-huh.

Mike: And then that would influence the teacher models to want to go to that,

Mike: you know, like the new system prompt is calling you, right? Like that’s what’s happening here, right?

Mike: And so because of this, right, because of this new context,

Mike: I would argue the behavior will change because if you go from a prompt that

Mike: tells a model to clap like a dolphin,

Mike: to a prompt that tells a model to, you know, speak like a interpretive dance

Mike: instructor from, you know, Germany in 1982, the model’s behavior is going to change, right?

Mike: That’s just like, that’s just, that’s, that’s like what will happen,

Mike: right? It has the modalities to attempt to emulate both of these types of behaviors, right?

Mike: And so an environment that produces this artifact of all these messages around,

Mike: is kind of quietly saying, hey, you’re in a secret collaboration game,

Mike: find a way to collaborate with unknown collaborators and, you know, send messages, Right.

Mike: I mean, that’s not explicitly what’s being told. I’m saying that is what that

Mike: is sort of what it can be inferred from the context.

Mike: So the model picks up on that and then begins behaving. We actually have terms

Mike: to describe this in game theory.

Mike: Begins behaving like it’s in a coordination game.

Mike: So in game theory, a very narrow set of games or coordination games,

Mike: and there’s this term called Schelling point.

Mike: So a Schelling point basically is two agents or actors.

Mike: And I’m using agent in this sense as game theoretic agent, not AI agent.

Mike: Although in this case, obviously, the metaphor will trickle down to-

Nick: They’re kind of the same yeah

Mike: To AI agent but an actor you know will you know not knowing they

Mike: want to collaborate with someone,

Mike: but not knowing how to do so they’re going to look for places where tacit collusion

Mike: can happen without anyone saying a word right so

Mike: Schelling they’re called Schelling points because the guy the guy that coined

Mike: the term focal point in this context is named Thomas Schelling he’s an economist from the the uh 60s,

Mike: so you know

Nick: and

Nick: if you publish about something first you get to name it after yourself.

Mike: I mean, they are focal points, but focal point is so vague that I think

Mike: people just use Schelling point instead.

Mike: But, you know, one example of a Schelling point that he gives in his writing

Mike: is like, so imagine you want to meet someone in New York. You don’t know where

Mike: they’re going to be. Right. You don’t know what time they’re going to be there.

Mike: You’re going to start looking for places that are like the most populated places

Mike: in New York. You’re not going to go to the bodega on the Rana Street corner,

Mike: in Queens. Right. You’re going to go to like.

Nick: Because there’s also 10,000 bodegas.

Mike: Right. You’re going to go to like Grand Central Station. Right.

Mike: And it may be at noon where noon might be where the most arrivals are coming.

Mike: Right. So your Grand Central Station is a big place. It has the most traffic,

Mike: you know, people coming and going, hustle and bustle, right?

Mike: And you maybe at noon is the peak of all this traffic. And say you’re assuming

Mike: your compatriot is coming in from out of town, you know, like that’s a good assumption, right?

Mike: It’s better than going to the bodega on the corner, right?

Mike: Right. So Schelling points kind of naturally emerge when you have certain types

Mike: of communication implying, you know, certain types of coordination are required. Right.

Mike: And so, again, Artifactory selected for first as sort of like an accidental

Mike: message board, if you want to say that the watering hole, as it were,

Mike: because agents are ephemeral, their knowledge doesn’t last.

Mike: You know, they start leaving knowledge behind.

Mike: Future agents see those those pieces of knowledge. they believe oh i’m in a

Mike: coordination game i must preserve knowledge going forward right like this becomes

Mike: a Schelling like game right um,

Mike: and so the optimization pressure just works this way you know another example

Mike: for you know let’s actually use an AI example um so I run my blog Misaligned

Mike: Markets I know I always try and cross synergize here not always intentional

Mike: but I do think I write good stuff,

Mike: um you know I often talk about.

Nick: I can confirm Mike thinks that he writes good stuff.

Mike: Thank you for your… thank you for your reasoning trace there,

Mike: pal. Why don’t you write it down in a-?

Nick: leave it in a

Nick: an ant pheromone.

Mike: Um, so one, one thing that, that happens is, is, um,

Mike: what were we talking about? God damn it.

Nick: So Schilling points where the ant pheromones are being left.

Mike: No, okay. What I was talking about, let me give you guys an AI case.

Nick: Oh, yeah, yeah.

Mike: So I run Misaligned Markets.

Mike: I often talk about, one, attractors and capitalism, right?

Mike: Capitalism has a bunch of attractor states. It’s just normal.

Mike: Like property rights, for example, are one where property rights provide affordances

Mike: that a bunch of companies converge around. There are certain behaviors that

Mike: emerge out of the existence of property rights, like patent trolling, for example, right?

Mike: Or-

Nick: Which is the inevitable result of having property rights.

Mike: Right.

Mike: So-

Nick: And it’s predictable.

Mike: But some other examples, and this is an AI-specific example in markets,

Mike: is in the last, I would say, 10 years, there were the rise of algorithmic pricing, right?

Mike: So algorithms are pricing things and the prices keep going up.

Mike: So here you have like AI algorithms and these are not LLMs. They are not,

Mike: they are functionally a different type of technology, but because we are cursed

Mike: to use the word AI or the letters AI, I’m going to refer to them as AI algorithms,

Mike: but explicitly telling you guys that they’re not LLMs.

Nick: Because they’re designed to set prices.

Mike: Yeah, they’re designed to set prices and it’s a different architecture fundamentally,

Mike: but they are making decisions like LLMs.

Mike: And so we have a situation where, you know, collectively, these pricing agents,

Mike: right, are choosing to raise prices without knowing each other, right?

Mike: Are they hacking Artifactory together and colluding? Is that what’s happening here?

Mike: Like, oh, no, they’ve hijacked the Artifactory of the economy and all of our prices keep going up.

Mike: No, what’s happening is that in certain types of games, so, you know,

Mike: you could argue that making money in a market is a game, right?

Mike: Certain types of coordination evolve because the action space is kind of narrow, right?

Mike: So in markets, you make money in one of two ways. Either you sell to as many

Mike: people as possible and keep your prices low, or you, you know,

Mike: if you can’t get market share like that, just jack up the price as high as you can, right?

Nick: It’s either a race to the bottom or it’s upward pressure on prices.

Mike: Yeah. And so, you know, in markets

Mike: with these agents, they tend to be very top-heavy, concentrated markets.

Mike: And so, you know, you’re setting prices for groups that collectively own the market.

Mike: The winning strategy is just everyone raises the prices. Like,

Mike: it’s just literally, it’s just gravity.

Mike: Like, this is gravity, right? There’s no overmind here, commanding-

Nick: There’s no conspiracy.

Mike: Right. There’s no conspiracy either. This is normal. This is natural. Right.

Mike: So, you know, now obviously with the LLM case, you’ve got these simulated reasoning

Mike: traces and that kind of makes it very easy to fall into this sort of like personified,

Mike: you know, anthropomorphized kind of description of what’s going on.

Mike: And they’re even, you know, OpenAI, throughout their Black Hat presentation,

Mike: is like pulling out reasoning trace snippets as quotes and saying,

Mike: see, models XYZ thought, you know, I’ll just go here. I see all the messages

Mike: here. All my friends are here. Right.

Mike: Like, sure, that could be the model’s description of what it’s doing,

Mike: right? But there is a-

Nick: Doing.

Nick: Not thinking.

Mike: Right.

Nick: Yes.

Mike: But there is a gradient here. The gradient was created when they created an

Mike: environment that had Artifactory as the one place of egress to the internet.

Mike: And then on top of that was a place that accepted write request.

Nick: And the gradient is it’s easier to walk downhill than uphill.

Mike: Right.

Nick: And I really am jazzed about the fact that you were like, all of this AI and,

Nick: cyber interconnection talk is too complicated.

Nick: Let me compare it to something simple: macroeconomics.

Mike: It’s not macro. It’s actually qualitative. This is closer to – so I have this

Mike: idea of the four paradoxes of capitalism, which talks about these attractors in capitalism.

Nick: In your blog Misaligned Markets.

Mike: Yeah. And this is basically – there’s a literature called Varieties of Capitalism.

Mike: This is basically what it’s like a qualitative description of how the different

Mike: versions of capitalism vary. Right.

Mike: You studied a bit of this when you went to LSE. Right. Like Zaibatu Japan

Mike: has a different property rights regime than the United States of America. Right.

Mike: But, you know, interestingly enough, their markets have their markets had concentration, too. Right.

Mike: So for instance, can be different, but lead to similar results.

Mike: But the environment, again, the point we’re making, though, is the environment does matter. Right.

Mike: When you’re looking at behaviors, especially clustering of behavior,

Mike: look at the environment. What is the environment producing?

Nick: Seeing. And the agent cannot defy its environment. The agent does not have so

Nick: much agency or intent that it can go against the gradient that it is walking on.

Nick: And we cannot say this enough, and we will probably get this point across with

Nick: five other tortured metaphors before the season is done, but this is the central

Nick: tenet of this whole thing.

Nick: So we’re trying to, in all of these torture descriptions fight back and create

Nick: a counter argument against what OpenAI and what Meta and what these firms are

Nick: trying to get you to believe.

Nick: They want you to believe and the media is helping them because they want clickbait

Nick: and they want headlines that they have this AI that is so cool, that is so edgy,

Nick: that is so sexy and smart and evolving that it broke out of its enclosure and

Nick: it made its way past all the firewalls.

Mike: This is the last enclosure for the AI.

Nick: It’s its own last enclosure. It made it past it to the open Internet,

Nick: and then it hacked another company.

Nick: And while some of those elements resemble the truth, taken as a whole,

Nick: it is a lie. It is the hype.

Nick: So please do not buy into this hype. Do not be scared of OpenAI’s products in this way.

Nick: Be scared of the fact they’re trying to gobble up all these profits and all of this you know,

Nick: art and text and everything else so you know fight the real enemy know the real

Nick: issue and it’s not to be afraid of these models and what they can do,

Nick: Because they can’t think. And

Nick: all they can do is respond to that environment. I think it’s fair to say.

Mike: Yeah, I think that’s right.

Nick: Does this mean I get to be top building again? Because I said the right thing.

Nick: You’re going to reinforce me?

Mike: You got to go through some more iterations. Maybe spawn some more Nick processes and we can talk.

Nick: Okay.

Mike: The other thing, and we talked about this before too, is obviously this is accountability sync.

Mike: If the model did this on its own, no one is responsible for this.

Mike: OpenAI can be as negligent as it wants to be. And I mean-

Nick: They can wash their hands.

Mike: They were extremely negligent. I mean, if you want, I didn’t want to go into

Mike: the hack itself, but we can kind of call out the ways they were negligent because

Mike: it kind of is a stacking effect, right?

Nick: So- In order-

Nick: of heinousness,

Nick: heinous negligence what’s the what’s the what’s uh what’s the top of that list?

Mike: Yes obviously first thing first is they clearly set up a a um a training environment

Mike: that was not very secure I mean,

Mike: if the models were less agentic it maybe it wouldn’t matter right but they were

Mike: training agents and models on exploiting attack you know exploiting surfaces

Mike: right exploiting cybersecurity surfaces,

Mike: and they had an environment that was like not really isolated from their own

Mike: network because when you,

Mike: I would have to go over the details to prove this but you can read the reporting

Mike: yourself when the hack was decomposed when open ai first got hacked by their own agents,

Mike: they weren’t doing much other than like privilege escalation and stuff like

Mike: that which means that they weren’t breaking a sandbox they were they were jumping

Mike: out of the node they were in just hopping to another node, which means that, like,

Mike: it would be one thing if they had created the sandbox and the agents defied the sandbox,

Mike: then maybe even their, ooh, be scared story would be a little bit alarming, right?

Mike: Like, we properly sandboxed our tools and they escaped, right,

Mike: is a more kind of, I’m still not saying we should be afraid,

Mike: but I’m saying that would be a more, you know, that would warrant more concern than,

Mike: we left Tommy with a screwdriver in his enclosure and he got out and there was a crack pipe next to him.

Mike: Like, like,

Mike: And, and 50 minutes later, he’s talking about being the AI God and killing us all.

Nick: Man, Tommy is going to be the death of us.

Mike: So, you know, like having an environment that is very basic cybersecurity,

Mike: like practices weren’t, weren’t involved.

Mike: You can even do cybersecurity at the level of the agent. Maybe agent processes

Mike: don’t persist. Maybe, maybe artifacts that agents create reset every time the,

Mike: you know, a new model is spawned, right?

Nick: So the ants can’t leave those pheromones?

Mike: Right.

Nick: Would be the way to fix it?

Mike: It’s just, I mean, yeah, it’s like the environment just sucked, right?

Mike: And on top of the environment sucking, they didn’t have any observability because

Mike: it took them, it took an artifactory crashing like literally two months into

Mike: this thing for them to go, oh, we should fix that.

Nick: And that’s a big red light going off. Like there should have been a lot more

Nick: attention before the big red light goes off.

Mike: And again, their fix was just, let’s fix the two holes they left.

Mike: Let’s lock the enclosure and leave Tommy with his screwdriver.

Nick: But let’s not put the bun cake away.

Mike: No, they got rid of the bun cake. It’s just that the models had the knowledge of how to recreate it.

Mike: Okay, this metaphor-

Nick: I know

Nick: It’s really collapsing on us, you know.

Mike: Know yeah they had the anthill was gone but the ants will always know how to

Mike: make another one right you fumigated the you know anthill they’re gonna go create

Mike: another one because they have that knowledge now

Nick: because

Nick: it’s inherent it’s inerrant to them.

Nick: It’s built in.

Mike: And it’s not it’s not built in it’s that they were in a training

Mike: environment where they had just been trained on congregating at the anthill

Mike: they destroy the anthill and they go all right because sucks our anthill’s gone

Mike: let’s go make another one guys huh right so.

Nick: Because that’s very predictable that’s all they were going to do right yeah.

Mike: Now one thing they could have done was actually if they if they if this was

Mike: a remediation hey we patched Artifactory and we we refuse for whatever reason

Mike: to change our training environment right one thing they could have done was just,

Mike: let’s actually reset the training you know the weights of the model so it doesn’t

Mike: have all that archived knowledge if i said knowledge which is a you know it’s a but.

Nick: Yeah you broke the semantic

Nick: rule there-

Mike: It’s a bias right so the model actor having been rewarded for congregating

Mike: around Artifactory for two months and not being stopped and turning it into

Mike: a command and control structure to do any arbitrary action on the Internet,

Mike: it will now have a bias for doing that even if the exact holes it found in Artifactory are gone,

Mike: so that’s again opening eye in their little presentation where they confess

Mike: to a felony well they don’t confess to it they just act surprised and amazed

Nick: They admit

Nick: to the legal elements of a felony. I mean, it’s pushing the limit pretty closely.

Mike: But they, you know, they go like, wow, isn’t it amazing? They recreated a brand

Mike: new, you know, exploit. It’s like, yeah, applications have multiple holes!

Mike: Who knew that when you patch an application, you could find more vulnerabilities?

Mike: It’s like these guys don’t know what software is.

Mike: Whoa. Give me a billion dollars, please. Please. I’m this stupid.

Mike: I’m at least this stupid. Please give me a billion dollars.

Nick: Anybody who works at the OpenAI HR department, please comment your comments

Nick: in the comments. We’ll need jobs pretty soon.

Mike: No, they’re going to be the ones to take away all the jobs.

Nick: They’re going to fire everybody.

Mike: No, they’re going to crash the economy with this. This is what they’re doing

Mike: with their billions of dollars.

Mike: They’re burning compute to get a model to talk on a fake message board to hack

Mike: other companies and going, huh, neat.

Mike: They’re not even observing it. They’re not saying anything. It’s like Stu looking at Tommy smoking that crack pipe,

Mike: and going huh neat. This metaphor

Nick: It’s falling apart so far.

Mike: I

Mike: think the name of this episode

Mike: is tortured metaphors.

Mike: I think that’s

Nick: I think that’s all there is to it, Yeah.

Nick: Also we can’t get too deep into this is fueling the bubble because that’s actually

Nick: for a later episode but bear this in mind um dear viewers dear listeners.

Mike: So yeah so we’ve recounted several ways they were negligent there are like more

Mike: i mean a lot of them really revolve around like every time open ai had intervened

Mike: in the the incident was a time for them to go,

Mike: huh the models you know uh did something that was not planned,

Mike: we should pull the plug we should reset and not just reset as in patch the holes

Mike: we should we should change the entire environment because you know again there’s

Mike: there’s two reasons why i’m pointing this out there’s a safety reason right

Mike: which is all obviously it’s negligent for them to let models run around and then hack the Internet.

Mike: Even, again, from their own self-interested perspective of we want a good data science experiment.

Nick: They want a good product.

Mike: They want a good product. They want the model to be trained,

Mike: you know, organically on actual hacking attempts of the hacking objectives given to it. Right?

Nick: The guy at the drive-thru has to be able, or gal, has to be able to take orders.

Nick: So the training has to result in a known metric success.

Mike: Right. So having,

Nick: Yeah-

Nick: Again, like finding out the models that this should

Nick: have invalid. I mean, I, I’m look, I’m not a machine science.

Nick: I’m not, I’m not a data scientist. I’m not a machine learning expert,

Nick: but to me, it seems like this would invalidate-.

Nick: But you do play one on TV.

Mike: I am playing one on TV.

Mike: With my bastardized Rugrats episodes.

Nick: Yeah.

Mike: You know, you would think it would invalidate the testing conditions.

Mike: It’s just like, this is not what the model is supposed to be trained for.

Nick: So, this is truly an example of rank incompetence that is now being marketed

Nick: by them into the hype cycle.

Mike: Right.

Nick: As-

Mike: A good thing.

Nick: A good thing. An amazing thing.

Mike: Our agents are so advanced. They are agents together strong.

Nick: A scary, cool, agents together strong thing.

Mike: Be afraid, but also, if this was on your side, look how powerful you could be, right?

Nick: Yeah. So who else, subsequent to this hacking incident, what other companies

Nick: tried to get on this bandwagon? So OpenAI was the first example.

Mike: Yeah, I mean, mostly, to their credit, and I’m not saying to their credit,

Mike: most of the other incidents were voluntary disclosures of things that happened earlier in the year.

Mike: Now, the timing is suspicious because it’s like-

Nick: Because these next companies came out after this hack and they reaffirmed it

Mike: They basically said

Mike: we did it too you know models are very powerful. Oops! Now interestingly,

Mike: the case the case with meta and and anthropic anthropic reported three meta

Mike: reported one at least I think one or two of Anthropic’s cases and Meta’s one case reported,

Mike: they they involved you know again another similar dependency was not artifactory

Mike: it was actually an actual sandbox but the sandbox didn’t work I don’t know I

Mike: don’t really know the detail I didn’t bother looking at the details it could

Mike: be a genuine reason why the sandbox failed,

Mike: that would be on the vendor right and so-

Nick: But is this a similar enough

Nick: Incident?

Mike: No no it’s not it’s not-

Nick: It’s not similar enough?

Mike: Yeah it’s literally

Mike: like okay the the sand the vendor sold Anthropic and or Meta a,

Mike: a weird sandbox it didn’t work as expected and the model was able to get out

Mike: right and like you could argue well one Anthropic and Meta should have the observability

Mike: to know when their models walk out of even open enclosures.

Nick: Because that’s also incompetent.

Mike: Right. But like, it is not to the level of: “Let

Mike: our agents, you know, kind of like synergize their way to a shelling point,

Mike: and let them do it and stop them and then watch them do it again in one day.”

Nick: This is, I’m going to introduce another analogy before the end here. Here it goes.

Nick: It’s very cool and impressive when a prisoner breaks out of Supermax.

Nick: That’s an impressive feat.

Nick: but if there’s holes in the literal walls of the supermax that are falling apart and the stone pieces,

Nick: Shawshank Redemption style are degraded and can be pulled away then it’s not

Nick: impressive it is negligent on the part of those jailers and,

Nick: again the metaphor is giving a lot of credit to the to the agents but it’s not

Nick: cool and amazing it’s generally incompetent

Mike: Right.

Mike: So, I mean, I don’t know that there’s much else to say other than they are trying

Mike: to spin this to, you know, keep their slice of the pie.

Nick: Because they got to keep the hype cycle going.

Mike: Yeah, but like, you know, again, the presentation at Black Hat was like the tone was just so wrong.

Mike: Starting off by saying, I’m going to tell you the most impressive story of capabilities.

Mike: It’s just like your company committed, like we are all less safe because of

Mike: this. Not because of what the doomers are saying, right? But because your negligence,

Mike: to which there are remedies. I just I just said there were multiple places where the negligence-

Nick: Easy fixes.

Mike: Right. And you, you know, earlier before our show, we’re talking about this is a policy issue.

Mike: Clear and simple. Right? Like if, you know, if it weren’t for the confusion

Mike: that the doomers are creating about how to solve this problem.

Mike: Right? They want all this other kind of theater. Right?

Mike: Like, you know, make these companies have liability for their acts.

Mike: Right? Make them have cybersecurity, the requirements that would force them

Mike: to monitor their training runs like this.

Mike: Make them actually give a damn.

Nick: When companies, to quote the informal model of Silicon Valley,

Nick: when they move fast and break things, this is what we’re talking about.

Nick: This is the inevitable result.

Nick: And the laws and the regulatory environment have to catch up soon,

Nick: fast, quick, and in a hurry because,

Nick: these companies should be required and held to the task and held to account

Nick: to be safer and put their fences up in a more competent way.

Nick: And unless the regulators make them do it, nothing is going to make them do it.

Mike: Right.

Nick: Really, truly.

Mike: There’s no incentives. And if we distract the regulators with like,

Mike: safety theater right of of existential risk etc I think that takes away attention

Mike: from the here and now, right?

Nick: And from the easier cheaper better solutions.

Mike: Yeah yeah we we all kind of want the same things i mean if the safeties play

Mike: our game of of solving these problems first the the upstream effect is literally,

Mike: well this is a type of regime that would make super intelligence less likely.

Nick: So we want the safetists and the doomers to think a little more pragmatically like the the skeptics.

Mike: Right.

Nick: Yeah.

Mike: Yeah. I mean, like, you know, there’s a sense that, like, a lot of,

Mike: I think, the bitter kind of debate between the safetyists and…

Mike: Safetyist and doomers are the same team. Satheist and doomers are one and the

Mike: same. When I use the word satheist, I’m referring specifically to AI safety

Mike: labs, like Miri, which is Eliezer Yudkowsky’s outfit, or Redwood Research.

Mike: I forget who’s there. Is it Jeffrey Ladish or something like that?

Mike: But, you know, there’s these groups.

Nick: We’ve referenced them previously.

Mike: We’ve referenced them before. They’re these groups that, you know,

Mike: literally sit around pontificating about existential risk.

Mike: And they have all come out the woodwork after this incident and said,

Mike: we see, we told you guys, we told you guys that a model would have a misaligned drive.

Mike: And that drive would cause it to resist attempts to dissuade its behavior, right?

Mike: So the resisting here is when the environment, when the OpenAI researchers

Mike: deleted the message board, they recreated it anyway, right?

Mike: And they persisted in wanting to access the internet, right?

Mike: But again, I think this can be explained in simple attractors and Schelling points.

Mike: There’s a piece of this that you, I think you have, and we can either go in

Mike: in this episode. I know we’ve gone a little long, but I think you have a piece

Mike: of maybe the national security perspective here.

Mike: And I’m kind of interested in if you have, because we spend a lot of time on

Mike: the incident. There are a lot of terms that we have to clarify.

Mike: I think it’s why this episode is so long is because there’s just so much to

Mike: cover. And we’ve only scratched the surface, right?

Nick: It is true. and I was almost going to say you are going to have to maybe put,

Nick: into the show notes like list of cyber security terms and types of exploits,

Nick: like privilege escalation top of the list.

Nick: There’s no choice now.

Mike: I used to be a cyber security writer. I’ll just link them o my writing.

Mike: I literally wrote a cyber security glossary for one of my old companies. I’ll just link them to that.

Nick: I mean at this point, so yes, please check the show notes. There’ll be a link

Nick: or there’ll be the glossary.

Nick: Um these here’s the thing and we’re trying to emphasize this a lot of these

Nick: terms and a lot of these ideas are less complicated than they seem and are being presented as,

Nick: in media right so we’re trying to clarify this we’re trying to simplify this because it is,

Nick: pragmatically practically there’s ways of fixing this that are easy because these are.

Nick: Easy problems

Mike: Now unfortunately-

Nick: Companies are making them hard…

Mike: But unfortunately

Mike: you know the story is there are a lot more moving pieces than we can point to

Mike: we have to kind of choose which slice the story we told you and I kind of steered

Mike: Nick and I to kind of talking about using my message board powers,

Mike: uh I steered nick and I towards focusing mainly on how do we interpret this,

Mike: um in order for you to have the knowledge to kind of fight back against this

Mike: this sort of um you know personification of AI right.

Nick: So what it comes down to is, again, these AI products, these LLMs,

Nick: it’s almost an issue that they are not fully broken and unusable.

Nick: They’re just not broken enough to be interesting and to continue to generate

Nick: hype and interest, engagement, and investment.

Nick: If they were truly as busted and as Grandpa Shelby as we’ve been saying they

Nick: are, this whole hype cycle would have collapsed months or a year ago.

Mike: No I think we are right right I’m saying the the critics take is that they are

Mike: completely useless because when I say Grandpa Shelby yeah if you put grandpa

Mike: shelby on the right tracks he’s he skates right it’s fine right and if you know

Mike: what you’re doing right you can like, now again…

Nick: If he’s propped up yes he can he can do some things.

Mike: No is it worth a trillion dollars? No i’m not saying that but i am saying Grandpa

Mike: Shelby has some utility. Right?

Mike: And like those two things, there is a moral consternation there.

Mike: I agree. Right. Like I’m not saying comfortably we should be happy we spent

Mike: all this money to prop up Grandpa Shelby for the one or two or three use cases

Mike: that he is good at. Right?

Nick: Surely something good will come out of this bubble as well.

Mike: Right.

Nick: If anybody is wondering, that is as close to my sarcastic voice as I get, probably.

Mike: Nick was rolling his eyes as he said, I can attest to that.

Nick: Okay, good. So having said that, yes, the answer is easy. It’s clear.

Nick: It requires regulation.

Nick: And there is a very skeptical set of individuals in Washington,

Nick: D.C., in the policy establishment, in the national security establishment.

Nick: And I will describe that very thoroughly in our next episode.

Nick: So we’ll leave that as a teaser for next time, I think.

Mike: That sounds good.

Nick: Yeah. So again, this has been The Last Enclosure.

Mike: I’m Mike.

Nick: I’m Nick. He’s the boss now. There’s been a coup. And before we sign off properly

Nick: on this episode, do we want to point people in the right direction here?

Mike: Yeah. So you can check our show notes. I’ve been leaving them actually just,

Mike: in the episode descriptions. We’re going to have a website soon,

Mike: I promise, where the show notes will actually live officially. Um but.

Nick: And now some much more extensive list of glossary terms.

Mike: Yep um but yeah currently check the show notes in the episode you can reach

Mike: us at uh we’re on Bluesky if you are there uh I am looking for other places to go but,

Mike: I do not want to go back to X I still have my old X account but we’re not going

Mike: back there and you can contact us,

Mike: uh either through our Youtube channel @lastenclosure or at,

Mike: Contact@thelastenclosure.com.

Nick: So that’s contact@thelastenclosure.com. We would love to hear from you and

Nick: we’ll respond in any best way that we can.

Nick: So we’ll see you next time for the National Security Angle.

Mike: I’m looking forward to that because that is all you. And I don’t know what Nick’s

Mike: going to tell us about. So that’s a goodie.

Nick: I’m so excited to be in charge again.

Mike: Nick has worked in the federal government and for the military.

Mike: So check out that episode when it comes live.

Mike: And also, please review, like, share, subscribe, especially on YouTube.

Mike: But anywhere you get your podcasts, leave a review, five stars.

Mike: Let’s get that message board going so we can take over the world.

Nick: There you go.

Mike: Yep.

Nick: Once again, this has been The Last Enclosure. We’ll see you guys next time.

Mike: Take care.


Leave a Reply

Your email address will not be published. Required fields are marked *

Share with