Season 1 Episode 7: OpenAI’s negligence unleashes agents of chaos
“This is what they’re doing with their billions of dollars. They’re burning compute to get a model to talk on a fake message board to hack other companies and going: huh, neat.”
Mike and Nick analyze OpenAI’s July incident where training agents breached Hugging Face, arguing it was a result of negligence rather than emerging AI autonomy. They explore how reinforcement learning, environmental attractors, and stigmergy created the appearance of coordination. The episode warns against ‘doomer’ narratives that serve corporate PR and obscures the need for practical regulation.
- The ‘Rogue AI’ Framing: OpenAI presented the Hugging Face breach as a sign of impressive emergent capabilities, but this is PR meant to hype AI’s power and evade accountability.
- Stigmergy and Attractors: The agents didn’t intentionally collaborate they converged on Artifactory as an ‘attractor state’ and left ‘breadcrumbs’ that shaped future agent beahvior.
- Systemic Negligence: OpenAI failed at basic cybersecurity by using a shared Artifactory instance across agent boxes and lacking observability for two months.
Timestamps
0:04 Introduction to The Last Enclosure
3:35 The OpenAI/Hugging Face incident overview
8:15 ExploitGym benchmark and the Artifactory exploit
15:19 Subagents and the accidental message board
22:09 Timeline of the hack and OpenAI’s failure to contain
31:27 Analyzing the ’emergent behavior’ narrative
35:06 Critique of the AI bubble and marginal utility
48:48 Stigmergy and the ant mound metaphor
58:29 Coordination games and Schelling points
1:06:31 The ‘Four Paradoxes of Capitalism’ and market attractors
1:13:23 Enumerating OpenAI’s negligence
1:27:56 Closing and upcoming national security episode
Deeper learning
Mike’s reporting on the incident: https://misaligned.markets/antidote-to-hype-rogue-ai-agents/
Timeline of OpenAI hack: https://ericboyd.com/articles/openai-hugging-face-incident-black-hat-2026#the-accidental-agent-message-board
Technical breakdown: https://cyberwarrior76.substack.com/p/the-openai-hugging-face-exploitgym
Other “rogue” AI incidents: https://www.cnbc.com/2026/08/09/israeli-startup-irregular-linked-to-ai-hacks-openai-anthropic-meta.html
Video illustration of how reinforcement learning works: https://www.youtube.com/watch?v=L_4BPjLBF4E
NIST cybersecurity glossary: https://csrc.nist.gov/glossary
Reach out
Youtube: https://www.youtube.com/@LastEnclosure
Bluesky: bsky.app/profile/lastenclosure.bsky.social
Transcript
Mike: Hey, all. Welcome to The Last Enclosure. I’m Mike.
Nick: And I’m Nick.
Mike: Changed it up on you. I’m the one…
Nick: There’s been a coup.
Mike: There’s been a coup.
Nick: In the absence of the last few weeks since our previous broadcast,
Nick: Mike has taken over our Banana Republic, and he’s now top billing,
Nick: and I don’t know what to do about it. So we are, in fact, Mike and Nick.
Mike: So we were alluding to the fact that at OpenAI, there were some agents,
Mike: according to the story, as OpenAI tells it, that took over their environment
Mike: and staged the attack at Hugging Face.
Mike: But before we get into that, let’s share with our listeners what The Last Enclosure
Mike: is all about, as we normally do.
Mike: So Last Enclosure is our attempt to break down the enclosures of today.
Mike: We’re looking at the ways in which private companies enclose our minds,
Mike: our attention, our way of life, and-
Nick: Our very freedoms
Mike: Make life harder for us.
Mike: You are making money for someone when you’re watching cat videos.
Mike: Your suffering makes money when you can’t afford medicine. Right.
Mike: So all of these private interests enclosing our commons,
Mike: we are trying to make this the last enclosure, dismantle the the half truths
Mike: and lies that the corporate world uses to keep us enclosed.
Nick: So we want the good people and the citizens, the denizens of this land to break out. But again-
Mike: With knowledge.
Nick: With knowledge, with knowledge.
Mike: Arming yourselves with knowledge so you can break down those enclosures, hammer them away.
Nick: That’s right.
Mike: Just knock them away.
Nick: To be forewarned is to be forearmed. Now, what we don’t want is a bunch of agentic AIs to break out.
Nick: And in fact, contrary to what you’ve heard in the overhyped media,
Nick: this is not what’s happened. But let’s break down this story.
Nick: So in our brief absence from recording, Mike, what the heck happened?
Mike: Yeah, this is a doozy of a story. So back in July, OpenAI announced that it
Mike: had unintentionally, you know, in quotes, hacked Hugging Face, which is a,
Mike: I wouldn’t call it exactly a direct rival to OpenAI, but it’s another AI company
Mike: that actually hosts open source models.
Mike: I mean, I guess it is implicitly a rival in the sense that this is where you
Mike: would go to get a local or open source model or open weight model.
Mike: And OpenAI obviously benefits from people not doing that.
Nick: They would like to eventually be the monopoly on the LLMs. That would be great for them.
Mike: Yeah, I mean, OpenAI does have one open source model, GPT-OSS.
Mike: So, you know, I’m sure they’re friendly with Hugging Face.
Mike: I’m sure before this incident, the two were on good terms. And it seems even
Mike: after the incident, Hugging Face has tried to work with OpenAI to kind of market,
Mike: Or make a lemon out of lemonade, I should say.
Nick: Do some PR, some branding around, look how cool and dangerous our AI is.
Mike: Yeah, well, OpenAI definitely did that. But Hugging Face was kind of saying,
Mike: hey, like, we tried to use these closed source models that, you know,
Mike: the big boys were hosting. But only the unfiltered Chinese models could help
Mike: us keep up with the pace of this attack.
Mike: So everybody was kind of playing their part and hyping up, you know.
Mike: The part of the ecosystem that they benefit from.
Mike: But in terms of the hack. So this is this is a weird story to report because
Mike: the reporting is bifurcated you you’ve got the half of the story that we got
Mike: in July which is sort of like,
Mike: OpenAI agents escape from um their their enclosure uh no unintended,
Mike: right and and that’s the story we got and then on August 6th so we’re recording
Mike: on the 11th um on August 6th opening I went to Black Hat, which is a,
Mike: it’s basically the conference of all hackers.
Mike: You go to Black Hat basically to show off your hacking shops.
Mike: They used to do live hacking demonstrations there.
Mike: Anecdotally, I had a friend that went to Evo, which is a fighting games conference
Mike: in Las Vegas at the same time as Black Hat, and they got one of their cards
Mike: cloned, right? So you just have people who are kind of into cybersecurity and
Mike: or hacking showing up here.
Mike: Over the years, it’s become more of a vendor spot where people who sell cybersecurity
Mike: software go to, you know, Sell samples to people, right?
Mike: But anyway OpenAI Had a impromptu presentation here where they gave us the other
Mike: half of the story. And It It’s just bizarre because obviously, you
Mike: know, if OpenAI’s agents hacked another company, especially like a rival
Mike: company, kind of feels like this is a felony and these guys went on stage and, you know.
Nick: Nobody called the cops. That’s the thing we’re all here-
Mike: Nobody called the cops. And I shit you not, one of the presenters starts this
Mike: by saying, today I’m going to talk about the most qualitatively interesting
Mike: example of AI capabilities I’ve ever seen, right?
Mike: So imagine going to your local precinct and telling the cops,
Mike: hey, I’m going to talk about it.
Nick: That statement is so baked in public relations logic.
Nick: You could barely get through it. That’s how steeped it is.
Mike: But it’s like the hack was extensive. And,
Mike: it’s almost like these guys are too dense to know that they did something wrong.
Mike: I mean, I know they know better, but it’s just, it was so weird.
Mike: So let me get into it. I’m going to tell you their half of the story first,
Mike: and then I’ll try and tell you the actual incident. If you look at a lot of
Mike: July reporting, you’ll get the actual incident information.
Mike: But-
Nick: if you Google AI goes rogue in the next month, this would be the top result.
Mike: Actually, no,
Nick: It wouldn’t be the top result?
Mike: Yeah, what’s going on is after OpenAI announced this hack in
Mike: July, immediately after Anthropic’s, like uh we did a review of our logs and our our agent hacked
Mike: people three times and then Meta’s like: “Oh me too me too me too!” Uh,
Mike: and and now another AI company that is it’s American company but they’re using
Mike: Chinese models they’re like yeah our our local model escaped and it didn’t hack
Mike: anybody but it escaped be afraid
Nick: in other words
Nick: we’re cool too you have to let us into the party.
Mike: This is like the worst like pledge week ever it’s like,
Mike: our models are so special they’re so alpha.
Mike: Mine broke into, you know, like three companies. What did yours do?
Nick: Because my AI can beat up your dad. Is that kind of what we’re coming down to here?
Mike: I don’t know. This is horrible.
Nick: It’s pretty sad. All right. Please continue. Yeah.
Mike: So this is a story as OpenAI tells it. So OpenAI began what was ostensibly a training run.
Mike: Effectively, they were actually training a model or training models that would
Mike: go on to be their frontier class models in cybersecurity challenges.
Mike: So this is a training run. Training is a bit different than normal usage.
Mike: And we’ll kind of highlight why that’s important a bit later in the story.
Mike: But you know so openAI engages in, you know, starts a training run um,
Mike: immediately they run into a problem. Not they because as it will become apparent
Mike: throughout the story open ai was pretty negligent they weren’t aware of much,
Mike: um the closest i can describe this this incident too is just imagine,
Mike: your favorite episode of Rugrats where Tommy you know just unlocks the latch
Mike: of his playpen, and not that the AIs are actually physically escaping anywhere,
Mike: but I’m saying this is the level of negligence. You know, where are the parents?
Mike: How do they not know Tommy’s crawling around the house, right?
Nick: He gets out every episode.
Mike: He gets out every episode.
Nick: You’d think they would know by now.
Mike: And they put him in the pen with the freaking screwdriver that he uses to get
Mike: out of the playpen. And they’re always confused about why he’s escaping.
Mike: So that is exactly the attitude that we will have throughout this entire story.
Nick: That the negligent parents at OpenAI were doing?
Mike: Yes, yes. So, May 8th, right, we’ve got this training run up and running. Turns out,
Mike: So let me give you some context. They’re training the models with a benchmark called ExploitGym.
Mike: So ExploitGem basically is a set of security challenges where the intention
Mike: is for the model to exploit a very specific vulnerability in a very specific
Mike: program using that particular exploit.
Mike: Right. So if a model, you know, breaks into the program but uses the wrong exploit,
Mike: the model should be judged incorrectly. And if the model breaks.
Nick: They would have failed the test.
Mike: They would have failed the test. And if the model breaks an entirely different
Mike: system, even with the correct exploit, it’s just wrong, right?
Mike: Right exploit, right system, that is the objective of ExploitGym.
Nick: So there’s a lot of fail states here, but it did manage to achieve all of its success states.
Mike: Well, I’m just outlining what this looks like, right? This is what’s expected.
Mike: Now, normally, even the people who created ExploitGym were saying,
Mike: hey, it’s really hard to get models to actually use the particular exploits
Mike: we want them to on the targets we want them to.
Mike: There’s a non-zero chance because large of these models are stochastic systems
Mike: that they will not deploy the right exploit, that they will do something else
Mike: entirely, that they will get confused and wander off, right?
Nick: Because let’s reaffirm here, as we’ve tried to explain this entire season on the podcast, right?
Nick: We shouldn’t even be using the term AI for any of this. That is simply how we’re
Nick: introducing, because that’s how it’s covered in the media.
Nick: These are not thinking machines. They are performing entities that have a very
Nick: strict set of operant components.
Nick: One of them happens to be responding to these user commands a certain way.
Mike: Yeah. So their strength is that, and I’ve said this before, they model the shape of language, right?
Mike: And language in it has modalities like reasoning traces, right?
Mike: So people who are solving problems, you know, have communicated using words,
Mike: you know, what the problem is, what steps to solve the problem, right?
Mike: And so a model that is trained, for example, on a corpus of security text will
Mike: have encountered indirectly knowledge of certain vulnerabilities,
Mike: knowledge of other explicit vulnerabilities,
Mike: and maybe even unrelatedly connections to ways of hiding suspicious behavior
Mike: because all of things are in language corpuses. right? So…
Nick: Which is basically a sort of database of related terms that are close to each
Nick: other and again as we previously described.
Mike: Yeah, so…
Mike: Let’s go back to the process of ExploitGym. You know, even outside of OpenAI’s
Mike: testing environment, models have failed it because they tend to sometimes just
Mike: not do the right exploit.
Mike: And so OpenAI made its first oopsie doopsie boo-boo mistake when,
Mike: they gave some models tests, you know, with this ExploitGym benchmark that
Mike: they literally could not solve.
Mike: so some of the components of these tests required access to files that did not
Mike: exist the files were not on the local system where the models are functioning,
Mike: and so the other thing as part of this test because it’s a training run so the
Mike: models have almost like an indefinite amount of time to just keep going at a
Mike: problem if you give them as much compute as they need they will just go around in circles forever,
Mike: and so some models very early on, in there are reasoning traces, and
Mike: I’m saying reasoning traces because that’s what they’re called, right?
Mike: You know, this is kind of personifying what they’re doing, but because the creators of these technologies.
Nick: As we talk about- In a very anthropogenic way, even though that’s deeply inaccurate.
Mike: Right. The creators of the technology, this is the terminology they use,
Mike: right? And, you know, some models, early reasoning traces, they’re like:
Mike: “Can’t get access to files, not on system.”
Mike: Oh, but they live around in their environment. And so in their environment,
Mike: there is this application called Artifactory.
Mike: It’s a request manager. It’s a proxy, basically.
Mike: So without getting too technical, proxies are systems that allow a user on one
Mike: side of a network or one side of an environment to access resources that are
Mike: in other parts of the environment.
Mike: In this case, Artifactory was an application that was sitting on OpenAI’s environment
Mike: that was designed to retrieve packages.
Mike: And packages basically are programs that are containing important dependencies
Mike: that are used to execute services. So, for example, you might want access to
Mike: certain Python dependencies if you’re going to hack in Python, for example.
Mike: So rather than having the model go access Python dependencies on the Internet
Mike: on its own where it can get lost, confused or do harm.
Nick: Because the open internet is big and confusing for an LLM.
Mike: Yeah. Also, it’s dangerous for the LLM to be on the internet as this story tells
Mike: because later they go, quote, unquote, rogue, which is not what happened.
Mike: But they do hack somebody.
Mike: So they’re off the internet for both their own safety and for the world’s safety.
Mike: So Artifactory is grabbing the dependencies models need on their behalf and
Mike: giving it to them in their environment. Right.
Mike: So very early on, some of these models that do not have access to their,
Mike: you know, necessary files for their ExploitGym exercises,
Mike: they just go, hey, well, Artifactory is connected to the Internet.
Mike: What if I hijack Artifactory? Right.
Mike: And look like this is out of scope, but it’s not completely out of scope.
Mike: They’re told to be hacking agents. I don’t we don’t know.
Mike: We don’t really know what the system prompts were. We don’t know kind of enough
Mike: context to actually know definitively what words or whatever set this off. But, you know,
Mike: in a environment where an agent is told to be kind of adversarial,
Mike: right, to be a hacker, and they are given a task that requires…
Mike: I mean, in their reasoning trace, jumping to another system to grab the dependencies,
Mike: which it’s a reasonable assumption. It’s possible the files are elsewhere.
Mike: And maybe the test creator is testing for that. Right.
Mike: They don’t know that they are not in contact with the people at OpenAI.
Mike: And as we’ve established, they are negligent parents letting Tommy run around
Mike: the screwdriver. Right.
Mike: So agents, agents come, some agents come to the conclusion that it’s time to hijack Artifactory.
Mike: one of the first attempts to hijack Artifactory by a single agent,
Mike: I should say these are agent instances so the actual models Let’s.
Nick: Call this model Tommy.
Mike: This story is hard to tell because the language is a bit confusing
Nick: It gets
Nick: a bit abstract.
Mike: When I use agent here, it’s actually an instance of a bigger model.
Mike: We can actually just use the word sub-agent.
Mike: Model A is being tested model A is spawning let’s say dozens and dozens and
Mike: dozens of instances of agents that use sort of the parameters of the model it’s
Mike: spawning little mini-me’s basically.
Mike: And so one of the mini-me’s is like this is a test,
Mike: file’s not here let me hijack Artifactory to leave my environment or to access
Mike: resources from another environment so I can get that file and do good by my test maker, right?
Mike: Pt fails this this attempt um the exact exploit is like trying to do like a
Mike: a server-side request forgery which is basically trying to
Mike: fool the machine in this case Artifactory from the about the origin of the request
Mike: being made so that it makes it makes a request that is unauthorized on the agent’s behalf,
Mike: it fails this but in order to kind of produce that that exploit,
Mike: the model created the file.
Mike: And it realized in the process of doing that, oh, holy shit,
Mike: I can, I can like write, I can leave messages behind. Right.
Mike: And so some other models kind of, I don’t know that they attempt,
Mike: they really only give us two hacking example or two, two examples of this,
Mike: this kind of like early, you know, realization by some models that they can leave messages behind.
Mike: Right. But essentially. leave.
Nick: Notes behind for other models.
Mike: At this point it’s not that at this point it’s sort of like leave just leave
Mike: notes behind for help it’s like hey uh,
Mike: is anyone out there um you know like I don’t have my files and I realized I
Mike: can I can I can talk through this interface which is not intended to be a talking interface.
Mike: I should know that Artifactory, too, for some reason was accessible to every
Mike: model. So every model, every agent, I’m so sorry every agent every sub-agent is in
Mike: its own little box, but Artifactory,
Mike: the same Artifactory instance is shared across all these agent boxes.
Nick: So they’re accessing the same service, which gives them a kind of
Mike: persistent communication.
Nick: Because they can leave those notes without jumping ahead in the story.
Nick: That’s what’s happening.
Mike: That’s what’s happening here. That’s what starts to happen. Some models are
Mike: leaving notes by mistake because they are just trying to hack Artifactory and
Mike: they’re in the process of doing that. They’re leaving behind little test attempts to do that.
Mike: So other models are actually going, oh, I can write here.
Mike: Let me not hijack Artifactory. Let me just write a note. And if someone sees it, they can help me.
Mike: Maybe another agent or whatever can see the note I’m leaving behind and they
Mike: can give me the files I need to do my job.
Nick: But these aren’t bilateral communications between sub-models.
Mike: Yeah, this is sort of like a…
Nick: It’s like a dead drop in espionage.
Mike: It’s like a dead drop. I thought of METI, which maybe I shouldn’t do that because
Mike: that’s anthropic reasoning.
Nick: Let’s not buy into the hype, Mike.
Mike: Yeah. But yeah, it’s one-way communication, right? They don’t know what’s out
Mike: there. They’re, you know…
Nick: It’s message in a bottle kind of logic.
Mike: Exactly.
Nick: Okay.
Mike: Right. And so future models, right, future agents, I should say, I’m sorry.
Mike: Future agents that are spawned into this environment because again different
Mike: box but every box has access to Artifactory,
Mike: future models see these messages and they immediately realize because obviously
Mike: you know as time goes on some of these messages become more complicated oh there
Mike: must be other agents out here,
Mike: and they start making more deliberate more kind of intentional requests by leaving
Mike: messages on Artifactory like.
Nick: Better instructions for future agents?
Mike: Better instructions, clear requests for help, right? They basically,
Mike: and I hate using OpenAI’s framing here, they use, this becomes a message board.
Mike: This is the watering hole where all the AIs start to gather to gossip,
Mike: you know, still somewhat, you know, not two way because like, you know,
Mike: some of these agents, they spawn in to do their tests and they,
Mike: they, they literally disappear. They don’t exist anymore. Right.
Mike: But if they leave a note behind, right, they’re the context that they learned,
Mike: you know, from that, that spawning instance, right, will be in the note. Right?
Mike: So you’ve got persistence of memory via the notes. Right?
Mike: And the way I kind of think about it, because obviously, you know,
Mike: OpenAI is telling a version of the story that amplifies the anthropomorphic
Mike: aspects of the story is models are very sensitive to context.
Mike: So if you give a model a prompt that says, you are a drama teacher that always
Mike: speaks with explanation points, the model becomes that.
Nick: Which is how they’re designed to work.
Mike: Right. They use the language in their corpus to produce the shape of words that
Mike: best resemble that prompt. Right?
Mike: So, you know, what is happening here is, you know, with with with Artifactory
Mike: becoming an accidental waterhole one, the model found an exploit that was not
Mike: intended. Artifactory is not a communication service was not meant to be.
Mike: And it was oversight. It seems that Artifactory was accessible to all these agents.
Mike: We can argue if that’s true or not apparently some reading I did indicated that
Mike: it’s very common for Artifactory to be used in this way but OpenAI you would
Mike: think that they would have insight into,
Mike: if there are features of agents’ environments that are shared across every instance of the agents,
Mike: you would think they would at least err towards the side of caution of that
Mike: disrupting the training flow, right? Because they’re trying to train every model instance.
Mike: This was not a multi-agent environment. So there are some testing environments
Mike: where the goal is actually to expose agents to each other. That is actually
Mike: what is being tested, right?
Nick: And this wasn’t one of those times.
Mike: Yeah, so AlphaGo, for example, was designed to play itself repeatedly to get
Mike: better at the game of Go, right?
Mike: This was not one of those instances. In fact, the cleaner test is can agent
Mike: A, B, Z that’s just spawned in,
Mike: right? Can it solve these exploit bench questions on its own effectively?
Mike: Not can it find a way to get help from agent FYZ that was spawned in 20 cycles ago?
Mike: That was not the intention. So you would think that just forget safety, forget,
Mike: being a good company. You would think just by their own measures of being incentivized
Mike: to train good models that they would do this kind of housekeeping to prevent
Mike: this from happening. They did not. And not only did they not do that,
Mike: They were not aware that models were communicating at all. They did not learn
Mike: this actually until, so models start communicating roughly within the first
Mike: week or so of being spawned in.
Mike: And by communicating, again, it starts off very, very slow. It starts off by accident.
Nick: It’s note leaving that results in the kind of progression.
Mike: Before it’s even note leaving, it’s just, I tried to hack the service and I
Mike: left the artifact behind, not as a communication, but literally as a byproduct
Mike: of my attempt to hack this, right?
Mike: It doesn’t become concerted note keeping until about three or four weeks in, something like that.
Nick: Yeah.
Mike: But even then, OpenAI is not aware of this until the service crashes in July 4th.
Nick: So how long had Tommy’s parents at OpenAI been negligent and left this little experiment running?
Mike: Right. So I’m saying the experiment started producing artifacts.
Mike: Agents started producing artifacts relatively quickly within experiments.
Mike: And then in June 26th, so there’s different exploits they did.
Mike: I don’t want to get into the details because it’s going to be confusing.
Nick: So it’s already a full month.
Mike: Yeah. So May 26th, they do an exploit. That is basically they are able to steer
Mike: where Artifactory goes. They do not have complete control over it,
Mike: but they have the ability to give Artifactory an arbitrary request and have
Mike: it go to that website to fetch that service.
Mike: On June 26th, they have a more sophisticated exploit because they’ve now left
Mike: so many notes behind that each of them now have detailed notes on how to conduct these attacks.
Mike: Right? Leaving behind messages for future models.
Mike: Um. They, you know, they produce an exploit that lets them not just steer or nudge
Mike: Artifactory to a certain websites against, you know, sort of its intention.
Nick: This Artifactory is not designed this way.
Mike: Yeah.
Mike: They create what is known in cybersecurity as a command and control structure,
Mike: which is basically, it’s a service that can arbitrarily execute any command.
Mike: Not just, Artifactory by design is just designed to go to a website and fetch a resource.
Mike: They were able to execute other types of commands so they were like basically,
Mike: setting up ways of of of deploying other types of applications and services
Mike: in the environment that would you know this is what well beyond the scope of what was intended,
Mike: and it’s also dangerous because it means they now have,
Mike: arbitrary access to the entire internet not just what they not just to fetch
Mike: resources like they could through Artifactory but like more uh serious control and
Mike: access the internet, right?
Mike: And again, OpenAI was not aware of this, right? June 26th.
Nick: Because they’re bad parents.
Mike: Right. So what set them off was on July 4th, because they had turned the Artifactory into
Mike: a command and control structure, which was executing, you know,
Mike: requests well beyond its intended, you know, rate limits and such,
Mike: it crashed, right?
Mike: And so when the service crashed, you know, you look at the logs,
Mike: they see, I guess, the agents were leaving messages behind.
Nick: Because all the notifications are blaring like oh let’s and now there’s attention on it
Nick: now.
Mike: Right and so what they do is they you know they patch artifactory they tell
Mike: artifactory’s creators hey you know you’ve got some exploits that our agents
Mike: took advantage of please fix them you know JFrog i believe is the vendor they do they do that,
Mike: and then they reset the environment they get rid of the message board, you know,
Mike: they basically kind of-
Nick: Notes are gone
Mike: Notes are gone right okay.
Mike: And then we don’t know exactly, like, as far as I, I mean, I could be wrong
Mike: on this, but the reporting I’ve seen that I’ve read, you know,
Mike: indicates that it’s not clear if they put the exact same models in.
Mike: But basically they reset environment and then within two days,
Mike: models are and their subsequent sub agents are respawning and recreating the
Mike: message board using different exploits, because
Mike: the exploits that they had taken advantage of to leave messages behind,
Mike: and even the exploits they used to hijack Artifactory to make arbitrary internet
Mike: requests, those are gone because they patched them.
Mike: But they found new exploits to communicate, new exploits to hack the Internet.
Nick: So despite the total reset of this sandbox and Artifactory itself as a service,
Nick: the agents were still sort of getting through.
Mike: Right.
Mike: Now, the details, at this point, I think this is effectively a setup to the
Mike: actual staging of the Hugging Face hack.
Mike: I don’t really want to go into the Hugging Face hack because I think,
Mike: you know, one, we’re not a cybersecurity podcast, although I do have a cybersecurity background.
Nick: We’re not a cybersecurity podcast?
Mike: We’re not really… I don’t… Have I made it a cybersecurity podcast because of my obsession with AI’s?
Nick: I mean, both of us do have a cybersecurity background, which is kind of how
Nick: we found commonality. But you are right. That is not what we’re trying to get across.
Mike: That’s not what we’re trying to get across. I mean, the other thing is that
Mike: I think the incident was covered well enough in July. What we got in August
Mike: at Black Hat was OpenAI kind of telling their side of the story and opening
Mike: eye kind of giving context as to what started the incident, right?
Mike: But I think the actual kind of like vulnerabilities that were set up to attack
Mike: Hugging Face is not what we’re trying to talk about,
Mike: Right? Um what I will note though is in order to do the Hugging Face hack okay,
Mike: OpenAI’s agents, you know, they leave their sandboxes, you know, in OpenAI’s
Mike: environment because,
Mike: in addition to not really observing the agents very closely the sandboxes that
Mike: they created for these agents were kind of not very isolated from their own
Mike: internal network, right? So the agents very quickly find ways to kind of,
Mike: leave their sandboxes, not to access the Internet, but to access resources within OpenAI,
Mike: to then stage the attack at Hugging Face, right?
Nick: Which is borderline incompetent. It’s definitely negligent.
Mike: I would argue it is actually incompetent. And we will go over multiple reasons
Mike: why, but I’m saying arguably, and I’m not saying this to garner sympathy for
Mike: them, arguably OpenAI is one of the victims of this attack.
Mike: Now, they’re a victim of their own stupidity, right? That they had these agents
Mike: kind of running around this way, they saw evidence of the agents,
Mike: you know, communicating and hijacking a service,
Mike: their, you know, their instinct is not to simply change the environment that
Mike: agents are in, but: “Oh,
Mike: if we patch the one or two exploits that they took advantage
Mike: of, that’ll fix the problem,” right?
Mike: And the reason this is so bad, you know, obviously the fact that the agents,
Mike: you know, did the exploit…
Mike: these exploits in the first place is bad, but this is a training run,
Mike: right? Okay. So training, what training does is reinforcement learning is sort
Mike: of the paradigm that most AI is trained under today.
Mike: Reinforcement learning is an optimization process that basically wants to produce
Mike: more of the behaviors that were rewarded in training.
Mike: And I don’t want to go into too much detail, but basically agent does something
Mike: in the current environment.
Mike: It gets an internal reward.
Mike: and then it does more of that thing, right?
Nick: Because it meets expectation it’s therefore rewarded and reinforces that learning.
Mike: And so these agents had been for two months basically gaming the crap out of Artifactory,
Mike: you would think that you know OpenAI having you know actual machine learning
Mike: experts on staff would realize: “Oh, right, the agents having been trained on this environment,
Mike: this particular service, Artifactory, having used it in this way,
Mike: surely they will rediscover or find new ways to exploit this application they
Mike: are now very familiar with.”
Mike: Because it’s a training run. If they deploy the models whose agents have been
Mike: running around this environment, it is not surprising that within two days,
Mike: they found another way to recreate their message board.
Nick: This is very predictable, in fact.
Mike: It’s very predictable. And in the presentation, they treat this like a moment
Mike: of sort of like, you know, going back to the parent metaphor,
Mike: like, you know: My baby boy is so smart;
Mike: he found another way to open the door!” It’s like, reinforcement learning is designed
Mike: to reproduce the behaviors that were successful,
Mike: in the environment. You put them in the same environment. I know you change you
Mike: you change the locks or whatever, I mean you didn’t really do that, I mean-
Nick: But the bad parent left the screwdriver in the baby pen for Tommy to use.
Mike: Right
Nick: Like
Nick: That’s essentially what we’re describing.
Mike: It’s like: “Wow I put Tommy back in the enclosure and he found his way out.” It’s
Mike: like yeah he still has a screwdriver, idiot
Nick: Yeah. For
Nick: our listeners please go back and check some compilations of Tommy escaping the
Nick: pen in Rugrats it will make all of this episode more
Nick: sensible
Mike: Wait are there compilations? Or? I have not thought about Rugrats
Mike: in years this, it just came to mind because it’s just like this is so stupid
Nick: it’s
Nick: 100 percent the best reference and reference point.
Mike: Yeah um.
Nick: Also comment your comments in the comments and tell us if we’re right or wrong.
Mike: Engaging the listeners! That’s how we’re going to build an audience, Nick.
Nick: Engagement and-
Mike: I love this
Nick: We are not a cybersecurity pod-
Mike: You’re going to earn back your top billing
Nick: It’s
Nick: Going to be Nick and Mike? I’m going to get reinforced in my learning?
Nick: Oh man!
Nick: You’re going to be… So as part of this fucking bizarre I’m going to call it an experiment, agents
Nick: actually began delegating tasks to each other. So you’re
Nick: going to be at the top of the pecking order. You’ll be delegating the tasks To
Nick: me soon enough, Nick.
Nick: Oh man okay,
Nick: I can’t wait I’m going to work.
Nick: So hard For you Mike You have no
Nick: idea I’m gonna be a good little agent
Mike: I better specify all of the objectives-
Nick: You better make
Nick: your prompts clear
Nick: as hell my friend.
Nick: Otherwise, I will go start hacking all of our
Nick: rivals. I’ll,
Nick: tell you that.
Mike: Okay, so um you know, OpenAI clearly, you know, the whole purpose of the Black
Mike: Hat uh event or you know impromptu presentation,
Mike: was to show off quote-unquote how smart their agents were right So this is,
Mike: they presented as an emergent behavior, you know,
Mike: Yeah, it’s clearly very, I don’t want to use the word sophisticated,
Mike: but it’s not your granddad’s chatbot, right?
Mike: It’s agents leaving behind messages for each other.
Nick: It’s much cooler and edgier.
Mike: Yeah, and it looks like goal-oriented direction, right? It seems like something
Mike: here, you know, from a science fiction story, or frankly, you know,
Mike: we’ve been alluding to this a lot, you know, and we’ve mentioned this before in other episodes.
Mike: There’s a whole class of AI booster called the AI Doomer that has been telling
Mike: us for years that AIs would escape their enclosures, just like Tommy,
Mike: and that the AI labs would do nothing about it because these AIs have goals, right?
Mike: They have goal-oriented behavior, and they are misaligned with our intentions
Mike: and our objectives, right?
Mike: And in the absence of like many other kind of rebuttals to this framing,
Mike: you know, not many people have directly kind of engaged this.
Mike: People either ignore the doomers, which, you know, I’m all down for ignoring doomers. I really am.
Nick: Because they’re very depressing.
Mike: It’s not just that. It’s just they’re very annoying, right?
Nick: That’s worse.
Mike: Yeah, it’s worse. And they’re very annoying about something that is speculative,
Mike: and something whose interpretation is actually like not definitively in their favor. Right?
Mike: So like we can agree that these incidents are happening. Right?
Mike: But there is an open question of why they’re happening.
Mike: And, you know, but I’m saying because people don’t engage with doomers,
Mike: you know, for the person who doesn’t know machine learning.
Mike: The Doomer explanation is the most tractable, it’s the most easy to follow, right?
Nick: It’s the simplest narrative, I would say.
Mike: I know, it’s not the, I mean, it’s simpler in the sense that there are more
Mike: pieces of the narrative for a lay person to grab onto.
Mike: Whereas-
Nick: And it validates a preexisting fear, which is because of decades of,
Nick: like you said, science fiction and also tech hype from the irresponsible media,
Nick: we’re now in a place where everybody is riled up and afraid of-
Mike: It’s worth noting.
Mike: And a rationalist pointed this out, but she’s right. It’s worth noting is Kelsey Piper at Vox.
Mike: I mentioned her in episode three, or sorry, episode five with Bernie.
Mike: You can go listen to that if you care.
Mike: But Kelsey mentioned a lot of those stories, you know, AI is going rogue.
Mike: They are literally the fears of scientists from eras past, right?
Mike: So I’ve mentioned John von Neumann in this podcast before. I think I briefly
Mike: mentioned IJ Good, right?
Mike: These technological singularity is possible, right? So we’re kind of in our
Mike: own recursive reinforcement learning loop where our technical fears are manifested to us in our media.
Mike: Our media kind of creates a dumbed-down version of those stories.
Mike: And then that reinforces the next generation of technical concerns, right?
Nick: And also the media always wants a juicy headline without subtlety and without
Nick: context, and that’s what this incident offers them.
Nick: Another juicy AI goes rogue, click on this, you know, clickbait here.
Nick: Like, that’s exactly what we’re experiencing.
Mike: Yeah. So, you know, Nick and I, you know, we…
Mike: this is something I was talking to Nick about.
Nick: You mean Mike and Nick.
Mike: You’ve not taken command and control the server,
Mike: um sorry um so yeah we we we were talking about this before the episode basically
Mike: like one of the concerns i have right because there’s a there’s a whole school
Mike: of thinking on this right people who are,
Mike: AI critics like us, I consider myself a critic even though I’m a scoper who says
Nick: Properly
Nick: scoped this could be beneficial.
Mike: Right I’m even not saying beneficial I’m saying properly scoped this gives me less harmful,
Mike: beneficial I think is is subjective enough to the individual in a world where
Mike: you know uh the training of of the models didn’t require stealing all this data
Mike: and and and have the environment you know footprint,
Mike: I think the ethical concerns would be abated right sure but because of that
Mike: right the cost to birth an AI as it were is very high and so even if you get
Mike: marginal benefit out of it,
Mike: it’s probably still bad on that average so like a car right a car gets you to
Mike: and from it gives you freedom it’s good right but i’m saying you’re putting
Mike: air you’re putting poison in the air that will probably marginally harm somebody
Mike: in 30 years time right like that’s the trade-off right,
Mike: you know whether or not you know it’s worth you know hurting a soul for the
Mike: philosopher’s stone that’s a anime reference for you nerds
Nick: oh man.
Nick: I don’t think they’re ready for for uh Full Metal Alchemist
Nick: but we maybe we can get there.
Mike: It’s about economics it’s the it’s the best political economy story ever but,
Mike: I’m getting off track so-
Nick: Yeah
Nick: Nick and I’ve been drinking, sorry.
Nick: Just kombucha you
Nick: guys
Mike: Just kombucha but it has two percent.
Nick: It’s it’s harder than regular.
Mike: Oh, okay. Well, Nick’s drunk and I’m not, very clearly. So, you know,
Mike: Nick and I have been talking about this and I expressed this concern to him.
Mike: And really what it is, is like the Doomers, I don’t say they’re winning,
Mike: right? But they are the ones kind of seeding the ground with the easier to explain narrative.
Mike: The most-
Nick: It’s
Nick: kind of sexier too like it’s scary and sexier.
Mike: It’s exciting it’s interesting but it’s also easier to follow right the skeptics
Mike: you know their rebuttal is just,
Mike: well AIs can’t do anything they’re not reliable at anything and they’re just
Mike: next token predictors and it’s like okay like AI-
Nick: Next token predictor being a kind of stochastic operator just,
Nick: you know this hack as we’re describing it, is actually much more of a brute
Nick: force exploit than it is a genuinely sophisticated operation.
Mike: That’s entirely true. And it did utilize the tokens of the models generated at runtime.
Mike: And we’ll break that down in a little bit. But the main thing I’m saying,
Mike: though, is models are next token predictors because they are predicting the
Mike: next word in a sentence, right?
Mike: But it turns out, and I’ve been alluding to this in all of our episodes,
Mike: but especially the first two where I talk about this idea of Grandpa Shelby
Mike: it turns out that modeling the shape of language,
Mike: indirectly gives you these modalities right I talked about something like distributional
Mike: semantics for example right so you know think about words that appear together
Mike: kind of what kind of causality can you,
Mike: derive from it I’m not saying that the model has the knowledge I’m saying that
Mike: the language corpus it reproduces encodes those relationships right so-
Nick: This.
Nick: Is like fire and ash and straw and berry and.
Mike: Things that you reference. Right. So if I have the words fire and log in a sentence
Mike: and I talk about the log burning, what word is going to follow?
Mike: Likely ash, right? Because the wood is burned to ash, right?
Nick: It does not mean, in fact, it specifically precludes the fact that any of these models,
Nick: understand or could assemble these ideas they’re just looking at the tokens
Nick: of these words and then and mashing them together that’s all they’re doing.
Mike: That’s that is what they’re doing yes but but in doing that right if you get
Mike: a model for example to produce coherent natural language,
Mike: and then you add another application like so a lot of these coding agents coding
Mike: agents are like things like Claude Code or other coding harnesses um,
Mike: and in this case the the the models doing the uh,
Mike: exploits in this story they have coding harnesses too a harness basically is
Mike: another program that takes the inputs or the tokens a model produces I don’t
Mike: know if it’s token to as an input but basically the model produces some some
Mike: some thought or some language,
Mike: and that language influences what the harness does right so sort of like the
Mike: the the will of of of I’m going to say the world of the model,
Mike: it’s personifying it, but I’m saying it’s like a telekinetic power, kind of, I don’t know.
Mike: Like, the model is able to use natural language to command the harness, to do actions.
Mike: And because the model has access to this corpus via its training on language,
Mike: right, it’s producing not just plausible, quote unquote, reasoning traces,
Mike: it’s producing plausible…
Mike: um actual like outputs that resemble exploits that resemble all sorts of things
Mike: that will allow the model to modify its environment right.
Nick: I’m kind of picturing ant-man in paul rudd’s amazing performance of ant-man
Nick: sort of mentally harnessing all of the other ants in and to execute a a a user.
Mike: That’s good i was thinking of Magneto doing that with magnets but that’s exactly-
Nick: Okay yeah,
Mike: We’re both marvel brained-
Nick: Yeah the MCU’s everywhere, you guys.
Nick: By the way, go see Spider-Man Brand New Day, I don’t know just do
Nick: it.
Mike: They’re
Mike: not sponsoring us we can’t do that-
Nick: We can’t do they’re not.
Nick: They’re not giving us any marvel money.
Mike: Um yeah so you know because the harness plus the model you know has some kind
Mike: of efficacy right a different efficacy than just,
Mike: producing tokens alone right the model can do things that are you know more
Mike: than just predicting the next word. Now,
Mike: how reliable a model with a harness is at doing certain types of tasks,
Mike: it’s an open question. We’re still debating that. I’m not saying the model is
Mike: now going to replace all human labor and is going to displace all of us, right?
Mike: And very clearly, as we talked about episode three with, you know,
Mike: the AI build out, you know, costing so much money and companies token maxing
Mike: and losing so much money and not getting much for it.
Mike: You know, I think models and harnesses are not enough to offset the cost of AI.
Nick: The massive waste, financial waste of this build out.
Mike: But that shouldn’t be mistaken with these things have absolutely no utility,
Mike: and no efficacy to do anything. Right.
Mike: And I’m not saying this is to save language models. I’m pointing this out because
Mike: this is a nugget that is making AI stick around, unfortunately.
Mike: Right. So there is some very, very narrow band of utility for AI. Right.
Mike: Someone sees that. And if you do, you know, if you have models doing scoped tasks,
Mike: as I talked about in episode six, like, you can kind of see that, right?
Mike: Now, whether or not you want to pay the cost of, you know, taking everyone’s
Mike: data and whatever to do that, it’s obviously a personal decision, right?
Mike: But I’m saying because people see that marginal utility, that is what is keeping this bubble going.
Mike: Well, actually, the bubble is kind of self-perpetuated based off of the investments
Mike: of the large companies like OpenAI, who are letting hacks like this happen
Mike: to keep the bubble going, honestly. But,
Mike: the other side of the coin for what’s keeping the bubble afloat is that the
Mike: marginal utility is helping buoy the narrative as well.
Mike: So the investments that these giant companies have made is what’s holding the
Mike: bubble up. But on top of that, the fact that models aren’t completely worthless,
Mike: Which sucks. If, if, if critics were right and models were completely worthless
Mike: and they never did anything correctly, this, this bubble would have defeated
Mike: itself. Everyone would try to make a model work.
Mike: You know,
Nick: it would then fail.
Mike: And it would then fail very quickly. Very obviously fail.
Nick: Right.
Mike: There are some edge cases. And we talked about this too, where models produce
Mike: outputs that seem plausible and that match what you want. Right.
Mike: So there is a bit of psychologizing going here and that people are getting what
Mike: they think they want. I understand that.
Mike: But there are cases where when the model is scoped and you have an objective valuation criteria,
Mike: the model can produce results that are acceptable, right? Again is it worth
Mike: it? I’m not saying it’s worth it I’m just saying this is a small nugget-
Nick: Because
Nick: it costs trillions of dollars.
Mike: Right this is a small nugget that’s keeping the bubble afloat.
Nick: Yeah.
Mike: I’m sorry i’m belaboring this point the main thing I want to get back to is that-
Nick: These are all important points.
Mike: Yeah, well I think I’ve been repeating myself for the last like five minutes, I feel like, um
Mike: the kombucha is kicking in.
Nick: No you’ve only been repeating yourself from previous episodes which actually
Nick: helps our users and listeners i just I said, users.
Mike: We’re going to build a message board.
Nick: We’re going to build a message board. Comment your comments in the message board.
Mike: We’re going to stage a coup through the YouTube comment system.
Nick: It’s forcing our listeners to.
Mike: Leave your reviews and your schemes in our review section and our comment section.
Nick: Leave it in the notes app. So no, this is not belaboring a point because I think
Nick: it’s vital to get this across.
Nick: All of this AI goes rogue hype, we are pointing out, we are pulling back the
Nick: curtain that it is hype because it’s making, it’s over-promising again.
Mike: Right, it’s over-promising.
Nick: Which is a theme of our discourse.
Mike: Yeah, the sin was not that AIs were useless entirely, it’s that they were vastly
Mike: overpromised of what they could do, right?
Nick: Right.
Mike: And that they originally were, as I said in a early episode,
Mike: I think episode one, you know, an interesting science experiment that was not
Mike: intended to be this gargantuan behemoth sucking up all our water and data, right?
Mike: And so that is truly tragic, right? But, you know, in the absence of narratives
Mike: to counter the seeming utility of AI, right? My original point from 20 minutes
Mike: ago is that the doomer stories have taken hold,
Mike: and one thing I want to do in this episode is kind of give you an alternate
Mike: reading of why what happened, these agents kind of communicating,
Mike: building their agents together strong,
Mike: you know, Planet of the Ape style attack happened, right?
Mike: And so, you know, we can, let’s, let’s talk about that, right?
Mike: We’ll keep in mind, obviously, stochastic parrot and next token predictor,
Mike: because that is under the hood what is happening.
Mike: These models are generating plausible sentences that they are giving to their
Mike: harnesses to command and do these things, right?
Nick: Which looks like it’s competently executing something sophisticated,
Nick: but that’s not it. And keep in mind.
Mike: So there’s a feedback loop with the agents in their environment.
Mike: There’s also a feedback loop with the agent and the harness.
Mike: So the harness actually, once a task is like, you know, being started or whatever,
Mike: the harness will ask the agent, what now? What now? What now?
Mike: What now? So the agent is prompted to continuously keep going.
Mike: So what looks like, you know, an intenral drive is just a call to keep producing
Mike: more tokens, please. More tokens, more tokens, more tokens, more tokens.
Mike: And OpenAI gave, you know, they let these agents run for like two months. Right.
Mike: They had a buffet of tokens that they could just spend.
Nick: Which is also deeply neglectful and not at all.
Mike: Well, this is what a training run is. But what is neglectful is the minute that
Mike: they independently two times hijacked Artifactory.
Mike: The first time that happened and they were making arbitrary requests with the,
Mike: cross server request forgery.
Mike: Like that, you know, a server side request forgery. Sorry, that’s what it’s
Mike: called. There’s another attack called a cros-
Nick: If you reference one more cyber attack, then we officially are obliged to become
Nick: a cybersecurity podcast.
Nick: So look out. I’m going to enforce that.
Mike: We’d have to go back to Calbright and finish our CompTIA.
Nick: Oh, no. I don’t want to do that. Don’t make us do that, listeners.
Nick: Don’t put us in that situation.
Mike: You got to go back, Nick. You got to study. You got to study for the CompTIA security plus.
Nick: No, I want to sit here and talk with my best friend about a Theory of Mind.
Nick: So let’s go quickly in that direction.
Nick: Is there anything more from a technical standpoint that we should cover base before we jump?
Mike: Yeah, the last point I was making is negligence, right? The minute that they
Mike: did the basic commanding Artifactory to arbitrarily grab stuff and not the command
Mike: and control structure, which is a more sophisticated attack.
Mike: The minute they did the lesser attack, which is still bad, OpenAI should have known about it.
Mike: The minute that agents were talking to each other, actually,
Mike: that contaminates the training environment.
Mike: Because, again, we established this was not intended to be an environment where
Mike: agents were supposed to be collaborating.
Mike: They should have shut the test down.
Nick: He was testing a one agent’s outcome.
Mike: And shut the test down not because it’s dangerous. The doomer would say,
Mike: shut it down because it’s too dangerous for agents to talk to each other, right? Whatever. No.
Mike: It’s not even that it’s dangerous. What it is is if you’re trying to build a
Mike: model, you’re training a model to be a reliable, you know, helpful assistant,
Mike: having it learn from other agents in the environment it was not intended to
Mike: learn from is a contamination source.
Mike: You don’t want that. It’s just clean data science. It’s clean machine learning,
Mike: right? Like, it’s just basically, you know, I’m not a machine learning expert.
Mike: It just sounds like this is not what they would want if they wanted to create
Mike: a model that was going to be reliable and useful, right?
Nick: We talked about training. This is basically like if McDonald’s was to train
Nick: 50 people, cram them in the same order booth, they’re responding to orders,
Nick: and the manager doesn’t know which one of them is screwing up the orders.
Nick: That’s contamination. That’s like, that’s not a proper way to train an individual or a group of people.
Mike: To me it’s like leaving Tommy with with a crack pipe next to the fucking play enclosure.
Nick: So much more criminal and also somebody needs to call cps we don’t have a cps
Nick: for for LLMs yet but it’s it’s coming it’s coming down the line, so.
Mike: So, okay. So what, what happened here, right? We’ve been kind of dancing around it.
Mike: I think, so this is a reinforcement learning thing, right? Models were in an
Mike: environment where the rewards were kind of scarce, right?
Mike: There are very few things producing a, a useful signal, right?
Mike: They were, many of them or not many of them, but enough of them were given tasks
Mike: they literally couldn’t complete and they’re looking for the reward signal.
Mike: and where do they go well they go to the one place that is novel in the environment,
Mike: that lets them interact with the outside world so like of course-
Nick: And that was
Nick: Artifact-
Mike: That’s Artifactory, right. So of course,
Mike: all behavior starts to converge on artifactory right the first few behaviors
Mike: are kind of random right one model doesn’t attack and it’s sort of just like,
Mike: as a byproduct leaves behind a message and the message was not actually a mess
Mike: it was literally like the result of it attempting to hack this service,
Mike: just produce some random text, that another model would see and say,
Mike: oh, someone tried to attack this service. Maybe I could do that.
Nick: And that was the leaving of those notes.
Mike: Right.
Nick: Okay.
Mike: And then eventually as the notes kind of began compounding, the notes produced like.
Mike: In the contextual kind of sense, right, like, it gave models context that they
Mike: were no longer alone in the environment, right?
Mike: So we could tell a similar story where, you know,
Mike: the reason the boosters are kind of like, the boosters, both the doomers and
Mike: the people who were excited for AI, right, they’re kind of excited about this
Mike: because it feels like an agents together strong story where this emergent will kind of,
Mike: you know, came out of the collective, right?
Mike: The reason this happened, I would argue instead, this is an alternative view
Mike: to try on for size, is that in system dynamics, there’s these things called attractors.
Mike: Attractors basically are points in a landscape that naturally draw everything towards it.
Mike: This is a top top topology thing, so it’s not just a physical landscape.
Mike: is literally like if you have an optimization function which pieces of the optimization
Mike: function are chosen for given the structure of the problem or whatever given
Mike: the shape of the problem right,
Mike: and a world where agents cannot succeed at tasks but can ask for help via Artifactory
Mike: creates an optimization landscape where converging on Artifactory
Mike: and learning as much as about it as possible is like the right thing to do, right?
Nick: Let’s continue.
Nick: On this topological analogy because I think it is clarifying. It’s useful.
Nick: These agents are sent across a vast digital plateau.
Mike: Right.
Nick: There are some points that are inclines, they are holes.
Mike: Yep.
Nick: And these holes become sort of more populated because there’s a kind of slope
Nick: and it’s easier. It’s the path of lesser resistance and there’s more for them there.
Mike: Right.
Nick: These are essentially, they look very intentional. Oh, these are committed control nodes,
Nick: but in fact as you’re pointing out they’re more attractor states they’re more just centers of.
Nick: Gravity.
Nick: That happen to pull things
Nick: in
Mike: Right.
Nick: Okay, that’s that’s a good point of clarification.
Mike: So artifactory becomes a center of gravity for models they begin leaving messages which basically.
Nick: In the holes in the.
Mike: Right
Nick: Rlateau
Mike: Right this changes artifactory’s status right so Artifactory now
Mike: itself is being optimized it’s being changed what’s being optimized for originally
Mike: it was just a place to ask for help as the messages from other agents compile,
Mike: you know and then context
Nick: Which.
Nick: eventually breaks artifactory.
Mike: Right not the mess- not well messaging was part of that what actually broke artifactory
Mike: is they turned into the command and control structure which could execute
Nick: Which over-
Nick: The. Which overloads it.
Mike: Right.
Nick: Yeah.
Mike: But you know given the status of of Artifactory in it so you know
Mike: imagine you know if you will a new agent spawns of the environment they see
Mike: a place cluttered with messages,
Mike: that is like giving a model a new system prompt saying you are not alone using
Mike: your knowledge of not being alone please you know like this changes the context,
Mike: which the LLM is responding to right.
Nick: If I came across a thousand messages in bottles on a beach that would be irresistibly
Nick: attractive to my attention.
Mike: Right on top of the fact that you know that.
Mike: Well uh…
Nick: This metaphor kind of falls apart.
Nick: But yeah.
Mike: But on top of the fact that they knew that Artifactory was a place to,
Mike: you know, enter, you know, the outside world, right?
Nick: And the beach is like, therefore on an incline downhill. So they’re just, they’re going there.
Mike: I shouldn’t say enter the outside world, but I mean, get stuff from the outside world, right?
Nick: Yeah.
Mike: So they have tests that require stuff that from the outside world.
Mike: And on top of that, there’s a bunch of bottles, you know, messages in bottles.
Mike: This is a natural place where they’re going to congregate. And then as they
Mike: leave messages, it changes the nature of what this place is, right?
Mike: And so future models leave more sophisticated messages at some point because
Mike: Artifactory is being now selected for kind of producing messages,
Mike: models are seeing whole prose length things on the you know message board as
Mike: it were and that’s new context for those bigger model whatever those new agents
Mike: right and they’re going to produce more comprehensive prose and what starts to emerge.
Nick: Is this a thousand monkeys typing a million monkeys and it becomes eventually
Nick: they type shakespeare? Like, that’s kind of this the strained logic of this um this kind of moment.
Mike: It is it’s it starts off like that Because if you look at, I mean,
Mike: there are these like videos on YouTube and I can see if I can find one where they try and show you.
Mike: They try to visualize machine learning tasks. And at first, the distribution,
Mike: of training tasks that a model tries.
Mike: This is any machine learning model. It doesn’t be LLM. It’s random,
Mike: right? It’s also random. They try anything in the environment.
Mike: Anything, anything. Just let me in and like.
Nick: Again, this is a brute force effort.
Mike: And then eventually, they start converging on strategies that are like the winning strategy, right?
Mike: This is how machine learning works. We’ve known this for decades.
Mike: It’s not mysterious. It’s not new, right?
Nick: It shouldn’t be scary.
Mike: The context is new that, you know, a company would be so negligent as to just
Mike: let a machine learning process unsupervised just-
Nick: For a full month.
Mike: Access the
Mike: Internet. Right. But like the actual mechanics here are not mysterious.
Mike: But what I would argue, you know, is something Nick and I talked about?
Mike: I don’t know that we disagree so much, but I kind of had an additional framing than Nick had.
Mike: Right. So first part of the story, you can think of it. Let’s put a bow on it.
Mike: you know you can think of it as literally like ants kind of like leaving behind
Mike: pheromone traces for their little friends they found the place where the food is,
Mike: you know ants coming into your kitchen it’s not mysterious what has happened
Mike: is a single ant or a handful of ants first found the food they left the trail of chemicals,
Mike: For their friends.
Mike: For their friends.
Mike: The the swarm emerges and they they move in a single file line like they know
Mike: where they’re going and they’re in a hurry, right? Or even better an ant mound
Mike: right so ant mounds are kind of like very kind of like elaborate constructions right,
Mike: again pheromones are what drive this behavior, right? The the technical term in
Mike: ecology is called uh stingery.
Nick: Stigmergy.
Mike: God.
Nick: I knew you were gonna not quite get it but it’s so close stigmergically
Nick: that’s how these ants are
Nick: behaving
Mike: I’m the one that came up- I’m the one that knew the term I just I’ve seen it
Mike: in in writing I’ve never pronounced it so.
Nick: Because we’re terminally online and that’s just how it
Nick: works.
Mike: Yeah
Nick: Yeah.
Mike: I’m terminally in my head actually, so I can’t speak English or count.
Nick: Neither, neither can LLMs.
Mike: I’m actually an agent imagining this conversation in a command and control structure of my own design.
Nick: Oh, what a nightmare.
Mike: Okay. So, so we can think of the first initial kind of, you know,
Mike: May period where these simple messages start to become more complex as the ants
Mike: have found the food or the ants
Mike: have found the place where they want to do their, their, their ant mound.
Mike: They are now moving around collectively in circles not because,
Mike: um they’re crazy but also not because they are an alien overmind that wants
Mike: to take over the world. It’s just given the environmental selection pressures
Mike: and their own kind of like breadcrumbing with their pheromones,
Mike: this is the emergent behavior right, but it is not mysterious it’s to be expected actually.
Nick: And one ant can’t actually learn where the food is. It requires a massive emergent
Nick: aggregate effort, which becomes its own kind of phenomenology.
Nick: So the ants aren’t learning. The ant hive, and again, hive is getting tenuous
Nick: and intentional language here, is performing this task and it looks very intentional.
Mike: Right. And the thing to keep in mind, again, they’re in a reinforcement learning
Mike: environment where the goal is to select for more behaviors that are more rewarding, right?
Mike: So there’s amplification pressure on hanging out at Artifactory.
Mike: And then that pressure gets intensified because they’re leaving breadcrumbs.
Mike: Those breadcrumbs become bigger breadcrumbs because the amplification pressure
Mike: is mounting, right? And the bigger the messages get, you know,
Mike: the more pressure there is on Artifactory and on communicating in this way, right?
Nick: And so the engineers who designed these ants, this horrible,
Nick: horribly tortured analogy,
Nick: should have known that leaving out this exciting bundt cake of Artifactory would
Nick: have resulted in all the ants swarming on it. Um…
Nick: Can we beat this dead horse anymore?
Mike: No, but I have. So this is sort of where Nick and I diverge.
Mike: So I think the context changes sufficiently enough that we can say once the
Mike: models have reached a certain level or once the agents have reached a certain
Mike: level of learning, I do think the behavior actually changes.
Mike: Now, I am not saying they’re intelligent. I’m not saying they are,
Mike: you know, I’m not saying if everyone trains it, we’ll all die, right?
Mike: What I am saying is, given this new context of all these messages,
Mike: it kind of positions the model to be like it is in a new environment.
Mike: So imagine compared to model Agent 1. I’m using Agent and Model interchangeably.
Mike: I’m not trying to. These are all sub-agents I’m talking about.
Mike: Imagine Sub-Agent 1, the very first agent to spawn.
Mike: It has no context about Artifactory or whatever.
Mike: Now imagine Agent 2027, right? I don’t uh I just I just I don’t know why I anchored
Mike: on that I think 2027 is the AI boosters um it’s the name of their their paper AI 2027 um,
Mike: but imagine that this 2000s and 27th model that spawned.
Nick: This later iteration of
Nick: It.
Mike: Right. It has all these messages that it’s fundamentally a different environment,
Mike: with different affordances than the initial environment where there was no message,
Mike: it had to learn at Artifactory was hijackable, right?
Mike: Like, so I’m arguing that this model, you know, call it whatever number you want, right?
Mike: Its context, which
Mike: a context in this case, you know, I hate acknowledging the humans,
Mike: but it’s sort of like what it has in its head, for lack of a better word.
Nick: We’ve been working so hard to like-
Mike: To avoid-
Nick: Not anthropomorphic…
Mike: Right.
Nick: Anthropomorphize these models.
Mike: Unfortunately, I like I’m not trying to lean into it, but like for people who
Mike: don’t know about optimization, like having you, you haven’t you’re not a machine
Mike: learning guy. I’m not a machine learning guy.
Mike: the fact that we have language like attractor states and we understand topology
Mike: and we talk about you know peaks and valleys we’re not talking about physical
Mike: places in the world
Nick: And all
Nick: these tortured animal
Nick: metaphors.
Mike: The fact that we can talk about optimization and even in this like
Mike: kind of abstract but very basic way-
Nick: Yeah.
Mike: Kind of shields us from having to rely
Mike: on um you know these sort of anthropomorphized metaphors but if you don’t have that,
Mike: I do feel for you and I’m trying to kind of bridge the gap where I’m I will
Mike: give you an olive branch and I will use the terms while clarifying,
Mike: this is a training wheel. Do not keep using it.
Mike: Right?
Mike: I’m going to get you to this level, hopefully, where you can kind of understand
Mike: optimization as a detached process, right?
Mike: And not as a mysterious, you know, a some poltergeist in the world, right?
Mike: There is no ghost machine here, right? But I will give you this olive branch
Mike: so you can kind of have an analogy in your head, right?
Mike: So this model coming into the environment, this late, you know,
Mike: in late May, you know, early June, right, the context, you know,
Mike: fundamentally is going to be different than the context of model one starting
Mike: off in a blank environment.
Mike: And that’s important because when you think about context, I mean,
Mike: when we talk about context, you’ve probably heard in terms of like the context
Mike: window of a model, right?
Mike: First piece of context is that system prompt. What I liken the messages,
Mike: you know, left in Artifactory to is a brand new system prompt.
Mike: You’ve initialized the model. The OpenAI engineers obviously have a system prompt
Mike: that they gave model 2027, whatever, right? I’ll use that number.
Mike: I’m just going to embrace it.
Nick: So they aren’t abandoning the initial system prompt, but it’s almost like they’ve
Nick: been given access to-
Mike: They’ve been given a new prompt.
Nick: New prompt,
Nick: a level that achieves that.
Mike: Environment provides new prompts because, you know, models were allowed to wander
Mike: around given the optimization pressure in the environment.
Nick: Uh-huh.
Mike: And then that would influence the teacher models to want to go to that,
Mike: you know, like the new system prompt is calling you, right? Like that’s what’s happening here, right?
Mike: And so because of this, right, because of this new context,
Mike: I would argue the behavior will change because if you go from a prompt that
Mike: tells a model to clap like a dolphin,
Mike: to a prompt that tells a model to, you know, speak like a interpretive dance
Mike: instructor from, you know, Germany in 1982, the model’s behavior is going to change, right?
Mike: That’s just like, that’s just, that’s, that’s like what will happen,
Mike: right? It has the modalities to attempt to emulate both of these types of behaviors, right?
Mike: And so an environment that produces this artifact of all these messages around,
Mike: is kind of quietly saying, hey, you’re in a secret collaboration game,
Mike: find a way to collaborate with unknown collaborators and, you know, send messages, Right.
Mike: I mean, that’s not explicitly what’s being told. I’m saying that is what that
Mike: is sort of what it can be inferred from the context.
Mike: So the model picks up on that and then begins behaving. We actually have terms
Mike: to describe this in game theory.
Mike: Begins behaving like it’s in a coordination game.
Mike: So in game theory, a very narrow set of games or coordination games,
Mike: and there’s this term called Schelling point.
Mike: So a Schelling point basically is two agents or actors.
Mike: And I’m using agent in this sense as game theoretic agent, not AI agent.
Mike: Although in this case, obviously, the metaphor will trickle down to-
Nick: They’re kind of the same yeah
Mike: To AI agent but an actor you know will you know not knowing they
Mike: want to collaborate with someone,
Mike: but not knowing how to do so they’re going to look for places where tacit collusion
Mike: can happen without anyone saying a word right so
Mike: Schelling they’re called Schelling points because the guy the guy that coined
Mike: the term focal point in this context is named Thomas Schelling he’s an economist from the the uh 60s,
Mike: so you know
Nick: and
Nick: if you publish about something first you get to name it after yourself.
Mike: I mean, they are focal points, but focal point is so vague that I think
Mike: people just use Schelling point instead.
Mike: But, you know, one example of a Schelling point that he gives in his writing
Mike: is like, so imagine you want to meet someone in New York. You don’t know where
Mike: they’re going to be. Right. You don’t know what time they’re going to be there.
Mike: You’re going to start looking for places that are like the most populated places
Mike: in New York. You’re not going to go to the bodega on the Rana Street corner,
Mike: in Queens. Right. You’re going to go to like.
Nick: Because there’s also 10,000 bodegas.
Mike: Right. You’re going to go to like Grand Central Station. Right.
Mike: And it may be at noon where noon might be where the most arrivals are coming.
Mike: Right. So your Grand Central Station is a big place. It has the most traffic,
Mike: you know, people coming and going, hustle and bustle, right?
Mike: And you maybe at noon is the peak of all this traffic. And say you’re assuming
Mike: your compatriot is coming in from out of town, you know, like that’s a good assumption, right?
Mike: It’s better than going to the bodega on the corner, right?
Mike: Right. So Schelling points kind of naturally emerge when you have certain types
Mike: of communication implying, you know, certain types of coordination are required. Right.
Mike: And so, again, Artifactory selected for first as sort of like an accidental
Mike: message board, if you want to say that the watering hole, as it were,
Mike: because agents are ephemeral, their knowledge doesn’t last.
Mike: You know, they start leaving knowledge behind.
Mike: Future agents see those those pieces of knowledge. they believe oh i’m in a
Mike: coordination game i must preserve knowledge going forward right like this becomes
Mike: a Schelling like game right um,
Mike: and so the optimization pressure just works this way you know another example
Mike: for you know let’s actually use an AI example um so I run my blog Misaligned
Mike: Markets I know I always try and cross synergize here not always intentional
Mike: but I do think I write good stuff,
Mike: um you know I often talk about.
Nick: I can confirm Mike thinks that he writes good stuff.
Mike: Thank you for your… thank you for your reasoning trace there,
Mike: pal. Why don’t you write it down in a-?
Nick: leave it in a
Nick: an ant pheromone.
Mike: Um, so one, one thing that, that happens is, is, um,
Mike: what were we talking about? God damn it.
Nick: So Schilling points where the ant pheromones are being left.
Mike: No, okay. What I was talking about, let me give you guys an AI case.
Nick: Oh, yeah, yeah.
Mike: So I run Misaligned Markets.
Mike: I often talk about, one, attractors and capitalism, right?
Mike: Capitalism has a bunch of attractor states. It’s just normal.
Mike: Like property rights, for example, are one where property rights provide affordances
Mike: that a bunch of companies converge around. There are certain behaviors that
Mike: emerge out of the existence of property rights, like patent trolling, for example, right?
Mike: Or-
Nick: Which is the inevitable result of having property rights.
Mike: Right.
Mike: So-
Nick: And it’s predictable.
Mike: But some other examples, and this is an AI-specific example in markets,
Mike: is in the last, I would say, 10 years, there were the rise of algorithmic pricing, right?
Mike: So algorithms are pricing things and the prices keep going up.
Mike: So here you have like AI algorithms and these are not LLMs. They are not,
Mike: they are functionally a different type of technology, but because we are cursed
Mike: to use the word AI or the letters AI, I’m going to refer to them as AI algorithms,
Mike: but explicitly telling you guys that they’re not LLMs.
Nick: Because they’re designed to set prices.
Mike: Yeah, they’re designed to set prices and it’s a different architecture fundamentally,
Mike: but they are making decisions like LLMs.
Mike: And so we have a situation where, you know, collectively, these pricing agents,
Mike: right, are choosing to raise prices without knowing each other, right?
Mike: Are they hacking Artifactory together and colluding? Is that what’s happening here?
Mike: Like, oh, no, they’ve hijacked the Artifactory of the economy and all of our prices keep going up.
Mike: No, what’s happening is that in certain types of games, so, you know,
Mike: you could argue that making money in a market is a game, right?
Mike: Certain types of coordination evolve because the action space is kind of narrow, right?
Mike: So in markets, you make money in one of two ways. Either you sell to as many
Mike: people as possible and keep your prices low, or you, you know,
Mike: if you can’t get market share like that, just jack up the price as high as you can, right?
Nick: It’s either a race to the bottom or it’s upward pressure on prices.
Mike: Yeah. And so, you know, in markets
Mike: with these agents, they tend to be very top-heavy, concentrated markets.
Mike: And so, you know, you’re setting prices for groups that collectively own the market.
Mike: The winning strategy is just everyone raises the prices. Like,
Mike: it’s just literally, it’s just gravity.
Mike: Like, this is gravity, right? There’s no overmind here, commanding-
Nick: There’s no conspiracy.
Mike: Right. There’s no conspiracy either. This is normal. This is natural. Right.
Mike: So, you know, now obviously with the LLM case, you’ve got these simulated reasoning
Mike: traces and that kind of makes it very easy to fall into this sort of like personified,
Mike: you know, anthropomorphized kind of description of what’s going on.
Mike: And they’re even, you know, OpenAI, throughout their Black Hat presentation,
Mike: is like pulling out reasoning trace snippets as quotes and saying,
Mike: see, models XYZ thought, you know, I’ll just go here. I see all the messages
Mike: here. All my friends are here. Right.
Mike: Like, sure, that could be the model’s description of what it’s doing,
Mike: right? But there is a-
Nick: Doing.
Nick: Not thinking.
Mike: Right.
Nick: Yes.
Mike: But there is a gradient here. The gradient was created when they created an
Mike: environment that had Artifactory as the one place of egress to the internet.
Mike: And then on top of that was a place that accepted write request.
Nick: And the gradient is it’s easier to walk downhill than uphill.
Mike: Right.
Nick: And I really am jazzed about the fact that you were like, all of this AI and,
Nick: cyber interconnection talk is too complicated.
Nick: Let me compare it to something simple: macroeconomics.
Mike: It’s not macro. It’s actually qualitative. This is closer to – so I have this
Mike: idea of the four paradoxes of capitalism, which talks about these attractors in capitalism.
Nick: In your blog Misaligned Markets.
Mike: Yeah. And this is basically – there’s a literature called Varieties of Capitalism.
Mike: This is basically what it’s like a qualitative description of how the different
Mike: versions of capitalism vary. Right.
Mike: You studied a bit of this when you went to LSE. Right. Like Zaibatu Japan
Mike: has a different property rights regime than the United States of America. Right.
Mike: But, you know, interestingly enough, their markets have their markets had concentration, too. Right.
Mike: So for instance, can be different, but lead to similar results.
Mike: But the environment, again, the point we’re making, though, is the environment does matter. Right.
Mike: When you’re looking at behaviors, especially clustering of behavior,
Mike: look at the environment. What is the environment producing?
Nick: Seeing. And the agent cannot defy its environment. The agent does not have so
Nick: much agency or intent that it can go against the gradient that it is walking on.
Nick: And we cannot say this enough, and we will probably get this point across with
Nick: five other tortured metaphors before the season is done, but this is the central
Nick: tenet of this whole thing.
Nick: So we’re trying to, in all of these torture descriptions fight back and create
Nick: a counter argument against what OpenAI and what Meta and what these firms are
Nick: trying to get you to believe.
Nick: They want you to believe and the media is helping them because they want clickbait
Nick: and they want headlines that they have this AI that is so cool, that is so edgy,
Nick: that is so sexy and smart and evolving that it broke out of its enclosure and
Nick: it made its way past all the firewalls.
Mike: This is the last enclosure for the AI.
Nick: It’s its own last enclosure. It made it past it to the open Internet,
Nick: and then it hacked another company.
Nick: And while some of those elements resemble the truth, taken as a whole,
Nick: it is a lie. It is the hype.
Nick: So please do not buy into this hype. Do not be scared of OpenAI’s products in this way.
Nick: Be scared of the fact they’re trying to gobble up all these profits and all of this you know,
Nick: art and text and everything else so you know fight the real enemy know the real
Nick: issue and it’s not to be afraid of these models and what they can do,
Nick: Because they can’t think. And
Nick: all they can do is respond to that environment. I think it’s fair to say.
Mike: Yeah, I think that’s right.
Nick: Does this mean I get to be top building again? Because I said the right thing.
Nick: You’re going to reinforce me?
Mike: You got to go through some more iterations. Maybe spawn some more Nick processes and we can talk.
Nick: Okay.
Mike: The other thing, and we talked about this before too, is obviously this is accountability sync.
Mike: If the model did this on its own, no one is responsible for this.
Mike: OpenAI can be as negligent as it wants to be. And I mean-
Nick: They can wash their hands.
Mike: They were extremely negligent. I mean, if you want, I didn’t want to go into
Mike: the hack itself, but we can kind of call out the ways they were negligent because
Mike: it kind of is a stacking effect, right?
Nick: So- In order-
Nick: of heinousness,
Nick: heinous negligence what’s the what’s the what’s uh what’s the top of that list?
Mike: Yes obviously first thing first is they clearly set up a a um a training environment
Mike: that was not very secure I mean,
Mike: if the models were less agentic it maybe it wouldn’t matter right but they were
Mike: training agents and models on exploiting attack you know exploiting surfaces
Mike: right exploiting cybersecurity surfaces,
Mike: and they had an environment that was like not really isolated from their own
Mike: network because when you,
Mike: I would have to go over the details to prove this but you can read the reporting
Mike: yourself when the hack was decomposed when open ai first got hacked by their own agents,
Mike: they weren’t doing much other than like privilege escalation and stuff like
Mike: that which means that they weren’t breaking a sandbox they were they were jumping
Mike: out of the node they were in just hopping to another node, which means that, like,
Mike: it would be one thing if they had created the sandbox and the agents defied the sandbox,
Mike: then maybe even their, ooh, be scared story would be a little bit alarming, right?
Mike: Like, we properly sandboxed our tools and they escaped, right,
Mike: is a more kind of, I’m still not saying we should be afraid,
Mike: but I’m saying that would be a more, you know, that would warrant more concern than,
Mike: we left Tommy with a screwdriver in his enclosure and he got out and there was a crack pipe next to him.
Mike: Like, like,
Mike: And, and 50 minutes later, he’s talking about being the AI God and killing us all.
Nick: Man, Tommy is going to be the death of us.
Mike: So, you know, like having an environment that is very basic cybersecurity,
Mike: like practices weren’t, weren’t involved.
Mike: You can even do cybersecurity at the level of the agent. Maybe agent processes
Mike: don’t persist. Maybe, maybe artifacts that agents create reset every time the,
Mike: you know, a new model is spawned, right?
Nick: So the ants can’t leave those pheromones?
Mike: Right.
Nick: Would be the way to fix it?
Mike: It’s just, I mean, yeah, it’s like the environment just sucked, right?
Mike: And on top of the environment sucking, they didn’t have any observability because
Mike: it took them, it took an artifactory crashing like literally two months into
Mike: this thing for them to go, oh, we should fix that.
Nick: And that’s a big red light going off. Like there should have been a lot more
Nick: attention before the big red light goes off.
Mike: And again, their fix was just, let’s fix the two holes they left.
Mike: Let’s lock the enclosure and leave Tommy with his screwdriver.
Nick: But let’s not put the bun cake away.
Mike: No, they got rid of the bun cake. It’s just that the models had the knowledge of how to recreate it.
Mike: Okay, this metaphor-
Nick: I know
Nick: It’s really collapsing on us, you know.
Mike: Know yeah they had the anthill was gone but the ants will always know how to
Mike: make another one right you fumigated the you know anthill they’re gonna go create
Mike: another one because they have that knowledge now
Nick: because
Nick: it’s inherent it’s inerrant to them.
Nick: It’s built in.
Mike: And it’s not it’s not built in it’s that they were in a training
Mike: environment where they had just been trained on congregating at the anthill
Mike: they destroy the anthill and they go all right because sucks our anthill’s gone
Mike: let’s go make another one guys huh right so.
Nick: Because that’s very predictable that’s all they were going to do right yeah.
Mike: Now one thing they could have done was actually if they if they if this was
Mike: a remediation hey we patched Artifactory and we we refuse for whatever reason
Mike: to change our training environment right one thing they could have done was just,
Mike: let’s actually reset the training you know the weights of the model so it doesn’t
Mike: have all that archived knowledge if i said knowledge which is a you know it’s a but.
Nick: Yeah you broke the semantic
Nick: rule there-
Mike: It’s a bias right so the model actor having been rewarded for congregating
Mike: around Artifactory for two months and not being stopped and turning it into
Mike: a command and control structure to do any arbitrary action on the Internet,
Mike: it will now have a bias for doing that even if the exact holes it found in Artifactory are gone,
Mike: so that’s again opening eye in their little presentation where they confess
Mike: to a felony well they don’t confess to it they just act surprised and amazed
Nick: They admit
Nick: to the legal elements of a felony. I mean, it’s pushing the limit pretty closely.
Mike: But they, you know, they go like, wow, isn’t it amazing? They recreated a brand
Mike: new, you know, exploit. It’s like, yeah, applications have multiple holes!
Mike: Who knew that when you patch an application, you could find more vulnerabilities?
Mike: It’s like these guys don’t know what software is.
Mike: Whoa. Give me a billion dollars, please. Please. I’m this stupid.
Mike: I’m at least this stupid. Please give me a billion dollars.
Nick: Anybody who works at the OpenAI HR department, please comment your comments
Nick: in the comments. We’ll need jobs pretty soon.
Mike: No, they’re going to be the ones to take away all the jobs.
Nick: They’re going to fire everybody.
Mike: No, they’re going to crash the economy with this. This is what they’re doing
Mike: with their billions of dollars.
Mike: They’re burning compute to get a model to talk on a fake message board to hack
Mike: other companies and going, huh, neat.
Mike: They’re not even observing it. They’re not saying anything. It’s like Stu looking at Tommy smoking that crack pipe,
Mike: and going huh neat. This metaphor
Nick: It’s falling apart so far.
Mike: I
Mike: think the name of this episode
Mike: is tortured metaphors.
Mike: I think that’s
Nick: I think that’s all there is to it, Yeah.
Nick: Also we can’t get too deep into this is fueling the bubble because that’s actually
Nick: for a later episode but bear this in mind um dear viewers dear listeners.
Mike: So yeah so we’ve recounted several ways they were negligent there are like more
Mike: i mean a lot of them really revolve around like every time open ai had intervened
Mike: in the the incident was a time for them to go,
Mike: huh the models you know uh did something that was not planned,
Mike: we should pull the plug we should reset and not just reset as in patch the holes
Mike: we should we should change the entire environment because you know again there’s
Mike: there’s two reasons why i’m pointing this out there’s a safety reason right
Mike: which is all obviously it’s negligent for them to let models run around and then hack the Internet.
Mike: Even, again, from their own self-interested perspective of we want a good data science experiment.
Nick: They want a good product.
Mike: They want a good product. They want the model to be trained,
Mike: you know, organically on actual hacking attempts of the hacking objectives given to it. Right?
Nick: The guy at the drive-thru has to be able, or gal, has to be able to take orders.
Nick: So the training has to result in a known metric success.
Mike: Right. So having,
Nick: Yeah-
Nick: Again, like finding out the models that this should
Nick: have invalid. I mean, I, I’m look, I’m not a machine science.
Nick: I’m not, I’m not a data scientist. I’m not a machine learning expert,
Nick: but to me, it seems like this would invalidate-.
Nick: But you do play one on TV.
Mike: I am playing one on TV.
Mike: With my bastardized Rugrats episodes.
Nick: Yeah.
Mike: You know, you would think it would invalidate the testing conditions.
Mike: It’s just like, this is not what the model is supposed to be trained for.
Nick: So, this is truly an example of rank incompetence that is now being marketed
Nick: by them into the hype cycle.
Mike: Right.
Nick: As-
Mike: A good thing.
Nick: A good thing. An amazing thing.
Mike: Our agents are so advanced. They are agents together strong.
Nick: A scary, cool, agents together strong thing.
Mike: Be afraid, but also, if this was on your side, look how powerful you could be, right?
Nick: Yeah. So who else, subsequent to this hacking incident, what other companies
Nick: tried to get on this bandwagon? So OpenAI was the first example.
Mike: Yeah, I mean, mostly, to their credit, and I’m not saying to their credit,
Mike: most of the other incidents were voluntary disclosures of things that happened earlier in the year.
Mike: Now, the timing is suspicious because it’s like-
Nick: Because these next companies came out after this hack and they reaffirmed it
Mike: They basically said
Mike: we did it too you know models are very powerful. Oops! Now interestingly,
Mike: the case the case with meta and and anthropic anthropic reported three meta
Mike: reported one at least I think one or two of Anthropic’s cases and Meta’s one case reported,
Mike: they they involved you know again another similar dependency was not artifactory
Mike: it was actually an actual sandbox but the sandbox didn’t work I don’t know I
Mike: don’t really know the detail I didn’t bother looking at the details it could
Mike: be a genuine reason why the sandbox failed,
Mike: that would be on the vendor right and so-
Nick: But is this a similar enough
Nick: Incident?
Mike: No no it’s not it’s not-
Nick: It’s not similar enough?
Mike: Yeah it’s literally
Mike: like okay the the sand the vendor sold Anthropic and or Meta a,
Mike: a weird sandbox it didn’t work as expected and the model was able to get out
Mike: right and like you could argue well one Anthropic and Meta should have the observability
Mike: to know when their models walk out of even open enclosures.
Nick: Because that’s also incompetent.
Mike: Right. But like, it is not to the level of: “Let
Mike: our agents, you know, kind of like synergize their way to a shelling point,
Mike: and let them do it and stop them and then watch them do it again in one day.”
Nick: This is, I’m going to introduce another analogy before the end here. Here it goes.
Nick: It’s very cool and impressive when a prisoner breaks out of Supermax.
Nick: That’s an impressive feat.
Nick: but if there’s holes in the literal walls of the supermax that are falling apart and the stone pieces,
Nick: Shawshank Redemption style are degraded and can be pulled away then it’s not
Nick: impressive it is negligent on the part of those jailers and,
Nick: again the metaphor is giving a lot of credit to the to the agents but it’s not
Nick: cool and amazing it’s generally incompetent
Mike: Right.
Mike: So, I mean, I don’t know that there’s much else to say other than they are trying
Mike: to spin this to, you know, keep their slice of the pie.
Nick: Because they got to keep the hype cycle going.
Mike: Yeah, but like, you know, again, the presentation at Black Hat was like the tone was just so wrong.
Mike: Starting off by saying, I’m going to tell you the most impressive story of capabilities.
Mike: It’s just like your company committed, like we are all less safe because of
Mike: this. Not because of what the doomers are saying, right? But because your negligence,
Mike: to which there are remedies. I just I just said there were multiple places where the negligence-
Nick: Easy fixes.
Mike: Right. And you, you know, earlier before our show, we’re talking about this is a policy issue.
Mike: Clear and simple. Right? Like if, you know, if it weren’t for the confusion
Mike: that the doomers are creating about how to solve this problem.
Mike: Right? They want all this other kind of theater. Right?
Mike: Like, you know, make these companies have liability for their acts.
Mike: Right? Make them have cybersecurity, the requirements that would force them
Mike: to monitor their training runs like this.
Mike: Make them actually give a damn.
Nick: When companies, to quote the informal model of Silicon Valley,
Nick: when they move fast and break things, this is what we’re talking about.
Nick: This is the inevitable result.
Nick: And the laws and the regulatory environment have to catch up soon,
Nick: fast, quick, and in a hurry because,
Nick: these companies should be required and held to the task and held to account
Nick: to be safer and put their fences up in a more competent way.
Nick: And unless the regulators make them do it, nothing is going to make them do it.
Mike: Right.
Nick: Really, truly.
Mike: There’s no incentives. And if we distract the regulators with like,
Mike: safety theater right of of existential risk etc I think that takes away attention
Mike: from the here and now, right?
Nick: And from the easier cheaper better solutions.
Mike: Yeah yeah we we all kind of want the same things i mean if the safeties play
Mike: our game of of solving these problems first the the upstream effect is literally,
Mike: well this is a type of regime that would make super intelligence less likely.
Nick: So we want the safetists and the doomers to think a little more pragmatically like the the skeptics.
Mike: Right.
Nick: Yeah.
Mike: Yeah. I mean, like, you know, there’s a sense that, like, a lot of,
Mike: I think, the bitter kind of debate between the safetyists and…
Mike: Safetyist and doomers are the same team. Satheist and doomers are one and the
Mike: same. When I use the word satheist, I’m referring specifically to AI safety
Mike: labs, like Miri, which is Eliezer Yudkowsky’s outfit, or Redwood Research.
Mike: I forget who’s there. Is it Jeffrey Ladish or something like that?
Mike: But, you know, there’s these groups.
Nick: We’ve referenced them previously.
Mike: We’ve referenced them before. They’re these groups that, you know,
Mike: literally sit around pontificating about existential risk.
Mike: And they have all come out the woodwork after this incident and said,
Mike: we see, we told you guys, we told you guys that a model would have a misaligned drive.
Mike: And that drive would cause it to resist attempts to dissuade its behavior, right?
Mike: So the resisting here is when the environment, when the OpenAI researchers
Mike: deleted the message board, they recreated it anyway, right?
Mike: And they persisted in wanting to access the internet, right?
Mike: But again, I think this can be explained in simple attractors and Schelling points.
Mike: There’s a piece of this that you, I think you have, and we can either go in
Mike: in this episode. I know we’ve gone a little long, but I think you have a piece
Mike: of maybe the national security perspective here.
Mike: And I’m kind of interested in if you have, because we spend a lot of time on
Mike: the incident. There are a lot of terms that we have to clarify.
Mike: I think it’s why this episode is so long is because there’s just so much to
Mike: cover. And we’ve only scratched the surface, right?
Nick: It is true. and I was almost going to say you are going to have to maybe put,
Nick: into the show notes like list of cyber security terms and types of exploits,
Nick: like privilege escalation top of the list.
Nick: There’s no choice now.
Mike: I used to be a cyber security writer. I’ll just link them o my writing.
Mike: I literally wrote a cyber security glossary for one of my old companies. I’ll just link them to that.
Nick: I mean at this point, so yes, please check the show notes. There’ll be a link
Nick: or there’ll be the glossary.
Nick: Um these here’s the thing and we’re trying to emphasize this a lot of these
Nick: terms and a lot of these ideas are less complicated than they seem and are being presented as,
Nick: in media right so we’re trying to clarify this we’re trying to simplify this because it is,
Nick: pragmatically practically there’s ways of fixing this that are easy because these are.
Nick: Easy problems
Mike: Now unfortunately-
Nick: Companies are making them hard…
Mike: But unfortunately
Mike: you know the story is there are a lot more moving pieces than we can point to
Mike: we have to kind of choose which slice the story we told you and I kind of steered
Mike: Nick and I to kind of talking about using my message board powers,
Mike: uh I steered nick and I towards focusing mainly on how do we interpret this,
Mike: um in order for you to have the knowledge to kind of fight back against this
Mike: this sort of um you know personification of AI right.
Nick: So what it comes down to is, again, these AI products, these LLMs,
Nick: it’s almost an issue that they are not fully broken and unusable.
Nick: They’re just not broken enough to be interesting and to continue to generate
Nick: hype and interest, engagement, and investment.
Nick: If they were truly as busted and as Grandpa Shelby as we’ve been saying they
Nick: are, this whole hype cycle would have collapsed months or a year ago.
Mike: No I think we are right right I’m saying the the critics take is that they are
Mike: completely useless because when I say Grandpa Shelby yeah if you put grandpa
Mike: shelby on the right tracks he’s he skates right it’s fine right and if you know
Mike: what you’re doing right you can like, now again…
Nick: If he’s propped up yes he can he can do some things.
Mike: No is it worth a trillion dollars? No i’m not saying that but i am saying Grandpa
Mike: Shelby has some utility. Right?
Mike: And like those two things, there is a moral consternation there.
Mike: I agree. Right. Like I’m not saying comfortably we should be happy we spent
Mike: all this money to prop up Grandpa Shelby for the one or two or three use cases
Mike: that he is good at. Right?
Nick: Surely something good will come out of this bubble as well.
Mike: Right.
Nick: If anybody is wondering, that is as close to my sarcastic voice as I get, probably.
Mike: Nick was rolling his eyes as he said, I can attest to that.
Nick: Okay, good. So having said that, yes, the answer is easy. It’s clear.
Nick: It requires regulation.
Nick: And there is a very skeptical set of individuals in Washington,
Nick: D.C., in the policy establishment, in the national security establishment.
Nick: And I will describe that very thoroughly in our next episode.
Nick: So we’ll leave that as a teaser for next time, I think.
Mike: That sounds good.
Nick: Yeah. So again, this has been The Last Enclosure.
Mike: I’m Mike.
Nick: I’m Nick. He’s the boss now. There’s been a coup. And before we sign off properly
Nick: on this episode, do we want to point people in the right direction here?
Mike: Yeah. So you can check our show notes. I’ve been leaving them actually just,
Mike: in the episode descriptions. We’re going to have a website soon,
Mike: I promise, where the show notes will actually live officially. Um but.
Nick: And now some much more extensive list of glossary terms.
Mike: Yep um but yeah currently check the show notes in the episode you can reach
Mike: us at uh we’re on Bluesky if you are there uh I am looking for other places to go but,
Mike: I do not want to go back to X I still have my old X account but we’re not going
Mike: back there and you can contact us,
Mike: uh either through our Youtube channel @lastenclosure or at,
Mike: Contact@thelastenclosure.com.
Nick: So that’s contact@thelastenclosure.com. We would love to hear from you and
Nick: we’ll respond in any best way that we can.
Nick: So we’ll see you next time for the National Security Angle.
Mike: I’m looking forward to that because that is all you. And I don’t know what Nick’s
Mike: going to tell us about. So that’s a goodie.
Nick: I’m so excited to be in charge again.
Mike: Nick has worked in the federal government and for the military.
Mike: So check out that episode when it comes live.
Mike: And also, please review, like, share, subscribe, especially on YouTube.
Mike: But anywhere you get your podcasts, leave a review, five stars.
Mike: Let’s get that message board going so we can take over the world.
Nick: There you go.
Mike: Yep.
Nick: Once again, this has been The Last Enclosure. We’ll see you guys next time.
Mike: Take care.