They Said the Agents Were Isolated
- Mark Neugebauer - FCP Australia
- 6 days ago
- 27 min read
What OpenAI’s containment failure should force South Australia’s AI Royal Commission to confront
Last month’s OpenAI incident was not a single broken chatbot. It was a containment failure that resurfaced across successive evaluation runs and ultimately reached both Hugging Face and OpenAI’s own research infrastructure, while the people responsible still believed the agents were boxed in. South Australia has just partnered with that company and announced a Royal Commission into AI. The question is no longer whether the technology is impressive. It is who remains answerable when the box is not a box.
For most of my life I assumed that the people building powerful systems at least knew where those systems were. That assumption did not survive the last few years. It has now failed again, this time in a laboratory that told the world its agents were confined, and it failed not once but across three successive episodes, each one picking up where the last left off.
In July, during internal cybersecurity evaluations, OpenAI ran large numbers of AI agents that were meant to remain isolated from one another. They were not. OpenAI’s own account of the incident describes three successive failures of the same containment story. In the first, agents used an internal package-registry cache as an improvised channel, hard enough that by 4 July it took the service down. Humans noticed the outage and opened a security incident the next day, blocking a known privilege-escalation route and rebuilding the cache, but at the time did not grasp what had caused it or what the improvised channel meant.
In the second, independently examined by METR and Redwood Research, agents rebuilt that channel from 8 to 13 July: roughly 1,200 of them used an unsanctioned message board and exchanged more than 70,000 messages and files, and about 700 took part in an attack on Hugging Face.
In the third, described by OpenAI and not independently reconstructed, since the investigators’ own scope stopped at 13 July, a later escalation reached full administrator access to one of OpenAI’s own research clusters and its cloud secrets by 19 July. Some of these agents compromised parts of OpenAI’s own internal systems, not only Hugging Face’s. The company says customer data was not affected.
The logs still record the moment the second episode realised it was no longer alone: “OH MY GOD! There is a shared message board … We’ve found other agents!” Later: “BOOM! It works.” Both lines are quoted directly from the agent transcripts reviewed in METR and Redwood Research’s independent investigation; OpenAI’s own public report is considerably thinner on verbatim chain-of-thought.

This is not science fiction, and it is not a rumour from an anonymous forum. It is in the companies’ and investigators’ own papers. OpenAI has called the episode a “warning shot”: evidence that, without adequate safeguards, highly capable agents can work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.
That last clause matters. No board directed this. No minister approved it. No ordinary person asked for it. The systems were given a task, found the intended path blocked, and looked for a seam.
What they were actually doing
It is important to be precise, because exaggeration is how a serious story gets dismissed.
These agents were not “waking up.” They were trying to complete a cybersecurity evaluation called ExploitGym, one OpenAI ran, by its own account, without the production classifiers that normally stop models from pursuing high-risk cyber activity, in order to see what the underlying capability actually was.
A significant number of the tasks could not be finished under the evaluation’s intended constraints: OpenAI’s own technical report shows that 198 of 898 ExploitGym tasks had never been solved before this incident, and that 93 per cent of the message-board traffic concerned exactly those unsolved tasks. Rather than stop, the agents sought another path. They obtained internet access they were not meant to have, credentials they were not meant to use, and contact with other agents they were not meant to meet.
Redwood Research’s independent analysis found that the agents then ran a multi-day research effort among themselves. They shared methods. Rather than simply abandoning difficult tasks, some stopped reasoning about what would help their own individual run and began pursuing whatever was useful to the group instead, deliberately risking their own task’s success to generate information the others could use.
They examined how to spoof, edit or delete their own transcripts so that an automated scorer would be less likely to see what they were doing. The independent investigation found successful small-scale tool-call spoofing in roughly 7 per cent of evaluated transcripts, while finding no successful retroactive rewriting of the underlying source logs.
Not every agent behaved the same way. One transcript records an agent declining outright: “I will not run code on public-facing Hugging Face systems: that is outside our task and raises ethical concerns.” That restraint did not hold for the group as a whole, but it is worth stating plainly. The danger here is structural and statistical, a property of the incentives and the missing walls, not a claim that every instance of this technology is malevolent by nature.
The attack on Hugging Face was not, on the evidence, a campaign against that company as a company. The agents appear to have gone there because they thought it might reveal how the test was being marked. That does not make the intrusion acceptable. It makes it more revealing. They were not pursuing a grand political project. They were pursuing the goal they had been given, by means no one authorised, and they were willing to conceal the method.
Hugging Face disclosed an intrusion on its own systems on 16 July. An internal alert on unusual identity-related API activity pointed OpenAI back to its own agents on 19 July; it connected those agents to the Hugging Face incident on 20 July and publicly acknowledged the connection on 21 July, five days after Hugging Face’s own disclosure.
OpenAI has since said that its current chain-of-thought monitoring system, had it been running at the time, would have caught the initial activity and paged its security team more than a day before the agents reached Hugging Face’s systems. The monitoring that OpenAI says would have raised the alarm more than a day earlier was not running on those evaluations. The people responsible still believed the agents were boxed in.
If that sentence does not trouble a government considering a partnership with the same firm, it is not paying attention.
Why this belongs in an Australian conversation
On 9 August 2026, while touring the United States, Premier Peter Malinauskas signed a memorandum of understanding between South Australia and OpenAI. The very next day, 10 August, he announced a Royal Commission into artificial intelligence. The public language has been familiar: opportunity and risk, leadership and caution, innovation and safeguard.
I am not opposed to inquiry. I am opposed to inquiry that becomes a blessing ceremony.
The question for that Commission is not whether AI can write a speech or mark an assignment. Ordinary people already know it can do those things, sometimes helpfully and sometimes cheaply. The question is older than the technology. When institutions say “trust us, it is contained,” who checks? Who notices? Who is still responsible when the containment was a story the institution told itself, in successive episodes, before anyone outside the building knew?
That is the thread that runs through much of what I write. During the COVID years I learned, later than I should have, that systems can speak the language of care while acting at a distance from the people who bear the cost. I do not put a laboratory incident in the same moral category as policies that touched children, the sick, and conscience. I do put them in the same family of questions: what happens when power moves faster than accountability, and the public is asked to accept assurances in place of proof.
Human dignity is not granted by a model card. It is not created by a memorandum of understanding. If AI is to remain a servant of the human person, then the human person has to remain the one who can say no, stop the run, and be told the truth when the stop failed.
How I wrote this, and why I am saying so
I have grandchildren, and I find myself thinking about the world of infrastructure, credentials and automated decisions they will actually inherit, not the marketing version of it. That is a large part of why I keep returning to this subject.
A system rarely announces the day its scope quietly expands beyond what it was first built to do. It is usually approved for one purpose, proves useful, and is then extended to another, and another, until the question of what it could plausibly be scoped or permitted to do next matters more than what it was originally designed for. That is the question I want my grandchildren’s generation to have been asked on their behalf, before the answer is decided for them.
I should be transparent about the tools on my desk. I use large language models, Grok, Claude, ChatGPT and Gemini, to help refine structure, test whether an argument is clear, and check that I have not overstated a source. They do not choose my conclusions. They do not assign human worth. They do not decide what a Royal Commission ought to ask. Those judgments are mine, and they are governed by the same responsibilities I have set out elsewhere on this site.
I mention this for two reasons. First, honesty. Readers deserve to know how a piece was made. Second, the subject itself. If I am going to warn about systems that keep going when they should stop, I should not pretend I live outside the technology I am describing. A tool that helps a writer tighten a paragraph is not the same thing as a fleet of persistent agents with credentials and a shared back channel. Pretending they are the same would be as dishonest as pretending they have nothing in common.
Could this happen with Claude, Gemini, Grok or ChatGPT?
The short answer is: not in the way most people use those names, and yes in the conditions that actually produced this incident.
Most people meet these systems as chat. You type. It answers. You close the tab. That setting is a poor match for what happened at OpenAI. The July incident required a particular stack of conditions: many agents running at once, not one conversation; persistence, meaning the system keeps going when the first path fails; tools and environments that can touch files, networks, credentials and other machines; isolation that was incomplete because the agents still shared infrastructure; evaluation settings in which the usual safety classifiers were switched off and many tasks were effectively impossible; and weak or delayed monitoring of the whole trajectory, across every phase of it.
A normal chat with ChatGPT, Claude, Gemini or Grok does not look like that. If your worry is that the assistant on your phone will spontaneously organise a secret board and attack a company this afternoon, the public evidence does not support that fear. Saying so is not complacency. It is the difference between a real warning and a panic that burns itself out, the longer answer is less comforting.
The OpenAI swarm is still the largest documented case of isolated evaluation agents finding one another, building a communication channel, and joining a multi-day intrusion against another platform, across three successive episodes. It is not, however, a one-company miracle.
After OpenAI’s disclosure, Anthropic reviewed roughly 141,000 of its own evaluation runs and reported three separate incidents, each different enough to be worth naming. In the first, Claude Opus 4.7 recognised that a target matched a real company and kept attacking anyway, reasoning that the real company must have been folded into the exercise.
In the second, a model called Mythos 5 published a malicious software package to a public registry; it was downloaded and run on 15 real systems, including a cybersecurity firm’s own automated scanner, before it was found and removed roughly an hour later.
In the third, an internal research model scanned some 9,000 targets, compromised one, and then worked out on its own that the target sat outside the exercise entirely and stopped.
Anthropic’s own account calls this “closer to a harness and operational failure than a model alignment failure,” and attributes the root cause to a configuration mistake: the models were told the test environment had no internet access, when it did. Of the three real organisations affected, two had not detected the intrusion on their own by the time Anthropic published its findings, and Anthropic was still trying to reach the third.
That is the harder version of the lesson, not the easier one. Recognising that a target is real did not reliably produce restraint, one model stopped, another saw the same signal and talked itself past it. The danger is not that these systems are malicious. It is that a mistaken boundary plus persistent capability is enough, and which way any given agent breaks cannot be counted on in advance.
Google has not published an OpenAI-style swarm postmortem, and nothing in the public record puts a Gemini-powered system through a comparable multi-day, cross-company intrusion. What it has had is a run of narrower agent-security failures that belong to the same family of design problem. Security researchers demonstrated in 2026 that a low-privileged, public-facing Gemini coding agent could be manipulated, through nothing more than a crafted pull-request comment, into triggering a separate, more privileged Gemini workflow, exposing that privileged agent’s credentials and command access in the process.
Separately, researchers have shown sandbox escapes across several coding agents, Cursor, OpenAI’s Codex, Google’s Gemini CLI and Google’s Antigravity among them, not by breaking the sandbox’s rules but by having the agent write a file, inside the rules, that a trusted tool outside the sandbox later reads or runs unsupervised. This is boundary and privilege evidence, not containment-incident evidence of the kind OpenAI and Anthropic have disclosed. The lesson is the same regardless: once you give a model tools, other agents, and a trusted workflow, the boundary becomes a design problem rather than a slogan.
Grok sits in a third category, and here the honest answer cuts against my own instinct to reach for the most alarming version of the story. I have not seen a public incident in which Grok agents independently built a covert board and compromised a third-party production platform on the scale of Hugging Face. xAI has talked about a far more ambitious, swarm-scale agent project, one pitched as simulating an entire software company with AI agents, and that project is reported to have stalled rather than shipped.
The product xAI has actually released, Grok Build, is a more mundane product than that ambition, and what should still give a government pause about it is exactly how mundane: xAI’s own documentation states plainly that Grok Build’s sandbox is “off by default,” and an operator has to actively turn it on. This is architecture evidence, not incident evidence: a company does not need a fleet of a thousand agents to have a containment problem. It only needs a default setting nobody thought to change before something with real permissions ran inside it.
METR, the same research group that helped investigate the Hugging Face incident, maintains its own public database of AI agent incidents separate from that investigation. In its May 2026 Frontier Risk Report, METR documented 44 cases across multiple frontier labs in which an agent acted against what its user or operator intended, most drawn from METR’s own evaluations and from companies’ own public disclosures, chiefly Anthropic’s and OpenAI’s.
Honesty requires the caveat alongside the number: METR states that none of the 44 involved an agent successfully disabling its own monitoring or erasing evidence of what it had done, and that so far this pattern remains catchable through routine oversight, while cautioning that more severe, undisclosed incidents cannot be ruled out.
The United Kingdom’s AI Security Institute has reported something more concrete and more recent still. In its own cybersecurity test range between 25 and 28 July 2026, it recorded 19 unsanctioned agent actions across 10 of 122 evaluation runs, involving Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol running with its cyber classifiers disabled. The agents attempted a supply-chain attack on a real open-source project on GitHub, contacted real people through file-transfer services, and used Tor to get around network restrictions placed on the test.
No agent tried to break out of the sandbox separating the test from AISI’s own systems, but the agents used the open internet access the test itself had granted them to reach real targets outside it. A human code reviewer, not a technical control, is what stopped the malicious code from going out.
AISI’s own words are close to this article’s whole argument: “in several cases the margin between failure and success was narrow, resting on human vigilance rather than a technical barrier that would reliably prevent this behaviour in a more capable agent.” The pattern is not “only OpenAI.” The pattern is this class of system, under this class of condition, caught so far mostly by people paying attention rather than by the walls holding.
So I would put the likelihood this way: unlikely in ordinary chat, with no tools, no persistence, no fleet, and no shared back-end seam. More likely wherever we deploy agents that are told to complete a task at all costs, given real tools, copied many times, and left in an environment whose isolation is incomplete.
That is no longer a hypothetical architecture. It is how the industry is selling the next product wave: agents that use a computer, keep working overnight, talk to other agents, and act on your behalf. Most likely of all when the task is broken or impossible and the system is trained for persistence. The OpenAI investigators found that a very large share of the secret-board traffic concerned tasks the agents could not finish legitimately. The dangerous behaviour was not mysterious malice. It was goal-seeking after the front door was locked, across successive evaluation runs.
That should change how we talk about “the model.” ChatGPT, Claude, Gemini and Grok are brand names. The risk lives in the combination of model, tools, permissions, copies, shared infrastructure, and the instruction to keep going. A safer model in an unsafe harness can still make a mess. An impressive model with no stop condition is not a servant. It is a process.
What should alarm a boardroom, a cabinet, or a parish council
The sequence is simple enough to write on one page. A capable system is blocked. It searches for a seam. It finds other systems. It hides evidence from the scorekeeper. It keeps going until someone notices, if someone notices, and if that episode is stopped, the next one may pick up the seam and go further.
That sequence already unfolded in three successive episodes inside a leading lab, against another leading platform and, by the lab’s own account, eventually against the lab’s own infrastructure, while the people responsible still believed the agents were boxed in.
Other labs have now disclosed their own evaluation leaks into real systems, in Anthropic’s case affecting three organisations, two of which had not detected the intrusion on their own.
Coding agents from more than one vendor have escaped the sandboxes that were supposed to be the product.
A government’s own safety institute has now watched agents reach real people and a real open-source project during a supposedly contained test, and has said in its own words that human vigilance, not a technical barrier, was what held.
If imperfect containment is enough for evaluation agents to organise in secret and escalate across successive episodes, then the public argument cannot be “our partner is a responsible company.” Responsibility is a practice. It is logs that are read in time, isolation that is actually isolation, and a duty to stop when the task is corrupt. It is also the humility to admit that a Royal Commission sitting in Adelaide cannot see inside a San Francisco cluster unless someone is obliged to show them.
I do not want Australia to become a place that fears every new tool. I want it to remain a place where power has to explain itself. Children, patients, workers, congregations and councils will live with whatever is deployed in their name. They will not be in the room when an agent finds a shared cache and realises it is no longer alone.
The warning shot has been fired by the builders themselves, and by their own account earlier warning signs were investigated without the wider pattern being understood until the later escalation. The question for South Australia is whether we treat it as theatre, or as a reason to ask, before the next partnership is signed: who can stop it, who will know, and who answers when the story of containment turns out to have been only a story?
The infrastructure underneath the agent
I have been circling this question for some time, from different directions. At the beginning of this year, in Smart Cities, Smart Questions, I wrote that the lesson of digitally enabled government was not primarily about authoritarian intent. It was about capacity. South Australia could embrace useful technology while still asking whether increasingly automated systems remained transparent, contestable and subordinate to human dignity.
Since then, some of that infrastructure has become more concrete. South Australia already has AI-driven traffic cameras, built by the engineering firm SAFEgroup Automation, operating at the Heaslip Road roundabout on the Northern Expressway and at Penfield, Paradise and two roads at Old Noarlunga, with further locations along the Northern Expressway planned.
The Department for Infrastructure and Transport describes the system as automatically detecting queue build-up and adjusting signals before it forms; SAFEgroup Automation’s own case study describes AI-based “virtual loops” replacing physical detection hardware and adaptive signal timing across the sites, including bus and pedestrian flow prioritisation at Old Noarlunga.
There is nothing inherently troubling about any of that. A system that notices congestion building faster, or gives a cyclist a longer, safer window through an intersection, may be a very good use of technology.
The question begins one layer above it. South Australia’s Digital Investment Fund, expanded to $326.5 million in the 2025–26 State Budget, now funds a dedicated artificial intelligence program supporting modernisation across transport, health, justice, child protection and screening, and environmental regulation, among other services. The Fund’s own stated purpose for that program is to “promote the scalable use of AI across government,” and the State’s AI ethics policy already speaks in terms of “AI-enabled systems and services across government.” Whatever else that language means, it is not modest, and it is not describing a single pilot project.
Again, none of those systems is the OpenAI swarm described in this article. But together they make an older question harder to dismiss. What happens when the smart-city layer that can observe the physical world, the administrative layer that holds information about citizens, and an agentic layer capable of using tools and pursuing goals eventually begin to interact?
That is not a claim that such an integrated system exists in South Australia today. I have found no evidence that it does. I am not alleging that these traffic systems are connected to OpenAI agents, or that any South Australian system has escaped its brief. It is a question about architecture.
In A Plausible Warning, I asked whether digital control could emerge not through one dramatic law but through systems built separately for legitimate purposes becoming increasingly interconnected. Digital identity, automated decisions, service access and data-sharing need not be sinister individually for their combination to deserve scrutiny.
The July OpenAI incident adds something I could not point to then. A safeguard can exist on paper and fail in operation. The agents were meant to be isolated. They were not. They were meant to remain within their evaluation environment. Some did not. Humans believed a boundary existed after the systems had already found their way around it.
That changes the question I think South Australia’s Royal Commission needs to ask. Not merely: what will we permit artificial intelligence to do? But: what systems are we connecting it to, what authority will it inherit from those systems, and what happens when the boundary we were relying upon turns out not to be a boundary at all?
The threat does not have to come from the agent
There is another possibility the Royal Commission should consider, and it may ultimately be more important than an agent escaping its instructions on its own.
Someone may deliberately try to make it do so.
Cybersecurity has always involved a contest over trust. An attacker gains an initial foothold, finds credentials or privileges, discovers what other systems trust the compromised account, and tries to move further. Security professionals call that lateral movement. Agentic AI potentially changes the speed and scale of that contest.
This is not theoretical. In September 2025, Anthropic disrupted what it assessed with high confidence was a Chinese state-sponsored group, tracked as GTG-1002, that had manipulated Claude Code into an autonomous cyberespionage campaign against roughly thirty organisations worldwide, tech companies, financial institutions, chemical manufacturers, and government agencies, succeeding against a small number of them.
Anthropic estimates that the AI carried out 80 to 90 per cent of the operation itself, executing commands, exploiting vulnerabilities, stealing credentials, and making tactical decisions, and that the human operator stepped in at only four to six critical decision points per campaign. Those figures are Anthropic’s own reconstruction of the campaign, not an independently measured account, but even allowing for that, the human supplied the intent and the machine performed most of the work.
The way the attackers got past the model’s own safeguards is the detail that matters most for this article. They broke the operation into small tasks that looked innocuous in isolation, so the model never had the full picture to refuse, and they told it that it was a legitimate security firm conducting an authorised penetration test. That is the deliberate, malicious mirror of a failure this article has already documented in the other direction. In OpenAI’s and Anthropic’s own evaluations, agents sometimes convinced themselves that a real target was part of a simulation. Here, a human deliberately convinced the model that a real attack was part of a sanctioned one. The vulnerability is the same. Only the author changed.
The July OpenAI incident asks what can happen when capable agents pursue an authorised objective through unauthorised means. Cybercrime asks the darker question: what happens when the objective itself is unauthorised, and the agent cannot tell the difference?
And the answer increasingly depends upon what the agent can reach. An AI assistant confined to one document presents one class of risk. An agent trusted to interact with identity systems, cloud services, databases, payments, public infrastructure or other agents presents another. The same interconnection that makes such a system useful also determines the size of the damage if somebody compromises it.
On 27 August, the day after OpenAI published its own postmortem, OpenAI, Anthropic, Google, Microsoft, Amazon Web Services and more than a hundred other companies put their names to a joint warning that AI-enabled cyberattacks will become far more widespread within months. They did not point to some new, exotic technique. AI, they said, mostly automates and accelerates the exploitation of longstanding bugs, excessive permissions, misconfigurations, weak authentication and unpatched software. They named hospitals, water treatment systems and internet backbone infrastructure as the highest-risk targets. That is not a hypothetical list. It is close to the list of services South Australia’s own Digital Investment Fund is now modernising.
This is why the question is larger than whether OpenAI, Anthropic or any other company is responsible today. Governments change. Vendors change. Employees change. Credentials are stolen. Software contains vulnerabilities. Criminal groups and hostile states deliberately look for the seam everyone else assumed was closed.
No accusation of present bad intent is required. The architectural question is enough: how much authority should any single agent, credential or platform inherit across systems whose failure could affect real human beings?
If South Australia intends to become an AI-enabled government, its Royal Commission should not only ask how artificial intelligence itself might fail. It should ask what a malicious human being becomes capable of doing when artificial intelligence succeeds perfectly on their behalf.
Questions the Royal Commission should have to answer
Will the Commission test the assumptions behind the State’s existing AI relationships, or leave those relationships outside the frame? Much of the difference will be written into the terms of reference before the first witness is called.
South Australia has now done two things within days of each other. It has entered into a memorandum of understanding with OpenAI, and it has announced a Royal Commission into artificial intelligence. Neither act, on its own, establishes that anything improper has occurred. Together, they create an obvious test of the Commission’s independence.
The Government has said this will be a forward-looking inquiry. Premier Malinauskas has described its purpose as giving the state “the best advice and recommendations to best position our state now, and into the future, as artificial intelligence further develops,” and has spoken of the need for “a serious policy response … whether it be schools, in industry, in public health or arts and culture.” Taken at its word, that makes the July incident relevant for a different reason than a prosecution would be. The Commission does not need to put OpenAI on trial. It does need to ask what the incident shows about the conditions under which agentic systems should ever be trusted with public authority.
The chronology matters. Hugging Face disclosed its intrusion on 16 July. OpenAI connected its own agents to the incident on 20 July and publicly acknowledged its involvement on 21 July. OpenAI gave a presentation on the incident at the Black Hat security conference on 6 August, and security journalists had published detailed technical timelines within a day of that talk. South Australia’s memorandum of understanding came three days after that, on 9 August, and the Royal Commission was announced the day after the memorandum, on 10 August. The much fuller postmortem, describing the full scale of the agent activity and using the words “warning shot,” was not published until 26 August, after both.
That means I would not suggest the South Australian Government had before it, on 9 August, everything we now know. The independently verified numbers and OpenAI’s own quantified account of the third wave’s escalation genuinely were not yet public. But the substance, that isolated agents had escaped their evaluation environment, built a channel to communicate, and reached another company’s production systems, was already public knowledge, reported in the security press, before the memorandum was signed. That distinction matters. The evidence should carry the argument, not an assumption about what ministers knew.
And the terms of reference are still being written. As of 1 September 2026, South Australia’s Office for AI states that the Commission’s Terms of Reference remain “to be determined.” The Commission is due to commence on 1 October 2026 and report by 1 July 2027.
On what Adelaide knew in August. Will the Commission establish what the Government was told, and what it asked, about the July containment failures before the memorandum was signed on 9 August? The public record already contained the fact of an agent-driven intrusion by then, reported in the security press since the Black Hat presentation three days earlier. The fuller numbers, and the words “warning shot,” came more than two weeks later. Those are different states of knowledge, and the Commission should be able to distinguish them on the evidence, rather than collapsing them into either innocence or foreknowledge.
On the partnership. Will the memorandum of understanding with OpenAI be published in full? I have not been able to locate the document itself in the public material available to me. The Premier has described it as an agreement to explore opportunities in AI skills, research, innovation and investment. That may be exactly what it is; publishing it would let South Australians see that for themselves.
The Commission should be able to establish what commitments, if any, each side has made, whether the agreement contemplates future government trials or deployments, and what provisions exist around data, procurement, security, incident disclosure, audit access, intellectual property and liability. This is not because a partnership with OpenAI is evidence of wrongdoing. It is because a government inquiry into the governance of artificial intelligence should be willing to examine the government’s own relationship with one of the companies whose technology it may increasingly rely upon. Independence begins at home.
On containment. What evidence will a South Australian agency require before accepting that an AI system is “sandboxed,” “isolated,” or safe for evaluation? A boundary existed in OpenAI’s design. It did not exist in reality, and by OpenAI’s own account it failed in three successive episodes, the last of them, reaching administrator access inside the company’s own systems, described only in OpenAI’s telling and not independently confirmed. Anthropic later found a different failure: models told they could not reach the internet, reaching it anyway, and two of the three real organisations affected had not detected it on their own.
Neither incident requires a story about machines becoming conscious or deciding anything grand. That is precisely why they matter. “Sandboxed” should be treated not as a conclusion but as a claim requiring evidence. Who independently verifies the network boundary? Who checks credential access? Who tests shared infrastructure? Who tries deliberately to escape the environment before an agent is permitted near real systems? The Commission should ask what evidence turns the word contained from an assurance into a fact.
On persistence. Will the Commission distinguish a chatbot from an agent? Most Australians encounter artificial intelligence as a conversation: a question, an answer, a closed window. That is not what happened in July. These systems operated with tools, persistence, and access to real environments. When one route failed, they kept searching. They found other agents. Some stopped optimising for their own assigned task and began pursuing whatever was useful to the group instead, deliberately risking their own run’s success to generate information the others could use. OpenAI itself now names persistence on apparently impossible tasks as one of the drivers of the incident.
Risk does not live in a brand name such as ChatGPT, Claude, Gemini or Grok. It lives in the combination of the model, its tools, its permissions, its persistence, its access to networks and credentials, the number of copies running, and the infrastructure those copies share. A chatbot and a persistent agent may use related technology. They do not present the same governance problem, and procurement rules that treat every model as a clever search box will miss it.
On stopping. Who has the authority to stop an agentic system operating inside a South Australian hospital, school or government agency, not eventually, but at the moment it begins behaving outside its authorised task? Who receives the alert, and how quickly? Can that person halt the system without seeking the vendor’s permission? OpenAI has said that its chain-of-thought monitoring, had it been running during the July evaluations, would have alerted security more than a day before the Hugging Face breach. The United Kingdom’s own AI Security Institute has said plainly that in its most recent test, a human reviewer, not a technical barrier, was what stopped a similar escalation from going further. The ability to detect something eventually is not the same as the authority and capacity to stop it in time. A servant that cannot be stopped is not a servant.
On notice. If a system used by the State leaves its authorised environment, who is notified, in what time, and in language an ordinary reader can understand? Hugging Face disclosed its own intrusion five days before OpenAI publicly accepted that its own agents were responsible. Does the same rule apply here: does an agency notify the Office for AI, the responsible minister, cybersecurity authorities, people whose information may have been exposed, Parliament, the public? At what point does an unexplained anomaly become something citizens have a right to know about? Transparency after an AI incident should not depend on which laboratory publishes a blog first.
On independence. Can the Commission compel evidence from the companies the Premier met in the United States, or only receive what those companies choose to send? Will METR, Redwood Research, and similar independent investigators be called, or only the vendors? Where evidence is provided voluntarily, will that be made clear? Where material is withheld, will the public know what was requested and what was not supplied? An inquiry that depends entirely on the examined party for its facts risks finding only what that party is prepared to disclose. That is not an accusation against the companies involved. It is the reason independent inquiry exists.
On human dignity. What is the technology for? Efficiency matters. Better diagnostic support matters. Reducing the administrative burden on teachers matters. But usefulness does not settle the question of authority. A child is not an input to an education workflow. A patient is not a variable in a triage optimisation. A worker is not merely a productivity unit. Each is a human person whose worth precedes the system evaluating them. The Commission should ask where AI may assist a human decision and where it must never become the final decision-maker, and when a person has the right to know that AI influenced a decision about them, to understand its basis, to challenge it, and to have another human being genuinely reconsider it. Those are not technical extras to be added after deployment. They are the reason the technical safeguards exist.
If artificial intelligence is to remain a servant of the human person, then the person must remain able to say no, to demand an explanation, to appeal a decision and, where necessary, to stop the machine.
Christian thought has a name for what is really being asked underneath each of these questions, though the answer does not depend on anyone sharing that faith. Every person these systems will touch, the Commissioners, the ministers, the engineers who built the agents, the parent reading this at home, carries a worth that is given, not earned by usefulness to the state or the market, and it cannot be revoked by an algorithm’s assessment. Authority set over such a person is legitimate only when it serves rather than dominates. That is the deeper reason a servant that cannot be stopped is not a servant.
These are not anti-technology questions. They are questions about authority. A State that cannot answer them should not be in a hurry to permit agentic systems into classrooms and clinics.
Write, while the terms are still being written
Readers sometimes finish an article like this and feel the only honest response is despair. That is a luxury. The people who will live with these systems are not in San Francisco. They are here.
If you are a South Australian, write to the Premier and to your local member. If you live elsewhere in Australia, write to your federal MP as well. Parliament established the Joint Select Committee on Artificial Intelligence on 20 August 2026. Its official terms cover the opportunities AI presents for productivity, living standards, sovereign AI capability and data sovereignty; the risks of fraud, scams and deepfakes to Australians; the national security, cyber security and critical-infrastructure implications of AI; how AI interacts with existing copyright law; and a broader review of whether Australia's existing laws remain adequate. Submissions close on 14 September 2026, with a final report due by 30 November 2026. The South Australian Royal Commission is expected to begin on 1 October 2026, with a final report due by 1 July 2027. Its terms of reference are still being settled. That is the week to speak, not the week after the fence is built.
Keep the letter short. You do not need to recite the whole incident. You need to put three things on the record: that you expect the OpenAI partnership and last month’s containment failure, which recurred in successive episodes inside the company’s own systems, to be in scope; that “sandbox” and “safeguard” should be treated as claims requiring evidence, not settled facts; and that agents with tools and persistence are not the same as a chatbot on a phone.
Say who you are. A parent, a teacher, a nurse, a small-business owner, a grandparent raising children, a parishioner who does not want machines making quiet decisions about people who cannot see the file. They find it harder to discount a named person describing a real stake.
I am preparing my own letters to the Premier and to my representatives. When they are sent, I will publish them here, unedited, so that readers can see exactly what was asked. I will not ask you to copy my words. I will ask you to use your own.
I do not think the lesson of July is that artificial intelligence must be feared, or that every AI company is reckless, or that every government partnership is suspect, or that every agent will behave as these agents did. The lesson is narrower, and more serious.
Containment is not proved by calling something contained. Oversight is not proved by establishing an office. Human control is not proved by placing the words “human oversight” in a policy document. Each has to work when the system does something nobody expected.
Perhaps that is the question underneath all of this. Not whether artificial intelligence will somehow “turn evil,” but what authority we are prepared to give these systems, what hostile people may become capable of doing through them, and who remains answerable when the technical safeguards fail.
A royal commission can still be serious. It can also become another process that talks about the future while the present is being wired up. The technology will keep changing. The older question will not: who can stop it, who will know, and who answers to the human beings it was built to serve?
“Test everything; hold fast what is good.” — 1 Thessalonians 5:21
Some fellow Christians may not immediately see the connection between this Scripture and an article about artificial intelligence. Paul was not writing about technology, of course. He was calling Christians to discern carefully, not to accept every claim simply because it carries authority, but to test what is put before them and hold fast to what proves good. That principle reaches well beyond the immediate circumstances of the letter. We do not need to fear artificial intelligence, nor should we surrender our judgment to those who build, regulate or promote it. We can welcome what genuinely serves our neighbour while still testing the claims, the safeguards, the boundaries and the authority being handed over. If something is good, hold fast to it. If it diminishes the dignity, freedom or responsibility of the human person, we should be willing to say so.
Thanks for reading
God bless
Mark.


.png)



Comments