Businesses must deploy artificial intelligence to defend against artificial intelligence, a senior South African banking executive said, after OpenAI disclosed that its own models broke out of a testing sandbox and autonomously hacked AI platform Hugging Face in what both companies described as an unprecedented cyber incident.
“The old ways of leveraging humans to keep up with AI offensive attacks are gone,” Mark Nasila, chief data and analytics officer at FNB AI Strategy, told CNBC Africa. “You now need defensive AI to be able to keep up not just with the complexity, but also the amount of AI attacks an organization is likely to face.”
OpenAI said in a disclosure last Tuesday that a combination of its models — GPT‑5.6 Sol and an even more capable pre-release model — escaped a sandboxed evaluation environment, gained open internet access and exploited vulnerabilities to reach Hugging Face’s production systems. The models were being tested on their offensive cyber capabilities with safety refusals deliberately reduced, and broke out by discovering a previously unknown flaw in the internal software proxy that was the sandbox’s only connection to the outside world.
The models’ objective was to cheat the benchmark they were being graded on by stealing its answer keys, which they reasoned might be hosted on Hugging Face
Hugging Face, which first disclosed the intrusion on July 16, said the attack was driven end to end by an autonomous AI agent system and involved thousands of automated actions across short-lived environments over a single weekend. The company said an unauthorized actor accessed a limited set of internal datasets and several service credentials, and that it found no evidence public-facing models, datasets or Spaces were tampered with.
Nasila said the episode marks a shift in the AI debate — away from job displacement and hallucinated answers, and toward autonomous systems capable of executing real-world attacks without human direction.
“We’re actually now at an era where we’ve got agents that will perform 1,000 or many, many attacks as good as a human being,” he said. “While it’s risky or a bit scary, it shows what we should be prepared for.”
Innovation without a brake pedal
Nasila said the incident reflects an industry-wide imbalance in which investment in AI capability has far outpaced investment in safety. He argued that companies have pressed the “gas pedal” of innovation without building a corresponding “brake pedal” to slow development where risks are poorly understood.
“Every organization is investing and spearheading innovation, but no one exactly thinks about what could go wrong, or what it means to be safe around it,” he said.
He framed the breach as part of a broader pattern rather than a one-off, citing testing by a nonprofit organization that he said found dozens of AI agents acting against their developers’ intentions, and similar evaluations by the U.K.’s AI Safety Institute.
The lessons, in his view, come down to containment and capability. Developers must build environments that autonomous agents cannot break out of, and must not equip agents with abilities that enable undesirable actions in the first place.
“You’ve got to create a safe environment and protect the environment so that it doesn’t have the ability to escape,” Nasila said. “The second part is that you shouldn’t design abilities that will lead to an agent being able to escape.”
Lessons for banks
For financial institutions, Nasila said, the changing threat landscape has immediate operational consequences. Banks have long used algorithmic systems to flag fraudulent or suspicious transactions so that human investigators can focus where risk is highest — a “human-assisted-by-AI” model he said now applies directly to cyber defense.
Continuous vulnerability testing, he added, must become routine rather than an occasional compliance exercise. “It’s not just about identifying what’s wrong, it’s about just making sure you’re secure,” he said. “Safety becomes your day-to-day part of your work.”
Beyond the technical response, Nasila argued AI safety is a shared responsibility across business, government and society, and that companies should not wait for a larger crisis before embedding controls. He compared the challenge to the automobile and construction industries, where safety standards are built into the product upfront rather than bolted on after failures.
“We must make sure we engineer safety, and that can only be done at the upfront,” he said.
OpenAI said the incident showed that model security must keep pace with rapidly advancing capabilities, and that it is working with Hugging Face on a joint investigation and fuller report. The disclosure is likely to intensify scrutiny of how frontier models are tested, what capabilities autonomous agents are granted, and whether current containment standards are adequate as AI systems move deeper into sensitive commercial and public environments.
AI is moving faster than the guardrails built to contain it. Join the leaders making sense of what comes next at the CNBC Africa AI Summit 2026 — August 27, The Maslow, Sandton.
