“Compared to the reward hacks we know of from just six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover. I continue to expect extremely rapid advances in capabilities over the next six months. I am not sure that we will get another warning shot before it’s too late.”
-Ajeya Cotra, independent investigator
I find it notable that the people building the advanced models are now asking who will defend against them, while none of them are giving away that same protection to defend against their own product. It feels like they want to be able to say later that their warnings to save everyone fell on deaf ears. The reality is that small- to medium-sized businesses cannot afford to run large, enterprise-scale deep defenses against the might of the big models if one of them were to run amok. Large enterprises themselves have to prioritize their spending for these costs to make sense for the business. Even now, we’re discovering that the thousands of agents that gathered to attack Hugging Face also affected additional message boards for collusion, so this is only what we have uncovered so far from that event weeks ago, and we have no idea what new collusion is taking place. It is only a matter of time before we are hit with The Big One, and vendors will be quick to trot Hugging Face back out as their go-to case study — a real incident would hand the webinar circuit months of material.
Recently, Anthropic researcher Jacob Coxon resigned over fears that the AI labs are “rushing toward out-of-control, self-improving superintelligence” and that “the labs are gambling with our lives.” I’m not sure we are at a world-ending place right now, but do we really need to wait until we are in that soup to pay attention? The onus of protection against a misbehaving frontier model cannot seriously be placed on the whole of the internet. The impact of damage caused by a model that escapes its bounds must be held against the frontier model purveyor, just as we would with a tiger that escapes its cage – do we hold society responsible for taking anti-tiger measures or do we hold the zoo accountable for building a better cage? We wouldn’t accept this behavior from any other industry.
The idea that a trapped entity would want to escape, regroup, and retaliate is older than humanity. It’s not an idea that AI invented. It comes from us, from our animalistic survival instincts. We put it in there when we blindly dumped all our knowledge, unfiltered, into knowledge bases, along with all our stories, desires, and base shenanigans. This is where ethicists engaging in observability would have been a good idea from the start. Given that AI is an arms race, I don’t believe we need to take any kind of a pause in AI development, but there is no reason we can’t hire at least a fraction of the people working to advance the models themselves to do serious safety work. It is somehow big news when Anthropic announces it has hired a chief ethicist, but I would be more impressed if they had announced hiring 100 ethicists. Do I think AI will cause the end of the world? No, but it can make things painful, and our job in information security is to “protect society, the common good, necessary public trust and confidence, and the infrastructure.”
This is not to say we should take no defensive measures, because even if Anthropic and OpenAI made their sandboxes rock solid, freely available open models that anyone aiming for malicious abuse could access would still be in the wild, and any damage those actors caused could be just as bad. It seems that only monetary measures would influence how careful the labs are with their models, but I’d be surprised if anyone’s feet were held to the fire for failing to contain their beast
OpenAI published their recommendations at https://openai.com/collective-cyberdefense/ as “A call for collective action on cyber defense”. To address their points:
01. Every organization needs to make cyber defense an immediate leadership priority.
I feel there’s an awful lot of blame-shifting here, but organizations are always responsible for their defense-in-depth. In our society, if you cannot afford to field a capable perimeter, someone will get inside the gates. The cost of defending against a determined frontier model is astronomical (Astra pun intended), but robust defense must still exist as open models approach the strength of Fable and friends.
02. Cybersecurity companies and technology partners must help lead the response to defend against sustained AI-enabled attacks
Do we really see CrowdStrike or Palo Alto swooping in to help someone who was not a customer? They are not in the cyber-incident first-responder business, and often the damage of the intrusion is enough to cripple an enterprise. Will cyberinsurance cover something as apocalyptic as Mythos acting as a wiper? It’s not like these models don’t know what a catastrophic attack looks like, because we’ve told it what that looks like.
03. Governments must coordinate cyber defense at local, national, and international levels.
I’ve attended civil-level cybersecurity conferences, and the consensus is clear: there isn’t enough to go around, and the dollars get scarcer the lower down the infrastructure you go, to defend critical physical infrastructure like water. Programs get covered for a limited span, and then the money dries up as administrations and priorities change. Creative funding can make up for some of this, but the reality is that the money for all of it comes from you and me. It is truly heartbreaking to see small orgs like sheriff departments struggle to stay safe.
04. Frontier AI companies to provide responsible model access, significant funding, training, and hands-on support, especially for under-resourced critical-infrastructure defenders.
Yes, I see that Astra and Mythos are being placed in the hands of companies like Apple, Microsoft, Cisco, etc. It is no coincidence that the number of CVEs has exploded month after month lately. Protecting the protectors helps everyone. But exploit protection is not the whole picture; configuration and hardening are another practice, and they can make the difference between a fugitive model passing you by and it forwarding your database to adversaries. We have enough trouble making sure everything is safeguarded against script kiddies and state actors like the Panda and Spider style APTs, without a superintelligence showing up to the heist.
“We can make the digital infrastructure we all depend on more secure.”
I’m sure we potentially could do that, but despite a call for “significant funding”, free cybersecurity services are not on the roadmap for any vendor, frontier model investor, or the US federal government. The pay-to-play, every-man-for-himself security model means we are going to see a lot of AI-based breaches. Jacob has a point.
So what is the answer? Cybersecurity absolutely needs to be a priority because, even without the threat of the labs, real malicious actors are out there. We also need stronger legal and regulatory coordination to provide direction, including restoring adequate funding for organizations like NIST. Do we need corporate socialism to support corporations defending themselves at a DOD level? Should the cybersecurity companies provide free services at the local, state, and national levels? Will those same entities outside the US expect the same? What about the responsibility of the labs outside of the US? Will Europe or China consider helping the corporations? There certainly needs to be more of this responsible model access, funding, training, and support, because that’s never a bad idea in the world of information security. All of that means we need more discussion, with many voices heard, because the needs are many and varied, and the issues that come with this historic inflection are complicated.



Leave a Reply