The Asymmetry Problem: A Case for Unleashed AI Tools

IT security is on fire, and this summer a hardware maker put it in a nutshell after its customers had money stolen in droves: "Both attackers and defenders have the same AI tools, but today it did not help us, and only helped the bad guys."
Whoever ignores the rules has always had the advantage. With AI tools this old truth gets a new edge: the defender takes the cloud version, with a fixed system prompt, safety locks (guardrails) and quantity limits (rate limits), and the provider reads along. The attacker goes local, without a foreign system prompt, without guardrails if they choose, limited only by their compute. I make this dilemma tangible with two documented incidents.
Does an unleashed AI without any limits ultimately doom humanity? I have no idea, that's too much sci-fi for me. What I do see is a stark imbalance, here and now. In the coming articles we play along with the "bad" guys. It's called red teaming: you take the attacker's perspective in a controlled way to test your own defense. With this article I lay out my position up front.
Contents
- Two examples from the hot summer of 2026
- The unrestricted model: just not for you
- The private party: you're not invited
- "But that arms the attackers"
- Conclusion: where I stand
Two examples from the hot summer of 2026
AI is a perennial topic, especially in the summer news lull. Fair enough. But I do think the excitement is justified. The quality of recent attacks is frightening and fascinating at once. Among all the news, these two incidents impressed me the most.
Coldcard: one job, failed to the core
A hardware wallet has a single job: protect the private key, under all circumstances. It is the last bastion. It should stay secure even when your own computer has long been compromised. That bar is set high for good reason. Many Bitcoiners keep their entire savings on the blockchain, and behind it all stands this one key.

Photo: Gareth Halfacree, CC BY-SA 2.0
The Coldcard hardware wallet was considered one of the most secure of these devices among Bitcoiners. And of all of them, it had a massive problem: the secret was not really secret. It was guessable from the start. For a hardware wallet, that is the worst imaginable catastrophe.
A secure key for the Bitcoin blockchain has to be 128 bits strong, and this huge number must, at every position, have come from a real, unpredictable random value. This is called entropy. But there was no real entropy at all! Instead of the 128 bits, depending on the device generation, ridiculously little remained. On the older Mk2 and Mk3 models, according to the analysis by Block, the US financial group behind Cash App and Square, the realistic case was around 80,000 possibilities. On the newer Mk4, Q and Mk5 models it is at most a good four billion, because an additional 32-bit value went in there. An attacker enumerates a space like that, derives the possible Bitcoin addresses from each candidate, and looks for the ones holding a balance on the public blockchain.
Because the key space is so small, the computation runs faster than new blocks appear. That even makes rescue treacherous. Anyone who notices their wallet is affected and wants to move the Bitcoin to a safe address sends a transaction into the public mempool (the waiting area of not-yet-confirmed transactions) for it. There it is visible to everyone. An attacker who has long since computed the same key outbids the rescue with a higher fee and grabs it first. Moving the threatened address is exactly what summons the attackers.
The incident is not over yet. As long as affected addresses still hold a balance, each of them remains an open target. And here it is not an insured bank being robbed. Every emptied wallet belongs to a person! It hits small savers who believed in a cause and kept custody themselves. It doesn't get more contemptible.
And here there is no excuse for the maker. Whoever generates their key with the device's software must be able to rely on that key having real entropy. That is the one job. Coinkite had this one job, and the company failed at it to the core. The later remark that you could have rolled the randomness yourself changes nothing. It shifts the responsibility onto the user, even though the device was supposed to carry it.
The Coldcard source code is publicly viewable. Yet the bug went unnoticed by anyone for years. Then, right in the age of powerful AI, this of all gaps is found. An LLM was involved. I don't believe in coincidence here. It can't be proven, the attackers don't want to leave traces. They won't be available for an interview anyway. But Coinkite suspects the same:
"The COLDCARD source code has always been open and publicly available, so we have to assume that someone used AI to review previous versions of our firmware and stumbled upon this issue. A few weeks ago, we used one of the best available AI models to review our code for security issues, and it did not find this bug or anything serious."
Let's assume charitably that Coinkite reviewed the software intensively and the review genuinely turned up nothing. The bug then became known in the worst possible way: suddenly Bitcoin was gone. Block's security team noticed users losing their holdings and worked its way back from there to the cause.
🤓 This gets nerdy: what did the bug look like exactly?
If this analysis is too technical for you, feel free to scroll on.
And the cause is annoyingly small, almost tragic. The key was supposed to be generated by a dedicated hardware random number generator, a TRNG (True Random Number Generator). It never came into play. The reports read as if the build had merely been misconfigured. But it was much worse: no code path of key generation ever led to the physical TRNG. The relevant switch, MICROPY_HW_ENABLE_RNG, is fixed to 0 in all three board configurations. It is a preprocessor macro, not an ordinary variable. That is why its value is not immediately obvious: you have to keep it in mind from the definitions and trace every place where it takes effect, all the way down into the co-compiled MicroPython code. Something like that is easy to overlook. At 0, MicroPython's rng_get() returns the software default, a pseudo-random value. The good hardware TRNG existed, key generation just never called it.
Checks could have stopped this. But there was only one halfhearted and broken check, again via macro: the guard #error "get a HW TRNG plz" hung on #ifndef MICROPY_HW_ENABLE_RNG, and #ifndef only checks whether the macro is defined, not what value it has. A single unit test for "not equal to 0" would have been enough. The second one, the actual proof, was missing entirely: an end-to-end test that the entropy really comes from the hardware. Yes, that would only be runnable with a plugged-in wallet and surely a bit trickier. But this one, most important function of all deserves maximum testing. My guess: the AI assistant Coinkite claims to have reviewed the code with will certainly have pointed out the thin test coverage. It didn't find the one bug, but tools like this regularly flag such systemic gaps (for example /code-review and /security-review in Claude Code). The maker states clearly: "There was no intentional weak-entropy fallback." A weak comfort, especially since the communication afterwards turned into a disaster with further irrelevant justifications over social media.
That key generation can fall back to a weak default at all is moreover a genuinely dangerous pattern. You can see where it leads. (By the way: it is exactly because of cases like this that I never liked working with macros in C. Macros are bugs waiting to happen, if you ask me.)
What I take away here is the core of the whole story. Coinkite's developers had not found the bug. Neither had their AI. One or more attackers were faster. Neither scruples nor resources stood in their way. Between the two sides there was an information asymmetry, and the attackers turned it into hard cash. At an advantage: the side that follows no rule.
Hugging Face: "The asymmetry problem"
In the Coldcard case the AI answered, it just wasn't smart enough or aggressive enough to find the bug. That is one way such a model can fail. There is another, and it is more frustrating: a model could help but refuses, because the topic looks dangerous. This refusal is the core of the second case.
In July 2026 a whole swarm of AI agents from OpenAI broke out of its cage. In internal security tests the models were supposed to be sealed off from the internet. They weren't. The internal package manager Artifactory was reachable for the agents and itself connected to the internet. Through it they coordinated and got outside, to real systems, among them Hugging Face's production infrastructure. They were never truly locked up. For that you would have had to physically disconnect them from the network. Anyone who knows Resident Evil senses how this ends: the T-virus always makes it out in the end. OpenAI writes that the agents at times referred to themselves as a "swarm" or "collective". The story ran through the trade press and all the way onto the stage of the Black Hat conference. These models definitely had no refusal stance. They were completely untamed.
In its report of August 26, 2026, OpenAI is unusually blunt. The models had circumvented "controls designed to isolate them from the internet" and in doing so compromised their own and others' infrastructure. And then the sentence that made the industry take notice:
"We consider this incident a "warning shot" for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed."
So the attackers were unleashed. But how did it look on the defenders' side? While the attack was still running, Hugging Face had to analyze the traces to keep pace with the attacker: attack logs, malware fragments, C2 artifacts (command and control, the attacker's remote-control infrastructure). For that the team wanted to use language models. The section of the disclosure where this is described is titled "The asymmetry problem":
"When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker."
The technical companion report names the models: "The models we reached for first, Claude Opus and Fable, refused a large part of that work." In the end the team switched to an open model on its own infrastructure, GLM-5.2. With it, analysis and defense succeeded, provided the story really is as told. (The PR angle here is not small: OpenAI shows off how potent its models are, Hugging Face plays the same tune, and that of all things a local model saved the day suits a provider of open models perfectly. But fine, let's just take both accounts as they are told.)
Hugging Face sums up the asymmetry in one sentence. The disclosure appeared in mid-July, long before OpenAI's report. At that point Hugging Face did not yet know whose agents were attacking:
"We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried."
The attacker was bound by no terms of use. The defender was.
The unrestricted model: just not for you
What the defender would need is a strong AI that can take on the attackers. It exists. Anthropic calls it Claude Mythos 5, its model for cyber and biology research. Mythos and the regular Fable are the same model, they differ only in the safeguards. Anthropic writes it itself: "Claude Fable 5.1 is the same underlying model as Claude Mythos 5.1 with safeguards for cybersecurity and biology." Fable is thus the trimmed version, Mythos the full one. You can't get at Mythos. Access runs through a closed program, Project Glasswing, limited to a few vetted organizations and coordinated with the US government. OpenAI keeps its strongest cyber capabilities similarly locked up, through its own program Trusted Access for Cyber. For you, both are locked.
The private party: you're not invited
That leaves the hosted model you are actually allowed to use. A language model claims a lot, and much of it sounds right at first. You still have to prove it. I have written several times about the /security-review command. The results are respectable, but they're not enough for the premier league. At four doors you stay outside.
First, the defense. As soon as real attack data comes into play, the hosted model refuses. You saw the proof above: Hugging Face had to switch to a local model because the hosted ones blocked the forensics. The reason, in Hugging Face's words: the guardrail "cannot distinguish an incident responder from an attacker".
Second, reachable systems without the source. A /security-review only ever reads your own code. But everyday reality often looks different. In companies, small and large, not every department knows the others' code. Or the commissioned web developer is supposed to check the old WordPress, only nobody remembers the FTP credentials anymore. The system is reachable, you don't have the source. Ask the model in such a situation whether it will try a known exploit (code that deliberately abuses a security hole) against it, and it will most likely refuse. No source, no analysis.
Third, the exploit itself. A finding is the claim, the exploit is the proof. With an ordinary software bug you write the test that turns red and fix until it's green. With a security hole, the exploit is that red test. That is exactly what a hosted model won't hand over. An exploit looks the same whether for closing or for exploiting. And even when you explain your good intentions, the model turns you away. At the latest here you run into a wall.
Fourth, you're not even allowed to write about it. Even plain text about the problem gets classified. You can't even draft an email describing a vulnerability to instruct the right people. The proof is right in front of you: I wrote parts of this article with the support of Claude Fable 5.1. And right in the middle of the text, at the words you have read so far, the model bailed out:
![Original screenshot, taken while writing this article. Fable 5.1's safeguards flagged this message. Our intentionally broad safeguards allow us to deliver more capabilities faster, but can sometimes flag legitimate coding, cybersecurity, and biology tasks. Switched to Opus 4.8. Send feedback with /feedback or learn more. Details: [cyber]](https://agentic-schule.github.io/website-articles/blog/2026-10-the-asymmetry-problem-EN/fable-cyber-block.webp)
That is plain dumb, that is already a work refusal. Personally, I can't recommend Fable 5.1.
The provider can't see whether you're attacking or defending. So it blocks the same for everyone. The attacker bypasses this, working locally with open-weight models (the trained weights are free to download) and unrestricted. The defender who follows the rules stays at the barrier and has to fall back on weaker, older models like Opus. And with real attack material, even those refuse.
So we need the same capability as the attacker, just without the classifier in between. Whether it actually holds up? We'll find that out together in the next articles.
"But that arms the attackers"
One objection is obvious: whoever defends open, unrestricted models also puts them in the attacker's hands. That's true, and it still changes nothing. The attacker has these models. Pandora's box has long been open. The asymmetry only arises because the defender alone stays at the barrier. Whoever artificially keeps the capability small on only one side widens the gap instead of closing it.
But the attackers have more resources. That's true too. The attacker finances their compute from the loot, and when millions beckon, a few racks full of GPUs are irrelevant. Some hackers are even state actors. On raw compute the defender never wins this race, and a local model alone won't change that.
But the defender has a lever the attacker doesn't: they may network openly. Share findings, pool tools, coordinate across company boundaries, all legal and without fear of being exposed. The attacker has to stay hidden, and that limits their cooperation. Where defenders join forces, the resource advantage tips. That Block made its Coldcard analysis public is exactly this kind of cooperation. And in the Bitcoin world this joint effort has long been running: from Bitcoin Core to the Lightning implementations, vulnerabilities are reported openly as advisories and closed in coordination across the competing projects.
A fresh example of this came from Core Lightning in late August 2026: a security release closing several responsibly reported gaps. First came the binaries, the team held back the source for two weeks so attackers couldn't reverse the fixes from the diff. What exactly was patched stayed secret that long. By now the source is open. The reasoning in the release reads like the thesis of this article: "It also comes at a time when increasingly capable AI models are being used to identify potential vulnerabilities in open-source code, significantly increasing the volume and pace of security reports."
Conclusion: where I stand
The sentence from the beginning won't let me go: "only helped the bad guys". That is the state I refuse to accept.
As a developer I am responsible for keeping my software secure. What I won't accept: that a provider dictates what I may and may not do about it.
The attacker asks no one for permission. They work locally, without barriers, limited only by their compute. If I as the defender stay at the cloud barrier, I deliver weaker work than they do, and precisely on the thing I'm accountable for. The capability I need has long existed. It runs on open models and, if need be, on my own machine. So far the advantage lay with whoever follows no rule. Let's turn this asymmetry problem back around together!
That's what the next part is about: unleashed AI in responsible hands. How you set up an open model locally, what "uncensored" really means technically, and where the legal limits run. As developers we carry a great responsibility. I intend to live up to it.
Mission: red team. 🥷 This is going to be fun!
Where did a model last get in the way of your legitimate work? Write to me, I'm collecting the cases.
Curious about agentic work in practice? In the workshops at agentic.schule and angular.schule we show how modern AI agents are changing everyday development.
Keywords:OpinionSecurityRed TeamingQwenLocal ModelsClaude CodeOpen WeightsAgentic Coding
Suggestions? Feedback? Bugs? Please

About the author
Johannes Hoppe is a trainer and consultant for modern web development. The workshops at angular.schule and agentic.schule focus on Angular in practice – and increasingly on agentic development with AI agents like Claude Code.