Part two of a series. Read part one: The Second Hacker Never Sleeps.
Strip away the noise and a real question sits underneath this year's autonomous hacking boom: who can actually do it, and who is just saying the word?
Mathieu Huysman and Tim De Wachter have spent the better part of a decade finding the things other people's tests miss, work they began as students convincing studios that the backdoors they were turning up in their spare time were a serious liability. Today they lead Cyrex, a team of more than twenty penetration testers, load testers and engineers who secure some of the most ambitious launches in games and enterprise, from Dune: Awakening to Path of Exile 2, for the studios that cannot afford to be wrong. As Mat put it recently, Cyrex once worked to stay ahead of the curve. Now, "we are the curve."
So when the team that set the pair hacking standard decides the second seat in an engagement should belong to an autonomous partner, it is worth understanding why. I sat down with Mat and Tim for the thinking behind it, the questions I get asked, and the ones a sceptic should ask.
The conversation
Take us back to the origin of this. You set the pair hacking standard with two humans in the room. Where did the idea of handing that second seat to an autonomous partner actually come from, and was there a moment it clicked?
Mat: Everything we do at Cyrex starts with a single question: how do we deliver the highest possible quality?
When we designed our penetration testing methodology, assigning a single engineer to a project simply didn’t make sense. Real-world attackers rarely operate alone, they collaborate, challenge each other’s assumptions, and combine different perspectives to find weaknesses. We wanted to replicate that dynamic in a professional and ethical way.
After experimenting with different team structures, we found that two senior engineers per engagement consistently delivered the best results. It became the sweet spot: enough collaboration to significantly increase depth and coverage, without introducing unnecessary overhead.
Pair Hacking was born from that philosophy. The only thing that has changed today is who occupies the second seat. Increasingly, that’s an autonomous AI partner that works alongside the engineer, enabling the same collaborative dynamic at a speed and scale that simply wasn’t possible before.
Tim: Pair hacking was always based on the idea that two different ways of thinking produce better offensive security work than one. Different backgrounds, different experiences, different ideas. You challenge each other’s assumptions and bounce ideas back and forth until you find something neither person might have found alone.
That made AI a very natural evolution for us. Frontier models have an enormous breadth of knowledge across attack techniques, vulnerabilities and technologies, but more importantly, you can actually reason with them. You can give an agent what you’ve found, challenge an assumption, explore an attack path and iterate together.
The moment it really clicked was when it stopped feeling like using AI as a tool and started feeling like having another hacker working alongside you. Except this partner can investigate several directions extremely quickly, come back with what it found and change its approach based on your feedback.
That combination is incredibly powerful, and importantly, defenders aren’t the only ones discovering it. Attackers are gaining exactly the same leverage.
Mat, you wrote that in 2016 Cyrex was trying to stay ahead of the curve, and that now "we are the curve." Where does an autonomous partner sit in that arc, is it the next evolution, or a break from how you've always worked?
Mat: It’s very much an evolution of how we’ve always worked at Cyrex, not a break from it.
We’re simply going faster and deeper than ever before. The autonomous AI partner is a major part of that, but it’s only one piece of the equation.
What has always made Cyrex different is the harness: our proprietary setup for testing software across platforms, architectures, and technology stacks. That is where we’ve consistently outperformed the market.
What’s changed now is that we’ve codified that expertise and embedded it into an autonomous partner. So in that sense, the AI partner doesn’t replace our methodology, it amplifies it. It takes the way Cyrex has always worked and pushes it to the next level.
Everyone is shouting "AI pentesting" this year. As the person building it, what is the real difference between an autonomous agent doing offensive work and a scanner with good marketing?
Tim: A scanner essentially asks: “Does this target match something I already know how to look for?” It’s generally driven by predefined rules, signatures and checks against a large database of known vulnerabilities.
An autonomous agent can ask a much broader question: “How can I compromise this target?”
That’s a fundamentally different problem. The agent can interpret what it discovers, form a hypothesis, try something, observe how the target responds and adapt its next action. It can also connect information across different parts of the environment and pursue an attack path over multiple steps.
That’s particularly important when it comes to false positives. A scanner might tell you that something looks vulnerable. An offensive agent can go further and attempt to establish whether it is actually exploitable and produce evidence of that exploitation.
So for me, autonomy isn’t about putting an LLM in front of a scanner. The important difference is the reasoning loop: observe, reason, act, evaluate and adapt.
How do the human and the agent split the work today, and which way is that line moving?
Tim: Automation has always been part of penetration testing. We’ve used it for years to accelerate repetitive work. What’s changing with agents is the amount of the thinking loop that can now be automated as well.
Reconnaissance is a good example. An agent can investigate a huge codebase, inspect services across an environment, analyze screenshots and correlate information far faster than a person could manually.
The same applies once you start attacking something. If you think about breaking into a house, traditional automation was very good at trying thousands of keys against the front door. An agent can try the keys, look at what happens, decide the front door isn’t promising, then go and check the back door and the toilet window instead.
That adaptive loop is the important difference. Fuzzing and other automation have always allowed us to test huge numbers of inputs, but they don’t inherently understand the response and rethink the attack strategy in the way an agent can.
Today, though, the human is still extremely important. LLMs hallucinate. We’ve also found that they can exaggerate findings or assign a higher severity than an experienced tester would. So humans validate exploitability, impact and severity, and critically, make sure the agent remains inside the authorized scope.
There’s also the physical boundary. If we’re testing something involving NFC cards, hardware or another physical system, the agent doesn’t currently have the embodiment needed to perform every action itself. The human becomes its interface to that environment.
So the line is definitely moving toward more autonomy. I think the human role increasingly moves away from performing every individual action and toward directing, validating and handling the parts of the environment the agent cannot yet reach.
When a client first hears that an autonomous agent is part of the engagement, what do they actually worry about, and how do you answer it?
Mat: The reactions are mixed, and they usually depend on the buyer and the vertical.
For some clients, the first concern is privacy and data security. Giving an AI agent access to their IP or source code naturally raises questions. Others are more skeptical about the actual capabilities of AI. And some simply prefer to stay with an expert-led engagement for now, rather than move directly into standalone autonomous testing.
That’s exactly why we support two models.
The first is our renewed Pair Hacking model: a senior engineer working alongside an autonomous agent, with the human still firmly in the loop. That’s ideal for clients who want the benefits of AI, but still want expert oversight.
The second is Deus Rex: our autonomous security testing product for partners who want to be in the driver’s seat and run tests at unprecedented speed, coverage, and independence
So we’re not forcing the market into one model. We’re serving both buyer segments. But directionally, we believe every Cyrex engagement will become increasingly autonomous. That is the future we’re betting on.
You've called pair hacking a standard much of the industry still hasn't caught up with. Does putting an autonomous partner in the second seat widen that gap, or does it just hand competitors a script to copy?
Mat: Pair Hacking is not just a methodology. It’s a mindset, and that is much harder to copy.
We’ve lived this way of working for more than ten years. It shaped how we hire, train, test, report, and think about quality. The industry had a decade to follow, and most didn’t.
So yes, competitors can copy the language. They can say “human plus AI” or “agent-assisted testing.” But they cannot easily replicate the depth behind it: the operational experience, the proprietary harness, and the judgment we’ve built across thousands of real engagements.
Putting an autonomous partner in the second seat doesn’t make the model easier to copy. It widens the gap, because it builds on a foundation that was already difficult to replicate.
You've said 2026 studios learn from simulations, not meltdowns. Two or three years out, what does offensive security look like if you two are right?
Mat: We’re entering an era where the number of attacks, and therefore the number of incidents , will be higher than ever before. I expect this to increase exponentially over the next few years, and in many ways, that trend has already started.
The capabilities of malicious actors have increased significantly. With AI now widely accessible, the barrier to entry has dropped dramatically. It feels like more people than ever can start attacking software, probing systems, and automating parts of the exploitation process. That is genuinely concerning.
At the same time, sophisticated attackers are operating with better tooling, more coordination, and at a scale, frequency, and volume we have not seen before.
Defenders need to match that velocity, ideally exceed it. Software is being written and shipped faster than ever, but traditional penetration testing simply cannot keep up with that pace.
That is where autonomous and AI-driven penetration testing comes in.
Security testing can no longer be treated as a one-off engagement or a final checkpoint before release. It needs to become part of the development lifecycle: running continuously, every time code is written, changed, and deployed.
One thing you want people to take away before we show them what's next?
Mat: The world is changing fast.
From autonomous vehicles in San Francisco to smart cities in Korea, software is becoming the operating layer of the world. It no longer just supports businesses in the background - it increasingly controls the systems people depend on every day.
With that level of dependence, security becomes existential. We need to be ready to protect businesses, infrastructure, and ultimately ourselves as humans, at all costs.
Tim: Hacking has always been a real skill. It takes years of learning and practice to become genuinely good at it, and historically that created a natural constraint on attackers. There were only so many skilled people and only so many targets they could spend their time attacking.
AI is starting to remove both of those constraints.
Someone with relatively little experience can now tap into an enormous amount of offensive security knowledge and have an AI guide them through techniques they wouldn’t previously have understood. At the other end of the spectrum, experienced attackers can automate and parallelize large parts of their workflow.
I remember game studios sometimes saying, “If we get hacked, at least it means we’re popular.” There was some logic to that because attackers had limited time and tended to concentrate on valuable targets.
I don’t think that assumption holds anymore. An attacker increasingly doesn’t have to choose between target A and target B. Their agents can work on both.
Attack is becoming scalable, so defense and security testing have to become scalable too.
If there is one thread running through everything Mat and Tim said, it is that none of this is a departure. It is the bet Cyrex was built on, that two adversaries beat one, taken to its logical end. The craft has not changed. The second hacker has just stopped needing sleep.
That is the work the whole team does, every engagement, every launch: find what others miss before anyone with worse intentions does. Cyrex was the curve a decade ago. It intends to stay there. And we are not finished showing you what that looks like.
Cyrex is an offensive security firm. We find the things other people's tests miss, by hand, for the biggest names in games and enterprise.
Talk to Mat directly · Book a security and load review · hello@cyrex.tech · Cyrex on LinkedIn · Cyrex on X

Written by Ian Matthews
Ian is Head of Marketing at Cyrex, responsible for the full marketing picture - from brand strategy and campaigns to the copy and content that brings Cyrex's work to life. With 8+ years across B2B and B2C, his speciality is bridging the technical and the human with sharp, results-driven messaging.



