In our last piece, we looked at how AI agents have started running real cyberattacks with very little human help. This time we want to make it concrete.
Forget the headlines for a minute. What does an autonomous attacker actually do when it turns its attention to a company like yours? Not a bank with a thousand-person security team. A normal, growing business with a cloud setup, a stack of SaaS tools, a handful of APIs and a team that ships every week.
Here's the walkthrough, step by step, from the agent's point of view. None of it needs a genius. It just needs patience, speed and no need for sleep, which is exactly what agents bring.
Step 1: It finds everything you forgot you had
The first thing an agent does is build a map. Every domain and subdomain, every cloud service with a public address, every login page, every API endpoint it can find. It reads your public code repositories, your job ads, your documentation and even old versions of your website to fill in the gaps.
This is where most companies get their first surprise. The map is almost always bigger than the one the IT team keeps. The staging server from a project two years ago. The marketing microsite an agency set up and nobody switched off. The test API that was only supposed to be live for a week.
A human attacker might find some of these over a few days. An agent finds them in hours and keeps checking, so the moment something new goes live, it shows up on the map too.
What it means for you: you can't protect what you don't know you're exposing, and attackers now have a better inventory of your estate than you might.
Step 2: It goes after the boring stuff
With the map built, the agent starts poking at every point on it. And it isn't looking for movie-style zero days. It's looking for the ordinary mistakes every company makes.
A storage bucket left readable to the whole internet. An admin panel still on its default settings. A software component two versions behind with a known flaw. An API that checks whether you're logged in, but not whether you should be seeing that particular record. An access key someone pasted into a public repo and deleted the next day, forgetting it lives on in the history.
None of these are exciting. That's the point. They're easy to miss, they build up quietly with every release, and checking for all of them across a large estate is exactly the kind of tireless, repetitive work agents are built for.
What it means for you: the risk isn't one catastrophic bug. It's dozens of small ones, each easy to fix, sitting there long enough for something to find them.
Step 3: It tries every key in every door
Sooner or later, the agent picks up credentials. Maybe from an old data breach involving your staff, maybe from a config file it found in step two, maybe from a phishing email it wrote itself after reading your team's LinkedIn profiles.
A human attacker with a stolen password usually tries it in a few obvious places. An agent tries it everywhere, at once. Your email, your VPN and internal network, your cloud console, your code repositories, every internal tool it has mapped. Then it notes exactly what each login opens and what that access leads to next.
This is where small hygiene gaps start to matter a lot. A reused password, a service account with more permissions than it needs, a system that never got multi-factor authentication because it was "only internal".
What it means for you: one weak login used to be one weak spot. Now it's a starting point the agent will explore in every direction it can.
Step 4: It joins the dots
This is the step that turns a list of minor issues into a breach.
On their own, most of the findings from steps one to three would sit near the bottom of a risk report. A forgotten staging server. A slightly outdated component. A support account with a weak password. But put them together and a path appears: in through the staging server, onto an internal network that trusts it, across with the support account's access, and into the customer database it was never meant to reach.
Connecting findings like this used to be where skilled human attackers earned their reputation. Agents are getting better at it fast, because they can hold the whole map in view and test every combination without getting tired or distracted.
What it means for you: severity ratings on individual findings don't tell the full story anymore. What matters is where they lead when someone strings them together.
How to keep pace
The good news is that every step above can be run by defenders too, and run first. The trick is matching the attacker's rhythm without losing the depth that only experienced people bring.
For the enterprises we work with, that comes down to three habits.
Deep assessments when it counts. Before major releases, big infrastructure changes or a new product going live, nothing replaces experienced penetration testers going through your systems properly. That's where business logic flaws and the cleverest attack chains get caught.
Continuous testing in between. Your map changes every time someone ships, spins up a service or adds an integration. Autonomous testing can keep watching it the way an attacker would, so the forgotten staging server gets flagged the week it appears rather than the year after.
Humans on the findings. An agent can flag a hundred issues. Someone who understands your business needs to confirm which ones are real, which ones chain together, and which ones to fix first. That's the difference between a noisy report and a safer company.
Put those together and you're running the same four steps as the attacker, on your own terms, with a head start.
See your company the way an agent does
Everything in this walkthrough is happening somewhere right now. The only real choice is whether you see the map before an attacker does.
We've spent over a decade testing for studios and enterprises, and more than 200 clients trust us to think like attackers. We built Deus Rex to bring that same thinking to autonomous testing: pair hacking, made autonomous, so the four steps above get run on your side first.
See how we're taking this on with Deus Rex at deusrex.ai

Written by Tim De Wachter
Tim is CTO and Co-Founder of Cyrex, with a decade of experience shaping the company's technical direction across application security, game security, and high-load architectures. He leads Cyrex's security and load testing engineering teams with a focus on reverse engineering, complex systems, and the application of AI to security and testing workflows.



