PROVE PENAGENT™

Agentic pentesting
for software products

PenAgent's 42 agents attack your product the way hackers do. Regularly. The agent swarm makes continuous security possible.

523 findings across 22 projects · Built on 400+ penetration testing rounds

Prove's pentesters at work

PROVE EXPERTISE · PenAgent™
PenAgent™ is built by Prove's pentesters, on the experience of more than 400 penetration testing rounds.

RESULTS FROM CLIENT PROJECTS

What has the agent swarm found?

We compared PenAgent™'s findings with expert-led penetration testing in five client projects. Over the same period, the agent swarm has been in use across 22 projects.

Comparison · 5 client projects
6

vulnerabilities that expert-led testing missed

1

of them was critical

91%

of the same findings as the experts

PenAgent in use · since 18 Aug 2026As of 3 Oct 2026
523

findings reported in seven weeks

22

projects tested by the agent swarm

111

critical or high-severity findings (21%)

Comparison: five client projects in late summer 2026. Usage figures: findings reported by PenAgent between 18 Aug and 3 Oct 2026; merged and withdrawn findings are excluded. The results do not guarantee future findings.

Sample finding · Access control
Severity breakdown · 523
Weekly trend

What does a typical finding look like, how do severities break down, and why did experts miss some gaps? It's all in the PDF.

Download sample findings (PDF) →
WHAT ARE YOU LOOKING FOR?

Why agentic pentesting?

Looking for a pentest provider?

Built by pentesters, run by agents

The experience of 400+ penetration testing rounds, encoded into 42 agents. A Prove pentester verifies every finding before it is reported.

Tell me more →
Can AI do pentesting?

Yes, when it actually attacks

A language model can describe an attack but can't execute it. PenAgent™'s specialist agents attack your running product and verify every finding.

See the comparison →
What about between pentests?

Test every change, not just the release candidate

PenAgent™ tests each week's code in the same week, so there are no blind weeks between releases.

Why this matters →
WHY SPEED MATTERS

Time from vulnerability disclosure to exploitation has collapsed
– from years to hours.


The average time from discovering a vulnerability to exploiting it in real attacks has collapsed in the AI era.

A version released untested can be compromised over a single weekend.

≈ 2.3 yrs
828
days
2018
684
2019
468
2020
324
2021
291
2022
147
2023
56
2024
23.2
2025
≈ 10 h
0.42
days
2026
When you release untested, you're betting attackers won't find the gaps before your next scheduled testing round. With AI, they can be found in hours.
Source: DNV Cyber, published by Yle; the 2026 figure is a forecast.
See what the agent swarm has found →
A CHANGE OF MINDSET IS NEEDED

Testing only the release is a thing of the past.
In security, reactive is too late.

Booking a penetration test for a release candidate is reactive by nature. Problems surface only once the version is already frozen. That was enough when exploiting a disclosed vulnerability took two years: a single snapshot gave a long safety margin.

In a ten-hour window there is no margin, and the gate still protects only one version. Proactive is the opposite: each week's code is tested in the same week, so the release candidate is already clean. And testing continues after the release too.

TODAY'S DEFAULT

Test the release candidate

One intensive pentest round, then months of releases in the dark until the next round is booked.

tested here risk grows every week →
1/52 weeks tested: 51 weeks of untested changes and growing risk
WITH THE SWARM

The candidate is already clean

Each week's code is tested in the same week, so the round confirms rather than reveals.

tested every week risk can't accumulate
52/52 weeks tested: not a single blind week
Illustrative. Risk accumulates in every untested week, before a round as much as after it. A round resets it for that week, after which it starts rising again straight away.
WHY THIS MATTERS

The release candidate is the worst possible place to find an authentication bypass.

At that point the options are delaying the release or shipping with a known risk. Weekly testing moves the finding earlier, when fixing it is cheap — and turns the agent swarm from a booked inspection into an always-on safeguard.

HOW IT WORKS

The agent swarm, every week.

HOW IT WORKS TODAY

Run by our pentesters, every week

01 / YOU PREPARE

You keep the target ready

Make sure the test environment and test data are in place for each weekly run.

→
02 / WE LAUNCH

Our pentesters launch the run

A Prove pentester launches the agent swarm. No setup work on your side.

→
03 / THE SWARM TESTS

The agent swarm does the testing

PenAgent™ works through the target in parallel and verifies its findings.

→
04 / WE DELIVER

Our pentesters review & deliver

A Prove pentester reviews the findings, removes the noise and delivers the report to your team.

Read more about how PenAgent works →
THE AGENT SWARM

Specialised agents – in use now, with more on the way.

Each agent is an expert in one attack area. 42 agents are now in production. They cover authentication, identity and access control most deeply, because that is where the most serious vulnerabilities hide. The agents also cover the full injection family, file uploads, AI features and the modern web surface.

IN USE NOW

Available today

42 production agents

Can an attacker log in as one of your users?

Every realistic route into an account that isn't theirs.

Can registration and password reset be abused?

The doors on both sides of the login screen.

Can login be brute-forced without anything stopping it?

Do automated attacks hit a wall, or keep going all night?

Can a user reach someone else's data or admin functions?

The rules on who can see and do what – OWASP's number one.

Can another site act through your logged-in users?

A user's own browser, borrowed by someone else's page.

Can malicious code run in your users' browsers?

Can a visitor's own session be turned against them?

Can your server be used to reach systems it shouldn't?

Your own backend as an attacker's proxy – or its file system.

Can your database be reached through a form field?

The oldest break-in method that still works. We prove it without reading or changing your business data.

Can your server be made to run something it shouldn't?

Input the server treats as commands rather than text.

Can file uploads be abused?

Accepting a file is a bug. Serving it back as code is a security incident.

Can your modern APIs be abused?

GraphQL and WebSockets follow different rules than your REST API.

Can your business rules be gamed?

Not a broken control, but a workflow used in an order nobody designed.

Can your AI features be turned against you?

If your assistant can call something, a chat message can try to reach it.

Is your infrastructure exposed or outdated?

The plumbing your application runs on.

Can the results be trusted and acted on?

A short, actionable list – not a wall of maybes.

42 specialists covering the highest-rated categories of the OWASP Top 10:2025, and most deeply the authentication, identity and access-control layers. Overlapping findings are merged automatically, so you get one entry per issue.
See sample findings (PDF) →
FAIR QUESTIONS WE GET ASKED

"Why not just use a language model?"
"We already review our code with AI."

Both are reasonable – but neither is penetration testing. A single language model has no specialists, no memory and no evidence.

AI-assisted code review reads source code, but never attacks what you actually shipped. Only one of these three tells you what an attacker can reach in your running system.

○

General-purpose language model

"Why not just ask Claude or GPT to hack it?"

  • ×One generalist does everything — shallow in every attack area.
  • ×Loses earlier findings and context when the context window fills up.
  • ×Works one step at a time — painfully slow on a large attack surface.
  • ×Can describe an attack in detail, but has no tools to execute it.
  • ×Reports plausible-sounding bugs it has never verified.
  • ×No structure and no audit trail of what was actually tested.
◐

AI-assisted code review

"We already review every PR with AI."

  • ◐Reads source code — never attacks the system you actually shipped.
  • ◐Can't see production configuration: TLS, exposed admin pages, outdated servers.
  • ◐Blind to workflows that leave the codebase — such as SSO through an external provider.
  • ◐Flags possible bugs without proof that they can be reached or exploited.
  • ◐Reviewed code isn't the same as production — environment and configuration differences go unnoticed.
  • ◐Can't confirm that a fix actually closed the gap after deployment.
Both still have their place. AI-assisted code review catches bugs in the editor, where they are cheapest to fix — it just can't tell you what an attacker can reach in what you deployed. Model-agnostic — currently runs on Claude and Gemini. The difference isn't the model, but the specialist agents, shared memory, real attacker tooling and verification through exploitation built around it.
Download sample findings (PDF) →
WHY AN AGENT SWARM

Beyond the shallowness of scanners and the limits of manual penetration testing.

PenAgent™ reasons like a human tester, but starts in minutes and works at machine scale.

◇ PenAgent™ Traditional scanner (DAST) Manual penetration testing
Availability Every week — no waiting list On demand Booked months in advance
Verifies & exploits findings Yes — verified, not guessed No — signature-based guesses Yes
Authentication & identity logic Yes — dedicated agents Largely blind Yes
False-positive noise Low — duplicates removed automatically High Low
Time to first result Minutes to hours Hours Days to weeks
Cost of repeat runs Fixed monthly price Moderate Hourly / daily rates
Most agents report within minutes. Some session checks intentionally run for up to 24 hours.
HOW TO GET STARTED

PenAgent™ is now available. The agent swarm can start this week.

PenAgent™ is now available from Prove. The agent swarm starts on your target right away, even when our testers' calendars are full.

Want to hear more?

Built by pentesters

Proven methods, encoded into an agent swarm.

PenAgent™ is built by Prove's experienced penetration testing team — experts who have delivered more than 400 penetration testing rounds. Every agent encodes the craft, attack chains and judgement of those real engagements, so the swarm tests the way our experts do — not like a checklist.

400+

Penetration testing rounds delivered for clients

42

Production agents in use now — and growing

100%

A Prove pentester verifies every reported finding.

NO WAITING

Our experts are booked months in advance.

Your testing doesn't have to wait.

Good pentesters are booked two to three months ahead. That is the honest reality of expert demand – and also two to three months of releases without oversight.

Read more about how PenAgent works →
SAMPLE FINDINGS

See what PenAgent has found.

Download the PDF: a sample finding from a real client project, the results of the comparison with expert-led testing, and the severity breakdown and weekly trend of 523 findings.

If you like, you can also book a 20-minute intro call with our expert on the thank-you page.

By submitting the form, you agree that we process your data in accordance with our privacy policy.

Download sample findings

We'll send the PDF to your email.

FAQ

Practical questions, straight answers.

SECURITY AND DATA PROTECTION
THE SERVICE
EVIDENCE AND REPORTS
CONTRACT

Prove Expertise Oy

Kirkkokatu 8A3, 90100 Oulu, Finland

Prove PenAgent™


PenAgent™ is a product built by Prove's penetration testing team: the 42 agents marked as in use already work in real client engagements, and new specialist agents are actively being developed.
Some tests add harmless, uniquely marked test values to the application. These are agreed in the kick-off meeting, logged and removed after the run. Security testing may only be performed on systems you own or have explicit permission to test.