Agentic AI-Assisted Penetration Testing

Cyber

September 24, 2026

Agentic AI-Assisted Penetration Testing: AI Scale. Human Precision

When Mythos launched, headlines warning of its potential dangers were prolific. For security risk leaders, it felt like a watershed moment, one that signaled that AI had arrived, but simultaneously raised widespread concern about risk. Many wanted to understand how real the risk was for their organization and what should be done about it.

For one large financial institution, it became an opportunity to assess whether AI-assisted penetration testing could uncover new insights in systems that had already undergone extensive expert-led testing. In partnership with Kroll, the organization undertook a data-driven risk evaluation, using AI-assisted penetration testing to better understand its exposure and establish a baseline for frontier AI risk.

Overview

 

Industry
  • Financial services
 
Challenges
  • Advancing cyber capabilities of frontier AI models
  • Understanding additional value of AI-assisted penetration testing

 

 

Kroll Services
  • Penetration Testing
  • Advisory
 
Impact
  • Role of AI-assisted penetration testing assessed
  • High- and critical-risk vulnerabilities identified
  • Expert-led approach to overcoming limitations of AI tests undertaken

The Challenge

A large multinational financial institution asked Kroll to assess whether rapid advances in frontier AI could materially change the cyber risk for its public-facing applications and infrastructure.

In May 2026, amid heightened attention being paid to advancing cyber capabilities of frontier AI models, especially Mythos, the client’s senior leadership became concerned that AI-enabled threat actors could use these tools to identify and exploit weaknesses in internet-facing systems more quickly and effectively.

The client asked Kroll’s Offensive Security team to evaluate several of its most prominent frontline banking applications and supporting infrastructure using an agentic AI harness designed to make extensive use of frontier AI capabilities.

The selected applications were among the client’s highest-priority systems. Kroll had assessed them regularly through the client’s annual penetration testing program, and the applications were also governed by a mature and rigorous application security program. This context was central to the engagement: The objective was to determine what additional insight an AI-assisted approach could provide when applied to systems that had already undergone repeated expert testing.

Kroll’s Solution

Kroll applied an agentic AI-assisted penetration testing harness within a consultant-led assessment. The harness was designed to augment established penetration testing practices rather than replace professional judgment.

The solution was model-agnostic, allowing the testing team to use appropriate frontier models as capabilities evolved. Agents were provided with a powerful set of penetration testing tools through the Model Context Protocol, enabling them access to the same set of tools used by human penetration testers.

Kroll is part of the verification programs offered by frontier AI labs, which allows it access to models with reduced guardrails to perform authorized penetration testing. Kroll’s agentic harness is not effective without access to these models because the standard guardrails consistently shut down our tests.

Integration with web security tool Burp Suite Professional allowed experienced consultants to monitor agent activity, examine AI-generated test cases, extend coverage and validate potential findings. This human oversight was essential to ensuring that Kroll’s professional consultants remained in control of the assessment and providing assurance about conclusions and the coverage.

Customized skill files aligned agent behavior with Kroll’s penetration testing methodology. The harness can also incorporate knowledge from previous assessments, allowing it to revisit known conditions, discover attack paths and test whether previously identified weaknesses appeared elsewhere.

All activity was logged to support review and auditability. This created a record of the agents’ actions and enabled consultants to examine how potential findings had been generated.

The resulting operating model combined the scale and persistence of AI-assisted testing with the contextual understanding, judgment and accountability of experienced penetration testers.

Assessment Results

Agentic AI-Assisted Penetration Testing

Agentic AI-Assisted Penetration Testing

Representative Findings

Multi-Factor Authentication Bypass

The assessment identified a medium-risk vulnerability that could allow certain users to bypass a required multi-factor authentication step.

Exploitation depended on a narrow combination of circumstances. A specific HTTP POST parameter, one of many parameters included in the request, had to be present but empty. The target account also had to meet a particular type and configuration.

In the AI-assisted approach, a broader combination of practical parameters and account conditions than would typically be practical within a time-constrained manual workflow, were tested. The testing also identified an indication that the relevant parameter warranted closer investigation.

The finding illustrates how agentic AI can extend test coverage by persistently exploring more combinations than reasonable for a human consultant, while still requiring a consultant to verify the finding and its risk.

One-Click Cross-Site Scripting

Testing identified a high-risk cross-site scripting vulnerability that could be used to exfiltrate personally identifiable information (PII).

A related cross-site scripting vulnerability had been identified at another endpoint during previous penetration tests. The AI-assisted assessment discovered additional affected endpoints and developed a complex payload that leveraged another known design weakness to access PII data through a separate application endpoint.

The payload demonstrated how multiple weaknesses could be chained into a more consequential attack path. It included a cross-origin resource sharing bypass and used an external interaction service to demonstrate data exfiltration within the authorized test environment.

This proof of impact materially improved the quality of the risk demonstration. A conventional proof of concept showing only script execution would not have communicated the combined ramifications of the affected endpoints and application design as clearly.

Unauthenticated Endpoint Leading to Host Compromise

The infrastructure assessment identified an unauthenticated REST interface that exposed endpoints containing highly sensitive information, including credentials, logs and configuration data. The condition ultimately enabled complete compromise of the affected host and was rated a critical risk.

The application’s HTTP root returned an error page. A human analyst had performed brief endpoint enumeration unsuccessfully before moving on to other areas of the assessment.

The AI agent recognized clues in the error message indicating the presence of a less common open-source application. Drawing on knowledge of that application’s default and potentially sensitive interfaces, it identified the unauthenticated endpoint that led to the information and credential disclosures.

This finding demonstrated the value of an LLM’s inherent knowledge base to identify an unfamiliar technology. The agent was able to identify relevant test cases for a less common tool and direct attention to an attack path that initial manual enumeration had not revealed.

Impact

AI-Assisted Penetration Testing Can Identify High- and Critical-Risk Vulnerabilities

Using AI to expand testing coverage and produce stronger demonstrations of potential impact was an effective way to examine the value of frontier AI, uncovering meaningful risks in highly scrutinized applications and infrastructure.

AI offered three clear advantages:

  • AI-assisted testing can increase coverage. Agents can explore large numbers of parameter combinations, configurations and related endpoints without losing persistence or consistency. This can help surface narrow conditions that may be difficult to prioritize during a conventional time-boxed assessment. 
  • AI can connect technical clues with broad knowledge of technologies, default interfaces and known design weaknesses. This is particularly valuable when assessing less common technologies. 
  • AI can help demonstrate combined risk. By chaining conditions and developing more complete proofs of concept, the harness helped show how individual weaknesses could produce a more serious business impact.

An Expert-Led Approach, Alongside the Benefits of AI-Assisted Penetration Testing, Can Offer Assurance

Limitations of AI-assisted penetration testing—particularly the volume of false-positive and low-value output generated before dedicated validation was introduced—were revealed. Raw AI output should not be treated as a finished penetration test. Without effective validation, the AI generated substantial noise. Human consultants were essential for verifying exploitability, removing low-value observations, managing scope and translating technical evidence into meaningful risk conclusions. This expert governance is required for safe and reliable offensive security work.

Organizations With Mature Security Programs Can Harness Advantages of AI-Assisted Penetration Testing

AI-assisted testing should not replace established testing practices, but it can strengthen them. In this scenario, with applications that had already benefited from years of recurring penetration testing and rigorous application security control, agentic approaches could identify important issues as part of exercises to pressure-test existing controls, investigate broader attack paths and prioritize the exposures that matter most. Those with similar established security fundamentals may be better positioned to use AI-assisted penetration testing to absorb advances in attacker capability than organizations with less mature programs.

Read more about our Offensive Security work and find out how AI can provide meaningful value in penetration testing when it is applied within a controlled, expert-led process.

Stay Ahead with Kroll

Cyber and Data Resilience

Kroll merges elite security and data risk expertise with frontline intelligence from thousands of incident responses and regulatory compliance, financial crime and due diligence engagements to make our clients more cyber- resilient.

Threat Exposure Management

Kroll’s field-proven cyber security assessment and testing solutions help identify, evaluate and prioritize risks to people, data, operations and technologies worldwide.