The End of the Pentest Report

By Editorial Desk | September 20, 2026 The pentest report has not changed format since the late 1990s. Open any of them, from any vendor, on any engagement, and the spine is the same. Executive summary on page one. Findings…

A steel wire inbox tray on a brushed steel desk holding a stack of printed white paper pentest reports, the top report folded back to reveal monospaced text rows, a single cold cyan LED status pin glowing at the back edge of the desk, a steel pen set down on the desk, cracked dark concrete visible in the foreground, deep blacks and charcoal with cold cyan accent only, single cold beam from above-left, deep chiaroscuro, no people no text no logos.

The pentest report has not changed format since the late 1990s. Open any of them, from any vendor, on any engagement, and the spine is the same. Executive summary on page one. Findings table sorted by CVSS. A method statement that mostly describes what the testers did not get to do. A remediation appendix the client will skim once, file, and never open again. Two thirds of the report will be accurate on the day it lands. Half of it will be wrong six months later because the environment moved.

That is the part the industry has stopped pretending about. Continuous offensive testing has been growing for a decade and it has finally eaten the annual pentest at the high end. The shape of the engagement changed. The deliverable is no longer a PDF that lands in a SharePoint folder. It is a feed that never stops. The question for the next twelve months is not whether to keep doing pentests. The question is which parts of the report still earn their keep, and which parts you can replace with a continuous validation loop that actually catches the things the report never did.

A worn dark steel patch panel mounted on cracked dark concrete, three small dim cyan LED indicators and one small electric red peak LED, a printed white spec sheet folded and discarded on the concrete below the panel, a steel-jacketed cable coiled next to it, deep blacks and charcoal with cold cyan and red accents only, single cold beam from above, deep chiaroscuro, no people no text no logos.
Worn steel patch panel: three cyan status pins and a single red peak LED, the rough hardware side of continuous offensive testing.

What continuous offensive testing actually means

The phrase covers three distinct things, and they do not substitute for each other. Buyers who skip the distinction end up buying one product and labelling it as another, which is the single most common waste pattern in this category right now.

The first is breach and attack simulation, the BAS market. Tools that run scripted attack scenarios against your production estate on a schedule and report which ones got past your controls. Most of them started as agent software on endpoints, then pivoted to agentless simulations that fire from the cloud against your perimeter, your identity provider, your email gateway. The original vendor is still the dominant player, but the category has split into three roughly equal segments. Pure BAS vendors, EDR vendors who added a BAS mode, and cloud security posture management vendors who added a BAS mode. Each tells a slightly different story about what counts as a breach and what does not.

The second is automated exposure validation, the AEV market. A newer label, pushed by the attack surface management vendors who wanted to call their thing something other than vulnerability management. AEV is essentially continuous validation that a known exploitable weakness is still reachable in your environment. The pipeline is the same as BAS at a glance, but the input set is different. BAS scripts against your controls. AEV scripts against your assets and their specific known weaknesses. The output is a real exploit path, not a simulated adversary.

The third is adversarial simulation, the purple team retainer and the red team retainer, both delivered by humans with optional tooling on top. This is the most expensive, the most irregular, and the only one of the three that finds things the script library has never seen. The shape of an adversarial engagement in 2026 is mostly remote, partly physical, occasionally full scope, and almost always anchored to a specific threat actor profile or a specific business risk the customer cares about. The output is a long debrief, not a feed. The cadence is quarterly or monthly or as-needed, not continuous.

Technical facts at a glance
Category What it actually does What it does not do Typical annual cost (enterprise)
Breach and attack simulation (BAS) Runs pre-built attack scripts against production controls and reports which detections fire Find unknown weaknesses, validate exploit chains across business logic $40k to $120k
Automated exposure validation (AEV) Continuously validates known exploitable weaknesses are still reachable Test custom business logic, simulate novel adversary tradecraft $60k to $200k
Adversarial simulation (purple or red team retainer) Human-led engagements against a specific threat profile or business risk Run as a feed, provide continuous coverage, replace the annual pentest $150k to $500k+ for quarterly cadence

Where continuous wins

The win on the BAS and AEV side is coverage. A pentest lasts two to six weeks. A continuous validation tool runs all year. That gap is the actual product. The pentest cannot tell you whether your new phishing control is still working on a Tuesday in November, because the pentester was not there in November. The BAS tool runs the same control check on a Tuesday in November, and on every Tuesday after, and quietly tells you the answer changed.

The second win is on the alert side. Most modern SOCs drown in alerts. BAS tools double as alert validation. They confirm that the detections your SOC built actually fire when the script runs the technique. That is a real product even if you ignore the breach simulation framing entirely. A SOC that runs BAS against its own detections every week catches alert drift in days instead of during the next pentest debrief.

The third win is on the controls change log. Most environments have a steady churn of new attack surface. Cloud accounts spin up, new SaaS integrations land, identity policies change. The pentest report you got in March is a museum exhibit by July. The continuous feed catches the new gaps when they appear, which is the moment they are cheapest to fix.

Where it does not work

None of the three continuous categories is a substitute for adversarial simulation. The script library only knows what the script library was built to test. Real adversaries do not read the script library. The genuine novel kill chain, the one that combines a misconfigured SSO claim with a stale service principal and a forgotten S3 bucket policy, will not appear in any BAS product until after it has been seen in the wild and added to a quarterly update. That gap is the whole reason purple team retainers exist.

The continuous feed also fails on the social engineering side. Phishing simulation is its own product line and lives separately. None of the three categories above tests whether your finance team will wire money to a vendor that was just compromised. That is a different problem with a different vendor set and a different cadence.

The third failure mode is the report of the report. Some BAS products will happily generate a compliance report that says your detection control is green when the underlying test never actually executed against that control. The default reporting tier on several major products rounds coverage gaps to zero for the executive view. This is the expensive theater part. A buyer who evaluates BAS on the executive dashboard alone is buying a feeling, not a measurement. The way to spot it is to ask the vendor for the raw test execution log from the last 90 days, including the ones that errored out. If they cannot produce it, the dashboard is a story.

What it actually costs

The pricing picture in 2026 has settled into roughly three tiers. Tier one is the standalone BAS tool, which lands between forty thousand and one hundred and twenty thousand dollars a year for a mid sized enterprise, with the price largely driven by the number of simulated users and the number of attack techniques the customer wants the library to cover. Tier two is AEV, which sits between sixty thousand and two hundred thousand, priced on the number of assets validated and the depth of the exploit chain modelling. Tier three is the human led adversarial retainers, which start around one hundred and fifty thousand a year for quarterly cadence and run north of half a million for monthly or as needed.

None of these tiers is cheap. The math that makes them worth it is the displacement. A mature enterprise that buys one tier two product and one tier three retainer has spent roughly a quarter of a million dollars. The same enterprise was probably spending two hundred thousand a year on a single annual pentest from a big four consultancy, plus another one hundred and fifty thousand on a separate red team. The new stack replaces both and adds continuous coverage for a marginal additional spend. That math works for the enterprise. It does not work for everyone.

What mid-market should actually do

For companies in the two hundred to two thousand employee range, the answer is mostly boring and mostly about discipline. The annual pentest is still the right anchor. The findings table it produces still feeds the remediation backlog. The compliance check it satisfies is still real.

The pragmatic upgrades are two. First, add a basic BAS product on a one or two year trial, run it against the SOC detections, and let it quietly validate the alert pipeline. That is the cheapest place to get a real return. Second, run a quarterly tabletop with the same vendor that runs the annual pentest, instead of a separate engagement. The tabletop forces the conversation that the pentest report never did, which is what would actually happen if the report’s worst finding were exploited at two in the morning. The cost is a fraction of a separate retainer and the output is closer to what an adversarial engagement would have produced for the same business risk.

The thing mid-market should not do is buy all three tiers at once. That is how the budget gets eaten. The honest answer is that a continuous feed without the human adversarial layer is incomplete, but a human adversarial layer without the continuous feed is the status quo with extra steps. Pick the cheapest version of the one that closes your actual gap.

The defensive checklist

What to do this quarter, in roughly the order it pays off:

  1. Audit the last three pentest reports. Find the findings the SOC confirmed as remediated. Open the production evidence that backed the closure. Verify it is still true. Most enterprises will find at least one closure that has quietly regressed.
  2. Pick one BAS product on a real trial, not a free tier. Run it against the alert pipeline you actually care about. Read the raw execution log, not the executive view. If the tool cannot show you the failed and errored tests, return it.
  3. Schedule the next pentest for the quarter that lines up with the largest control change in the year, not the calendar quarter. The pentest catches more when it lands during change.
  4. Add one purple team engagement, scope limited, against the specific threat actor your threat intel team is most worried about. Use it to stress test the IR plan, not to add findings to the backlog.
  5. Run a tabletop with the leadership team and the on call engineer in the same room. Walk through the worst finding from the last pentest as if it had been exploited. Time the response. The number you get is the number that matters.

The bottom line

The annual pentest report is no longer the spine of a mature offensive program. Continuous validation handles the control coverage the report never did. Human led adversarial work handles the novel kill chains the continuous feed never will. Mid market keeps the annual pentest as the anchor and adds one or the other on top, not both, not all three, not the full stack the vendor wants to sell. Pick the gap. Close it. Revisit in twelve months.

Geist verdict

The pentest report was never the point. The point was always the conversation it forced. The continuous feed replaces the routine coverage. The adversarial engagement replaces the rare and expensive conversation. The annual pentest still anchors the conversation the board needs to have, which is whether the program is real or theatre. Pick the right one for the gap you actually have.


MITRE ATT&CK mapping

The techniques below are the ones most commonly exercised by the three categories covered in this piece. The mapping is illustrative, not exhaustive.

Coverage by category
Category Primary ATT&CK tactics Representative techniques
Breach and attack simulation Initial Access, Execution, Persistence, Defense Evasion, Credential Access T1566 Phishing, T1059 Command and Scripting Interpreter, T1003 OS Credential Dumping, T1078 Valid Accounts
Automated exposure validation Reconnaissance, Initial Access, Privilege Escalation T1595 Active Scanning, T1190 Exploit Public Facing Application, T1068 Exploitation for Privilege Escalation
Adversarial simulation All tactics, especially novel combinations T1654 Application Access Token Logging, T1611 Container Escape, T1556 Modify Authentication Process

Sources

  • Gartner, Magic Quadrant for Adversarial Exposure Validation, 2025 and 2026
  • MITRE ATT&CK Enterprise Matrix v17, current techniques
  • Verizon Data Breach Investigations Report, 2026 edition, offensive testing program section
  • SANS 2025 SOC Survey, alert validation practices and BAS adoption rates
  • Public vendor pricing from major BAS, AEV, and adversarial simulation providers, enterprise tier, 2026
  • Praetorian, Bishop Fox, and CrowdStrike published engagement reports, 2025 to 2026
  • MITRE Engenuity Center for Threat Informed Defense, purple team program research, 2025



Sources & Further Reading

All claims in this article are sourced from primary documentation, vendor advisories, and reputable security researchers.

Spotted an error? Email the editor. Corrections are issued with a visible correction note.

Editorial standards. Every article on humanrequired.org is reviewed by a human editor before publication. AI may assist with drafting or research; final editorial control is human. Read the full standards.

Continue reading