Solusec: Solutions for Cyber Security

Operated by Solusec Ltd
CREST accredited · IASME Certification Body

Judging what you bought

Reading a cheap report: did you buy a test or a scan?

The report is the only part of a penetration test most buyers ever see. It is also the part that tells you, if you know where to look, exactly how it was produced.

Cheap CREST penetration testing › Reading a cheap report

You cannot watch a penetration test happen. You hand over a scope, some access and a purchase order, and a document comes back. If the document is forty pages with a professional cover and a severity chart, it looks like the thing you bought.

Frequently it is. Sometimes it is the output of an automated vulnerability scanner exported into a template, and the buyer has no way of telling. Since that product commonly sells at £250 to £500 a day against £800 to £1,200 for accredited manual testing, the incentive to blur the line is obvious.

This is not a judgement on scanning, which is a legitimate and useful product sold honestly by plenty of firms. It is about knowing which one is in your hands, because if a contract clause or an insurer asked for a penetration test, a scan will not satisfy it and you will buy the test a second time.

The report itself gives it away. Here is how to read it.

The single question that decides it

Pick the most serious finding in the report and ask one thing of it: could a developer who has never seen this document reproduce the issue from what is written here, and would they understand what it means for this organisation specifically?

A manual finding survives that question. It describes a request, a parameter, a response and a consequence in your environment. A scanner finding does not, because it was never produced by someone who tried it. It was produced by a signature matching a version string or a response pattern, and there is nothing behind it to write down.

Everything below is a variation of that question.

The tells, in the order they are easiest to check

Plugin identifiers standing in for evidence

If findings carry a numeric identifier from a scanning tool, a plugin name, or a reference into a vendor knowledge base, and that identifier is the evidence rather than an aside, the finding came from a tool and nobody verified it. Manual evidence looks like a request and a response, a screenshot with the relevant part indicated, or a short transcript. A tool reference alongside real evidence is fine and normal. A tool reference instead of evidence is the tell.

Everything rated medium

Look at the distribution of severities. A real engagement produces an uneven spread: a few things that matter a great deal, a lot that matter slightly, and usually several informational items. A report where almost everything sits in one band, and that band is medium, is a report where nobody applied judgement to the ratings. Uniform ratings are a default, not a conclusion.

The same applies to scores lifted whole from a public vulnerability database. A severity rating is supposed to reflect the issue in your environment, which means the tester should have adjusted it for how exposed the affected system is and what it holds. Unmodified base scores throughout mean no adjustment happened.

No reproduction steps

Check whether you could follow the finding. A real one tells you where to go, what to send, what comes back and how to recognise it. If findings stop at a description of the class of issue, whoever wrote it did not perform it, and nor will your developer, who will come back and ask what exactly they are supposed to fix.

Remediation text that would fit any organisation

Read the remediation advice for three findings in a row. If it could be pasted into any report for any client without changing a word, it was. Genuine remediation guidance names your technology, your version, the specific file or setting, and frequently acknowledges a constraint that only applies to you. Generic guidance is not useless, but a whole report of it means nobody looked at your system while writing.

No statement of what was not covered

A report produced by a person who did the work knows where the work stopped. It says which systems were in scope and which were excluded, whether testing was authenticated and against which roles, what could not be reached, what was agreed as out of bounds, and what conclusions the document does not support. A scan has no such awareness, because a scanner does not know what it did not scan. Its absence is one of the strongest signals in this list, and it is the one buyers notice least.

Findings organised by host instead of by risk

Automated output is naturally ordered by target, because that is how the tool walked the estate. Reports written by people are ordered by what matters, with related issues grouped into a single finding that explains how they combine. Ten separate entries for the same expired certificate on ten addresses is a tool's view. One finding covering all ten is a person's.

Nothing that required being logged in

If the target had user accounts and you supplied credentials, look for findings that could only have come from inside: something about what one user could see of another, about a function reachable by the wrong role, about a workflow that could be completed out of order. If the entire findings list could have been produced by someone who never signed in, the credentials you supplied were probably never used.

False positives left standing

Findings for software you do not run, for an operating system that is not on the host, or for a vulnerability that plainly does not apply to your configuration. These appear in any scan. Their presence in a delivered report means nobody validated the output, which is the entire difference between the two products.

No sign of a second reader

CREST accreditation obliges a provider to have a second person review a report before delivery. That review shows up as internally consistent severities, an executive summary that matches the technical findings, and an absence of the small errors that survive one pair of eyes. Contradictions between the summary and the detail, or a severity in the table that differs from the same finding in the body, mean nobody checked.

What good looks like on the page

For contrast, a finding written by someone who did the work usually contains: what the issue is in one sentence; where it is, precisely; how it was confirmed, with the actual evidence; what an attacker could do with it against your organisation; a severity with a line of reasoning for why it is that and not the band above or below; how to fix it, referencing your stack; and how to verify the fix.

That is perhaps half a page per finding. It is also why writing and reviewing are a real part of what a day buys, and why a report of thirty short entries is often worth less than one of eight proper ones.

What to do if the report fails these tests

Go back to the provider before you go anywhere else, and be specific rather than accusatory. Three questions, in writing, get you a long way.

  1. How many hours of manual testing were performed, by whom, and on what dates? An honest answer here resolves most cases in one direction or the other.
  2. Can you provide the evidence supporting findings three, seven and eleven? Pick findings rated high or medium. If evidence exists it will arrive quickly. If it does not, the reply will be about methodology rather than about those findings.
  3. Who reviewed this report before delivery? A provider with a review step will name a role and a date without hesitation.

If the answers do not hold up, your position depends on what you actually contracted for. This is the moment a purchase order saying only "security testing" costs you, and it is the argument for specifying the deliverable in writing before you buy: evidence per finding, reproduction steps, a statement of exclusions, and a named reviewer. Those four lines in an order are what turn a disappointing report into a remedy rather than a lesson.

And if it turns out you bought a scan honestly sold as a scan, the fault may be in the purchase rather than the product. A scan is the right buy for some purposes. It is simply not the thing a requirement naming a penetration test is asking for.

Is a short report necessarily a bad report?

No, and long reports are often the weaker ones. A test of a small, well maintained scope should produce few findings, and padding it out with informational entries and appendices is the usual way a thin engagement is made to look substantial. Judge the quality of individual findings rather than the page count.

Our report has no critical or high findings at all. Should we be suspicious?

Not by itself. A well maintained system genuinely can come back with nothing serious, particularly if you cleared out the obvious problems beforehand. What matters is whether the report shows evidence of having looked: reproduction steps on what it did find, a clear statement of scope and exclusions, and findings that could only have come from someone using the system.

Can we ask to see the raw testing evidence?

Yes, and an accredited provider retains it as part of their obligations. Ask for the supporting evidence behind two or three specific findings rather than everything, which is a reasonable request that is quick to satisfy. A provider who cannot produce it for a finding they rated high has a problem.

The report looks automated but the provider is accredited. What then?

Raise it with them directly and in writing first, naming the specific findings and what is missing from them. Accredited providers are required to operate a complaints procedure, and the accrediting body is a further route if that procedure does not resolve it. Most cases turn out to be a report that was rushed rather than a scan in disguise, and they are usually fixed once asked.

Second opinion on a report you already have

Send us a report you are unsure about. We will tell you what it shows evidence of and what it does not, whether or not you buy anything from us.