Picture a company that gets a clean report from an automated scanner on a Friday. The next Tuesday a security researcher emails them. The researcher changed one number in an API request and downloaded invoices that belonged to another company. No scanner would have flagged it, because the server returned a perfectly normal response. Finding this kind of flaw is what a penetration test is for. A person, with your written permission, tries to break into the application the way a real attacker would, and then writes down exactly how they did it and how to close the gap.
Pentest or Vulnerability Scan?
A scanner such as Qualys, Nessus, or OWASP ZAP runs thousands of known checks against your server and application. It is good at missing patches, outdated libraries, weak TLS settings, and missing security headers. It has no idea that a branch admin should only see their own branch, that a voucher must not be redeemed twice, or that step three of an approval flow can be skipped by calling the API directly. A pentester sits in Burp Suite, reads every request, and chains small weaknesses into a real attack. You need both. Scans are cheap and should run every week or month. A pentest is a deeper check you buy at specific moments.
| Aspect | Vulnerability scan | Penetration test |
|---|---|---|
| Done by | Automated tool, configured by a person | Experienced tester using tools and manual work |
| Finds | Known CVEs, misconfiguration, missing headers | Access control flaws, business logic abuse, chained attacks |
| False positives | Common, needs triage | Rare, each finding is reproduced |
| Frequency | Weekly or monthly | Yearly and after major changes |
| Duration | Hours | Usually 5 to 15 working days |
| Output | Long list sorted by tool severity | Report with evidence, CVSS score, and fix advice |
Black Box, Grey Box, or White Box
In a black box test the tester gets a URL and nothing else, like an outsider. It sounds realistic, but a good part of the paid days goes into reconnaissance that your team could have answered in an email. In a grey box test the tester gets accounts for each user role and the API documentation. For most business applications this gives the most findings per day, and it is what we usually recommend. A white box test adds source code and architecture diagrams. It suits high-risk parts such as payment, lending, or authentication, where you want the tester to read the code that decides who gets money.
What Goes in the Scope
- Web application: the URLs, every user role, and which features matter most to the business.
- API: an OpenAPI file or Postman collection, all versions still reachable, and partner endpoints that the main app never calls.
- Mobile app: Android and iOS builds, tested against OWASP MASVS for local storage, certificate pinning, root or jailbreak detection, and secrets left in the binary.
- Internal network and cloud: VPN access for the tester, server ranges, and a read-only review of cloud settings such as AWS IAM and storage buckets.
- Usually out of scope: denial of service, social engineering, and third-party services such as payment gateways, unless they are agreed in writing.
Vendors price by scope, so a vague request like test our website gets a vague quote and a thin report. Count the roles, the endpoints, and the apps before you ask for prices.
When You Need One
- Before the first launch of an application that handles money or personal data.
- After major changes: a new payment flow, a rewrite of login, or an API opened to partners.
- At least once a year for production systems, since code, dependencies, and attackers all change.
- When an enterprise client asks for it during vendor onboarding. Banks and large companies now send security questionnaires that ask for a recent pentest report.
- For banks, payment providers, and fintech lenders supervised by OJK or Bank Indonesia, regular independent security testing is an expected part of IT risk management, and examiners ask to see the results.
- ISO 27001 auditors look for evidence that technical vulnerabilities are found and fixed. UU PDP (Law No. 27 of 2022) requires appropriate security for personal data. A pentest report with a retest is solid evidence for both.
How an Engagement Runs
- 1Scoping: a call to agree on targets, roles, environment, and test type, followed by a quote with the number of tester days.
- 2Rules of engagement: a signed document with the testing window, tester IP addresses, emergency contacts on both sides, what happens if a critical flaw or signs of a real breach are found, and how test data is handled.
- 3Preparation: your team sets up the environment and accounts, described in the next section.
- 4Testing window: typically 5 working days for a small web app with a few roles, and 10 to 15 days for web, API, and both mobile apps. Critical findings should be reported the same day instead of waiting for the final report.
- 5Report: an executive summary for management, then each finding with a CVSS score, affected URL or component, reproduction steps, screenshots or requests as evidence, and a recommended fix.
- 6Retest: after your team fixes the findings, usually within 30 to 90 days, the vendor checks them again in one to three days and issues an updated report. This retest letter is what auditors and clients actually want to see.
Preparing Your Environment
Test on staging that matches production: the same build, the same server configuration, and ideally the same WAF rules. Fill it with realistic but anonymised data, because an empty database hides bugs in search, export, and reporting. Create two accounts for every role so the tester can try to reach one user's data from another account. That is how IDOR flaws, the most common serious finding we see, get caught. Give testers a way past OTP and CAPTCHA, and whitelist their IPs in Cloudflare or your WAF, otherwise you pay for a week of testing the firewall. Take a backup before the window starts, and tell your hosting provider and on-call staff so an alert at 2 a.m. does not turn into an incident call.
Reading the Report and Prioritising Fixes
Most reports score findings with CVSS. Version 3.1 is still common, and version 4.0 has been out since late 2023. The bands are Critical 9.0 to 10.0, High 7.0 to 8.9, Medium 4.0 to 6.9, and Low 0.1 to 3.9, plus informational notes. The score is generic, so read it with your business in mind. A Medium finding that lets any logged-in user download customers' KTP scans is more urgent than a High on an internal page that only two admins can reach.
- Critical and High: fix within days, and block the release if the app is not live yet.
- Medium: plan into the next sprint, within about 30 days.
- Low and informational: batch them into the backlog and handle them with regular maintenance.
- Look for the root cause. If one endpoint is missing an authorization check, others probably are too. Fix it in shared middleware or a policy layer, then ask the tester to check the pattern, not only the single URL.
A pentest report with twenty open findings and no retest is a documented list of ways into your system.
Choosing a Vendor
- Ask for a redacted sample report. Good findings have clear reproduction steps and evidence. Pasted scanner output with a company logo is a warning sign.
- Ask about methodology: OWASP WSTG for web, OWASP MASVS and MASTG for mobile, the OWASP API Security Top 10 for APIs, and PTES or NIST SP 800-115 for network work.
- Check the certifications of the people who will actually test, such as OSCP, OSWE, Burp Suite Certified Practitioner, or CREST. Ask whether any of the work is subcontracted.
- Confirm that a retest is included in the price and how long the retest offer stays open.
- Agree on an NDA and on where evidence is stored, who can see it, and when it is deleted.
- Be wary of quotes that are too cheap for the scope. Two days for a large application with five roles is a scan, whatever the proposal calls it.
Budget your own team's time as well. Keep one developer on standby during the testing window to answer questions and reset accounts, and reserve a sprint after the report for fixes. A pentest without time to fix the findings is money spent on a document.
Scanners should run all year, and the pentest is for the questions only a person can answer. In the managed services we run for clients, scheduled scanning, an annual pentest, and the retest are part of the same routine, so findings go straight into the backlog of the team that maintains the code.
Key takeaways
- Vulnerability scans find known issues cheaply. A pentest finds access control and business logic flaws that only a person spots.
- Grey box testing with accounts for every role gives the most value per tester day for most business applications.
- Test before launch, after major changes, at least yearly, and whenever clients, regulators, or auditors ask for evidence.
- Prepare a realistic staging environment with two accounts per role, whitelisted tester IPs, and a developer on standby.
- Prioritise by business impact as well as CVSS, fix the root cause, and always finish with a retest.


