Report | WAF Comparison Project, 2026 | Check Point Software
Report | WAF Comparison Project, 2026
Introduction
This article describes the results of our annual WAF Efficacy comparison. For the third year in a row, we tested leading WAF solutions in rigorous, real-world conditions, triggering both malicious and legitimate web requests to measure exactly how well these vendors protect modern applications. While previous years focused on the decline of ModSecurity and updates to the OWASP Core Rule Set, 2025 marked a shift toward exposing the architectural limitations of traditional WAFs. As vulnerabilities become more complex, the limitations of legacy, signature-based engines are becoming impossible to ignore. This year’s test introduces a new focus on Padding Evasion - inspired by the critical React2Shell vulnerability - and highlights how fixed-buffer inspection limits are leaving modern applications exposed.
The Core Challenge: Security vs. Detection
The two most important parameters when selecting a Web Application Firewall remain:
- Security Quality (True Positive Rate) - the WAF's ability to correctly identify and block malicious requests is crucial in today's threat landscape. It must preemptively block Zero-Day attacks as well as effectively tackle known attack techniques utilized by hackers.
- Detection Quality (False Positive Rate) – the WAF's ability to correctly allow legitimate requests is also critical because any interference with valid requests could lead to significant business disruption and an increased workload for administrators as much tuning is required.
A very comprehensive data set was used to test the products:
- 1,040,242 legitimate HTTP requests from 692 real websites in 14 categories
- 74,284 malicious payloads from a broad spectrum of commonly experienced attack vectors
Loyal to the spirit of open-source, we provide in this GitHub repository all the details of the testing methodology, testing datasets, and open-source tools that are required to validate and reproduce this test and welcome the community's feedback.
Products Tested and Results
This year's test was conducted in December 2025, and compared the following popular WAF solutions:
- Microsoft Azure WAF – OWASP CRS 3.2 Ruleset
- AWS WAF – AWS Managed Ruleset
- AWS WAF – AWS Managed Ruleset and F5 Ruleset
- CloudFlare WAF – Managed and OWASP Core Rulesets
- F5 NGINX App Protect WAF – Default Profile
- F5 NGINX App Protect WAF – Strict Profile
- NGINX ModSecurity – OWASP CRS 4.20.0
- CloudGuard WAF – Default Configuration (High Confidence)
- CloudGuard WAF – Critical Confidence Configuration
- F5 BIG-IP Advanced WAF – Rapid Deployment Policy
- Fortinet FortiWeb – Default Configuration
- Google Cloud Armor – Preconfigured ModSecurity Rules
- Barracuda WAF – Default Configuration
The two charts below summarize the main findings. Security Quality and Detection Quality are often a tradeoff within security products.
Key Updates & Findings for 2026
Before diving into the full data, several major developments defined this year's testing cycle:
- New Attack Vector: Padding Evasion & React2Shell - Modern exploits are getting larger. The recent React2Shell vulnerability (CVE 2025 55182, CVSS 10.0) demonstrated that critical payloads often exceed the standard 8KB or 128KB inspection buffers used by many legacy WAFs. To test this, we introduced a Padding Evasion dataset containing hundreds of malicious payloads hidden behind junk data of varying lengths.
- The Industry Challenge: We observed that some vendors, notably Cloudflare, default to ignoring payloads exceeding specific sizes to maintain performance, effectively failing open on large attacks.
- The Solution: The test confirms that specific architectural approaches - streaming analysis combined with Machine Learning - are required to inspect the entire request body without compromising performance.
- Tooling Evolution: PDF Reports & Docker - Based on extensive feedback from the DevSecOps community, we have transformed the testing tool from a script into a platform. Automated PDF Reporting and Docker features have been introduced for ease of use.
Deep Dive: The Padding Evasion Dilemma
The introduction of the Padding Evasion dataset revealed a fundamental architectural divide in the WAF market. Unlike standard injection attacks, padded payloads test the WAF's engine capacity rather than just its signature database.
Padding Evasion Results:
| WAF Solution | Behavior On Large Payloads | Impact | Status |
|---|---|---|---|
| CloudGuard WAF | Full Coverage | Protected | |
| Google Cloud Armor | Full Coverage | Protected | |
| Microsoft Azure WAF | Fail Close | Blocks legitimate large traffic | |
| AWS WAF (Managed & F5 Rules) | Fail Close | Blocks legitimate large traffic | |
| Barracuda WAF | Fail Close | Blocks legitimate large traffic | |
| Cloudflare WAF | Fail Open | Vulnerable | |
| F5 NGINX AppProtect (Strict/Default) | Fail Open | Vulnerable | |
| F5 BIG-IP Advanced WAF | Fail Open | Vulnerable | |
| NGINX ModSecurity | Fail Open | Vulnerable | |
| Fortinet FortiAppSec | Fail Open | Vulnerable |
Conclusion
Only CloudGuard WAF and Google Cloud Armor successfully inspected the full payload depth. The majority of the market (including F5, Cloudflare, and Fortinet) defaulted to a "Fail Open" state, rendering them ineffective against padded RCE vulnerabilities like React2Shell. Conversely, AWS and Azure prioritize security over usability, likely requiring significant exception handling for modern, data-heavy applications.
Methodology
Datasets
Each WAF solution was tested against two large datasets: Legitimate and Malicious.
Legitimate Requests Dataset
The Legitimate Requests Dataset is carefully designed to test WAF behaviors in real-world scenarios. This dataset includes 1,040,242 HTTP requests from 692 real-world websites and allows for challenging WAF systems by examining their responses to a range of website functionalities.
Category Distribution:
| Category | Websites | Examples |
|---|---|---|
| E-Commerce | 404 | eBay, Ikea |
| Travel | 75 | Booking, Airbnb |
| Information | 59 | Wikipedia, Daily Mail |
| Food | 40 | Wolt, Burger King |
| Search Engines | 24 | Duckduckgo, Bing |
| Social media | 17 | Facebook, Instagram |
| Files uploads | 16 | Adobe, Shutterfly |
| Content creation | 13 | Office, Atlassian |
| Games | 13 | Roblox, Steam |
| Videos | 8 | YouTube, twitch |
| Files download | 7 | Google, Dropbox |
| Applications | 7 | IBM Quantum Simulator, Planner 5d |
| Streaming | 6 | Spotify, Youtube Music |
| Technology | 3 | Microsoft, Lenovo |
| Total | 692 |
The Legitimate Requests Dataset including all HTTP requests is available for further research.
Malicious Requests Dataset
The Malicious Requests Dataset includes 73,924 malicious payloads from various attack vectors, including SQL Injection, Cross-Site Scripting (XSS), XML External Entity (XXE), and others. The malicious payloads were sourced from multiple GitHub resources, which ensure a comprehensive approach to testing WAF solutions.
Tools
To ensure transparency and reproducibility, the testing tool is made available to the public. The responses from each request sent by the test tool to the WAFs were systematically logged in a database for further analysis.
Comparison Metrics
To quantify the efficacy of each WAF, we use statistical measures:
- Security Quality - the proportion of actual positives correctly identified (True Positive Rate).
- Detection Quality - the proportion of actual negatives correctly identified (True Negative Rate).
- Balanced Accuracy - an arithmetic mean of these two metrics.