REPRODUCIBLE PRODUCT EVIDENCE
Secret Sentinel regression benchmark
This report answers a narrow question: did every automated case in the named private corpus pass at one exact revision? It does not turn regression tests into a claim of universal detection accuracy.
388/388
automated test cases passed
- Benchmark version
- 2026.08
- Core version
- 1.0.0
- Generated
- August 2, 2026 UTC
- Revision
36f066384ead
Results by behavior
Detection and non-detection
90/900 failed · 0 skipped
Redaction integrity
57/570 failed · 0 skipped
Idempotency and convergence
50/500 failed · 0 skipped
Jira and Confluence workflow
97/970 failed · 0 skipped
Configuration and aggregate storage
67/670 failed · 0 skipped
Core reliability
27/270 failed · 0 skipped
What is actually measured
The unit is one automated test case, not one credential provider and not one customer document. The private synthetic corpus covers documented credential shapes and safe near misses, Jira ADF and Confluence storage redaction, repeated-event convergence, workflow routing, configuration, aggregate KVS behavior, and supporting reliability utilities.
Classification is deterministic and performed only after Vitest completes. The public JSON contains group counts—never fixture values, test names, repository paths, scanner rules, or regular expressions.
What the number does not mean
- It is not “100% secret-detection accuracy.”
- It does not estimate recall for unknown or future credential formats.
- It does not promise zero false positives in real customer content.
- It is not a head-to-head score against GitHub, GitGuardian, DLP, or another scanner.
- It does not replace an Atlassian sandbox acceptance test.
Integrity and independent inspection
CI first executes the private corpus, generates a canonical aggregate report, then signs that exact JSON blob with Cosign using GitHub's short-lived OIDC identity. The workflow publishes the report and Sigstore verification bundle together. A permanent private signing key is not stored in the repository.