LEAKS detector
noseyparker PII detection benchmark
Measured on 3 of 41 retained datasets under the same frozen scoring and masking protocol.
7.59%fully hidden annotations
8.33%detected / overlapped
1.65%extra masking
—CPU characters / second
Strengths
Highest complete-masking rates by normalized data type.
Passwords, keys & tokens18.16%205 / 1,129
Bank accounts & cards0.00%0 / 169
Documents & identifiers0.00%0 / 609
Watch list
Lowest observed complete-masking rates; inspect before deployment.
Addresses & locations0.00%0 / 176
Phone numbers & email0.00%0 / 390
People's names0.00%0 / 228
Language coverage
The overall leader can differ from the best choice for one language slice.
| Language | Datasets | Fully hidden ↑ | Detected ↑ | Extra masking ↓ |
|---|---|---|---|---|
| English | 2 | 18.18% | 20.12% | 1.74% |
| Russian | 1 | 1.02% | 1.02% | 0.00% |
Configuration, source and version
- Family
- LEAKS
- Revision
- 0.24.0
- Eligible datasets
- 3 / 41
- Character F1
- 0.320
- Untouched diagnostic
- 91.67% · 2,476 / 2,701
- Benchmark version
- 1.0.3 · 2026-09-09
- CPU seconds / 10k chars
- Not measured
- GPU seconds / 10k chars
- Not measured
Results by dataset
Complete masking is primary; detection shows whether any character overlapped.
| Dataset | Gold annotations | Fully hidden ↑ | Detected ↑ | Extra masking ↓ |
|---|---|---|---|---|
| hivetrace | 1,667 | 1.02% | 1.02% | 0.00% |
| secrets-rules | 746 | 16.22% | 16.22% | 10.97% |
| secrets-issues | 288 | 23.26% | 30.21% | 0.83% |