PII Arena
PPLX detector

pplx+homoglyph PII detection benchmark

Measured on 3 of 41 retained datasets under the same frozen scoring and masking protocol.

81.93%fully hidden annotations
89.11%detected / overlapped
7.62%extra masking
CPU characters / second

Strengths

Highest complete-masking rates by normalized data type.

Phone numbers & email90.24%703 / 779
People's names88.91%1,323 / 1,488
Documents & identifiers88.23%2,369 / 2,685

Watch list

Lowest observed complete-masking rates; inspect before deployment.

Network identifiers59.56%215 / 361
Addresses & locations64.81%919 / 1,418
Bank accounts & cards79.03%294 / 372

Language coverage

The overall leader can differ from the best choice for one language slice.

LanguageDatasetsFully hidden ↑Detected ↑Extra masking ↓
English175.69%95.83%10.28%
Russian282.18%88.84%1.56%

Configuration, source and version

Family
PPLX
Revision
f1f90a53823f5df0a1344c1e137d9fffdaab54d6
Eligible datasets
3 / 41
Character F1
0.552
Untouched diagnostic
10.89% · 815 / 7,486
Benchmark version
1.0.3 · 2026-09-09
CPU seconds / 10k chars
Not measured
GPU seconds / 10k chars
2.153

Results by dataset

Complete masking is primary; detection shows whether any character overlapped.

DatasetGold annotationsFully hidden ↑Detected ↑Extra masking ↓
corrupt-secrets-issues28875.69%95.83%10.28%
corrupt-redmadrobot5,53178.67%86.53%1.58%
corrupt-hivetrace1,66793.82%96.52%1.45%