HF detector
ru-legal-ner+cpu PII detection benchmark
Measured on 2 of 41 retained datasets under the same frozen scoring and masking protocol.
62.97%fully hidden annotations
79.94%detected / overlapped
1.70%extra masking
—CPU characters / second
Strengths
Highest complete-masking rates by normalized data type.
People's names91.88%5,360 / 5,834
Phone numbers & email86.39%1,339 / 1,550
Organizations73.60%538 / 731
Watch list
Lowest observed complete-masking rates; inspect before deployment.
Network identifiers2.62%34 / 1,297
Customer & employee IDs7.45%36 / 483
Other sensitive attributes10.71%24 / 224
Language coverage
The overall leader can differ from the best choice for one language slice.
| Language | Datasets | Fully hidden ↑ | Detected ↑ | Extra masking ↓ |
|---|---|---|---|---|
| Russian | 2 | 62.97% | 79.94% | 1.70% |
Configuration, source and version
- Family
- HF
- Revision
- c0313a42ac147ccc38f5c3b0b3e77cab53208683
- Eligible datasets
- 2 / 41
- Character F1
- 0.788
- Untouched diagnostic
- 20.06% · 3,652 / 18,205
- Benchmark version
- 1.0.3 · 2026-09-09
- CPU seconds / 10k chars
- Not measured
- GPU seconds / 10k chars
- Not measured
Results by dataset
Complete masking is primary; detection shows whether any character overlapped.