Ensembles
Explore the masking tradeoffs of fixed detector compositions.
Dataset slice
Composition outcomes
Manually selected compositions evaluated on this benchmark; no independent selection holdout.
| Composition | Untouched ↓ | Fully hidden ↑ | Extra masking ↓ | Char F1 ↑ |
|---|---|---|---|---|
| 20.95% | 74.92% | 11.42% | 0.716 | |
| 13.19% | 79.86% | 5.88% | 0.763 | |
| 3.46% | 91.94% | 15.27% | 0.736 | |
| 2.09% | 92.96% | 18.84% | 0.701 | |
| 1.87% | 94.08% | 17.10% | 0.714 | |
| 1.43% | 94.40% | 19.82% | 0.685 | |
| Partial coverage · 40/41 datasets | 27.17% | 64.40% | 6.37% | 0.713 |
Detection and full hiding
A larger union can hide more annotations while masking more unannotated text.
Ensemble costs in the source report are derived from member throughput, not measured end-to-end latency. Compositions containing BardsAI have a mixed-device GPU estimate.