The toolkit to test, validate, and evaluate your models and surface, curate, and prioritize the most valuable data for labeling.
-
Updated
May 23, 2025 - Python
The toolkit to test, validate, and evaluate your models and surface, curate, and prioritize the most valuable data for labeling.
🛠️ Corrected Test Sets for ImageNet, MNIST, CIFAR, Caltech-256, QuickDraw, IMDB, Amazon Reviews, 20News, and AudioSet
AQuA: A Benchmarking Tool for Label Quality Assessment, NeurIPS'23 D&B
Analysis of label prop errors in ML pipelines and their impact on accuracy+demographic parity
Can capture-recapture count the benchmark label errors nobody found? Two pre-registered experiments on MMLU-Redux with LLM judges and confident learning (negative result).
To associate your repository with the label-errors topic, visit your repo's landing page and select "manage topics."