Speech processing
How to Evaluate Audio Classification Models
Design audio classification tests with class balance, macro F1, confusion analysis, and real recording conditions.
Accuracy can look high when one audio class dominates the dataset. A useful test shows how the classifier performs on each event and how often it confuses events that matter.
Define labels before collecting clips
Write a short description and boundary examples for every class. Decide whether clips may have multiple labels, an unknown label, or background only. Collect recordings from the microphones, environments, and languages expected in use. Split by speaker, device, or recording session where leakage would otherwise make the held-out set too easy.
Report class-level performance
For a single-label classifier, report a confusion matrix, precision, recall, and F1 per class plus macro F1. For multi-label tasks, specify the decision threshold and report per-label precision and recall. Show class counts and uncertainty for small slices. Hugging Face Evaluate supports audio classification evaluation and combined accuracy, precision, recall, and F1 metrics.
Stress the full pipeline
Test noise, clipping, quiet events, overlapping sounds, and short versus long clips. Measure preprocessing and inference latency together. Review false alarms and missed events in terms of their product cost; an alerting system may prefer a different threshold from a music tagging tool. Publish label definitions, threshold, split rule, and model revision so a future comparison can be repeated.