Listening to
the Pantanal.
How do you recognize wildlife in a noisy soundscape—with only 90 minutes of CPU time? Our work brings together distillation and diverse audio models to tackle that constraint.
About the working note
Listening to the Pantanal: Runtime-Constrained Distillation and Diversity-Aware Ensembling for BirdCLEF+ 2026.
The ensemble combines sound event detection models with a Perch/ProtoSSM branch using per-class rank fusion. Proceedings are currently in staging.