BEANS-Next
37,718 examples across 45 tasks. Evaluate bioacoustic abilities with audio and text, using a common library for inference and scoring.
Broadening Audio-Language Capabilities
for Bioacoustics
A benchmark and training resource for audio-language models, spanning acoustic perception, biological category recognition, scene understanding, and in-context learning.
Paper and citation: coming soon on arXiv.
37,718 examples across 45 tasks. Evaluate bioacoustic abilities with audio and text, using a common library for inference and scoring.
44 million audio-language pairs from 6.5 million distinct audio clips. Training supervision built from bioacoustic archives, rich metadata, and synthetic audio.
Bioacoustic research requires more than identifying which species is present. BEANS-Next organizes evaluation into four tiers of abilities shared with ROOTS training tasks.
Describe vocalizations and reason about their acoustic properties, including pitch and recording quality.
Recognize species, call types, and behaviors, and describe the biological content of recordings.
Count events, track their order and overlap, and relate acoustic properties to individual species.
Use reference audio to identify unfamiliar species or call types and detect sounds in a query recording.
ROOTS combines tasks built from structured metadata, LLM-based text synthesis, and audio-derived information. Synthetic recordings supply controlled examples where strong labels are scarce.
These separately released resources provide source clips and annotated synthetic scenes used in ROOTS.
Extracted vocalization clips with species metadata, used as building blocks for synthetic tasks.
View on Hugging Face →Synthetic soundscapes with event timing and species annotations for detection and scene understanding.
View on Hugging Face →Synthetic scenes with event timing and source identities for detection and source-based tasks.
View on Hugging Face →Existing bioacoustic and general-purpose audio-language models struggle with abilities beyond familiar classification tasks.
Fine-tuning NatureLM-audio separately on each tier of ROOTS improves performance across all four tiers of BEANS-Next. Open-ended description, counting, and transfer to real-world recordings remain challenging.