BEANS-Next & ROOTS

BEANS-Next and ROOTS

Broadening Audio-Language Capabilities
for Bioacoustics

A benchmark and training resource for audio-language models, spanning acoustic perception, biological category recognition, scene understanding, and in-context learning.

Christos Plachouras1,2, David Robinson1, Marius Miron1, Gagan Narula1, Paul Laisné1, Anthony L. T. Fine1, Benno Weck1, Ellen Gilsenan-McMahon1, Diane Kim1, Laura Hay Mack1, Maddie Cusimano1, Sara Keen1, Lukas Rauch1, Benjamin Hoffman1, Emmanuel Chemla1, Emmanouil Benetos2, Johan Pauwels2, Milad Alizadeh1, Matthieu Geist1, Olivier Pietquin1

1 Earth Species Project2 Queen Mary University of London

Paper and citation: coming soon on arXiv.

BEANS-Next · The benchmark

Beyond species recognition

Bioacoustic research requires more than identifying which species is present. BEANS-Next organizes evaluation into four tiers of abilities shared with ROOTS training tasks.

BEANS-Next evaluates audio-language models across four tiers, while ROOTS supplies training data from bioacoustic archives, citizen science, synthetic data, and audio-derived information.
Evaluation and training follow a shared taxonomy, from acoustic perception to learning from reference recordings.
TIER 1

Acoustic perception

Describe vocalizations and reason about their acoustic properties, including pitch and recording quality.

TIER 2

Biological category recognition

Recognize species, call types, and behaviors, and describe the biological content of recordings.

TIER 3

Scene understanding

Count events, track their order and overlap, and relate acoustic properties to individual species.

TIER 4

In-context learning

Use reference audio to identify unfamiliar species or call types and detect sounds in a query recording.

ROOTS · The training resource

Broader supervision from animal sounds

ROOTS combines tasks built from structured metadata, LLM-based text synthesis, and audio-derived information. Synthetic recordings supply controlled examples where strong labels are scarce.

Data construction pipeline: extract information from audio, synthesize language using metadata, score and filter the resulting examples, and compute statistics.
Construction combines source metadata and audio-derived evidence with text synthesis, scoring, and filtering.

Audio components

These separately released resources provide source clips and annotated synthetic scenes used in ROOTS.

AnimalSpeak-PseudoVox

Extracted vocalization clips with species metadata, used as building blocks for synthetic tasks.

View on Hugging Face →

Synthetic Strong Detection

Synthetic soundscapes with event timing and species annotations for detection and scene understanding.

View on Hugging Face →

Synthetic Detect-Diarize

Synthetic scenes with event timing and source identities for detection and source-based tasks.

View on Hugging Face →
Research findings

Broader tasks reveal
room to improve

Existing bioacoustic and general-purpose audio-language models struggle with abilities beyond familiar classification tasks.

Fine-tuning NatureLM-audio separately on each tier of ROOTS improves performance across all four tiers of BEANS-Next. Open-ended description, counting, and transfer to real-world recordings remain challenging.