The goal of this protocol is to determine species- and site-specific confidence score thresholds using logistic regression to improve detection precision in large acoustic datasets processed with automated acoustic recognition software.
Method Article
The goal of this protocol is to determine species- and site-specific confidence score thresholds using logistic regression to improve detection precision in large acoustic datasets processed with automated acoustic recognition software.
Passive acoustic monitoring (PAM) has become an invaluable tool for biodiversity research, enabling the non-invasive collection of vast datasets. However, a significant challenge remains in efficiently and reliably processing this large volume of data to extract species-specific information across varying locations. This paper presents a detailed, step-by-step protocol to address this challenge using a machine learning detector module within a bioacoustics analysis software. The methodology is designed to accurately and confidently identify and validate bird vocalisations from raw acoustic recordings.
Our protocol details the process from initial data collection using autonomous recording units (ARUs) to the final generation of a high-quality annotated dataset. Key steps include configuring the machine learning detector module to generate initial detections, a manual validation procedure to calculate precision tables, and a logistic regression analysis to determine a species-specific and, where appropriate, a location-specific confidence score threshold. This statistically derived threshold is then used to refine the detector’s output, tested on two overlap configurations (0 s and 2 s). We show that applying the derived optimal confidence score thresholds substantially improves the machine learning-based avian sound recognition models detection precision across sites. For the three sites used to illustrate the process (Liverpool Park, Cairngorms, and Glasgow Suburban) precision increased by 26.1%, 17.7%, and 17% for an overlap of 0 s, and by 28.77%, 16.87%, and 15% for an overlap of 2 s. We suggest the resulting methodology is superior to manual counting methods in both speed and reliability. In summary, this paper provides a reproducible framework that facilitates the accessible and effective use of machine learning approaches in bioacoustics, enabling researchers to confidently leverage large acoustic datasets for ecological studies and parameter analysis.
Animal vocalisations offer valuable insights into the natural world and act as powerful biological tools allowing researchers to investigate various aspects of animals lives1. The field of bioacoustics has become increasingly important in biodiversity monitoring and conservation, aiding in species identification and in assessing population trends and ecosystem health2,3,4. This field of research also contributes to our understanding of animal behaviour, for example, in communication and habitat use5,6. More recently, it has been demonstrated as a promising approach for assessing animal welfare7,8.
One of the key advantages of using acoustic analysis as a research tool is the recent technological advances in autonomous recording units (ARUs), which enable researchers to non-invasively and remotely monitor animal vocalisations for extended periods of time, a technique known as passive acoustic monitoring (PAM)9. This has allowed for the collection of data from cryptic species without human intervention10,11. Additionally, ARUs are cost-effective and minimise the need for trained researchers to be in the field9. A study on the cryptic European nightjar (Caprimulgus europaeus) demonstrated these advantages, finding that the use of ARUs was significantly more effective than human observers in detecting the species12. Large-scale bioacoustics studies are increasingly becoming more popular because they can provide insights into a range of ecological and behavioural aspects, they are also highly applicable and well-suited for field studies. Yet ARUs running for long periods mean researchers are inevitably faced with the challenge of efficiently and reliably processing vast amounts of acoustic data13.
More recently, researchers have begun combining PAM with the use of automatic recognition tools to detect species within large datasets, enhancing and speeding up the overall processing of data. These techniques automatically recognise and classify species based on vocalisation type13, often using machine learning and deep learning models such as artificial neural networks, which have gained traction in recent years14. The reliability of automatic recognition software is often superior to manual methods of species detection, as demonstrated in studies on Giant Panda (Ailuropoda melanoleuca) vocalisations15 and Malaysian frog species16. Importantly, one of the primary drivers behind adoption of these approaches is processing speed, particularly as the volume of acoustic data generated through PAM continues to grow13. Manual analysis of large acoustic datasets is inefficient and impractical at scale, yet automated recognition tools can process recordings in a fraction of the time required by a human observer. For example, one study reported that 257 h of manual visual scanning was equivalent to 1 h of processing time using BirdNET (machine learning–based avian sound recognition model)17.
This machine learning–based avian sound recognition model is one such tool that has recently gained traction for its use in ornithology and other acoustics studies, it is an accessible automated software for bird species identification that relies on a deep learning model18,19. Most studies to date have used the machine learning–based avian sound recognition model to generate species lists for a given area, and with human validation it has been shown to detect 89% of species in a community, while requiring less than a quarter of the time used by visual scanning alone17. This machine learning–based avian sound recognition model has also been shown to effectively identify the presence of a specific target species in an area, such as the Eurasian Bittern (Botaurus stellaris), with a 93.7% success rate20. However, this model often has a higher success rate when combined with manual validation and isn’t yet reliable as a fully automatic detection tool21. There are also other challenges to consider according to Wood and Kahl21 regarding its production of unitless predictions.
The machine learning detector module assigns each detected bird vocalisation a quantitative confidence score between 0 and 1, which indicates how certain the software is that the detected sound has been correctly identified as a particular species18,19,21. Precision, defined as the proportion of detections that are true positives (i.e. correct identifications), is a key metric influenced by this score22. Consequently, a higher confidence score may produce less detections, but with higher precision, yet miss detections especially in noisy environments, and a lower score may produce more detections with less certain precision22,23. The relationship between the confidence scores and precision is complex, it is often found that some high confidence scores can be false positives, for example in locations with high levels of background noise, or with species that are likely to mimic other species vocalisations23,24. Due to the variability in precision among species and location, it is possible during the configuration process to assign individual confidence score thresholds to each bird species, whereby only detections with scores above that threshold will be labelled. A high universal confidence score threshold could be applied across all species and locations, however, doing so often increases the number of false negatives (missed identifications)22,25,26, while the number of false positives vary between species and with environmental differences between recordings.
Several studies and best-practice guidelines outline various methods for determining an optimal confidence score threshold at which to run the machine learning detector module for specific species and study locations, enabling detections to be accepted with greater confidence as true positives21,24,26,27,28. One approach involves calculating F-scores across a range of confidence thresholds and selecting the threshold that maximises the F-score, balancing both precision and recall29,30. An alternative approach, which allows for probability calibration, involves manually validating a random subset of predictions and fitting a logistic regression to establish a relationship between the binary validation outcome (true/false detection) and the associated confidence score21. This method can then be used to determine context-specific thresholds, which may be adjusted across species and environmental conditions (e.g., varying background noise)19. Such methods have been applied in passive acoustic monitoring studies on Yucatán black howler monkeys (Alouatta pigra), gray wolves (Canis lupus) and coyotes (C. latrans)27,28.
The aim of this methods paper is to demonstrate and validate a scalable methodology for accurately and reliably extracting species-specific vocal data from large acoustic datasets using the machine learning–based avian sound recognition model within Raven Pro-bioacoustics analysis software (ver 1.6)19,31. This approach incorporates the calculation of appropriate confidence score thresholds following Wood and Kahl’s21 best-practice guidelines. While these guidelines21 provide a comprehensive framework for applying confidence scores within the machine learning–based avian sound recognition model19, they are primarily presented as best-practice recommendations. Few studies have operationalised these procedures across distributed ARU study systems under realistic monitoring constraints, where acoustic conditions and background noise vary spatially22. This protocol (Figure 1) is particularly suitable for large-scale passive acoustic monitoring studies involving multiple sites, variable acoustic environments, or target species with high misclassification rates. Furthermore, the implementation of these guidelines within the bioacoustics analysis software (ver 1.6)19,31, a more accessible and user-friendly platform compared to coding-based environments32,33 has not been previously demonstrated.

Figure 1: Schematic representation of the entire workflow. Please click here to view a larger version of this figure.
To address this gap, we apply and validate the protocol using large-scale datasets recorded with AudioMoths - ARUs (ver 1.2.0)34,35, across a diverse range of habitats. Our focal species the European Robin (Erithacus rubecula), exhibits high song variability and structural complexity36,37, and demonstrated a susceptibility to misidentification in initial pilot analyses. By applying species- and site-specific confidence score thresholds to the machine learning–based avian sound recognition model configuration within the bioacoustics analysis software19,31, we evaluate the protocol under varying acoustic conditions, incorporating differing background noise levels and species assemblages, and demonstrate that tailoring threshold optimisation can substantially improve detection precision in comparison to default settings. This approach is designed to ultimately facilitate accessible, streamlined, and efficient application of the machine learning–based avian sound recognition model19 in large scale PAM studies, with minimal coding.
Access restricted. Please log in or start a trial to view this content.
Ethics statement
This protocol follows the ethical guidelines of Liverpool John Moores University. Methodology involved passive acoustic monitoring only, and did not include the handling, disturbance, or experimental manipulation of animals. No hazardous procedures were required to deploy the ARUs. ARUs were deployed in accordance with institutional research and fieldwork guidelines, and appropriate signage was installed at recording sites to inform the public of acoustic monitoring. However, it is highly advised to check weather forecasts before conducting field activity, and to conduct a risk assessment for each field site.
Software, licensing and usage
NOTE: ARU data should be processed using the machine learning–based avian sound recognition model within the bioacoustics analysis software (v1.6)19,31, with detector outputs also processed in the same software (v1.6)31. This is a licensed software and is not freely available, use of the workflow described here requires access via an institutional or individual license. In addition, the machine learning–based avian sound recognition model within the software (v1.6)19,31, does not currently run on Apple devices with ARM-based processor architecture (M1 or M2 chips), however this protocol is compatible with macOS-based computers with Intel CPUs and Windows operating systems
1. Collecting audio data
2. Using the machine learning detector module in the bioacoustics analysis software19,31 for automatic species detection
3. Create precision tables for target species at a representative ARU from each study location
4. Perform a logistic regression analysis on the validation tables to generate an optimal confidence score threshold

Access restricted. Please log in or start a trial to view this content.
To illustrate the effectiveness of the described protocol, we present representative results from the analysis of acoustic data for the European Robin (Erithacus rubecula) collected at three sites of varying environments and avian species assemblages, including a park in Liverpool (ARU 1), a naturally regenerating woodland site within The Cairngorms National Park (ARU 2), and a suburban site in Glasgow (ARU 3). These examples demonstrate how applying statistically derived confidence score thresholds can improve ...
Access restricted. Please log in or start a trial to view this content.
The use of the protocol outlined in this paper provides a reliable, streamlined method for generating species-specific and accurately classified datasets of bird vocalisations, which has proven useful when handling large audio datasets. We chose to test the protocol on the European Robin due to its well-known structural and frequency variability within its song repertoires43. Meaning initially, detections were found to have high mislabelling (false positive) rates (Table 1). We al...
Access restricted. Please log in or start a trial to view this content.
Authors have no competing financial interest.
We thank Liverpool John Moore’s University for guidance and resources. We also thank the Wild Animal Initiative (WAI) for funding support, this research was funded by WAI grants (S-2023–00038).
Access restricted. Please log in or start a trial to view this content.
| Name | Company | Catalog Number | Comments |
|---|---|---|---|
| Autonomous Recording Units v1.2.0 | Open Acoustic Devices | N/A | Autonomous recording unit (ARU) |
| Bioacoustics Analysis Software v1.6.5 | Cornell Lab of Ornithology | RRID:SCR_016190 | Sound analysis software |
| Data Transfer Cable | Various Manufacturers | N/A | Used to transfer data from devices |
| Device Configuration Applications | Open Acoustic Devices | N/A | Apps for configuring AudioMoths |
| Machine Learning Detector Module v2.4 | Cornell Lab of Ornithology | N/A | Machine learning detector |
| Rechargeable Alkaline Batteries | Various Manufacturers | N/A | Power supply for recording units |
| Removable Memory Cards | Various Manufacturers | N/A | Data storage for recordings |
| Statistical Computing Software | R Core Team | RRID:SCR_001905 | Statistical computing software |
Access restricted. Please log in or start a trial to view this content.
Request permission to reuse the text or figures of this JoVE article
Request Permission