Method Article

Enhancing an Avian Sound Recognition Model's Detection Precision via Logistic Regression of Large Acoustic Datasets: A Case Study of the European Robin (Erithacus rubecula)

437 views

DOI:

10.3791/70479

April 11th, 2026

In This Article

Summary

The goal of this protocol is to determine species- and site-specific confidence score thresholds using logistic regression to improve detection precision in large acoustic datasets processed with automated acoustic recognition software.

Abstract

Passive acoustic monitoring (PAM) has become an invaluable tool for biodiversity research, enabling the non-invasive collection of vast datasets. However, a significant challenge remains in efficiently and reliably processing this large volume of data to extract species-specific information across varying locations. This paper presents a detailed, step-by-step protocol to address this challenge using a machine learning detector module within a bioacoustics analysis software. The methodology is designed to accurately and confidently identify and validate bird vocalisations from raw acoustic recordings.

Our protocol details the process from initial data collection using autonomous recording units (ARUs) to the final generation of a high-quality annotated dataset. Key steps include configuring the machine learning detector module to generate initial detections, a manual validation procedure to calculate precision tables, and a logistic regression analysis to determine a species-specific and, where appropriate, a location-specific confidence score threshold. This statistically derived threshold is then used to refine the detector’s output, tested on two overlap configurations (0 s and 2 s). We show that applying the derived optimal confidence score thresholds substantially improves the machine learning-based avian sound recognition models detection precision across sites. For the three sites used to illustrate the process (Liverpool Park, Cairngorms, and Glasgow Suburban) precision increased by 26.1%, 17.7%, and 17% for an overlap of 0 s, and by 28.77%, 16.87%, and 15% for an overlap of 2 s. We suggest the resulting methodology is superior to manual counting methods in both speed and reliability. In summary, this paper provides a reproducible framework that facilitates the accessible and effective use of machine learning approaches in bioacoustics, enabling researchers to confidently leverage large acoustic datasets for ecological studies and parameter analysis.

Introduction

Animal vocalisations offer valuable insights into the natural world and act as powerful biological tools allowing researchers to investigate various aspects of animals lives1. The field of bioacoustics has become increasingly important in biodiversity monitoring and conservation, aiding in species identification and in assessing population trends and ecosystem health2,3,4. This field of research also contributes to our understanding of animal behaviour, for example, in communication and habitat use5,6. More recently, it has been demonstrated as a promising approach for assessing animal welfare7,8.

One of the key advantages of using acoustic analysis as a research tool is the recent technological advances in autonomous recording units (ARUs), which enable researchers to non-invasively and remotely monitor animal vocalisations for extended periods of time, a technique known as passive acoustic monitoring (PAM)9. This has allowed for the collection of data from cryptic species without human intervention10,11. Additionally, ARUs are cost-effective and minimise the need for trained researchers to be in the field9. A study on the cryptic European nightjar (Caprimulgus europaeus) demonstrated these advantages, finding that the use of ARUs was significantly more effective than human observers in detecting the species12. Large-scale bioacoustics studies are increasingly becoming more popular because they can provide insights into a range of ecological and behavioural aspects, they are also highly applicable and well-suited for field studies. Yet ARUs running for long periods mean researchers are inevitably faced with the challenge of efficiently and reliably processing vast amounts of acoustic data13

More recently, researchers have begun combining PAM with the use of automatic recognition tools to detect species within large datasets, enhancing and speeding up the overall processing of data. These techniques automatically recognise and classify species based on vocalisation type13, often using machine learning and deep learning models such as artificial neural networks, which have gained traction in recent years14. The reliability of automatic recognition software is often superior to manual methods of species detection, as demonstrated in studies on Giant Panda (Ailuropoda melanoleuca) vocalisations15 and Malaysian frog species16. Importantly, one of the primary drivers behind adoption of these approaches is processing speed, particularly as the volume of acoustic data generated through PAM continues to grow13. Manual analysis of large acoustic datasets is inefficient and impractical at scale, yet automated recognition tools can process recordings in a fraction of the time required by a human observer. For example, one study reported that 257 h of manual visual scanning was equivalent to 1 h of processing time using BirdNET (machine learning–based avian sound recognition model)17.

This machine learning–based avian sound recognition model is one such tool that has recently gained traction for its use in ornithology and other acoustics studies, it is an accessible automated software for bird species identification that relies on a deep learning model18,19. Most studies to date have used the machine learning–based avian sound recognition model to generate species lists for a given area, and with human validation it has been shown to detect 89% of species in a community, while requiring less than a quarter of the time used by visual scanning alone17. This machine learning–based avian sound recognition model has also been shown to effectively identify the presence of a specific target species in an area, such as the Eurasian Bittern (Botaurus stellaris), with a 93.7% success rate20. However, this model often has a higher success rate when combined with manual validation and isn’t yet reliable as a fully automatic detection tool21. There are also other challenges to consider according to Wood and Kahl21 regarding its production of unitless predictions.

The machine learning detector module assigns each detected bird vocalisation a quantitative confidence score between 0 and 1, which indicates how certain the software is that the detected sound has been correctly identified as a particular species18,19,21. Precision, defined as the proportion of detections that are true positives (i.e. correct identifications), is a key metric influenced by this score22. Consequently, a higher confidence score may produce less detections, but with higher precision, yet miss detections especially in noisy environments, and a lower score may produce more detections with less certain precision22,23. The relationship between the confidence scores and precision is complex, it is often found that some high confidence scores can be false positives, for example in locations with high levels of background noise, or with species that are likely to mimic other species vocalisations23,24. Due to the variability in precision among species and location, it is possible during the configuration process to assign individual confidence score thresholds to each bird species, whereby only detections with scores above that threshold will be labelled. A high universal confidence score threshold could be applied across all species and locations, however, doing so often increases the number of false negatives (missed identifications)22,25,26, while the number of false positives vary between species and with environmental differences between recordings.

Several studies and best-practice guidelines outline various methods for determining an optimal confidence score threshold at which to run the machine learning detector module for specific species and study locations, enabling detections to be accepted with greater confidence as true positives21,24,26,27,28. One approach involves calculating F-scores across a range of confidence thresholds and selecting the threshold that maximises the F-score, balancing both precision and recall29,30. An alternative approach, which allows for probability calibration, involves manually validating a random subset of predictions and fitting a logistic regression to establish a relationship between the binary validation outcome (true/false detection) and the associated confidence score21. This method can then be used to determine context-specific thresholds, which may be adjusted across species and environmental conditions (e.g., varying background noise)19. Such methods have been applied in passive acoustic monitoring studies on Yucatán black howler monkeys (Alouatta pigra), gray wolves (Canis lupus) and coyotes (C. latrans)27,28.

The aim of this methods paper is to demonstrate and validate a scalable methodology for accurately and reliably extracting species-specific vocal data from large acoustic datasets using the machine learning–based avian sound recognition model within Raven Pro-bioacoustics analysis software (ver 1.6)19,31. This approach incorporates the calculation of appropriate confidence score thresholds following Wood and Kahl’s21 best-practice guidelines. While these guidelines21 provide a comprehensive framework for applying confidence scores within the machine learning–based avian sound recognition model19, they are primarily presented as best-practice recommendations. Few studies have operationalised these procedures across distributed ARU study systems under realistic monitoring constraints, where acoustic conditions and background noise vary spatially22. This protocol (Figure 1) is particularly suitable for large-scale passive acoustic monitoring studies involving multiple sites, variable acoustic environments, or target species with high misclassification rates. Furthermore, the implementation of these guidelines within the bioacoustics analysis software (ver 1.6)19,31, a more accessible and user-friendly platform compared to coding-based environments32,33 has not been previously demonstrated.

Bioacoustics machine learning workflow diagram for detecting species; logistic regression, precision analysis.
Figure 1: Schematic representation of the entire workflow. Please click here to view a larger version of this figure.

To address this gap, we apply and validate the protocol using large-scale datasets recorded with AudioMoths - ARUs (ver 1.2.0)34,35, across a diverse range of habitats. Our focal species the European Robin (Erithacus rubecula), exhibits high song variability and structural complexity36,37, and demonstrated a susceptibility to misidentification in initial pilot analyses. By applying species- and site-specific confidence score thresholds to the machine learning–based avian sound recognition model configuration within the bioacoustics analysis software19,31, we evaluate the protocol under varying acoustic conditions, incorporating differing background noise levels and species assemblages, and demonstrate that tailoring threshold optimisation can substantially improve detection precision in comparison to default settings. This approach is designed to ultimately facilitate accessible, streamlined, and efficient application of the machine learning–based avian sound recognition model19 in large scale PAM studies, with minimal coding.

Access restricted. Please log in or start a trial to view this content.

Protocol

Ethics statement
This protocol follows the ethical guidelines of Liverpool John Moores University. Methodology involved passive acoustic monitoring only, and did not include the handling, disturbance, or experimental manipulation of animals. No hazardous procedures were required to deploy the ARUs. ARUs were deployed in accordance with institutional research and fieldwork guidelines, and appropriate signage was installed at recording sites to inform the public of acoustic monitoring. However, it is highly advised to check weather forecasts before conducting field activity, and to conduct a risk assessment for each field site.

Software, licensing and usage
NOTE: ARU data should be processed using the machine learning–based avian sound recognition model within the bioacoustics analysis software (v1.6)19,31, with detector outputs also processed in the same software (v1.6)31. This is a licensed software and is not freely available, use of the workflow described here requires access via an institutional or individual license. In addition, the machine learning–based avian sound recognition model within the software (v1.6)19,31, does not currently run on Apple devices with ARM-based processor architecture (M1 or M2 chips), however this protocol is compatible with macOS-based computers with Intel CPUs and Windows operating systems

1. Collecting audio data

  1. Prepare ARUs for deployment
    1. Insert batteries into the ARUs (rechargeable alkaline batteries were used during this experiment).
    2. Insert appropriately sized removable memory cards relative to the amount of data to be collected (2 GB removable memory cards were used during this experiment).
      Note: For recording schedules producing 108 1-min files per day over a 30 day period (see step 1.2.5.2), each file was approximately 5760 kB, resulting in a total data volume of 622 MB. With an estimated daily energy consumption of 26 mAh. Therefore, removable memory cards with a capacity of 2 GB were used for the representative results to provide sufficient storage. Additionally, rechargeable alkaline batteries retained usable power for the longest period, as opposed to lithium AA batteries and rechargeable NiMH, under the test conditions.
  2. Configure each ARU for deployment
    1. Download the ‘ARU device firmware update application’39 and the ‘ARU device configuration application’40 from the ‘App Download Page’ on the ARU website41.
    2. Connect the ARUs individually to a computer using a data transfer cable.
    3. Ensure the device is switched to the USB/Off mode.
    4. Use the ‘ARU device firmware update application’39 to update the devices to the latest firmware.
    5. Use the ‘ARU device configuration application’40 to set the ARUs to record at a target time of day for a selected number of days. Determine active recording periods and sleep intervals.
      1. Identify the time windows when target species are most active (e.g., dawn chorus for diurnal birds), and schedule recording periods accordingly.
      2. Set an appropriate duty cycle, specifying how long the ARUs are actively recording versus in sleep mode.
        Example: For these representative results, ARUs were programmed to record daily between 4:00–9:00 AM and 5:00–9:00 PM from 1st June-30th June 2023. These time windows were selected to capture both the dawn and dusk choruses. During active periods, devices recorded 1 min WAV files followed by a 4 min sleep interval, resulting in a total of 108 1 min recordings per day.
        Note: ARUs can record continually, but this rapidly increases battery consumption and removable memory card usage.
      3. Use the ‘ARU device configuration application’42 to set the devices to record at a sampling rate of 48 kHz as recommended for capturing a full range of bird vocalisations34,35 and with medium gain (typically left at default unless recording in noisy environments34,35). Other sampling rates can be used depending on the frequency of target species vocalisations, for example for ultrasonic wildlife, such as bats or amphibians, use a sampling rate of 384 kHz35.
    6. After configuring all ARUs, switch each device to the CUSTOM mode.
  3. Device protection and labelling
    1. Insert each ARU device into an Antistatic Ziplock bag. Use additional waterproof tape as needed to prevent water damage. Place Silica Gel packets inside the bags to absorb moisture.
    2. Label each ARU by number and its location.
  4. ARU deployment
    1. Determine locations for ARUs to be deployed within the target species habitat.
    2. Deploy ARUs at spacing distances appropriate to the study requirements and target species.
      Note: Detection distance depends on habitat, background noise, and species’ call volume. While loud species (e.g., Eurasian bittern) can be detected by ARUs up to 800 m, recordings degrade at longer ranges, limiting fine scale measurements20. Here, one ARU per site was analysed, with units about 100 m apart to ensure good spatial coverage and clear vocalisations. Adjust spacing according to species, habitat, and project goals.

2. Using the machine learning detector module in the bioacoustics analysis software19,31 for automatic species detection

  1. Transfer audio data
    1. After the deployment period has ended, remove removable memory cards from the ARUs, keeping note of which removable memory cards correspond to each numbered ARU.
    2. Transfer all WAV files from each removable memory cards onto a computer using a card reader.
    3. Organise all WAV files generated from each ARU into folders labelled with the corresponding ARU number and location.
    4. Save back up versions of all data onto additional devices. Pause point: Protocol can be paused here.
    5. Restrict access to this data to authorised project researchers and ensure data is password protected to maintain confidentiality, particularly where human activity may have been inadvertently recorded.
  2. Configure the machine learning detector module
    1. Open the bioacoustics analysis software31 and under the drop-down menu select Tools > Detector > Learning Detector.
    2. In the Choose Detector Inputs window, select Browse. In the subsequent Open Sound Files window, select all the WAV files within the relevant ARU folder and click Open (Supplementary Files S1 and S2).
    3. In the Configure New Sound Window that opens, under Paging > Page Sound, enter the appropriate Page Size for the length of each recording (e.g., 60 s for 1 min WAV files) (Supplementary File S3).
    4. In the same Configure New Sound Window, under Multiple Files select Open as file sequence in one window, then click OK (Supplementary File S3).
    5. In the returning Choose Detector Inputs window, select Waveform under Available Signals and Views on the right and click the << arrows button to move it over to Required Inputs. Confirm that the status under the Required Inputs box reads Ready!, and click OK (Supplementary File S4).
    6. In the Configure Machine Learning Detector window that appears, under the Inputs tab, select the desired model from Select Model. When analysing bird species use the global avian classification model, which among various updates include over 3000–6000 bird species19 (Supplementary File S5).
    7. Under the Inputs tab, leave the Overlap at the default of 0 s to prevent any interference with the subsequent confidence score threshold analysis (Supplementary File S5).
    8. Under the Outputs tab, select the drop-down menu Output Class File and select the output list that relates to the geographical location of the study (Supplementary File S6).
    9. Under the Outputs tab there will be a list of bird species relevant to the chosen Output Class File. Lower the Threshold to 0.1 and select Apply All, make sure that all values under Threshold are now at 0.1 in the list (Supplementary File S6).
    10. Select the Suppress All Species option and ticks will appear in all boxes in the species list. Then locate the target species and un-tick it in the Suppress column (Supplementary File S6).
    11. Select OK to begin the Learning Detector, this will process the files and create selections where it has detected the target species vocalisations.
    12. Monitor the Progress Manager to track Percentage Completed and Time Remaining (Supplementary File S7).
  3. Save the Learning Detector output
    Note: A Learning Detector table will be generated during processing and displayed beneath the spectrogram view. The numbered selections in this table correspond to the detector’s identified selections on the spectrogram view. Each selection is fixed at 3 s in length and may not always be repeated over the entire vocalisation being detected. Each selection is assigned an identification Label and Score in the selection table (Supplementary File S8). These scores are a unitless prediction and will be used in the logistic regression analysis in subsequent stages.
    1. Select File > Save Selection Table “Learning Detector” As and save the file using the ARU name and an identifier such as: AM1_LearningDetectorTable_Original.txt.

3. Create precision tables for target species at a representative ARU from each study location

  1. Prepare the validation dataset
    1. Open the Learning Detector selection table that was saved in step 2.3.1 and transfer to a spreadsheet software.
    2. Confirm that only the target species appears under the Label column.
    3. Delete all columns apart from the Selection and Score columns.
    4. Randomise the rows using the spreadsheet’s randomisation function (e.g., add a column using =RAND(), sort this column, then delete the randomisation column).
    5. Retain the first 300 randomised detections, if fewer than 300 exist, retain all detections.
      Note: Selecting 300 detections for validation is standard in similar studies using 1 month of acoustic data24. This sample size estimates precision and mislabelling rates reliably, captures a range of detection scores, and keeps workload manageable. More selections may be needed for longer datasets.
    6. Save the validation dataset using the ARU name being analysed and an identifier such as: AM1_ValidationDataset.xlsx.
  2. Prepare bioacoustics analysis software for manual validation
    1. Open the software31, select File > Open Sound Files… and in the Open Sound Files window select all the corresponding WAV files for the ARU being analysed, and click Open.
    2. In the Configure New Sound Window which opens, under Paging, select Page Sound and enter the Page size used in step 2.2.3. Under Multiple Files select Open as file sequence in one window and select OK.
    3. Navigate to File > Open Selection Table… and open the Learning Detection Table for the ARU being analysed saved in step 2.3.1.
    4. In the Panel on the left of the Sound view, under the Layouts tab, in Views:, unselect the Waveform view so all files are displayed in Spectrogram view to enable visual confirmation of vocalisation structure.
  3. Validate randomised selections
    1. Alongside the bioacoustics analysis software31 with the prepared data ready for manual validation, open the validation spreadsheet saved in step 3.1.6.
    2. Create a new column in the spreadsheet labelled Correctly Detected.
    3. Locate each selection number from the randomised list within the spectrogram view and input whether it has been correctly detected (1) or incorrectly detected (0) into the Correctly Detected column.
  4. Calculate detector precision
    1. Calculate precision of the detector (prior to the logistic regression) for the ARU and target species using the formula:
    2. Precision (%) = (Number of correct detections/Total validated detections) x 100
    3. For datasets with precision of 90% or greater, retain current detector settings.
    4. For datasets with precision lower than 90%, carry onto section 4 to fit a logistic regression analysis to determine the optimal threshold to set the detector to run at.
      Note: When precision falls below 90%, high mislabelling rates make reliable extraction of vocalisations difficult without threshold adjustment. A 90% cut-off balances detection reliability with adequate sample size. Previous machine learning–based avian sound recognition model calibration studies also regard precision rates of 90% and above as being sufficiently reliable42.

4. Perform a logistic regression analysis on the validation tables to generate an optimal confidence score threshold

  1. Generate confidence score thresholds, utilising Wood & Kahl’s21 guidelines
    1. Use the software R32 to import the validation dataset created in section 3.
    2. Run a logistic regression using the generalized linear modelling function with the family as binomial on the dataset.
    3. Set the desired probability of precision as 0.9, this means that the probability that a detection is a true positive will be 90% or above.
      Note: Setting the desired probability of precision to 0.9 ensures that most detections are true positives, however, even at this level false positives can occur in large acoustic datasets and users should be aware of this limitation. A 90% precision threshold allows users to maintain high numbers of correct detections, while retaining enough data for a robust analysis, aiming for 100% would limit usable data.
    4. Calculate the optimal threshold by solving the fitted logistic regression equation for the predictor value (confidence score) that yields a predicted probability of 0.9 using the Logistic Regression Analysis Script below:
      Logistic regression formula; threshold calculation; applied statistics equation; math concept diagram.
      p = 0.9, and β0 and β1 ​are the intercept and slope from the fitted model
      Logistic Regression Analysis Script for statistical computing software
      # Load validation table
      ThresholdRobin <-read.csv('ValidationTableRobin.csv')
      # Run Logistic Regression
      Model1<-glm(Correctly.Detected~Score, data=ThresholdRobin, family='binomial')
      Model1
      # Desired probability above 90% precision
      p <- 0.9 # the desired p (probability of true positive)
      threshold <- (log(p/(1-p))- Model1$coefficients[1]) / Model1$coefficients [2]
      Note: Instead of using the Logit Score as in previous studies24,27 as the independent variable, use the raw confidence scores as the independent variable. In Wood and Kahl’s21 guidelines it is mentioned that either the raw confidence scores or back-transformed versions can be used.
  2. Apply the optimal confidence score
    1. Open the bioacoustics analysis software31.
    2. Re-run the Learning Detector on the ARU being investigated using the steps in section 2. However, this time in the configuration process, the Threshold for the target species being detected must be changed to the threshold determined in section 4.1 of the protocol.
    3. (Optional) Increase Overlap to 2 s to increase detection resolution.
      Note: The Overlap value determines how much adjacent detected segments overlap. At 0 s vocalisations on segment borders might be missed (false negatives). Increasing overlap improves detection resolution and predictive power but slows analysis19,29. Here, for the logistic regression, a default overlap of 0 s was used. With the representative results and confidence score thresholds calculated in the logistic regression analysis, an overlap of both 0 and 2 s is tested.
    4. A new list of detections will be created that now should be more accurate, with less false positives (see Results).
    5. Save the new Learning Detector output with the list of detections that were generated using the derived confidence score threshold by selecting File > Save Selection Table “Learning Detector” As and save the file using the ARU name and an identifier such as: AM1_LearningDetectorTable_OptimisedThreshold.txt.
    6. This file represents the final validated dataset generated by the protocol. Use the saved optimised learning detector selection table as the final output of this protocol. Please note that subsequent analyses, such as annotation of vocalisations and acoustic parameter extraction, fall outside the scope of this method.

Access restricted. Please log in or start a trial to view this content.

Results

To illustrate the effectiveness of the described protocol, we present representative results from the analysis of acoustic data for the European Robin (Erithacus rubecula) collected at three sites of varying environments and avian species assemblages, including a park in Liverpool (ARU 1), a naturally regenerating woodland site within The Cairngorms National Park (ARU 2), and a suburban site in Glasgow (ARU 3). These examples demonstrate how applying statistically derived confidence score thresholds can improve ...

Access restricted. Please log in or start a trial to view this content.

Discussion

The use of the protocol outlined in this paper provides a reliable, streamlined method for generating species-specific and accurately classified datasets of bird vocalisations, which has proven useful when handling large audio datasets. We chose to test the protocol on the European Robin due to its well-known structural and frequency variability within its song repertoires43. Meaning initially, detections were found to have high mislabelling (false positive) rates (Table 1). We al...

Access restricted. Please log in or start a trial to view this content.

Disclosures

Authors have no competing financial interest.

Acknowledgements

We thank Liverpool John Moore’s University for guidance and resources. We also thank the Wild Animal Initiative (WAI) for funding support, this research was funded by WAI grants (S-2023–00038).

Access restricted. Please log in or start a trial to view this content.

Materials

List of materials used in this article
NameCompanyCatalog NumberComments
Autonomous Recording Units v1.2.0Open Acoustic DevicesN/AAutonomous recording unit (ARU)
Bioacoustics Analysis Software v1.6.5Cornell Lab of OrnithologyRRID:SCR_016190Sound analysis software
Data Transfer CableVarious ManufacturersN/AUsed to transfer data from devices
Device Configuration ApplicationsOpen Acoustic DevicesN/AApps for configuring AudioMoths
Machine Learning Detector Module v2.4Cornell Lab of OrnithologyN/AMachine learning detector 
Rechargeable Alkaline BatteriesVarious ManufacturersN/APower supply for recording units
Removable Memory CardsVarious ManufacturersN/AData storage for recordings
Statistical Computing SoftwareR Core TeamRRID:SCR_001905Statistical computing software

References

  1. Briefer, E. F. Vocal expression of emotions in mammals: Mechanisms of production and evidence. J Zool. 288 (1), 1-20 (2012).
  2. Changapur, M. W., Seema, S., Sowmya, B. J. Bioacoustics monitoring to improve conservation efforts for endangered species. 2023 7th International Conference on Computation System and Information Technology for Sustainable Solutions (CSITSS), , IEEE. (2023).
  3. Sebastián-González, E., Pang-Ching, J., Barbosa, J. M., Hart, P. Bioacoustics for species management: two case studies with a Hawaiian forest bird. Ecol Evol. 5 (20), 4696-4705 (2015).
  4. Raich, X., et al. Use of bioacoustics in species identification: Piranhas from genus Pygocentrus (Teleostei: Serrasalmidae) as a case study. PLOS One. 15 (10), e0241316(2020).
  5. Pillay, R., et al. Bioacoustic monitoring reveals shifts in breeding songbird populations and singing behaviour with selective logging in tropical forests. J Appl Ecol. 56 (11), 2482-2492 (2019).
  6. Teixeira, D., Maron, M., van Rensburg, B. J. Bioacoustic monitoring of animal vocal behaviour for conservation. Conserv Sci Pract. 1 (8), e72(2019).
  7. Coutant, M., Villain, A. S., Briefer, E. L. A scoping review of the use of bioacoustics to assess various components of farm animal welfare. Appl Anim Behav Sci. 275, 106286(2024).
  8. Mcloughlin, M. P., Stewart, R., McElligott, A. G. Automated bioacoustics: methods in ecology and conservation and their potential for animal welfare monitoring. J R Soc Interface. 16 (155), 20190225(2019).
  9. Shonfield, J., Bayne, E. M. Autonomous recording units in avian ecological research: Current use and future applications. Avian Conserv Ecol. 12 (1), 14(2017).
  10. Newson, S. E., Bas, Y., Murray, A., Gillings, S. Potential for coupling the monitoring of bush-crickets with established large-scale acoustic monitoring of bats. Methods Ecol Evol. 8 (9), 1051-1062 (2017).
  11. Wheeldon, A., et al. Comparison of acoustic and traditional point count methods to assess bird diversity and composition in the Aberdare National Park, Kenya. Afr J Ecol. 57 (2), 168-176 (2019).
  12. Zwart, M. C., Baker, A., McGowan, P. J. K., Whittingham, M. J. The Use of Automated Bioacoustic Recorders to Replace Human Wildlife Surveys: An Example Using Nightjars. PLoS ONE. 9 (7), e102770(2014).
  13. Xie, J., et al. Review of automatic recognition technology for bird vocalizations in the deep learning era. Ecol Inform. 73, 101927(2023).
  14. Yang, F., Jiang, Y., Xu, Y. Design of bird sound recognition model based on lightweight. IEEE Access. 10, 85189-85198 (2023).
  15. Liao, Z., et al. Automatic recognition of giant panda vocalizations using wide spectrum features and deep neural network. Math Biosci Eng. 20 (8), 15456-15475 (2023).
  16. Jaafar, H., Ramli, D. A. Automatic syllables segmentation for frog identification system. 2013 IEEE 9th International Colloquium on Signal Processing and its Applications, Kuala Lumpur, Malaysia, , https://ieeexplore.ieee.org/document/6530046 224-228 (2013).
  17. Ware, L., Mahon, C. L., McLeod, L., Jetté, J. F. Artificial intelligence (BirdNET) supplements manual methods to maximize bird species richness from acoustic data sets generated from regional monitoring. Can J Zool. 101 (12), 1031-1051 (2023).
  18. Kahl, S., Wood, C. M., Eibl, M., Klinck, H. BirdNET: A deep learning solution for avian diversity monitoring. Ecol Inform. 61, 101236(2021).
  19. Kahl, K., Wood, C., Klinck, M. Cornell Lab of Ornithology. BirdNET: Bird sound identification. [computer program]. , The Cornell Lab of Ornithology. Ithica (NY). https://birdnet.cornell.edu (2021).
  20. Manzano-Rubio, R., et al. Low-cost open-source recorders and ready-to-use machine learning approaches provide effective monitoring of threatened species. Ecol Inform. 72, 101910(2022).
  21. Wood, C. M., Kahl, S. Guidelines on the appropriate use of BirdNet scores and other detector outputs. J Ornithol. 165, 777-782 (2024).
  22. Pérez-Granados, C. BirdNET: applications, performance, pitfalls and future opportunities. Ibis. 165 (3), 1068-1075 (2023).
  23. Singer, D., et al. Aggregated time-series features boost species-specific differentiation of true and false positives in passive acoustic monitoring of bird assemblages. Remote Sens Ecol Conserv. 10 (4), 517-530 (2024).
  24. Bota, G., et al. Hearing to the Unseen: AudioMoth and BirdNET as a Cheap and Easy Method for Monitoring Cryptic Bird Species. Sensors. 23 (16), 7176(2023).
  25. Amorós-Ausina, A., Schuchmann, K. -L., Marques, M. I., Pérez-Granados, C. Living together, singing together: Revealing similar patterns of vocal activity in two tropical songbirds applying BirdNET. Sensors. 24 (17), 5780(2024).
  26. Tseng, S., Hodder, D. P., Otter, K. A. Setting BirdNET confidence thresholds: species-specific vs. universal approaches. J Ornithol. 166 (4), 1123-1135 (2025).
  27. Wood, C. M., Cruz, A. B., Kahl, S. Pairing a user-friendly machine-learning animal sound detector with passive acoustic surveys for occupancy modelling of an endangered primate. Am J Primatol. 85 (8), e23507(2023).
  28. Sossover, D., Burrows, K., Kahl, S., Wood, C. M. Using the BirdNET algorithm to identify wolves, coyotes, and potentially their interactions in a large audio dataset. Mamm Res. 69, 159-165 (2024).
  29. Funosas, D., et al. Assessing the potential of BirdNET to infer European bird communities from large-scale ecoacoustic data. Ecol Indic. 164, 112146(2023).
  30. Ziegenhorn, M. A., et al. ArcticSoundsNET: BirdNET embeddings facilitate improved bioacoustic classification of Arctic species. Ecol Inform. 90, 103270(2025).
  31. Raven Pro: Interactive sound analysis software. [Computer program]. , Version 1.6, The Cornell Lab of Ornithology. https://ravensoundsoftware.com/ (2025).
  32. R Core Team. R: A language and environment for statistical computing [computer program]. , R Foundation for Statistical Computing. https://www.R-project.org/ (2024).
  33. Python Language Reference [Internet]. , Version 3.14.3, Python Software Foundation. https://www.python.org (2026).
  34. Hill, A. P., et al. AudioMoth: Evaluation of a Smart open Acoustic Device for Monitoring Biodiversity and the Environment. Methods in Ecol and Evol. 9 (5), 1199-1211 (2017).
  35. Hill, A. P., et al. AudioMoth: A low-cost acoustic device for monitoring biodiversity and the environment. HardwareX. 6, e00073(2019).
  36. Bremond, J. C. Recherche sur la semantique et les elements vecteurs d'information dans les signaux aeoustiques du rouge-gorge (Erithacus rubecula). Terre Vie. 22, https://hal.science/hal-03531635v1 109-220 (1968).
  37. Hoelzel, A. R. Song characteristics and response to playback of male and female robins, Erithacus rubecula. Ibis. 128, 115-127 (1968).
  38. Sandbrook, C., et al. Principles for the socially responsible use of conservation monitoring technology and data. Conserv Sci Pract. 3 (5), e374(2021).
  39. AudioMoth Flash App [computer program]. , Version 1.7.0, Open Acoustic Devices. https://www.openacousticdevices.info/applications (2025).
  40. AudioMoth Configuration App [computer program]. , Version 1.12.1, Open Acoustic Devices. https://www.openacousticdevices.info/applications (2025).
  41. Applications. , Open Acoustic Devices. https://www.openacousticdevices.info/applications (2025).
  42. Sethi, S. S., et al. Large-scale avian vocalization detection delivers reliable global biodiversity insights. Proc Natl Acad Sci U S A. 121 (33), e2315933121(2024).
  43. Verboom, W. C. Bird vocalizations: songs of the European robin (Erithacus rubecula). JunoBioacoustics. 202107, (2021).

Access restricted. Please log in or start a trial to view this content.

Reprints and Permissions

Request permission to reuse the text or figures of this JoVE article

Request Permission

Tags

European Robin DetectionPassive Acoustic MonitoringMachine Learning DetectorBioacoustics AnalysisAutonomous Recording UnitsConfidence Score Threshold

Related Articles