Mice (Mus musculus) communicate using ultrasonic vocalizations (USVs) during social interactions. However, because mouse USVs are not associated with clear visual indicators and occur at frequencies outside the human auditory range, it remains challenging to identify individual vocalizing mice within a group. This protocol describes a method for recording synchronized video and ultrasonic audio data during multi-animal social interactions, enabling simultaneous capture of behavior and vocal activity. A computational pipeline is outlined that integrates multi-animal tracking, vocalization detection, sound-source localization, and assignment of vocalizations to individual animals. The validation procedures used to assess tracking accuracy, vocalization extraction, localization precision, and assignment confidence at each stage of the pipeline are further described. Application of this protocol to a four-animal demonstration dataset yields individual-resolved vocalization tracking, spatially precise sound-source estimates, and quantitative measures of vocal output and acoustic features. Together, these approaches provide a robust framework for linking vocal communication to individual behavior and enable downstream analyses of sex differences, individual variability, and social communication in freely interacting mice.