mAP (mean average precision, in %) is the standard active-speaker-detection score: for every annotated face at every frame the model outputs a speaking probability, and mAP is the area under the precision–recall curve of those scores against the ground-truth speaking labels — 100 means every speaking face is ranked above every silent one. The number next to each title is ROAM-ASD (ours) against SOTA — the state of the art, i.e. the best-performing published method on that dataset (WASD and UniTalk per category).
Data description: Movie clips; the standard active-speaker benchmark.
Data description: Broadcast and interview video, often many small faces on screen.
Data description: Short in-the-wild clips with many candidate faces and one speaker.
Data description: Optimal Conditions: people talking in an alternate manner, with minor interruptions, cooperative poses, and face availability.
Data description: Speech Impairment: frontal-pose subjects either talking via video-conference call (delayed speech) or in a heated discussion with potential talking overlap (speech overlap), face availability ensured.
Data description: Face Occlusion: at least one subject has partial facial occlusion, while keeping good speech quality.
Data description: Human Voice Noise: communication between speakers while another human voice plays in the background, with face availability and subject cooperation.
Data description: Surveillance Settings: speaker communication in video-surveillance scenarios, with varying audio and image quality and no guarantee of face access, speech quality, or subject cooperation.
Data description: Language: speech in languages other than English.
Data description: Noise: background sound competing with the speech (music, crowds, engines, other voices).
Data description: Visual: visually hard scenes, such as crowded frames, small or partly occluded faces, and inserts.
Data description: Mixed: several of the above difficulties at once.