ROAM-ASD: Robust Open-World Active Speaker Detection with Flexible Multimodal Fusion

mAP (mean average precision, in %) is the standard active-speaker-detection score: for every annotated face at every frame the model outputs a speaking probability, and mAP is the area under the precision–recall curve of those scores against the ground-truth speaking labels — 100 means every speaking face is ranked above every silent one. The number next to each title is ROAM-ASD (ours) against SOTA — the state of the art, i.e. the best-performing published method on that dataset (WASD and UniTalk per category).

green frame: predicted speakingred frame: predicted silentorange circle: prediction disagrees with ground truthnumber above a face: speaking probability

AVA-ActiveSpeakermAP: ROAM-ASD (ours) 96.5 · SOTA 95.6

Data description: Movie clips; the standard active-speaker benchmark.

ASWmAP: ROAM-ASD (ours) 99.28 · SOTA 98.3

Data description: Broadcast and interview video, often many small faces on screen.

TalkiesmAP: ROAM-ASD (ours) 98.19 · SOTA 96.1

Data description: Short in-the-wild clips with many candidate faces and one speaker.

WASDmAP over all five categories pooled: ROAM-ASD (ours) 98.79 · SOTA 93.7

WASD · Optimal ConditionsmAP: ROAM-ASD (ours) 99.62 · SOTA 97.8

Data description: Optimal Conditions: people talking in an alternate manner, with minor interruptions, cooperative poses, and face availability.

WASD · Speech ImpairmentmAP: ROAM-ASD (ours) 99.68 · SOTA 98.3

Data description: Speech Impairment: frontal-pose subjects either talking via video-conference call (delayed speech) or in a heated discussion with potential talking overlap (speech overlap), face availability ensured.

WASD · Face OcclusionmAP: ROAM-ASD (ours) 98.92 · SOTA 95.4

Data description: Face Occlusion: at least one subject has partial facial occlusion, while keeping good speech quality.

WASD · Human Voice NoisemAP: ROAM-ASD (ours) 96.29 · SOTA 84.7

Data description: Human Voice Noise: communication between speakers while another human voice plays in the background, with face availability and subject cooperation.

WASD · Surveillance SettingsmAP: ROAM-ASD (ours) 96.18 · SOTA 77.9

Data description: Surveillance Settings: speaker communication in video-surveillance scenarios, with varying audio and image quality and no guarantee of face access, speech quality, or subject cooperation.

UniTalkmAP overall: ROAM-ASD (ours) 87.92 · SOTA 83.2

UniTalk · languagemAP: ROAM-ASD (ours) 91.56 · SOTA 86.7

Data description: Language: speech in languages other than English.

UniTalk · noisemAP: ROAM-ASD (ours) 90.60 · SOTA 84.1

Data description: Noise: background sound competing with the speech (music, crowds, engines, other voices).

UniTalk · visual (crowded)mAP: ROAM-ASD (ours) 89.28 · SOTA 84.9

Data description: Visual: visually hard scenes, such as crowded frames, small or partly occluded faces, and inserts.

UniTalk · mixed (hard)mAP: ROAM-ASD (ours) 81.19 · SOTA 77.9

Data description: Mixed: several of the above difficulties at once.