ROAM-ASD with missing modalities

green frame: predicted speakingred frame: predicted silentorange circle: disagrees with ground truth% in each panel header: accuracy over the clip under that condition

One clip, left to right: all streams, no audio, mouth dropped, face dropped

Four separate videos of the same moment. No audio: the model was fed silence, and this clip carries no audio track. Mouth dropped: the model receives no mouth crop, so the mouth region of every face is masked. Face dropped: the model receives no face crop but still gets the mouth crop, so the face is masked except the mouth. Boxes and scores in each video are what the model output under that condition. Press play all to run the four together. Under each example the same clip is shown for the control model trained without modality dropout (same architecture, data and schedule, streams never dropped in training).

AVA-ActiveSpeaker — C25wkwAMB-w, 1038–1048 s

Same clip, model trained without modality dropout (control):

all streamsno audiomouth droppedface dropped
ROAM-ASD100.0%94.0%100.0%100.0%
no-dropout control99.8%89.8%71.4%99.8%

UniTalk - mixed (hard) — ABqo634gI5s, 11–21 s

Same clip, model trained without modality dropout (control):

all streamsno audiomouth droppedface dropped
ROAM-ASD99.2%97.8%99.3%99.3%
no-dropout control99.7%98.0%34.6%99.3%

WASD - Face Occlusion — ZQESmCZZnH4_185-215, 0–10 s

Same clip, model trained without modality dropout (control):

all streamsno audiomouth droppedface dropped
ROAM-ASD100.0%100.0%100.0%100.0%
no-dropout control100.0%98.6%75.0%100.0%

WASD - Human Voice Noise — Lh4dSfTGnGk_230-254, 12–22 s

Same clip, model trained without modality dropout (control):

all streamsno audiomouth droppedface dropped
ROAM-ASD96.8%97.2%96.8%97.0%
no-dropout control98.7%70.2%70.2%98.5%

WASD - Surveillance Settings — ZOIPEe4fcaM_824-854, 0–10 s

Same clip, model trained without modality dropout (control):

all streamsno audiomouth droppedface dropped
ROAM-ASD100.0%100.0%100.0%100.0%
no-dropout control100.0%72.6%96.4%100.0%

Temporary dropout: missing intervals in a stream (relative holes)

A stream is removed in random holes whose length is a fraction of the track's own length — 5%, 15% or 30% of the track — placed so that 30% of each track is missing in total (seeded, identical for all models), and the model is scored only inside the holes. Left panel: the clip with all streams. Right panel: the gap run — an amber border and banner mark every frame in which the stream was missing, and the boxes show what the model output at that moment. Each row shows the 12 s window of the clip that contains the most missing frames for that hole length.

WASD - Surveillance Settings — ZOIPEe4fcaM_824-854, face + mouth holes of 5% of the track

Clip accuracy: no gaps 100.0% · gap run, all frames 100.0% · inside the holes 100.0% (427 frames)

WASD - Surveillance Settings — ZOIPEe4fcaM_824-854, face + mouth holes of 15% of the track

Clip accuracy: no gaps 100.0% · gap run, all frames 100.0% · inside the holes 100.0% (590 frames)

WASD - Surveillance Settings — ZOIPEe4fcaM_824-854, face + mouth holes of 30% of the track

Clip accuracy: no gaps 100.0% · gap run, all frames 100.0% · inside the holes 100.0% (696 frames)

WASD - Human Voice Noise — Lh4dSfTGnGk_230-254, audio holes of 5% of the track

Clip accuracy: no gaps 96.1% · gap run, all frames 97.5% · inside the holes 100.0% (231 frames)

WASD - Human Voice Noise — Lh4dSfTGnGk_230-254, audio holes of 15% of the track

Clip accuracy: no gaps 97.6% · gap run, all frames 99.0% · inside the holes 100.0% (428 frames)

WASD - Human Voice Noise — Lh4dSfTGnGk_230-254, audio holes of 30% of the track

Clip accuracy: no gaps 98.2% · gap run, all frames 97.9% · inside the holes 99.1% (429 frames)

AVA-ActiveSpeaker — C25wkwAMB-w, face + mouth holes of 5% of the track

Clip accuracy: no gaps 97.9% · gap run, all frames 99.5% · inside the holes 99.6% (262 frames)

AVA-ActiveSpeaker — C25wkwAMB-w, face + mouth holes of 15% of the track

Clip accuracy: no gaps 97.9% · gap run, all frames 98.0% · inside the holes 100.0% (264 frames)

AVA-ActiveSpeaker — C25wkwAMB-w, face + mouth holes of 30% of the track

Clip accuracy: no gaps 91.5% · gap run, all frames 97.0% · inside the holes 95.1% (288 frames)

AVA-ActiveSpeaker — kMy-6RtoOVU, face + mouth holes of 5% of the track

Clip accuracy: no gaps 96.0% · gap run, all frames 94.7% · inside the holes 96.9% (322 frames)

AVA-ActiveSpeaker — kMy-6RtoOVU, face + mouth holes of 15% of the track

Clip accuracy: no gaps 96.2% · gap run, all frames 92.6% · inside the holes 88.1% (336 frames)

AVA-ActiveSpeaker — kMy-6RtoOVU, face + mouth holes of 30% of the track

Clip accuracy: no gaps 96.2% · gap run, all frames 96.6% · inside the holes 92.9% (296 frames)

UniTalk - mixed (hard) — ABqo634gI5s, face + mouth holes of 5% of the track

Clip accuracy: no gaps 84.6% · gap run, all frames 81.9% · inside the holes 73.5% (717 frames)

UniTalk - mixed (hard) — ABqo634gI5s, face + mouth holes of 15% of the track

Clip accuracy: no gaps 98.9% · gap run, all frames 89.6% · inside the holes 77.9% (789 frames)

UniTalk - mixed (hard) — ABqo634gI5s, face + mouth holes of 30% of the track

Clip accuracy: no gaps 74.6% · gap run, all frames 82.4% · inside the holes 90.6% (877 frames)