Four separate videos of the same moment. No audio: the model was fed silence, and this clip carries no audio track. Mouth dropped: the model receives no mouth crop, so the mouth region of every face is masked. Face dropped: the model receives no face crop but still gets the mouth crop, so the face is masked except the mouth. Boxes and scores in each video are what the model output under that condition. Press play all to run the four together. Under each example the same clip is shown for the control model trained without modality dropout (same architecture, data and schedule, streams never dropped in training).
Same clip, model trained without modality dropout (control):
| all streams | no audio | mouth dropped | face dropped | |
|---|---|---|---|---|
| ROAM-ASD | 100.0% | 94.0% | 100.0% | 100.0% |
| no-dropout control | 99.8% | 89.8% | 71.4% | 99.8% |
Same clip, model trained without modality dropout (control):
| all streams | no audio | mouth dropped | face dropped | |
|---|---|---|---|---|
| ROAM-ASD | 99.2% | 97.8% | 99.3% | 99.3% |
| no-dropout control | 99.7% | 98.0% | 34.6% | 99.3% |
Same clip, model trained without modality dropout (control):
| all streams | no audio | mouth dropped | face dropped | |
|---|---|---|---|---|
| ROAM-ASD | 100.0% | 100.0% | 100.0% | 100.0% |
| no-dropout control | 100.0% | 98.6% | 75.0% | 100.0% |
Same clip, model trained without modality dropout (control):
| all streams | no audio | mouth dropped | face dropped | |
|---|---|---|---|---|
| ROAM-ASD | 96.8% | 97.2% | 96.8% | 97.0% |
| no-dropout control | 98.7% | 70.2% | 70.2% | 98.5% |
Same clip, model trained without modality dropout (control):
| all streams | no audio | mouth dropped | face dropped | |
|---|---|---|---|---|
| ROAM-ASD | 100.0% | 100.0% | 100.0% | 100.0% |
| no-dropout control | 100.0% | 72.6% | 96.4% | 100.0% |
A stream is removed in random holes whose length is a fraction of the track's own length — 5%, 15% or 30% of the track — placed so that 30% of each track is missing in total (seeded, identical for all models), and the model is scored only inside the holes. Left panel: the clip with all streams. Right panel: the gap run — an amber border and banner mark every frame in which the stream was missing, and the boxes show what the model output at that moment. Each row shows the 12 s window of the clip that contains the most missing frames for that hole length.
Clip accuracy: no gaps 100.0% · gap run, all frames 100.0% · inside the holes 100.0% (427 frames)
Clip accuracy: no gaps 100.0% · gap run, all frames 100.0% · inside the holes 100.0% (590 frames)
Clip accuracy: no gaps 100.0% · gap run, all frames 100.0% · inside the holes 100.0% (696 frames)
Clip accuracy: no gaps 96.1% · gap run, all frames 97.5% · inside the holes 100.0% (231 frames)
Clip accuracy: no gaps 97.6% · gap run, all frames 99.0% · inside the holes 100.0% (428 frames)
Clip accuracy: no gaps 98.2% · gap run, all frames 97.9% · inside the holes 99.1% (429 frames)
Clip accuracy: no gaps 97.9% · gap run, all frames 99.5% · inside the holes 99.6% (262 frames)
Clip accuracy: no gaps 97.9% · gap run, all frames 98.0% · inside the holes 100.0% (264 frames)
Clip accuracy: no gaps 91.5% · gap run, all frames 97.0% · inside the holes 95.1% (288 frames)
Clip accuracy: no gaps 96.0% · gap run, all frames 94.7% · inside the holes 96.9% (322 frames)
Clip accuracy: no gaps 96.2% · gap run, all frames 92.6% · inside the holes 88.1% (336 frames)
Clip accuracy: no gaps 96.2% · gap run, all frames 96.6% · inside the holes 92.9% (296 frames)
Clip accuracy: no gaps 84.6% · gap run, all frames 81.9% · inside the holes 73.5% (717 frames)
Clip accuracy: no gaps 98.9% · gap run, all frames 89.6% · inside the holes 77.9% (789 frames)
Clip accuracy: no gaps 74.6% · gap run, all frames 82.4% · inside the holes 90.6% (877 frames)