multimodalopinionbullishExplicit, compositional and editable spatio-temporal scene graph representations can enable richer grounded activity understanding from first-person videoComputer Vision27 Jul 2026http://arxiv.org/abs/2607.02425v1