multimodal
fact
bearish
Vision-language models fail the majority of navigation tasks in 4DSynth-Nav benchmark and stall after early subtasks
Two vision-language models evaluated across three difficulty tiers both fail the majority of tasks and stall after early subtasks.
Computer Vision28 Aug 2026