multimodal
fact
bullish
CODE framework with OWL-ViT L/14 backbone achieves 21.7 U-mAP and 40.8 K-mAP in Task 1 of Real-World Detection benchmark, surpassing previous state of the art by 2.6 and 2.3 points respectively
with the OWL-ViT L/14 backbone, CODE achieves 21.7 U-mAP and 40.8 K-mAP in Task 1, surpassing the previous state of the art by 2.6 and 2.3 points, respectively
Computer Vision30 Aug 2026