3 packages found
[CVPR 2025 🔥]A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
[CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural langu
Code repo for the paper: Attacking Vision-Language Computer Agents via Pop-ups