MBZUAI/NADI-2025-Sub-task-3-test
Viewer • Updated • 365 • 20
Natural Language Processing, Machine Learning, and Computer Vision
Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding
Training-Free Speech-Centric Omni Understanding with Frozen VLMs