Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation Paper • 2211.06687 • Published Nov 12, 2022 • 7
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published 12 days ago • 82
Vidu S1: A Real-Time Interactive Video Generation Model Paper • 2607.03118 • Published 19 days ago • 138
Program-as-Weights: A Programming Paradigm for Fuzzy Functions Paper • 2607.02512 • Published 20 days ago • 122
Towards Automating Scientific Review with Google's Paper Assistant Tool Paper • 2606.28277 • Published 26 days ago • 9
Autodata: An agentic data scientist to create high quality synthetic data Paper • 2606.25996 • Published 28 days ago • 18
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance Paper • 2606.19195 • Published Jun 17 • 141
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding Paper • 2605.27365 • Published May 26 • 146
From Context to Skills: Can Language Models Learn from Context Skillfully? Paper • 2604.27660 • Published May 3 • 171
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents Paper • 2604.26752 • Published Apr 29 • 112
3AM: Segment Anything with Geometric Consistency in Videos Paper • 2601.08831 • Published Jan 13 • 34