Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation Paper • 2211.06687 • Published Nov 12, 2022 • 7
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published 12 days ago • 82
Vidu S1: A Real-Time Interactive Video Generation Model Paper • 2607.03118 • Published 19 days ago • 138
CohereLabs/cohere-transcribe-arabic-07-2026 Automatic Speech Recognition • 2B • Updated 8 days ago • 23.5k • 125
Program-as-Weights: A Programming Paradigm for Fuzzy Functions Paper • 2607.02512 • Published 20 days ago • 122
Towards Automating Scientific Review with Google's Paper Assistant Tool Paper • 2606.28277 • Published 26 days ago • 9
Autodata: An agentic data scientist to create high quality synthetic data Paper • 2606.25996 • Published 28 days ago • 18
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance Paper • 2606.19195 • Published Jun 17 • 141