AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis Paper • 2607.28618 • Published 2 days ago • 286
PhiZero: A World Model Built Around Physical Language Paper • 2607.28624 • Published 2 days ago • 152
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System Paper • 2607.27380 • Published 3 days ago • 63
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 2 days ago • 275
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published 3 days ago • 125
Inflect v2 Collection Complete local text-to-waveform speech models at 3.96M and 9.36M parameters, with official PyTorch and ONNX Runtime releases. • 5 items • Updated 6 days ago • 9
Kimi-VL-A3B Collection Moonshot's efficient MoE VLMs, exceptional on agent, long-context, and thinking • 6 items • Updated Mar 2 • 84
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published 5 days ago • 27
Parallel Decoding Distillation for Fast Image and Video Generation Paper • 2607.26004 • Published 4 days ago • 12
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Paper • 2607.21553 • Published 9 days ago • 39
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD Paper • 2607.20145 • Published 10 days ago • 72
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published 16 days ago • 170
Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation Paper • 2607.09581 • Published 22 days ago • 6
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Paper • 2607.13125 • Published 14 days ago • 138