LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V7-GGUF Image-Text-to-Text • 35B • Updated about 14 hours ago • 333k • 429
When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation Paper • 2608.03632 • Published 4 days ago • 21
ST-WAM: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts Paper • 2607.28993 • Published 8 days ago • 6
CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Paper • 2608.02589 • Published 5 days ago • 25
Progressive Agent Skill Generation via Reinforcement Learning Paper • 2608.01678 • Published 5 days ago • 58
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published 13 days ago • 100