wangzixuan
wangzx1210
AI & ML interests
None yet
Recent Activity
upvoted a paper 3 days ago
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning upvoted a paper 24 days ago
TTPO: Test-Time Policy Optimization liked a dataset 26 days ago
xiamoent/Agent-G2-ALFWorld-Webshop-sft-data