Multi-hop video questions for RLVR, and the 8B model trained on them with Second-Wave Exploration.
Nguyen Quang Trung
ngqtrung
AI & ML interests
None yet
Recent Activity
updated a collection 3 days ago
Video HopChain updated a collection 3 days ago
Video HopChain updated a model 3 days ago
ngqtrung/video-hopchain-8bOrganizations
Qwen3-VL-8B video RLVR - GRPO checkpoint ladder
Every published checkpoint of the 8B video-QA GRPO campaign, ranked by core-3 mean_accuracy (5,645 rows). Untrained base = 0.4426.
Qwen3-VL-8B RLVR — Datasets (v1)
Curated SFT + GRPO RL datasets (video MC-QA, OMR math-image, OpenMMReasoner-RL, Vero) for Qwen3-VL-8B post-training.
VMAR — Train-Ready (SFT + RL + Eval)
Train-ready VMAR datasets: teacher-distilled SFT corpus, RL prompt set, and the curated in-loop eval benchmark.
Video RLVR — final training data
The datasets behind our Qwen3-VL-8B video RLVR runs: the base 24f/100k mixture and the HopChain v5 multi-hop corpora.
Qwen3-VL-8B RLVR — Models (v1)
Qwen3-VL-8B GRPO RLVR checkpoints from a token-dropout exploration study. OMR ppexplore=winner (0.714); video ~0.485 dead-heat.
VMAR — Raw & Source
Raw multi-style distilled traces, the pre-distillation template seed, and the never-trained real-audio eval.
videorl
Video HopChain
Multi-hop video questions for RLVR, and the 8B model trained on them with Second-Wave Exploration.
Video RLVR — final training data
The datasets behind our Qwen3-VL-8B video RLVR runs: the base 24f/100k mixture and the HopChain v5 multi-hop corpora.
Qwen3-VL-8B video RLVR - GRPO checkpoint ladder
Every published checkpoint of the 8B video-QA GRPO campaign, ranked by core-3 mean_accuracy (5,645 rows). Untrained base = 0.4426.
Qwen3-VL-8B RLVR — Models (v1)
Qwen3-VL-8B GRPO RLVR checkpoints from a token-dropout exploration study. OMR ppexplore=winner (0.714); video ~0.485 dead-heat.
Qwen3-VL-8B RLVR — Datasets (v1)
Curated SFT + GRPO RL datasets (video MC-QA, OMR math-image, OpenMMReasoner-RL, Vero) for Qwen3-VL-8B post-training.
VMAR — Raw & Source
Raw multi-style distilled traces, the pre-distillation template seed, and the never-trained real-audio eval.
VMAR — Train-Ready (SFT + RL + Eval)
Train-ready VMAR datasets: teacher-distilled SFT corpus, RL prompt set, and the curated in-loop eval benchmark.
videorl