Multimodal Datasets liuhaotian/LLaVA-CC3M-Pretrain-595K Preview • Updated Jul 6, 2023 • 563 • 180 lmms-lab-encoder/COCO-Caption Viewer • Updated Mar 8, 2024 • 81.3k • 20.3k • 15
Multi-turn VQA alexshengzhili/SciGraphQA-295K-train Viewer • Updated Aug 8, 2023 • 296k • 182 • 12 HuggingFaceM4/VisDial Viewer • Updated Jun 9, 2023 • 133k • 490 • 1 jxu124/visdial Viewer • Updated May 20, 2023 • 133k • 389 • 1 HuggingFaceFV/finevideo Viewer • Updated Apr 30 • 39.5k • 16.2k • 375
Multi-turn VQA alexshengzhili/SciGraphQA-295K-train Viewer • Updated Aug 8, 2023 • 296k • 182 • 12 HuggingFaceM4/VisDial Viewer • Updated Jun 9, 2023 • 133k • 490 • 1 jxu124/visdial Viewer • Updated May 20, 2023 • 133k • 389 • 1 HuggingFaceFV/finevideo Viewer • Updated Apr 30 • 39.5k • 16.2k • 375
Multimodal Datasets liuhaotian/LLaVA-CC3M-Pretrain-595K Preview • Updated Jul 6, 2023 • 563 • 180 lmms-lab-encoder/COCO-Caption Viewer • Updated Mar 8, 2024 • 81.3k • 20.3k • 15