DataFlex-RL: An Evaluation Platform for RLVR Data Policies Paper • 2609.06107 • Published 16 days ago • 162
DataFlex-RL: An Evaluation Platform for RLVR Data Policies Paper • 2609.06107 • Published 16 days ago • 162
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios Paper • 2607.14989 • Published Jul 16 • 1
WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing Paper • 2607.25765 • Published Jul 28
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Paper • 2607.20465 • Published May 19 • 56
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators Paper • 2607.20465 • Published May 19 • 56
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Paper • 2605.09635 • Published Jul 23 • 63
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Paper • 2607.16617 • Published Jul 18 • 99
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning Paper • 2605.22012 • Published May 21 • 45
TraceAV-Bench: Benchmarking Multi-Hop Trajectory Reasoning over Long Audio-Visual Videos Paper • 2605.07593 • Published May 8 • 1
Towards Next-Generation LLM Training: From the Data-Centric Perspective Paper • 2603.14712 • Published Mar 16
OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Paper • 2604.04707 • Published Apr 6 • 200
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs Paper • 2605.09635 • Published Jul 23 • 63
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines Paper • 2607.16617 • Published Jul 18 • 99
PAS: Data-Efficient Plug-and-Play Prompt Augmentation System Paper • 2407.06027 • Published Jul 8, 2024 • 10
KeyVideoLLM: Towards Large-scale Video Keyframe Selection Paper • 2407.03104 • Published Jul 3, 2024 • 1
Synth-Empathy: Towards High-Quality Synthetic Empathy Data Paper • 2407.21669 • Published Jul 31, 2024
CFBench: A Comprehensive Constraints-Following Benchmark for LLMs Paper • 2408.01122 • Published May 5, 2025
MathScape: Evaluating MLLMs in multimodal Math Scenarios through a Hierarchical Benchmark Paper • 2408.07543 • Published Aug 14, 2024