SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published 23 days ago • 71
Multi-scale Predictive Representations for Goal-conditioned Reinforcement Learning Paper • 2605.09364 • Published May 10
VectorGym: A Multitask Benchmark for SVG Code Generation, Sketching, and Editing Paper • 2603.29852 • Published Feb 22 • 6
ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval Paper • 2511.00903 • Published Nov 2, 2025
BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning Paper • 2508.09804 • Published Aug 13, 2025
Do Enterprise Systems Need Learned World Models? The Importance of Context to Infer Dynamics Paper • 2605.12178 • Published May 12 • 65
BigDocs: An Open and Permissively-Licensed Dataset for Training Multimodal Models on Document and Code Tasks Paper • 2412.04626 • Published Dec 5, 2024 • 15
Chitrarth: Bridging Vision and Language for a Billion People Paper • 2502.15392 • Published Feb 21, 2025
LitLLMs, LLMs for Literature Review: Are we there yet? Paper • 2412.15249 • Published Dec 15, 2024 • 2
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs Paper • 2511.04727 • Published Nov 6, 2025
VoiceAgentBench: Are Voice Assistants ready for agentic tasks? Paper • 2510.07978 • Published Oct 9, 2025
Seeing Straight: Document Orientation Detection for Efficient OCR Paper • 2511.04161 • Published Nov 6, 2025
Designing Production-Scale OCR for India: Multilingual and Domain-Specific Systems Paper • 2602.16430 • Published Feb 18 • 1
Chitranuvad: Adapting Multi-Lingual LLMs for Multimodal Translation Paper • 2502.20420 • Published Feb 27, 2025
CUA-Suite: Massive Human-annotated Video Demonstrations for Computer-Use Agents Paper • 2603.24440 • Published Mar 25 • 99