MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions Paper • 2507.10859 • Published Jul 14, 2025
Why Stop at Words? Unveiling the Bigger Picture through Line-Level OCR Paper • 2508.21693 • Published Aug 29, 2025
TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification Paper • 2501.00398 • Published Apr 3, 2025
Aurelia: Test-time Reasoning Distillation in Audio-Visual LLMs Paper • 2503.23219 • Published Mar 29, 2025
Gencho: Room Impulse Response Generation from Reverberant Speech and Text via Diffusion Transformers Paper • 2602.09233 • Published Feb 9
Dynamic Multi-Byte Prediction With Hierarchical Language Models Paper • 2608.15454 • Published about 1 month ago • 19
Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos Paper • 2607.16107 • Published Jul 17 • 12
Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models Paper • 2603.29263 • Published Mar 31
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding Paper • 2508.12687 • Published Aug 23, 2025
Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music Paper • 2604.10905 • Published Apr 13 • 29