view article Article Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers +1 tomaarsen, NohTow, raphaelsty • 8 days ago • 93
view article Article Lattice: an 8 MB static retriever that embeds Wikipedia in 7 minutes erikkaum • 18 days ago • 19
view article Article NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval nvidia • Jul 16 • 59
ShadowPEFT: Shadow Network for Parameter-Efficient Fine-Tuning Paper • 2604.19254 • Published Apr 21 • 30
KoViDoRe Benchmark Collection Korean Visual Document Retrieval Benchmark • 10 items • Updated 6 days ago • 6
view article Article Nano-BEIR: A Multilingual Information Retrieval Benchmark with Quality-Enhanced Queries sionic-ai • Dec 22, 2025 • 12
Fantastic (small) Retrievers and How to Train Them: mxbai-edge-colbert-v0 Tech Report Paper • 2510.14880 • Published Oct 16, 2025 • 20
Embeddings datasets ⚡️ Collection This collection gather datasets for embeddings pre-training and fine-tuning. • 19 items • Updated 26 days ago • 5
NanoBEIR-fr 🍺 Collection French translation of zeta-alpha-ai's NanoBEIR collection • 14 items • Updated Mar 2 • 2
Fixing Data That Hurts Performance: Cascading LLMs to Relabel Hard Negatives for Robust Information Retrieval Paper • 2505.16967 • Published May 22, 2025 • 24
RLHN Datasets Collection RLHN: Cleaned Training Datasets with False Negatives Identified & Relabeled as ground truth. • 5 items • Updated May 23, 2025 • 4
NanoBEIR 🍺 Collection A collection of smaller versions of BEIR datasets with 50 queries and up to 10K documents each. • 13 items • Updated Sep 11, 2024 • 27
Simple linear attention language models balance the recall-throughput tradeoff Paper • 2402.18668 • Published Feb 28, 2024 • 20