AI & ML interests

None defined yet.

Recent Activity

Articles

Organization Card

LILT

We build the multilingual layer for English-first AI. Custom evals, benchmarks, and RL environments across 200+ languages.

Most agent and coding benchmarks ship in English. We build the audited non-English counterparts — and the multilingual environments models train on — so labs and enterprises can measure and improve what their models actually do in the languages their users speak.

Why we publish here

Open releases make it easier for the community to stress-test our work, reproduce our scores, and extend our benchmarks to new languages. Every artifact is paired with a paper, a scoring script, and explicit limitations.

What you'll find here

  • Benchmarks & datasets — multilingual evaluations across coding, agents, tool use, long context, instruction following, and domain QA. Audited splits across our priority languages, scalable to 200+.
  • RL environments — multilingual training environments for agentic and tool-using models, with reproducible scoring.
  • Leaderboards & scoring — Gradio Spaces with reproducible submission flows.
  • Baselines — frontier-model scores published with exact prompts, decoding params, and dated snapshots.
  • Papers — methodology, audit workflow, and findings.

Currently featured

📌 Terminal-Bench-LILT — multilingual coding-agent benchmark. 300 tasks across AR / CS / DE / ES / HI / JA / KO / SR / TR / ZH, authored by native-speaking engineers rather than translated, verified deterministically under Harbor. Best frontier model resolves 63.1%; a subset is solved by none of the seven models evaluated. Replacing native-language instructions with English moves pass rates ≤7pp — comprehension is not the bottleneck. Dataset, paper, and code linked in the pinned collection.

📌 GAIA-v2-LILT — multilingual agent benchmark across AR / DE / HI / KO / PT-BR. +20.7pp average gain post human-audit on frontier agents. Dataset, paper, and leaderboard linked in the pinned collection.

Links

Citation

If you use one of our datasets or benchmarks, please cite the corresponding paper linked on each dataset card.

models 0

None public yet

datasets 0

None public yet