multilingual-e5-large-swe-16384
This model is a 42.73% smaller version of intfloat/multilingual-e5-large optimized for Swedish language via vocabulary size reduction using the trimming method.
This trimmed model should perform similarly to the original model with only 16,384 tokens and a much smaller memory footprint. However, it may not perform well for other languages as tokens not commonly used in the selected languages were removed from the vocabulary.
Model Statistics
| Metric |
Original |
Trimmed |
Reduction |
| Vocabulary size |
250,037 tokens |
16,384 tokens |
93.44% |
| Model size |
559,890,432 params |
320,665,600 params |
42.73% |

Mining Dataset Statistics
Usage
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("alphaedge-ai/multilingual-e5-large-swe-16384")
query = "My query in Swedish"
documents = [
"Chunk in Swedish",
"Chunk in Swedish",
"Chunk in Swedish",
]
query_embeddings = model.encode_query(query)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
Citations
Multilingual E5
@article{wang2024multilingual,
title={Multilingual E5 Text Embeddings: A Technical Report},
author={Wang, Liang and Yang, Nan and Huang, Xiaolong and Yang, Linjun and Majumder, Rangan and Wei, Furu},
journal={arXiv preprint arXiv:2402.05672},
year={2024}
}
Trimming blog post
@misc{hf_blogpost_trimming,
title={Introduction to Trimming},
author={Loïck BOURDOIS and Tom AARSEN and Bram VANROY and Christopher AKIKI and Woojun JUNG and Manuel ROMERO and Prithiv SAKTHI},
year={2026},
url={https://huggingface.co/blog/lbourdois/introduction-to-trimming},
}