Instructions to use TuTuCSF/WeMM-Embedding-2B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use TuTuCSF/WeMM-Embedding-2B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use TuTuCSF/WeMM-Embedding-2B-GGUF with Ollama:
ollama run hf.co/TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use TuTuCSF/WeMM-Embedding-2B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use TuTuCSF/WeMM-Embedding-2B-GGUF with Docker Model Runner:
docker model run hf.co/TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M
- Lemonade
How to use TuTuCSF/WeMM-Embedding-2B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.WeMM-Embedding-2B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use TuTuCSF/WeMM-Embedding-2B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use TuTuCSF/WeMM-Embedding-2B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "TuTuCSF/WeMM-Embedding-2B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- WeMM-Embedding-2B-GGUF
- Benchmark result at a glance
- Full evaluated task matrix
- Comprehensive comparison: WeMM-Embedding vs Qwen3-VL-Embedding
- Evaluation environment
- Evaluation methodology
- Important tokenizer / prompt requirements
- llama.cpp server
- Text embedding example
- Visual inputs
- Benchmark artifacts
- Available GGUF files
- Limitations
- Result summary
- Upstream
- License
- Benchmark result at a glance
- WeMM-Embedding-2B-GGUF
156 MMEB-v3 tasks · ~54 hours local evaluation · Q4_K_M + Q8_0 mmproj
GGUF conversion of tencent/WeMM-Embedding-2B, focused on local multimodal embedding inference with llama.cpp.
Recommended validated pair
- Main model:
WeMM-Embedding-2B-Q4_K_M.gguf - Visual projector:
mmproj-WeMM-Embedding-2B-Q8_0.gguf - Output dimension: 2048
- Pooling: last token
- Output normalization: L2
- Matryoshka dimensions:
64, 128, 256, 512, 1024, 2048
This repository is not only a conversion. The Q4_K_M + Q8_0 projector pair was evaluated end-to-end on 156 MMEB-v3 tasks, requiring approximately 54 hours of local evaluation.
Benchmark result at a glance
Evaluation date: 2026-09-04
| Benchmark group | Metric | Tasks | Tencent official WeMM-2B | This GGUF | Δ | Score retained |
|---|---|---|---|---|---|---|
| MMEB-v2 Image | Hit@1 | 36 | 79.60 | 79.18 | -0.42 | 99.47% |
| MMEB-v2 VisDoc | NDCG@5 | 27 | 80.70 | 78.07 | -2.63 | 96.75% |
| MMEB-v3 Text | NDCG@5 | 53 | 45.30 | 43.65 | -1.65 | 96.37% |
| MMEB-v3 MCMR | Hit@1 | 1 | 42.50 | 41.01 | -1.49 | 96.49% |
The published Tencent scores above are the native WeMM-Embedding-2B aggregate results from the WeMM technical report. They are not a BF16 rerun on the same RTX 5080 machine, so the differences should be read as practical agreement with the published native baseline rather than a perfectly controlled quantization-only ablation.
The important result is that the Q4_K_M main model with a Q8_0 visual projector preserves roughly 96–99% of the published native score across the four directly comparable benchmark groups.
Full evaluated task matrix
This run covers 156 / 190 MMEB-v3 tasks:
| Modality / group | Tasks | Metric | Local result |
|---|---|---|---|
| Image | 36 | Hit@1 | 79.18 |
| VisDoc | 27 | NDCG@5 | 78.07 |
| Text | 53 | NDCG@5 | 43.65 |
| Tool | 35 | Hit@1 | 49.00 |
| Memory | 4 | Hit@1 | 35.67 |
| MCMR | 1 | Hit@1 | 41.01 |
| Tool + Memory, no-GUI Agent subset | 39 | Hit@1 | 47.63 |
Not evaluated in this run:
- 15 Video tasks
- 8 GUI tasks
- 11 Audio tasks
The original WeMM model supports video, but video was intentionally excluded from this GGUF benchmark. Audio is not supported by WeMM-Embedding. Because the 8 GUI tasks were excluded, the local 39-task Tool + Memory result must not be compared directly with Tencent's published 47-task Agent aggregate.
Comprehensive comparison: WeMM-Embedding vs Qwen3-VL-Embedding
The tables below separate published native-model results from this repository's local GGUF result. This matters because comparing a local Q4_K_M run directly with a native BF16/FP16 leaderboard submission is useful as a practical reference, but it is not a controlled quantization-only ablation.
Official MMEB-v2 family comparison
Tencent reports the following results on the 78-task MMEB-v2 benchmark.
| Model | Size | AVG | Image Hit@1 | Video Hit@1 | VisDoc NDCG@5 |
|---|---|---|---|---|---|
| Qwen3-VL-Embedding | 2B | 73.2 | 75.0 | 61.9 | 79.2 |
| WeMM-Embedding | 2B | 77.9 | 79.6 | 70.8 | 80.7 |
| WeMM-Embedding | 4B | 79.2 | 80.8 | 72.1 | 82.0 |
| Qwen3-VL-Embedding | 8B | 77.8 | 80.1 | 67.1 | 82.4 |
| WeMM-Embedding | 9B | 80.6 | 81.9 | 74.3 | 83.3 |
At roughly the same 2B scale, official WeMM-Embedding-2B leads Qwen3-VL-Embedding-2B by:
| MMEB-v2 metric | WeMM-2B advantage |
|---|---|
| AVG | +4.7 |
| Image | +4.6 |
| Video | +8.9 |
| VisDoc | +1.5 |
At the larger scale, WeMM-Embedding-9B compared with Qwen3-VL-Embedding-8B:
| MMEB-v2 metric | WeMM-9B advantage |
|---|---|
| AVG | +2.8 |
| Image | +1.8 |
| Video | +7.2 |
| VisDoc | +0.9 |
The largest consistent WeMM advantage on MMEB-v2 is video retrieval, while the gap on visual-document retrieval is much smaller.
Official MMEB-v3 family comparison
Tencent reports the following results on all 190 MMEB-v3 tasks.
| Model | Size | V3-All | Text NDCG@5 | Agent Hit@1 | MCMR Hit@1 | Audio |
|---|---|---|---|---|---|---|
| Qwen3-VL-Embedding | 2B | 50.9 | 39.2 | 39.3 | 42.0 | 0.0 |
| WeMM-Embedding | 2B | 56.0 | 45.3 | 45.1 | 42.5 | 0.0 |
| WeMM-Embedding | 4B | 58.2 | 47.9 | 49.0 | 41.9 | 0.0 |
| Qwen3-VL-Embedding | 8B | 53.5 | 42.5 | 38.4 | 38.0 | 0.0 |
| WeMM-Embedding | 9B | 59.5 | 48.8 | 51.0 | 49.3 | 0.0 |
At the 2B scale, official WeMM-Embedding-2B leads Qwen3-VL-Embedding-2B by:
| MMEB-v3 metric | WeMM-2B advantage |
|---|---|
| V3-All | +5.1 |
| Text | +6.1 |
| Agent | +5.8 |
| MCMR | +0.5 |
At the larger scale, WeMM-Embedding-9B compared with Qwen3-VL-Embedding-8B:
| MMEB-v3 metric | WeMM-9B advantage |
|---|---|
| V3-All | +6.0 |
| Text | +6.3 |
| Agent | +12.6 |
| MCMR | +11.3 |
This GGUF vs official WeMM-Embedding-2B
This is the most relevant table for evaluating the quality of the conversion itself.
| Benchmark group | Official WeMM-2B | This 2B Q4_K_M GGUF | Δ | Score retained |
|---|---|---|---|---|
| Image | 79.60 | 79.18 | -0.42 | 99.47% |
| VisDoc | 80.70 | 78.07 | -2.63 | 96.75% |
| Text | 45.30 | 43.65 | -1.65 | 96.37% |
| MCMR | 42.50 | 41.01 | -1.49 | 96.49% |
Across these four directly comparable published benchmark groups, the tested Q4_K_M + Q8_0 projector pair retains approximately 96–99% of the published native WeMM-Embedding-2B score.
This GGUF vs Qwen3-VL-Embedding-2B
This comparison answers a practical deployment question: how does the locally quantized WeMM-2B compare with the published native Qwen3-VL-Embedding-2B baseline?
| Benchmark group | Qwen3-VL-Embedding-2B native | WeMM-2B Q4_K_M GGUF | GGUF difference |
|---|---|---|---|
| Image | 75.00 | 79.18 | +4.18 |
| VisDoc | 79.20 | 78.07 | -1.13 |
| Text | 39.20 | 43.65 | +4.45 |
| MCMR | 42.00 | 41.01 | -0.99 |
Even after Q4_K_M quantization of the main model, this WeMM-2B GGUF remains clearly ahead of the published Qwen3-VL-Embedding-2B native baseline on Image and Text, while trailing by about one point on VisDoc and MCMR.
Cross-size reference: this 2B Q4 GGUF vs Qwen3-VL-Embedding-8B
This is not an apples-to-apples comparison, but it is useful for understanding the deployment efficiency of the 2B GGUF.
| Benchmark group | Qwen3-VL-Embedding-8B native | WeMM-2B Q4_K_M GGUF | GGUF difference |
|---|---|---|---|
| Image | 80.10 | 79.18 | -0.92 |
| VisDoc | 82.40 | 78.07 | -4.33 |
| Text | 42.50 | 43.65 | +1.15 |
| MCMR | 38.00 | 41.01 | +3.01 |
Notably, the tested 2B Q4_K_M WeMM GGUF exceeds the published Qwen3-VL-Embedding-8B score on Text and MCMR, while coming within one point on Image. Qwen3-VL-Embedding-8B remains substantially stronger on VisDoc.
What cannot be compared directly
This local run intentionally excludes Video and GUI, so several official aggregate columns cannot be reconstructed faithfully:
- MMEB-v2 AVG cannot be computed because the 15 Video tasks were not evaluated.
- MMEB-v3 V3-All cannot be computed because Video, GUI and Audio are not all present.
- Tencent's official Agent score covers 47 tasks. This local run contains 35 Tool + 4 Memory tasks but excludes 8 GUI tasks.
- The local Tool + Memory average (47.63 Hit@1 over 39 tasks) is therefore a useful no-GUI subset result, but it must not be presented as the official 47-task Agent score.
In short, use Image / VisDoc / Text / MCMR for direct published-score comparisons, and treat Tool / Memory as additional local evidence.
Evaluation environment
| Item | Configuration |
|---|---|
| GPU | NVIDIA GeForce RTX 5080 16 GB |
| CPU | AMD Ryzen 7 9800X3D |
| System RAM | 32 GB |
| OS | Windows 11 |
| Inference backend | llama.cpp llama-server |
| Main GGUF | WeMM-Embedding-2B-Q4_K_M.gguf |
| mmproj | mmproj-WeMM-Embedding-2B-Q8_0.gguf |
| Image / VisDoc / Tool / Memory context | 32768, KV F16 |
| Text context | 262144, KV Q8_0 |
| Pooling | last |
| Normalization | L2 |
| Parallel slots | -np 1 |
| Approx. evaluation time | 54 hours |
The Text benchmark intentionally uses 262K context. This matters for LongEmbed tasks and avoids turning long-document truncation into a fake "quantization loss".
The final result file records the two context profiles explicitly:
image/visdoc/tool/memory = ctx32768 kvf16text = ctx262144 kvq8_0
Evaluation methodology
The benchmark follows Tencent's released mmeb_v3_eval pipeline, which is based on the official VLM2Vec MMEB-v3 evaluator.
Pinned references used by this evaluation work:
- Tencent WeMM-Embedding evaluation code commit:
9ed7e2d7914cd67a031c3e2a4fef3faaac314721 - VLM2Vec evaluator commit:
2638a8413fda4b98668a29ea763b4898814bfea7 - MMEB-V3 data revision:
4a5560b2b64384204b6fea8a82ea986eba51f5aa
The GGUF backend replaces only model inference. Dataset parsing, instructions, query/candidate construction and RankingMetrics remain aligned with the released evaluator.
The tested modalities are:
- Image
- VisDoc
- Text
- Tool
- Memory
- MCMR
Dataset-defined sample caps are preserved where present. Other tasks use their full query/candidate sets.
Important tokenizer / prompt requirements
WeMM embedding extraction uses the final <embedding> token with last-token pooling.
For text input, the validated raw prompt surface is:
<|im_start|>user
YOUR_TEXT<|im_end|><embedding>
There is no newline between <|im_end|> and <embedding>.
The tokenizer must leave <embedding> as the unique final token. When using llama-server, disable automatic EOS insertion:
--override-kv tokenizer.ggml.add_eos_token=bool:false
The benchmark performs a /tokenize audit before evaluation and rejects a configuration where <embedding> is not the unique final token.
llama.cpp server
Validated 32K multimodal / general retrieval configuration:
llama-server \
-m WeMM-Embedding-2B-Q4_K_M.gguf \
--mmproj mmproj-WeMM-Embedding-2B-Q8_0.gguf \
--embedding \
--pooling last \
--embd-normalize 2 \
--override-kv tokenizer.ggml.add_eos_token=bool:false \
-ngl 99 \
-c 32768 \
-np 1 \
-b 512 \
-ub 512 \
--cache-type-k f16 \
--cache-type-v f16 \
--flash-attn on \
--image-min-tokens 64 \
--image-max-tokens 8192 \
--no-cache-prompt \
--no-webui \
--host 127.0.0.1 \
--port 8080
For the 262K Text benchmark profile:
llama-server \
-m WeMM-Embedding-2B-Q4_K_M.gguf \
--embedding \
--pooling last \
--embd-normalize 2 \
--override-kv tokenizer.ggml.add_eos_token=bool:false \
-ngl 99 \
-c 262144 \
-np 1 \
-b 512 \
-ub 512 \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--flash-attn on \
--no-cache-prompt \
--no-webui \
--host 127.0.0.1 \
--port 8080
Text embedding example
With the server running:
curl http://127.0.0.1:8080/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"input":"<|im_start|>user\nRepresent the meaning of this sentence.<|im_end|><embedding>","encoding_format":"float"}'
The returned vector should have 2048 dimensions and be L2 normalized.
Visual inputs
For image / visual-document embeddings, use the paired:
mmproj-WeMM-Embedding-2B-Q8_0.gguf
The benchmark uses WeMM-compatible image/text ordering, an image budget of 64–8192 visual tokens, last-token pooling and L2 normalization.
Raw multimodal HTTP request schemas can change between llama.cpp builds. If reproducing the benchmark, keep the exact prompt ordering and verify that <embedding> remains the final token after tokenization.
Benchmark artifacts
For reproducibility, this repository should include the final benchmark outputs:
benchmarks/
├── 2B_scores_summary.csv
└── 2B_scores_detail.txt
2B_scores_summary.csv contains one row per evaluated dataset.
2B_scores_detail.txt contains all reported ranking metrics for all 156 datasets.
Available GGUF files
The 2B conversion set uses these filenames:
WeMM-Embedding-2B-Q4_K_M.gguf
WeMM-Embedding-2B-BF16.gguf
mmproj-WeMM-Embedding-2B-Q8_0.gguf
mmproj-WeMM-Embedding-2B-BF16.gguf
For local deployment, the benchmarked pair is:
WeMM-Embedding-2B-Q4_K_M.gguf
+
mmproj-WeMM-Embedding-2B-Q8_0.gguf
Limitations
- This benchmark does not include Video, GUI or Audio.
- Tool + Memory is only a 39-task no-GUI subset of the official Agent benchmark.
- The published Tencent native scores are used as the reference baseline; a full native BF16 rerun under identical local conditions was not performed.
llama.cppmultimodal APIs evolve. Revalidate tokenizer ending, image handling and embedding dimensions when changing builds.- Long-text Text evaluation requires substantially more context memory than ordinary image / document retrieval.
Result summary
The practical conclusion from this run:
WeMM-Embedding-2B-Q4_K_M.gguf+mmproj-WeMM-Embedding-2B-Q8_0.ggufis a high-fidelity local quantization of WeMM-Embedding-2B for the evaluated image, visual-document, text, tool and memory retrieval workloads.
Across the four directly comparable published benchmark groups, the GGUF pair retains approximately 96–99% of the published native WeMM-Embedding-2B score, while reducing the main model to Q4_K_M for local deployment.
Upstream
- Original model:
tencent/WeMM-Embedding-2B - Project:
Tencent/WeMM-Embedding - Technical report: WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
- arXiv:
2608.24053
License
The upstream WeMM-Embedding model is released under the Apache License 2.0. This GGUF repository follows the upstream licensing terms. Third-party evaluation datasets and components retain their original licenses.
WeMM-Embedding-2B-GGUF
tencent/WeMM-Embedding-2B 的 GGUF 转换版本,面向使用 llama.cpp 的本地多模态 Embedding 推理。
推荐且已经完整验证的组合
- 主模型:
WeMM-Embedding-2B-Q4_K_M.gguf - 视觉投影器:
mmproj-WeMM-Embedding-2B-Q8_0.gguf - 输出维度:2048
- Pooling:最后一个 token(last token)
- 输出归一化:L2
- Matryoshka 维度:
64, 128, 256, 512, 1024, 2048
这个仓库不只是一次 GGUF 转换。
Q4_K_M + Q8_0 mmproj组合已经完成 156 个 MMEB-v3 任务的端到端实测,本地累计评测时间约 54 小时。
核心评测结果
评测日期:2026-09-04
| Benchmark 组 | 指标 | 任务数 | 腾讯官方 WeMM-2B | 本仓库 GGUF | Δ | 分数保留率 |
|---|---|---|---|---|---|---|
| MMEB-v2 Image | Hit@1 | 36 | 79.60 | 79.18 | -0.42 | 99.47% |
| MMEB-v2 VisDoc | NDCG@5 | 27 | 80.70 | 78.07 | -2.63 | 96.75% |
| MMEB-v3 Text | NDCG@5 | 53 | 45.30 | 43.65 | -1.65 | 96.37% |
| MMEB-v3 MCMR | Hit@1 | 1 | 42.50 | 41.01 | -1.49 | 96.49% |
上表中的腾讯分数来自 WeMM 技术报告公布的原生模型聚合成绩。它们不是在同一张 RTX 5080 上重新跑出来的 BF16 对照,因此这些差值应该理解为“本地 GGUF 结果与官方原生基线的一致程度”,而不是严格控制变量下的纯量化损失。
最重要的结论是:
Q4_K_M主模型配合Q8_0视觉投影器,在四个可直接比较的公开 benchmark 组上保留了大约 96%–99% 的官方原生分数。
完整评测任务矩阵
本次评测覆盖 156 / 190 个 MMEB-v3 任务:
| 模态 / 分组 | 任务数 | 指标 | 本地结果 |
|---|---|---|---|
| Image | 36 | Hit@1 | 79.18 |
| VisDoc | 27 | NDCG@5 | 78.07 |
| Text | 53 | NDCG@5 | 43.65 |
| Tool | 35 | Hit@1 | 49.00 |
| Memory | 4 | Hit@1 | 35.67 |
| MCMR | 1 | Hit@1 | 41.01 |
| Tool + Memory(不含 GUI 的 Agent 子集) | 39 | Hit@1 | 47.63 |
本轮未测试:
- 15 个 Video 任务
- 8 个 GUI 任务
- 11 个 Audio 任务
原始 WeMM 模型支持视频,但本次 GGUF benchmark 有意排除了 Video。WeMM-Embedding 不支持 Audio。
由于排除了 8 个 GUI 任务,本地 39 项 Tool + Memory 结果不能直接与腾讯官方公布的 47 项 Agent 聚合分数比较。
全面对比:WeMM-Embedding vs Qwen3-VL-Embedding
下面将官方原生模型结果与本仓库的本地 GGUF 实测结果分开列出。这样做很重要,因为拿本地 Q4_K_M 直接和原生 BF16/FP16 榜单提交比较,作为部署参考很有价值,但并不是严格意义上的纯量化消融实验。
官方 MMEB-v2 全系列对比
腾讯在 78 项 MMEB-v2 上公布的结果如下:
| 模型 | 尺寸 | AVG | Image Hit@1 | Video Hit@1 | VisDoc NDCG@5 |
|---|---|---|---|---|---|
| Qwen3-VL-Embedding | 2B | 73.2 | 75.0 | 61.9 | 79.2 |
| WeMM-Embedding | 2B | 77.9 | 79.6 | 70.8 | 80.7 |
| WeMM-Embedding | 4B | 79.2 | 80.8 | 72.1 | 82.0 |
| Qwen3-VL-Embedding | 8B | 77.8 | 80.1 | 67.1 | 82.4 |
| WeMM-Embedding | 9B | 80.6 | 81.9 | 74.3 | 83.3 |
同为约 2B 规模时,官方 WeMM-Embedding-2B 相比 Qwen3-VL-Embedding-2B:
| MMEB-v2 指标 | WeMM-2B 优势 |
|---|---|
| AVG | +4.7 |
| Image | +4.6 |
| Video | +8.9 |
| VisDoc | +1.5 |
大模型尺寸下,WeMM-Embedding-9B 相比 Qwen3-VL-Embedding-8B:
| MMEB-v2 指标 | WeMM-9B 优势 |
|---|---|
| AVG | +2.8 |
| Image | +1.8 |
| Video | +7.2 |
| VisDoc | +0.9 |
MMEB-v2 中 WeMM 最稳定、最明显的优势出现在视频检索;VisDoc 上的差距则小得多。
官方 MMEB-v3 全系列对比
腾讯在完整 190 项 MMEB-v3 上公布的结果:
| 模型 | 尺寸 | V3-All | Text NDCG@5 | Agent Hit@1 | MCMR Hit@1 | Audio |
|---|---|---|---|---|---|---|
| Qwen3-VL-Embedding | 2B | 50.9 | 39.2 | 39.3 | 42.0 | 0.0 |
| WeMM-Embedding | 2B | 56.0 | 45.3 | 45.1 | 42.5 | 0.0 |
| WeMM-Embedding | 4B | 58.2 | 47.9 | 49.0 | 41.9 | 0.0 |
| Qwen3-VL-Embedding | 8B | 53.5 | 42.5 | 38.4 | 38.0 | 0.0 |
| WeMM-Embedding | 9B | 59.5 | 48.8 | 51.0 | 49.3 | 0.0 |
同为 2B 规模时,官方 WeMM-Embedding-2B 相比 Qwen3-VL-Embedding-2B:
| MMEB-v3 指标 | WeMM-2B 优势 |
|---|---|
| V3-All | +5.1 |
| Text | +6.1 |
| Agent | +5.8 |
| MCMR | +0.5 |
大模型尺寸下,WeMM-Embedding-9B 相比 Qwen3-VL-Embedding-8B:
| MMEB-v3 指标 | WeMM-9B 优势 |
|---|---|
| V3-All | +6.0 |
| Text | +6.3 |
| Agent | +12.6 |
| MCMR | +11.3 |
本仓库 GGUF vs 官方 WeMM-Embedding-2B
这是判断本次 GGUF 转换质量最重要的一张表。
| Benchmark 组 | 官方 WeMM-2B | 本仓库 2B Q4_K_M GGUF | Δ | 分数保留率 |
|---|---|---|---|---|
| Image | 79.60 | 79.18 | -0.42 | 99.47% |
| VisDoc | 80.70 | 78.07 | -2.63 | 96.75% |
| Text | 45.30 | 43.65 | -1.65 | 96.37% |
| MCMR | 42.50 | 41.01 | -1.49 | 96.49% |
在四个可以直接和官方公开成绩比较的 benchmark 组上,Q4_K_M + Q8_0 mmproj 组合保留了约 96%–99% 的官方原生能力。
本仓库 GGUF vs Qwen3-VL-Embedding-2B
这张表更接近实际部署问题:已经量化到 Q4 的 WeMM-2B,与 Qwen3-VL-Embedding-2B 官方原生模型相比如何?
| Benchmark 组 | Qwen3-VL-Embedding-2B 原生 | WeMM-2B Q4_K_M GGUF | GGUF 差值 |
|---|---|---|---|
| Image | 75.00 | 79.18 | +4.18 |
| VisDoc | 79.20 | 78.07 | -1.13 |
| Text | 39.20 | 43.65 | +4.45 |
| MCMR | 42.00 | 41.01 | -0.99 |
即使主模型已经量化为 Q4_K_M,本仓库的 WeMM-2B 在 Image 和 Text 上仍明显高于官方 Qwen3-VL-Embedding-2B 原生基线;VisDoc 和 MCMR 则落后约 1 分。
跨尺寸参考:本仓库 2B Q4 vs Qwen3-VL-Embedding-8B
这不是严格同口径比较,但很适合衡量 2B GGUF 的部署效率。
| Benchmark 组 | Qwen3-VL-Embedding-8B 原生 | WeMM-2B Q4_K_M GGUF | GGUF 差值 |
|---|---|---|---|
| Image | 80.10 | 79.18 | -0.92 |
| VisDoc | 82.40 | 78.07 | -4.33 |
| Text | 42.50 | 43.65 | +1.15 |
| MCMR | 38.00 | 41.01 | +3.01 |
值得注意的是:
实测的 2B Q4_K_M WeMM GGUF 在 Text 和 MCMR 上超过了官方 Qwen3-VL-Embedding-8B,Image 也只差 0.92 分;Qwen3-VL-Embedding-8B 在 VisDoc 上则明显更强。
哪些数据不能直接比较
本地评测有意排除了 Video 与 GUI,因此无法严谨重建部分官方聚合指标:
- 因为未评测 15 个 Video 任务,所以不能重建 MMEB-v2 AVG。
- 因为没有完整覆盖 Video、GUI、Audio,所以不能计算官方口径的 MMEB-v3 V3-All。
- 腾讯官方 Agent 分数覆盖 47 项;本地只有 35 Tool + 4 Memory,没有 8 GUI。
- 因此本地
Tool + Memory = 47.63 Hit@1 / 39 tasks可以作为 no-GUI Agent 子集参考,但不能写成官方 47-task Agent 成绩。
因此,推荐使用 Image / VisDoc / Text / MCMR 做直接官方成绩对比,把 Tool / Memory 视为额外的本地验证证据。
评测环境
| 项目 | 配置 |
|---|---|
| GPU | NVIDIA GeForce RTX 5080 16 GB |
| CPU | AMD Ryzen 7 9800X3D |
| 系统内存 | 32 GB |
| 操作系统 | Windows 11 |
| 推理后端 | llama.cpp llama-server |
| 主 GGUF | WeMM-Embedding-2B-Q4_K_M.gguf |
| mmproj | mmproj-WeMM-Embedding-2B-Q8_0.gguf |
| Image / VisDoc / Tool / Memory context | 32768,KV F16 |
| Text context | 262144,KV Q8_0 |
| Pooling | last |
| Normalization | L2 |
| 并行槽位 | -np 1 |
| 累计评测时间 | 约 54 小时 |
Text benchmark 有意使用 262K context。这是 LongEmbed 等超长文本任务能够公平评测的关键,避免把长文截断造成的损失误判为量化损失。
最终结果文件明确记录了两套 context profile:
image/visdoc/tool/memory = ctx32768 kvf16text = ctx262144 kvq8_0
评测方法
评测沿用腾讯发布的 mmeb_v3_eval pipeline;它基于官方 VLM2Vec MMEB-v3 evaluator。
本轮评测使用的固定版本:
- Tencent WeMM-Embedding evaluation code:
9ed7e2d7914cd67a031c3e2a4fef3faaac314721 - VLM2Vec evaluator:
2638a8413fda4b98668a29ea763b4898814bfea7 - MMEB-V3 data revision:
4a5560b2b64384204b6fea8a82ea986eba51f5aa
GGUF backend 只替换模型推理部分。数据集解析、instructions、query/candidate 构造以及 RankingMetrics 均继续与腾讯发布的 evaluator 对齐。
已测试:
- Image
- VisDoc
- Text
- Tool
- Memory
- MCMR
数据集本身定义了 sample cap 的任务继续沿用该 cap;其他任务使用完整 query / candidate 集合。
tokenizer / prompt 的关键要求
WeMM 的 embedding 从最终 <embedding> token 提取,并使用 last-token pooling。
纯文本的验证 prompt 形式:
<|im_start|>user
YOUR_TEXT<|im_end|><embedding>
<|im_end|> 与 <embedding> 之间不能插入换行。
tokenizer 必须确保 <embedding> 是唯一的最后一个 token。使用 llama-server 时,需要关闭自动 EOS:
--override-kv tokenizer.ggml.add_eos_token=bool:false
benchmark 在正式评测前通过 /tokenize 检查这一点;如果 <embedding> 不是唯一的最终 token,评测会直接拒绝继续。
llama.cpp 服务配置
已经验证的 32K 多模态 / 通用检索配置:
llama-server \
-m WeMM-Embedding-2B-Q4_K_M.gguf \
--mmproj mmproj-WeMM-Embedding-2B-Q8_0.gguf \
--embedding \
--pooling last \
--embd-normalize 2 \
--override-kv tokenizer.ggml.add_eos_token=bool:false \
-ngl 99 \
-c 32768 \
-np 1 \
-b 512 \
-ub 512 \
--cache-type-k f16 \
--cache-type-v f16 \
--flash-attn on \
--image-min-tokens 64 \
--image-max-tokens 8192 \
--no-cache-prompt \
--no-webui \
--host 127.0.0.1 \
--port 8080
用于 262K Text benchmark 的配置:
llama-server \
-m WeMM-Embedding-2B-Q4_K_M.gguf \
--embedding \
--pooling last \
--embd-normalize 2 \
--override-kv tokenizer.ggml.add_eos_token=bool:false \
-ngl 99 \
-c 262144 \
-np 1 \
-b 512 \
-ub 512 \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--flash-attn on \
--no-cache-prompt \
--no-webui \
--host 127.0.0.1 \
--port 8080
文本 Embedding 示例
服务器启动后:
curl http://127.0.0.1:8080/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"input":"<|im_start|>user\nRepresent the meaning of this sentence.<|im_end|><embedding>","encoding_format":"float"}'
返回向量应为 2048 维,并已经完成 L2 normalization。
图像输入
Image / VisDoc embedding 需要配套:
mmproj-WeMM-Embedding-2B-Q8_0.gguf
本次 benchmark 使用与 WeMM 对齐的图像 / 文本顺序、64–8192 visual tokens 图像预算、last-token pooling 与 L2 normalization。
llama.cpp 的多模态 HTTP schema 可能随版本变化。复现 benchmark 时,应优先保证 prompt 顺序完全一致,并重新确认 <embedding> 在 tokenization 后仍是最后一个 token。
Benchmark 原始结果
建议仓库同时保留最终结果文件:
benchmarks/
├── 2B_scores_summary.csv
└── 2B_scores_detail.txt
2B_scores_summary.csv:每个评测数据集一行。
2B_scores_detail.txt:包含全部 156 个数据集的所有 ranking metrics。
GGUF 文件
2B 转换文件使用:
WeMM-Embedding-2B-Q4_K_M.gguf
WeMM-Embedding-2B-BF16.gguf
mmproj-WeMM-Embedding-2B-Q8_0.gguf
mmproj-WeMM-Embedding-2B-BF16.gguf
本轮完整 benchmark 使用的是:
WeMM-Embedding-2B-Q4_K_M.gguf
+
mmproj-WeMM-Embedding-2B-Q8_0.gguf
限制
- 本轮 benchmark 不包含 Video、GUI、Audio。
- Tool + Memory 只是官方 Agent benchmark 的 39 项 no-GUI 子集。
- 官方腾讯原生成绩作为参考基线,本地没有再用完全相同环境完整重跑一遍 BF16。
llama.cpp多模态 API 会持续变化。更换 build 后,应重新验证 tokenizer 尾部、图像处理和 embedding 维度。- 262K Text 长上下文评测的内存与 KV cache 成本明显高于普通图像和文档检索。
结论
本轮测试支持以下实际结论:
WeMM-Embedding-2B-Q4_K_M.gguf+mmproj-WeMM-Embedding-2B-Q8_0.gguf是 WeMM-Embedding-2B 在已测试 Image、VisDoc、Text、Tool、Memory retrieval 工作负载上的高保真本地量化版本。
在四个能够与腾讯官方公开成绩直接比较的 benchmark 组上,这套 GGUF 组合保留了大约 96%–99% 的官方原生分数,同时把主模型压缩到了 Q4_K_M,更适合本地部署。
上游项目
- 原始模型:
tencent/WeMM-Embedding-2B - 项目:
Tencent/WeMM-Embedding - 技术报告:WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
- arXiv:
2608.24053
License
上游 WeMM-Embedding 模型采用 Apache License 2.0。本 GGUF 仓库遵循上游许可条款。第三方评测数据集和组件继续遵循各自原始 License。
- Downloads last month
- 815
4-bit
16-bit