WeMM-Embedding-2B-GGUF

English · 简体中文

156 MMEB-v3 tasks · ~54 hours local evaluation · Q4_K_M + Q8_0 mmproj

GGUF conversion of tencent/WeMM-Embedding-2B, focused on local multimodal embedding inference with llama.cpp.

Recommended validated pair

  • Main model: WeMM-Embedding-2B-Q4_K_M.gguf
  • Visual projector: mmproj-WeMM-Embedding-2B-Q8_0.gguf
  • Output dimension: 2048
  • Pooling: last token
  • Output normalization: L2
  • Matryoshka dimensions: 64, 128, 256, 512, 1024, 2048

This repository is not only a conversion. The Q4_K_M + Q8_0 projector pair was evaluated end-to-end on 156 MMEB-v3 tasks, requiring approximately 54 hours of local evaluation.

Benchmark result at a glance

Evaluation date: 2026-09-04

Benchmark group Metric Tasks Tencent official WeMM-2B This GGUF Δ Score retained
MMEB-v2 Image Hit@1 36 79.60 79.18 -0.42 99.47%
MMEB-v2 VisDoc NDCG@5 27 80.70 78.07 -2.63 96.75%
MMEB-v3 Text NDCG@5 53 45.30 43.65 -1.65 96.37%
MMEB-v3 MCMR Hit@1 1 42.50 41.01 -1.49 96.49%

The published Tencent scores above are the native WeMM-Embedding-2B aggregate results from the WeMM technical report. They are not a BF16 rerun on the same RTX 5080 machine, so the differences should be read as practical agreement with the published native baseline rather than a perfectly controlled quantization-only ablation.

The important result is that the Q4_K_M main model with a Q8_0 visual projector preserves roughly 96–99% of the published native score across the four directly comparable benchmark groups.

Full evaluated task matrix

This run covers 156 / 190 MMEB-v3 tasks:

Modality / group Tasks Metric Local result
Image 36 Hit@1 79.18
VisDoc 27 NDCG@5 78.07
Text 53 NDCG@5 43.65
Tool 35 Hit@1 49.00
Memory 4 Hit@1 35.67
MCMR 1 Hit@1 41.01
Tool + Memory, no-GUI Agent subset 39 Hit@1 47.63

Not evaluated in this run:

  • 15 Video tasks
  • 8 GUI tasks
  • 11 Audio tasks

The original WeMM model supports video, but video was intentionally excluded from this GGUF benchmark. Audio is not supported by WeMM-Embedding. Because the 8 GUI tasks were excluded, the local 39-task Tool + Memory result must not be compared directly with Tencent's published 47-task Agent aggregate.

Comprehensive comparison: WeMM-Embedding vs Qwen3-VL-Embedding

The tables below separate published native-model results from this repository's local GGUF result. This matters because comparing a local Q4_K_M run directly with a native BF16/FP16 leaderboard submission is useful as a practical reference, but it is not a controlled quantization-only ablation.

Official MMEB-v2 family comparison

Tencent reports the following results on the 78-task MMEB-v2 benchmark.

Model Size AVG Image Hit@1 Video Hit@1 VisDoc NDCG@5
Qwen3-VL-Embedding 2B 73.2 75.0 61.9 79.2
WeMM-Embedding 2B 77.9 79.6 70.8 80.7
WeMM-Embedding 4B 79.2 80.8 72.1 82.0
Qwen3-VL-Embedding 8B 77.8 80.1 67.1 82.4
WeMM-Embedding 9B 80.6 81.9 74.3 83.3

At roughly the same 2B scale, official WeMM-Embedding-2B leads Qwen3-VL-Embedding-2B by:

MMEB-v2 metric WeMM-2B advantage
AVG +4.7
Image +4.6
Video +8.9
VisDoc +1.5

At the larger scale, WeMM-Embedding-9B compared with Qwen3-VL-Embedding-8B:

MMEB-v2 metric WeMM-9B advantage
AVG +2.8
Image +1.8
Video +7.2
VisDoc +0.9

The largest consistent WeMM advantage on MMEB-v2 is video retrieval, while the gap on visual-document retrieval is much smaller.

Official MMEB-v3 family comparison

Tencent reports the following results on all 190 MMEB-v3 tasks.

Model Size V3-All Text NDCG@5 Agent Hit@1 MCMR Hit@1 Audio
Qwen3-VL-Embedding 2B 50.9 39.2 39.3 42.0 0.0
WeMM-Embedding 2B 56.0 45.3 45.1 42.5 0.0
WeMM-Embedding 4B 58.2 47.9 49.0 41.9 0.0
Qwen3-VL-Embedding 8B 53.5 42.5 38.4 38.0 0.0
WeMM-Embedding 9B 59.5 48.8 51.0 49.3 0.0

At the 2B scale, official WeMM-Embedding-2B leads Qwen3-VL-Embedding-2B by:

MMEB-v3 metric WeMM-2B advantage
V3-All +5.1
Text +6.1
Agent +5.8
MCMR +0.5

At the larger scale, WeMM-Embedding-9B compared with Qwen3-VL-Embedding-8B:

MMEB-v3 metric WeMM-9B advantage
V3-All +6.0
Text +6.3
Agent +12.6
MCMR +11.3

This GGUF vs official WeMM-Embedding-2B

This is the most relevant table for evaluating the quality of the conversion itself.

Benchmark group Official WeMM-2B This 2B Q4_K_M GGUF Δ Score retained
Image 79.60 79.18 -0.42 99.47%
VisDoc 80.70 78.07 -2.63 96.75%
Text 45.30 43.65 -1.65 96.37%
MCMR 42.50 41.01 -1.49 96.49%

Across these four directly comparable published benchmark groups, the tested Q4_K_M + Q8_0 projector pair retains approximately 96–99% of the published native WeMM-Embedding-2B score.

This GGUF vs Qwen3-VL-Embedding-2B

This comparison answers a practical deployment question: how does the locally quantized WeMM-2B compare with the published native Qwen3-VL-Embedding-2B baseline?

Benchmark group Qwen3-VL-Embedding-2B native WeMM-2B Q4_K_M GGUF GGUF difference
Image 75.00 79.18 +4.18
VisDoc 79.20 78.07 -1.13
Text 39.20 43.65 +4.45
MCMR 42.00 41.01 -0.99

Even after Q4_K_M quantization of the main model, this WeMM-2B GGUF remains clearly ahead of the published Qwen3-VL-Embedding-2B native baseline on Image and Text, while trailing by about one point on VisDoc and MCMR.

Cross-size reference: this 2B Q4 GGUF vs Qwen3-VL-Embedding-8B

This is not an apples-to-apples comparison, but it is useful for understanding the deployment efficiency of the 2B GGUF.

Benchmark group Qwen3-VL-Embedding-8B native WeMM-2B Q4_K_M GGUF GGUF difference
Image 80.10 79.18 -0.92
VisDoc 82.40 78.07 -4.33
Text 42.50 43.65 +1.15
MCMR 38.00 41.01 +3.01

Notably, the tested 2B Q4_K_M WeMM GGUF exceeds the published Qwen3-VL-Embedding-8B score on Text and MCMR, while coming within one point on Image. Qwen3-VL-Embedding-8B remains substantially stronger on VisDoc.

What cannot be compared directly

This local run intentionally excludes Video and GUI, so several official aggregate columns cannot be reconstructed faithfully:

  • MMEB-v2 AVG cannot be computed because the 15 Video tasks were not evaluated.
  • MMEB-v3 V3-All cannot be computed because Video, GUI and Audio are not all present.
  • Tencent's official Agent score covers 47 tasks. This local run contains 35 Tool + 4 Memory tasks but excludes 8 GUI tasks.
  • The local Tool + Memory average (47.63 Hit@1 over 39 tasks) is therefore a useful no-GUI subset result, but it must not be presented as the official 47-task Agent score.

In short, use Image / VisDoc / Text / MCMR for direct published-score comparisons, and treat Tool / Memory as additional local evidence.

Evaluation environment

Item Configuration
GPU NVIDIA GeForce RTX 5080 16 GB
CPU AMD Ryzen 7 9800X3D
System RAM 32 GB
OS Windows 11
Inference backend llama.cpp llama-server
Main GGUF WeMM-Embedding-2B-Q4_K_M.gguf
mmproj mmproj-WeMM-Embedding-2B-Q8_0.gguf
Image / VisDoc / Tool / Memory context 32768, KV F16
Text context 262144, KV Q8_0
Pooling last
Normalization L2
Parallel slots -np 1
Approx. evaluation time 54 hours

The Text benchmark intentionally uses 262K context. This matters for LongEmbed tasks and avoids turning long-document truncation into a fake "quantization loss".

The final result file records the two context profiles explicitly:

  • image/visdoc/tool/memory = ctx32768 kvf16
  • text = ctx262144 kvq8_0

Evaluation methodology

The benchmark follows Tencent's released mmeb_v3_eval pipeline, which is based on the official VLM2Vec MMEB-v3 evaluator.

Pinned references used by this evaluation work:

  • Tencent WeMM-Embedding evaluation code commit: 9ed7e2d7914cd67a031c3e2a4fef3faaac314721
  • VLM2Vec evaluator commit: 2638a8413fda4b98668a29ea763b4898814bfea7
  • MMEB-V3 data revision: 4a5560b2b64384204b6fea8a82ea986eba51f5aa

The GGUF backend replaces only model inference. Dataset parsing, instructions, query/candidate construction and RankingMetrics remain aligned with the released evaluator.

The tested modalities are:

  • Image
  • VisDoc
  • Text
  • Tool
  • Memory
  • MCMR

Dataset-defined sample caps are preserved where present. Other tasks use their full query/candidate sets.

Important tokenizer / prompt requirements

WeMM embedding extraction uses the final <embedding> token with last-token pooling.

For text input, the validated raw prompt surface is:

<|im_start|>user
YOUR_TEXT<|im_end|><embedding>

There is no newline between <|im_end|> and <embedding>.

The tokenizer must leave <embedding> as the unique final token. When using llama-server, disable automatic EOS insertion:

--override-kv tokenizer.ggml.add_eos_token=bool:false

The benchmark performs a /tokenize audit before evaluation and rejects a configuration where <embedding> is not the unique final token.

llama.cpp server

Validated 32K multimodal / general retrieval configuration:

llama-server \
  -m WeMM-Embedding-2B-Q4_K_M.gguf \
  --mmproj mmproj-WeMM-Embedding-2B-Q8_0.gguf \
  --embedding \
  --pooling last \
  --embd-normalize 2 \
  --override-kv tokenizer.ggml.add_eos_token=bool:false \
  -ngl 99 \
  -c 32768 \
  -np 1 \
  -b 512 \
  -ub 512 \
  --cache-type-k f16 \
  --cache-type-v f16 \
  --flash-attn on \
  --image-min-tokens 64 \
  --image-max-tokens 8192 \
  --no-cache-prompt \
  --no-webui \
  --host 127.0.0.1 \
  --port 8080

For the 262K Text benchmark profile:

llama-server \
  -m WeMM-Embedding-2B-Q4_K_M.gguf \
  --embedding \
  --pooling last \
  --embd-normalize 2 \
  --override-kv tokenizer.ggml.add_eos_token=bool:false \
  -ngl 99 \
  -c 262144 \
  -np 1 \
  -b 512 \
  -ub 512 \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --flash-attn on \
  --no-cache-prompt \
  --no-webui \
  --host 127.0.0.1 \
  --port 8080

Text embedding example

With the server running:

curl http://127.0.0.1:8080/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{"input":"<|im_start|>user\nRepresent the meaning of this sentence.<|im_end|><embedding>","encoding_format":"float"}'

The returned vector should have 2048 dimensions and be L2 normalized.

Visual inputs

For image / visual-document embeddings, use the paired:

mmproj-WeMM-Embedding-2B-Q8_0.gguf

The benchmark uses WeMM-compatible image/text ordering, an image budget of 64–8192 visual tokens, last-token pooling and L2 normalization.

Raw multimodal HTTP request schemas can change between llama.cpp builds. If reproducing the benchmark, keep the exact prompt ordering and verify that <embedding> remains the final token after tokenization.

Benchmark artifacts

For reproducibility, this repository should include the final benchmark outputs:

benchmarks/
├── 2B_scores_summary.csv
└── 2B_scores_detail.txt

2B_scores_summary.csv contains one row per evaluated dataset.

2B_scores_detail.txt contains all reported ranking metrics for all 156 datasets.

Available GGUF files

The 2B conversion set uses these filenames:

WeMM-Embedding-2B-Q4_K_M.gguf
WeMM-Embedding-2B-BF16.gguf
mmproj-WeMM-Embedding-2B-Q8_0.gguf
mmproj-WeMM-Embedding-2B-BF16.gguf

For local deployment, the benchmarked pair is:

WeMM-Embedding-2B-Q4_K_M.gguf
+
mmproj-WeMM-Embedding-2B-Q8_0.gguf

Limitations

  • This benchmark does not include Video, GUI or Audio.
  • Tool + Memory is only a 39-task no-GUI subset of the official Agent benchmark.
  • The published Tencent native scores are used as the reference baseline; a full native BF16 rerun under identical local conditions was not performed.
  • llama.cpp multimodal APIs evolve. Revalidate tokenizer ending, image handling and embedding dimensions when changing builds.
  • Long-text Text evaluation requires substantially more context memory than ordinary image / document retrieval.

Result summary

The practical conclusion from this run:

WeMM-Embedding-2B-Q4_K_M.gguf + mmproj-WeMM-Embedding-2B-Q8_0.gguf is a high-fidelity local quantization of WeMM-Embedding-2B for the evaluated image, visual-document, text, tool and memory retrieval workloads.

Across the four directly comparable published benchmark groups, the GGUF pair retains approximately 96–99% of the published native WeMM-Embedding-2B score, while reducing the main model to Q4_K_M for local deployment.

Upstream

  • Original model: tencent/WeMM-Embedding-2B
  • Project: Tencent/WeMM-Embedding
  • Technical report: WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
  • arXiv: 2608.24053

License

The upstream WeMM-Embedding model is released under the Apache License 2.0. This GGUF repository follows the upstream licensing terms. Third-party evaluation datasets and components retain their original licenses.


WeMM-Embedding-2B-GGUF

tencent/WeMM-Embedding-2B 的 GGUF 转换版本,面向使用 llama.cpp 的本地多模态 Embedding 推理。

推荐且已经完整验证的组合

  • 主模型:WeMM-Embedding-2B-Q4_K_M.gguf
  • 视觉投影器:mmproj-WeMM-Embedding-2B-Q8_0.gguf
  • 输出维度:2048
  • Pooling:最后一个 token(last token)
  • 输出归一化:L2
  • Matryoshka 维度:64, 128, 256, 512, 1024, 2048

这个仓库不只是一次 GGUF 转换。Q4_K_M + Q8_0 mmproj 组合已经完成 156 个 MMEB-v3 任务的端到端实测,本地累计评测时间约 54 小时

核心评测结果

评测日期:2026-09-04

Benchmark 组 指标 任务数 腾讯官方 WeMM-2B 本仓库 GGUF Δ 分数保留率
MMEB-v2 Image Hit@1 36 79.60 79.18 -0.42 99.47%
MMEB-v2 VisDoc NDCG@5 27 80.70 78.07 -2.63 96.75%
MMEB-v3 Text NDCG@5 53 45.30 43.65 -1.65 96.37%
MMEB-v3 MCMR Hit@1 1 42.50 41.01 -1.49 96.49%

上表中的腾讯分数来自 WeMM 技术报告公布的原生模型聚合成绩。它们不是在同一张 RTX 5080 上重新跑出来的 BF16 对照,因此这些差值应该理解为“本地 GGUF 结果与官方原生基线的一致程度”,而不是严格控制变量下的纯量化损失。

最重要的结论是:

Q4_K_M 主模型配合 Q8_0 视觉投影器,在四个可直接比较的公开 benchmark 组上保留了大约 96%–99% 的官方原生分数。

完整评测任务矩阵

本次评测覆盖 156 / 190 个 MMEB-v3 任务

模态 / 分组 任务数 指标 本地结果
Image 36 Hit@1 79.18
VisDoc 27 NDCG@5 78.07
Text 53 NDCG@5 43.65
Tool 35 Hit@1 49.00
Memory 4 Hit@1 35.67
MCMR 1 Hit@1 41.01
Tool + Memory(不含 GUI 的 Agent 子集) 39 Hit@1 47.63

本轮未测试:

  • 15 个 Video 任务
  • 8 个 GUI 任务
  • 11 个 Audio 任务

原始 WeMM 模型支持视频,但本次 GGUF benchmark 有意排除了 Video。WeMM-Embedding 不支持 Audio。

由于排除了 8 个 GUI 任务,本地 39 项 Tool + Memory 结果不能直接与腾讯官方公布的 47 项 Agent 聚合分数比较。

全面对比:WeMM-Embedding vs Qwen3-VL-Embedding

下面将官方原生模型结果与本仓库的本地 GGUF 实测结果分开列出。这样做很重要,因为拿本地 Q4_K_M 直接和原生 BF16/FP16 榜单提交比较,作为部署参考很有价值,但并不是严格意义上的纯量化消融实验。

官方 MMEB-v2 全系列对比

腾讯在 78 项 MMEB-v2 上公布的结果如下:

模型 尺寸 AVG Image Hit@1 Video Hit@1 VisDoc NDCG@5
Qwen3-VL-Embedding 2B 73.2 75.0 61.9 79.2
WeMM-Embedding 2B 77.9 79.6 70.8 80.7
WeMM-Embedding 4B 79.2 80.8 72.1 82.0
Qwen3-VL-Embedding 8B 77.8 80.1 67.1 82.4
WeMM-Embedding 9B 80.6 81.9 74.3 83.3

同为约 2B 规模时,官方 WeMM-Embedding-2B 相比 Qwen3-VL-Embedding-2B:

MMEB-v2 指标 WeMM-2B 优势
AVG +4.7
Image +4.6
Video +8.9
VisDoc +1.5

大模型尺寸下,WeMM-Embedding-9B 相比 Qwen3-VL-Embedding-8B:

MMEB-v2 指标 WeMM-9B 优势
AVG +2.8
Image +1.8
Video +7.2
VisDoc +0.9

MMEB-v2 中 WeMM 最稳定、最明显的优势出现在视频检索;VisDoc 上的差距则小得多。

官方 MMEB-v3 全系列对比

腾讯在完整 190 项 MMEB-v3 上公布的结果:

模型 尺寸 V3-All Text NDCG@5 Agent Hit@1 MCMR Hit@1 Audio
Qwen3-VL-Embedding 2B 50.9 39.2 39.3 42.0 0.0
WeMM-Embedding 2B 56.0 45.3 45.1 42.5 0.0
WeMM-Embedding 4B 58.2 47.9 49.0 41.9 0.0
Qwen3-VL-Embedding 8B 53.5 42.5 38.4 38.0 0.0
WeMM-Embedding 9B 59.5 48.8 51.0 49.3 0.0

同为 2B 规模时,官方 WeMM-Embedding-2B 相比 Qwen3-VL-Embedding-2B:

MMEB-v3 指标 WeMM-2B 优势
V3-All +5.1
Text +6.1
Agent +5.8
MCMR +0.5

大模型尺寸下,WeMM-Embedding-9B 相比 Qwen3-VL-Embedding-8B:

MMEB-v3 指标 WeMM-9B 优势
V3-All +6.0
Text +6.3
Agent +12.6
MCMR +11.3

本仓库 GGUF vs 官方 WeMM-Embedding-2B

这是判断本次 GGUF 转换质量最重要的一张表。

Benchmark 组 官方 WeMM-2B 本仓库 2B Q4_K_M GGUF Δ 分数保留率
Image 79.60 79.18 -0.42 99.47%
VisDoc 80.70 78.07 -2.63 96.75%
Text 45.30 43.65 -1.65 96.37%
MCMR 42.50 41.01 -1.49 96.49%

在四个可以直接和官方公开成绩比较的 benchmark 组上,Q4_K_M + Q8_0 mmproj 组合保留了约 96%–99% 的官方原生能力。

本仓库 GGUF vs Qwen3-VL-Embedding-2B

这张表更接近实际部署问题:已经量化到 Q4 的 WeMM-2B,与 Qwen3-VL-Embedding-2B 官方原生模型相比如何?

Benchmark 组 Qwen3-VL-Embedding-2B 原生 WeMM-2B Q4_K_M GGUF GGUF 差值
Image 75.00 79.18 +4.18
VisDoc 79.20 78.07 -1.13
Text 39.20 43.65 +4.45
MCMR 42.00 41.01 -0.99

即使主模型已经量化为 Q4_K_M,本仓库的 WeMM-2B 在 ImageText 上仍明显高于官方 Qwen3-VL-Embedding-2B 原生基线;VisDoc 和 MCMR 则落后约 1 分。

跨尺寸参考:本仓库 2B Q4 vs Qwen3-VL-Embedding-8B

不是严格同口径比较,但很适合衡量 2B GGUF 的部署效率。

Benchmark 组 Qwen3-VL-Embedding-8B 原生 WeMM-2B Q4_K_M GGUF GGUF 差值
Image 80.10 79.18 -0.92
VisDoc 82.40 78.07 -4.33
Text 42.50 43.65 +1.15
MCMR 38.00 41.01 +3.01

值得注意的是:

实测的 2B Q4_K_M WeMM GGUF 在 Text 和 MCMR 上超过了官方 Qwen3-VL-Embedding-8B,Image 也只差 0.92 分;Qwen3-VL-Embedding-8B 在 VisDoc 上则明显更强。

哪些数据不能直接比较

本地评测有意排除了 Video 与 GUI,因此无法严谨重建部分官方聚合指标:

  • 因为未评测 15 个 Video 任务,所以不能重建 MMEB-v2 AVG
  • 因为没有完整覆盖 Video、GUI、Audio,所以不能计算官方口径的 MMEB-v3 V3-All
  • 腾讯官方 Agent 分数覆盖 47 项;本地只有 35 Tool + 4 Memory,没有 8 GUI。
  • 因此本地 Tool + Memory = 47.63 Hit@1 / 39 tasks 可以作为 no-GUI Agent 子集参考,但不能写成官方 47-task Agent 成绩。

因此,推荐使用 Image / VisDoc / Text / MCMR 做直接官方成绩对比,把 Tool / Memory 视为额外的本地验证证据。

评测环境

项目 配置
GPU NVIDIA GeForce RTX 5080 16 GB
CPU AMD Ryzen 7 9800X3D
系统内存 32 GB
操作系统 Windows 11
推理后端 llama.cpp llama-server
主 GGUF WeMM-Embedding-2B-Q4_K_M.gguf
mmproj mmproj-WeMM-Embedding-2B-Q8_0.gguf
Image / VisDoc / Tool / Memory context 32768,KV F16
Text context 262144,KV Q8_0
Pooling last
Normalization L2
并行槽位 -np 1
累计评测时间 约 54 小时

Text benchmark 有意使用 262K context。这是 LongEmbed 等超长文本任务能够公平评测的关键,避免把长文截断造成的损失误判为量化损失。

最终结果文件明确记录了两套 context profile:

  • image/visdoc/tool/memory = ctx32768 kvf16
  • text = ctx262144 kvq8_0

评测方法

评测沿用腾讯发布的 mmeb_v3_eval pipeline;它基于官方 VLM2Vec MMEB-v3 evaluator。

本轮评测使用的固定版本:

  • Tencent WeMM-Embedding evaluation code:9ed7e2d7914cd67a031c3e2a4fef3faaac314721
  • VLM2Vec evaluator:2638a8413fda4b98668a29ea763b4898814bfea7
  • MMEB-V3 data revision:4a5560b2b64384204b6fea8a82ea986eba51f5aa

GGUF backend 只替换模型推理部分。数据集解析、instructions、query/candidate 构造以及 RankingMetrics 均继续与腾讯发布的 evaluator 对齐。

已测试:

  • Image
  • VisDoc
  • Text
  • Tool
  • Memory
  • MCMR

数据集本身定义了 sample cap 的任务继续沿用该 cap;其他任务使用完整 query / candidate 集合。

tokenizer / prompt 的关键要求

WeMM 的 embedding 从最终 <embedding> token 提取,并使用 last-token pooling

纯文本的验证 prompt 形式:

<|im_start|>user
YOUR_TEXT<|im_end|><embedding>

<|im_end|><embedding> 之间不能插入换行

tokenizer 必须确保 <embedding> 是唯一的最后一个 token。使用 llama-server 时,需要关闭自动 EOS:

--override-kv tokenizer.ggml.add_eos_token=bool:false

benchmark 在正式评测前通过 /tokenize 检查这一点;如果 <embedding> 不是唯一的最终 token,评测会直接拒绝继续。

llama.cpp 服务配置

已经验证的 32K 多模态 / 通用检索配置:

llama-server \
  -m WeMM-Embedding-2B-Q4_K_M.gguf \
  --mmproj mmproj-WeMM-Embedding-2B-Q8_0.gguf \
  --embedding \
  --pooling last \
  --embd-normalize 2 \
  --override-kv tokenizer.ggml.add_eos_token=bool:false \
  -ngl 99 \
  -c 32768 \
  -np 1 \
  -b 512 \
  -ub 512 \
  --cache-type-k f16 \
  --cache-type-v f16 \
  --flash-attn on \
  --image-min-tokens 64 \
  --image-max-tokens 8192 \
  --no-cache-prompt \
  --no-webui \
  --host 127.0.0.1 \
  --port 8080

用于 262K Text benchmark 的配置:

llama-server \
  -m WeMM-Embedding-2B-Q4_K_M.gguf \
  --embedding \
  --pooling last \
  --embd-normalize 2 \
  --override-kv tokenizer.ggml.add_eos_token=bool:false \
  -ngl 99 \
  -c 262144 \
  -np 1 \
  -b 512 \
  -ub 512 \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --flash-attn on \
  --no-cache-prompt \
  --no-webui \
  --host 127.0.0.1 \
  --port 8080

文本 Embedding 示例

服务器启动后:

curl http://127.0.0.1:8080/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{"input":"<|im_start|>user\nRepresent the meaning of this sentence.<|im_end|><embedding>","encoding_format":"float"}'

返回向量应为 2048 维,并已经完成 L2 normalization。

图像输入

Image / VisDoc embedding 需要配套:

mmproj-WeMM-Embedding-2B-Q8_0.gguf

本次 benchmark 使用与 WeMM 对齐的图像 / 文本顺序、64–8192 visual tokens 图像预算、last-token pooling 与 L2 normalization。

llama.cpp 的多模态 HTTP schema 可能随版本变化。复现 benchmark 时,应优先保证 prompt 顺序完全一致,并重新确认 <embedding> 在 tokenization 后仍是最后一个 token。

Benchmark 原始结果

建议仓库同时保留最终结果文件:

benchmarks/
├── 2B_scores_summary.csv
└── 2B_scores_detail.txt

2B_scores_summary.csv:每个评测数据集一行。

2B_scores_detail.txt:包含全部 156 个数据集的所有 ranking metrics。

GGUF 文件

2B 转换文件使用:

WeMM-Embedding-2B-Q4_K_M.gguf
WeMM-Embedding-2B-BF16.gguf
mmproj-WeMM-Embedding-2B-Q8_0.gguf
mmproj-WeMM-Embedding-2B-BF16.gguf

本轮完整 benchmark 使用的是:

WeMM-Embedding-2B-Q4_K_M.gguf
+
mmproj-WeMM-Embedding-2B-Q8_0.gguf

限制

  • 本轮 benchmark 不包含 Video、GUI、Audio。
  • Tool + Memory 只是官方 Agent benchmark 的 39 项 no-GUI 子集
  • 官方腾讯原生成绩作为参考基线,本地没有再用完全相同环境完整重跑一遍 BF16。
  • llama.cpp 多模态 API 会持续变化。更换 build 后,应重新验证 tokenizer 尾部、图像处理和 embedding 维度。
  • 262K Text 长上下文评测的内存与 KV cache 成本明显高于普通图像和文档检索。

结论

本轮测试支持以下实际结论:

WeMM-Embedding-2B-Q4_K_M.gguf + mmproj-WeMM-Embedding-2B-Q8_0.gguf 是 WeMM-Embedding-2B 在已测试 Image、VisDoc、Text、Tool、Memory retrieval 工作负载上的高保真本地量化版本。

在四个能够与腾讯官方公开成绩直接比较的 benchmark 组上,这套 GGUF 组合保留了大约 96%–99% 的官方原生分数,同时把主模型压缩到了 Q4_K_M,更适合本地部署。

上游项目

  • 原始模型:tencent/WeMM-Embedding-2B
  • 项目:Tencent/WeMM-Embedding
  • 技术报告:WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report
  • arXiv:2608.24053

License

上游 WeMM-Embedding 模型采用 Apache License 2.0。本 GGUF 仓库遵循上游许可条款。第三方评测数据集和组件继续遵循各自原始 License。

Downloads last month
815
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TuTuCSF/WeMM-Embedding-2B-GGUF

Finetuned
Qwen/Qwen3.5-2B
Quantized
(7)
this model