AECMOS (ONNX) β€” non-intrusive MOS estimator for acoustic echo cancellation

AECMOS is a neural, non-intrusive (no clean reference needed) speech-quality estimator for acoustic echo cancellation. Given the three signals of an AEC scenario β€” the far-end loopback (lpb), the mic capture, and the AEC's enhanced output (enh) β€” it predicts two mean-opinion-score values:

  • Echo MOS β€” how well the echo was removed, and
  • Other / degradation MOS β€” how much the near-end speech was degraded.

This repository re-hosts the pretrained ONNX weights, unmodified, so they can be pulled programmatically (e.g. from a pure-Dart runtime) instead of cloning the full challenge repo.

Provenance & license

These models are from Microsoft's AEC-Challenge repository (AECMOS/AECMOS_local): https://github.com/microsoft/AEC-Challenge. They are released there under the MIT license (Copyright Β© Microsoft Corporation), and are re-hosted here unmodified under the same MIT license for convenience. Attribution and the AECMOS method belong to the original authors; see the AEC-Challenge repo and the AECMOS paper (AECMOS: A speech quality assessment metric for acoustic echo cancellation, ICASSP 2022) for details and citation.

If you use these models, please cite the original AECMOS work and Microsoft's AEC-Challenge, not this mirror.

Files

Each file's Microsoft "run id" is in its name; the scorer selects its configuration from that id.

file run id sampling rate STFT n_fft GRU hidden scenario marker
aecmos_1663915512.onnx 1663915512 16 kHz 513 [4, 1, 64] yes
aecmos_1668423760.onnx 1668423760 48 kHz 1537 [4, 1, 96] yes

Not included: run id 1663829550 (16 kHz, no scenario marker). It is available from the upstream AEC-Challenge repo; this mirror only carries the two files above.

Inputs are the three signals as mono PCM scaled to [-1, 1); the front-end is a librosa-parity log-mel spectrogram (160 mel bands, odd n_fft, periodic Hann, Slaney mel filterbank, power_to_db with ref=max/top_db=80, then (x+40)/40 normalization), stacked to a [1, 3, frames, 160] tensor with the GRU state zeroed.

Usage

Pure Dart β€” no ONNX Runtime / FFI

These run on the pure-Dart inference engine onnx_runtime_dart (conv + max-pool + bidirectional GRU, matched to the Python reference to ~1e-6), so scoring works on every target including web/WASM.

// See onnx_runtime_dart/example/aecmos/ for AecmosScorer + the mel front-end.
final scorer = AecmosScorer('aecmos_1663915512.onnx');
final (:echoMos, :otherMos) = scorer.score('dt', lpb, mic, enh); // st | nst | dt

The CometBeat music app ships a headless CLI wrapper for AEC-quality evaluation:

# Drop a model into the model cache, then score a scenario (headerless PCM16 mono):
mkdir -p ~/.cache/onnx_runtime_dart_models
huggingface-cli download cstr/aecmos-onnx aecmos_1663915512.onnx \
  --local-dir ~/.cache/onnx_runtime_dart_models
dart run bin/aecmos.dart ~/.cache/onnx_runtime_dart_models/aecmos_1663915512.onnx \
  lpb.raw mic.raw enh.raw dt

st = far-end single-talk, nst = near-end single-talk, dt = double-talk.

Download

from huggingface_hub import hf_hub_download
path = hf_hub_download("cstr/aecmos-onnx", "aecmos_1663915512.onnx")
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support