Textual Echo Cancellation (TEC) β Single Interfering Voice (TecSingleInterfering)
This is not an officially supported Google product.
Multi-source attention sequence-to-sequence Textual Echo Cancellation (TEC) model trained on 24 kHz LibriTTS user speech mixed at 0 dB SNR with reverberant LJ Speech (single-speaker TTS interference). Takes the noisy/reverberant microphone spectrogram (SpeechEncoderV1) and the lightweight TTS source transcript (TtsEncoderV2, < 0.1 KB side input) to reconstruct the clean user speech spectrogram.
- GitHub Repository: https://github.com/wq2012/tec
- PyPI Package:
textual-echo-cancellation(v0.2.0) - Interactive Hugging Face Space Demo: https://huggingface.co/spaces/wq2012/tec
- Paper: Textual Echo Cancellation (IEEE SLT 2021, arXiv:2008.06006v4)
- Audio Demo Page: https://google.github.io/speaker-id/publications/TEC/
Open-Source Reproduction Notice: This pretrained model was trained using the standalone open-source reproduction library (
wq2012/tec) built onlingvoandtensorflow. The original IEEE SLT 2021 paper was conducted using Google's internal codebase and 281k-utterance internal LibriTTS + TTS datasets rather than this open-source reproduction.
Training Configuration
- Condition: Single interfering voice (LibriTTS + LJ Speech)
- Dataset Splits: 400 training utterances (
train), 100 validation utterances (dev), 100test-cleanutterances, and 100test-otherutterances at 24 kHz mixed at 0 dB SNR with synthetic image-method room impulse responses (RT60 = 0.25 s). - Architecture: Speech encoder (
SpeechEncoderV1: 2Γ strided 3Γ3 Conv2D with stride(2, 2), 1Γ Bi-CLSTM with stride(1, 1), and 3Γ Bi-LSTM layers), text encoder (TtsEncoderV2: 98-symbol ASCII character vocabularyVOCAB_SIZE = 98, 512-dim embedding, 3Γ Conv1D + 1Γ Bi-LSTM, for TEC models), and multi-source GMM monotonic attention decoder (MultiSourceFbeDecoderV1/FbeDecoderV1: 2Γ256 Pre-Net, 2-layer autoregressive LSTM, frame reduction factorr = 4, and 5-layer 1D convolutionalPostEditConvNet). - Trainable Parameters:
18,206,463(18,214,084total model variables;218.51 MBcheckpoint shard;19.02 MBdynamic-range quantizedmodel.tflite) - Checkpoint Selection: Validation-loss early stopping (
best.ckptselected on the 100-utterancedevsplit over 500 training steps with batch size 8 and Adamlr = 1e-3).
Model Performance
1. Open-Source Reproduction Evaluation (TecSingleInterfering, N = 100 mixtures per split)
Evaluated on 24 kHz LibriTTS (test-clean: N = 100 mixtures, 1,041 reference words; test-other: N = 100 mixtures, 1,039 reference words) mixed at 0 dB SNR under the Single interfering voice (LibriTTS + LJ Speech) condition, using Qwen3-ASR-0.6B-F16 via audio.cpp for ASR WER scoring and 13-MFCC Dynamic Time Warping for MCD (with 10,000-resample non-parametric bootstrap 95% confidence intervals):
| Metric | test-clean (N = 100) |
test-other (N = 100) |
Paper Reference (TEC (proposed)) |
|---|---|---|---|
| WER (%) β | 64.07% (667/1041, 95% CI [56.32, 72.33]) | 85.37% (887/1039, 95% CI [78.48, 92.38]) | 15.5% / 39.8% |
| MCD (dB) β | 7.99 dB (95% CI [7.75, 8.24]) | 8.60 dB (95% CI [8.26, 8.96]) | 7.51 dB / 8.54 dB |
| Side Input Size (KB) β | 0.059 KB (padded: 0.059 KB) | 0.058 KB (padded: 0.058 KB) | 0.10 KB |
| Complexity (GFLOPS) β | 3.71 GFLOPS | 3.50 GFLOPS | 7.27 GFLOPS (401 frames) |
2. Full Comparison Across All Pretrained Models & Baselines (N = 100 per split)
| Condition | Method | Hugging Face Model | WER (%) test-clean β | WER (%) test-other β | MCD (dB) test-clean β | MCD (dB) test-other β | Side Input test-clean (KB) β | Avg Complexity (GFLOPS) β |
|---|---|---|---|---|---|---|---|---|
| Single interfering voice (LibriTTS + LJSpeech) | GroundTruth |
β | 2.59 [1.54, 3.79] | 6.16 [4.07, 8.47] | 0.00 | 0.00 | 0.000 | 0.00 |
MicrophoneSignal |
β | 80.79 [72.97, 89.04] | 91.15 [84.42, 98.09] | 8.93 [8.62, 9.24] | 9.56 [9.18, 9.93] | 0.000 | 0.00 | |
NlmsAec (AEC-NLMS) |
β | 70.32 [61.76, 79.26] | 87.49 [79.59, 95.49] | 8.76 [8.46, 9.07] | 9.41 [9.04, 9.79] | 189.494 | 0.05 | |
NoSideInputSingleInterfering |
wq2012/vanilla_seq2seq_single_interfering |
73.10 [65.02, 81.44] | 86.81 [79.79, 93.88] | 8.08 [7.84, 8.33] | 8.81 [8.45, 9.17] | 0.000 | 3.35 | |
AecSingleInterfering |
wq2012/aec_single_interfering |
8.17 [6.06, 10.42] | 19.35 [15.23, 23.96] | 5.83 [5.59, 6.07] | 6.20 [5.89, 6.53] | 189.494 | 5.11 | |
TecSingleInterfering |
wq2012/tec_single_interfering |
64.07 [56.32, 72.33] | 85.37 [78.48, 92.38] | 7.99 [7.75, 8.24] | 8.60 [8.26, 8.96] | 0.059 | 3.71 | |
| Multiple interfering voices (LibriTTS + VCTK) | GroundTruth |
β | 2.59 [1.54, 3.79] | 6.35 [4.22, 8.67] | 0.00 | 0.00 | 0.000 | 0.00 |
MicrophoneSignal |
β | 83.67 [75.45, 91.89] | 91.72 [85.43, 98.27] | 8.07 [7.86, 8.29] | 8.35 [8.06, 8.65] | 0.000 | 0.00 | |
NlmsAec (AEC-NLMS) |
β | 75.02 [66.40, 83.62] | 86.72 [79.65, 94.04] | 7.85 [7.65, 8.06] | 8.17 [7.88, 8.47] | 195.914 | 0.05 | |
NoSideInputMultiInterfering |
wq2012/vanilla_seq2seq_multi_interfering |
77.33 [67.79, 86.82] | 84.79 [77.53, 91.94] | 7.58 [7.36, 7.80] | 7.97 [7.65, 8.30] | 0.000 | 3.46 | |
AecMultiInterfering |
wq2012/aec_multi_interfering |
9.32 [6.42, 12.54] | 21.56 [16.63, 26.85] | 5.08 [4.92, 5.25] | 5.32 [5.07, 5.57] | 195.914 | 5.29 | |
TecMultiInterfering |
wq2012/tec_multi_interfering |
65.90 [57.38, 74.18] | 78.25 [71.17, 85.23] | 7.19 [7.02, 7.36] | 7.53 [7.26, 7.82] | 0.062 | 3.84 |
3. Text-Content Sensitivity Ablation (N = 100 per split)
Verifies that TtsEncoderV2 actively exploits the semantic and phonetic content of the interfering TTS transcript rather than acting as a content-independent bias:
| Split | Side-Input Condition | WER (%) β | MCD (dB) β | ΞWER vs. matched |
Paired p-value (Bootstrap / Wilcoxon) |
|---|---|---|---|---|---|
single_test_clean (TecSingleInterfering) |
matched (true TTS transcript) |
64.07% | 7.99 dB | β | β |
shuffled (mismatched transcript) |
71.66% | 8.06 dB | +7.59% | p < 0.0001 / p = 0.00099 | |
empty (no text / Vanilla-Seq2seq) |
73.10% | 8.08 dB | +9.03% | p < 0.0001 / p = 0.00014 | |
multi_test_clean (TecMultiInterfering) |
matched (true TTS transcript) |
65.90% | 7.19 dB | β | β |
shuffled (mismatched transcript) |
75.31% | 7.46 dB | +9.41% | p < 0.0001 / p = 0.00017 | |
empty (no text / Vanilla-Seq2seq) |
77.33% | 7.58 dB | +11.43% | p < 0.0001 / p = 0.00012 |
4. Input SNR Ablation on single_test_clean (N = 100 mixtures per SNR)
| Input SNR | MicrophoneSignal WER / MCD |
NlmsAec WER / MCD |
Vanilla-Seq2seq WER / MCD |
TecSingleInterfering WER / MCD |
AecSingleInterfering WER / MCD |
|---|---|---|---|---|---|
| +5 dB | 34.01% / 7.89 dB | 25.55% / 7.68 dB | 24.88% / 7.01 dB | 21.81% / 6.92 dB | 6.05% / 5.33 dB |
| 0 dB | 80.79% / 8.93 dB | 70.32% / 8.76 dB | 73.10% / 8.08 dB | 64.07% / 7.99 dB | 8.17% / 5.83 dB |
| -5 dB | 104.80% / 9.91 dB | 103.55% / 9.78 dB | 103.55% / 9.27 dB | 102.50% / 9.20 dB | 33.14% / 6.92 dB |
| -10 dB | 107.88% / 10.75 dB | 107.78% / 10.65 dB | 106.82% / 10.19 dB | 106.05% / 10.11 dB | 88.86% / 8.14 dB |
Files in This Repository
best.ckpt.data-00000-of-00001,best.ckpt.index,best.ckpt.meta,checkpoint: TensorFlow / Lingvo checkpoint forTecSingleInterfering(selected by minimum validation loss ondev).model.tflite: Dynamic-range quantized TensorFlow Lite (.tflite) FlatBuffer model forTecSingleInterfering(requires TFLite FlexSELECT_TF_OPSruntime for dynamic RNNTensorListoperations).evaluation_metrics.json: Verified N = 100 evaluation results (test-cleanandtest-other), 95% bootstrap confidence intervals, paired significance tests, and ablations forTecSingleInterfering.
How to Use
1. Install textual-echo-cancellation
pip3 install textual-echo-cancellation huggingface_hub
2. Download the Model from Hugging Face
from huggingface_hub import snapshot_download
model_dir = snapshot_download(repo_id="wq2012/tec_single_interfering")
print("Downloaded model to:", model_dir)
3. Run Inference via CLI (scripts/inference.py)
python3 -m scripts.inference \
--model TecSingleInterfering \
--checkpoint_path "${MODEL_DIR}/best.ckpt" \
--mixed_wav /path/to/mixed_input.wav \
--interfering_text "currently in mountain view it is 72 degrees" \
--output_wav /tmp/enhanced_clean.wav
4. Run Inference via Python API
import os
from huggingface_hub import snapshot_download
from tec import inference
model_dir = snapshot_download(repo_id="wq2012/tec_single_interfering")
ckpt_path = os.path.join(model_dir, "best.ckpt")
result = inference.run_inference_on_wav(
model_name="TecSingleInterfering",
mixed_wav_path="/path/to/mixed_input.wav",
interfering_text="currently in mountain view it is 72 degrees",
checkpoint_path=ckpt_path,
output_wav_path="/tmp/enhanced_clean.wav",
)
print("Enhanced log-Mel spectrogram shape:", result["predicted_mel"].shape)
5. On-Device Inference with Quantized TFLite (model.tflite)
Note: Because the multi-source attention decoder uses dynamic tf.while_loop / TensorList operations, model.tflite is exported with tf.lite.OpsSet.SELECT_TF_OPS (Flex delegate) enabled.
import os
import numpy as np
from huggingface_hub import snapshot_download
from tec import inference
model_dir = snapshot_download(repo_id="wq2012/tec_single_interfering")
tflite_path = os.path.join(model_dir, "model.tflite")
dummy_mel = np.zeros((100, 128), dtype=np.float32)
tflite_out = inference.run_tflite_inference(
tflite_path=tflite_path,
mixed_mel=dummy_mel,
interfering_text="currently in mountain view it is 72 degrees",
)
print("TFLite predicted log-Mel shape:", tflite_out["predicted_mel"].shape)
Expected console output:
TFLite predicted log-Mel shape: (100, 128)
Citation
If you use this model or the textual-echo-cancellation library in your research, please cite the original paper:
@inproceedings{ding2021textual,
title={Textual Echo Cancellation},
author={Ding, Shaojin and Jia, Ye and Hu, Ke and Wang, Quan},
booktitle={2021 IEEE Spoken Language Technology Workshop (SLT)},
pages={666--673},
year={2021},
organization={IEEE}
}
- Downloads last month
- 12