Textual Echo Cancellation (TEC) β€” Single Interfering Voice (TecSingleInterfering)

This is not an officially supported Google product.

Multi-source attention sequence-to-sequence Textual Echo Cancellation (TEC) model trained on 24 kHz LibriTTS user speech mixed at 0 dB SNR with reverberant LJ Speech (single-speaker TTS interference). Takes the noisy/reverberant microphone spectrogram (SpeechEncoderV1) and the lightweight TTS source transcript (TtsEncoderV2, < 0.1 KB side input) to reconstruct the clean user speech spectrogram.

Open-Source Reproduction Notice: This pretrained model was trained using the standalone open-source reproduction library (wq2012/tec) built on lingvo and tensorflow. The original IEEE SLT 2021 paper was conducted using Google's internal codebase and 281k-utterance internal LibriTTS + TTS datasets rather than this open-source reproduction.


Training Configuration

  • Condition: Single interfering voice (LibriTTS + LJ Speech)
  • Dataset Splits: 400 training utterances (train), 100 validation utterances (dev), 100 test-clean utterances, and 100 test-other utterances at 24 kHz mixed at 0 dB SNR with synthetic image-method room impulse responses (RT60 = 0.25 s).
  • Architecture: Speech encoder (SpeechEncoderV1: 2Γ— strided 3Γ—3 Conv2D with stride (2, 2), 1Γ— Bi-CLSTM with stride (1, 1), and 3Γ— Bi-LSTM layers), text encoder (TtsEncoderV2: 98-symbol ASCII character vocabulary VOCAB_SIZE = 98, 512-dim embedding, 3Γ— Conv1D + 1Γ— Bi-LSTM, for TEC models), and multi-source GMM monotonic attention decoder (MultiSourceFbeDecoderV1 / FbeDecoderV1: 2Γ—256 Pre-Net, 2-layer autoregressive LSTM, frame reduction factor r = 4, and 5-layer 1D convolutional PostEditConvNet).
  • Trainable Parameters: 18,206,463 (18,214,084 total model variables; 218.51 MB checkpoint shard; 19.02 MB dynamic-range quantized model.tflite)
  • Checkpoint Selection: Validation-loss early stopping (best.ckpt selected on the 100-utterance dev split over 500 training steps with batch size 8 and Adam lr = 1e-3).

Model Performance

1. Open-Source Reproduction Evaluation (TecSingleInterfering, N = 100 mixtures per split)

Evaluated on 24 kHz LibriTTS (test-clean: N = 100 mixtures, 1,041 reference words; test-other: N = 100 mixtures, 1,039 reference words) mixed at 0 dB SNR under the Single interfering voice (LibriTTS + LJ Speech) condition, using Qwen3-ASR-0.6B-F16 via audio.cpp for ASR WER scoring and 13-MFCC Dynamic Time Warping for MCD (with 10,000-resample non-parametric bootstrap 95% confidence intervals):

Metric test-clean (N = 100) test-other (N = 100) Paper Reference (TEC (proposed))
WER (%) ↓ 64.07% (667/1041, 95% CI [56.32, 72.33]) 85.37% (887/1039, 95% CI [78.48, 92.38]) 15.5% / 39.8%
MCD (dB) ↓ 7.99 dB (95% CI [7.75, 8.24]) 8.60 dB (95% CI [8.26, 8.96]) 7.51 dB / 8.54 dB
Side Input Size (KB) ↓ 0.059 KB (padded: 0.059 KB) 0.058 KB (padded: 0.058 KB) 0.10 KB
Complexity (GFLOPS) ↓ 3.71 GFLOPS 3.50 GFLOPS 7.27 GFLOPS (401 frames)

2. Full Comparison Across All Pretrained Models & Baselines (N = 100 per split)

Condition Method Hugging Face Model WER (%) test-clean ↓ WER (%) test-other ↓ MCD (dB) test-clean ↓ MCD (dB) test-other ↓ Side Input test-clean (KB) ↓ Avg Complexity (GFLOPS) ↓
Single interfering voice (LibriTTS + LJSpeech) GroundTruth β€” 2.59 [1.54, 3.79] 6.16 [4.07, 8.47] 0.00 0.00 0.000 0.00
MicrophoneSignal β€” 80.79 [72.97, 89.04] 91.15 [84.42, 98.09] 8.93 [8.62, 9.24] 9.56 [9.18, 9.93] 0.000 0.00
NlmsAec (AEC-NLMS) β€” 70.32 [61.76, 79.26] 87.49 [79.59, 95.49] 8.76 [8.46, 9.07] 9.41 [9.04, 9.79] 189.494 0.05
NoSideInputSingleInterfering wq2012/vanilla_seq2seq_single_interfering 73.10 [65.02, 81.44] 86.81 [79.79, 93.88] 8.08 [7.84, 8.33] 8.81 [8.45, 9.17] 0.000 3.35
AecSingleInterfering wq2012/aec_single_interfering 8.17 [6.06, 10.42] 19.35 [15.23, 23.96] 5.83 [5.59, 6.07] 6.20 [5.89, 6.53] 189.494 5.11
TecSingleInterfering wq2012/tec_single_interfering 64.07 [56.32, 72.33] 85.37 [78.48, 92.38] 7.99 [7.75, 8.24] 8.60 [8.26, 8.96] 0.059 3.71
Multiple interfering voices (LibriTTS + VCTK) GroundTruth β€” 2.59 [1.54, 3.79] 6.35 [4.22, 8.67] 0.00 0.00 0.000 0.00
MicrophoneSignal β€” 83.67 [75.45, 91.89] 91.72 [85.43, 98.27] 8.07 [7.86, 8.29] 8.35 [8.06, 8.65] 0.000 0.00
NlmsAec (AEC-NLMS) β€” 75.02 [66.40, 83.62] 86.72 [79.65, 94.04] 7.85 [7.65, 8.06] 8.17 [7.88, 8.47] 195.914 0.05
NoSideInputMultiInterfering wq2012/vanilla_seq2seq_multi_interfering 77.33 [67.79, 86.82] 84.79 [77.53, 91.94] 7.58 [7.36, 7.80] 7.97 [7.65, 8.30] 0.000 3.46
AecMultiInterfering wq2012/aec_multi_interfering 9.32 [6.42, 12.54] 21.56 [16.63, 26.85] 5.08 [4.92, 5.25] 5.32 [5.07, 5.57] 195.914 5.29
TecMultiInterfering wq2012/tec_multi_interfering 65.90 [57.38, 74.18] 78.25 [71.17, 85.23] 7.19 [7.02, 7.36] 7.53 [7.26, 7.82] 0.062 3.84

3. Text-Content Sensitivity Ablation (N = 100 per split)

Verifies that TtsEncoderV2 actively exploits the semantic and phonetic content of the interfering TTS transcript rather than acting as a content-independent bias:

Split Side-Input Condition WER (%) ↓ MCD (dB) ↓ Ξ”WER vs. matched Paired p-value (Bootstrap / Wilcoxon)
single_test_clean (TecSingleInterfering) matched (true TTS transcript) 64.07% 7.99 dB β€” β€”
shuffled (mismatched transcript) 71.66% 8.06 dB +7.59% p < 0.0001 / p = 0.00099
empty (no text / Vanilla-Seq2seq) 73.10% 8.08 dB +9.03% p < 0.0001 / p = 0.00014
multi_test_clean (TecMultiInterfering) matched (true TTS transcript) 65.90% 7.19 dB β€” β€”
shuffled (mismatched transcript) 75.31% 7.46 dB +9.41% p < 0.0001 / p = 0.00017
empty (no text / Vanilla-Seq2seq) 77.33% 7.58 dB +11.43% p < 0.0001 / p = 0.00012

4. Input SNR Ablation on single_test_clean (N = 100 mixtures per SNR)

Input SNR MicrophoneSignal WER / MCD NlmsAec WER / MCD Vanilla-Seq2seq WER / MCD TecSingleInterfering WER / MCD AecSingleInterfering WER / MCD
+5 dB 34.01% / 7.89 dB 25.55% / 7.68 dB 24.88% / 7.01 dB 21.81% / 6.92 dB 6.05% / 5.33 dB
0 dB 80.79% / 8.93 dB 70.32% / 8.76 dB 73.10% / 8.08 dB 64.07% / 7.99 dB 8.17% / 5.83 dB
-5 dB 104.80% / 9.91 dB 103.55% / 9.78 dB 103.55% / 9.27 dB 102.50% / 9.20 dB 33.14% / 6.92 dB
-10 dB 107.88% / 10.75 dB 107.78% / 10.65 dB 106.82% / 10.19 dB 106.05% / 10.11 dB 88.86% / 8.14 dB

Files in This Repository

  • best.ckpt.data-00000-of-00001, best.ckpt.index, best.ckpt.meta, checkpoint: TensorFlow / Lingvo checkpoint for TecSingleInterfering (selected by minimum validation loss on dev).
  • model.tflite: Dynamic-range quantized TensorFlow Lite (.tflite) FlatBuffer model for TecSingleInterfering (requires TFLite Flex SELECT_TF_OPS runtime for dynamic RNN TensorList operations).
  • evaluation_metrics.json: Verified N = 100 evaluation results (test-clean and test-other), 95% bootstrap confidence intervals, paired significance tests, and ablations for TecSingleInterfering.

How to Use

1. Install textual-echo-cancellation

pip3 install textual-echo-cancellation huggingface_hub

2. Download the Model from Hugging Face

from huggingface_hub import snapshot_download

model_dir = snapshot_download(repo_id="wq2012/tec_single_interfering")
print("Downloaded model to:", model_dir)

3. Run Inference via CLI (scripts/inference.py)

python3 -m scripts.inference \
  --model TecSingleInterfering \
  --checkpoint_path "${MODEL_DIR}/best.ckpt" \
  --mixed_wav /path/to/mixed_input.wav \
  --interfering_text "currently in mountain view it is 72 degrees" \
  --output_wav /tmp/enhanced_clean.wav

4. Run Inference via Python API

import os
from huggingface_hub import snapshot_download
from tec import inference

model_dir = snapshot_download(repo_id="wq2012/tec_single_interfering")
ckpt_path = os.path.join(model_dir, "best.ckpt")

result = inference.run_inference_on_wav(
    model_name="TecSingleInterfering",
    mixed_wav_path="/path/to/mixed_input.wav",
    interfering_text="currently in mountain view it is 72 degrees",
    checkpoint_path=ckpt_path,
    output_wav_path="/tmp/enhanced_clean.wav",
)
print("Enhanced log-Mel spectrogram shape:", result["predicted_mel"].shape)

5. On-Device Inference with Quantized TFLite (model.tflite)

Note: Because the multi-source attention decoder uses dynamic tf.while_loop / TensorList operations, model.tflite is exported with tf.lite.OpsSet.SELECT_TF_OPS (Flex delegate) enabled.

import os
import numpy as np
from huggingface_hub import snapshot_download
from tec import inference

model_dir = snapshot_download(repo_id="wq2012/tec_single_interfering")
tflite_path = os.path.join(model_dir, "model.tflite")

dummy_mel = np.zeros((100, 128), dtype=np.float32)
tflite_out = inference.run_tflite_inference(
    tflite_path=tflite_path,
    mixed_mel=dummy_mel,
    interfering_text="currently in mountain view it is 72 degrees",
)
print("TFLite predicted log-Mel shape:", tflite_out["predicted_mel"].shape)

Expected console output:

TFLite predicted log-Mel shape: (100, 128)

Citation

If you use this model or the textual-echo-cancellation library in your research, please cite the original paper:

@inproceedings{ding2021textual,
  title={Textual Echo Cancellation},
  author={Ding, Shaojin and Jia, Ye and Hu, Ke and Wang, Quan},
  booktitle={2021 IEEE Spoken Language Technology Workshop (SLT)},
  pages={666--673},
  year={2021},
  organization={IEEE}
}
Downloads last month
12
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Space using wq2012/tec_single_interfering 1

Collection including wq2012/tec_single_interfering

Paper for wq2012/tec_single_interfering