When a user speaks to a smart speaker or voice assistant while the device is playing back a Text-to-Speech (TTS) response, the microphone captures an overlapping mixture of the user's speech and the reverberant TTS playback echo. Textual Echo Cancellation (TEC) cancels the TTS echo using the source text of the TTS playback (< 0.1 KB side input) and reconstructs the clean user speech for Automatic Speech Recognition (ASR).
| Approach | Audio Input to ASR | Side Input Payload | ASR Recognition Result |
|---|---|---|---|
| Run the demo to compare ASR results before and after Textual Echo Cancellation. | |||
Real 0 dB SNR reverberant mixtures ($\mathrm{RT}_{60} = 0.25\text{ s}$) from the LibriTTS test-clean + LJ Speech evaluation set. Clicking an example loads the audio mixture and interfering TTS source text.
| Example | Interfering TTS Playback Text (Side Input to Cancel) | Ground-Truth Clean User Query (Target) | Action |
|---|---|---|---|
| Sample 1 | a table showing the figures for the year ending Michaelmas eighteen oh two. |
"I can't see you at all, anywhere." | Load & Run → |
| Sample 2 | and to approve or disapprove the public policy written into these laws. |
"Because the thing had been such a scare?" | Load & Run → |
| Sample 3 | was living in the city while the walls were still standing, though in a ruinous condition. |
"I must know about you." | Load & Run → |
| Sample 4 | a subsequent bullet, which was lethal, shattered the right side of his skull. |
"I should much prefer that you called in the aid of the police." | Load & Run → |