← All case studies

Running a neural audio model in real time on low-power, cost-effective embedded hardware

Limited power from connected equipment and the cost of large deployments shaped this embedded noise suppression system. We implemented a neural audio model for speech noise suppression on the selected hardware and retained its behavior in reference comparisons. The integrated audio configuration processed frames within a 16 ms budget.

Project
Neural noise suppression for a host-powered voice accessory
Key outcomes
  • A streaming C implementation for the embedded platform selected for the host-powered accessory.
  • Checked model outputs and streaming state matched the reference within numerical test tolerances.
  • The integrated audio configuration processed frames within its 16 ms budget in a hardware test.

Background

The voice accessory needed to improve speech in noise at a practical unit cost for large deployments. It also had to avoid a separate battery maintenance task. Charging additional accessories or replacing disposable batteries was not a practical operating model. Separate batteries would add cost, charging or replacement supplies, equipment setup, and user training. They would also add preparation steps that could be missed when the equipment was needed urgently.

The design therefore called for power from the equipment the accessory connects to, with no separate accessory battery. That limited supply drove the choice of a low-power embedded platform. The goal was to use a trained neural audio model to suppress noise in speech on this platform while keeping the equipment practical to deploy and use.

Challenge

The power requirement shaped the platform choice; the platform’s memory and processing limits then shaped the implementation. Neural inference had to share those resources with sample-rate conversion, gain control, and continuous audio input and output. Each frame had a 16 ms processing budget. Late processing or a gap in sample transfer could interrupt the audio.

The model’s learned behavior also had to survive the implementation changes. Its recurrent layers and streaming caches retained information across frames, and automatic conversion did not support all the required operations and state updates. An error in that retained state could affect later speech processing. The implementation had to preserve the reference computation while controlling memory use, execution time, and audio transfer.

Solution

We rebuilt the streaming neural computation in C using the trained weights and the original implementation as the reference. Comparisons covered intermediate outputs, retained state, and the final output across a sequence of frames. This made the learned network's behavior a requirement for each optimization.

Memory use became part of the design. We separated temporary workspace from persistent state and audio buffers, then placed frequently accessed weights in fast on-chip RAM. That change reduced standalone model processing from about 47.6 to 27.0 ms per frame. Reaching the 16 ms budget required further work: we reused normalization terms and skipped zero entries in sparse matrices to reduce repeated calculations. Reference comparisons checked that these changes preserved the model computation.

We also measured processor-specific routines at the operation sizes the model actually used. Fourier-transform routines helped, with their inverse scaling adjusted to preserve signal amplitude. For the small vector operations in this audio path, scalar C loops were faster than the corresponding general-purpose DSP calls in the controlled comparison. Those measurements guided the implementation of each part of the audio path.

The neural model then became part of a continuous system. Chained audio buffers removed a gap in sample transfer. A two-core design assigned neural processing to one core and audio transfer, preprocessing, and output processing to the other. This gave inference the processing time it needed while keeping the audio interfaces serviced.

Result

The C implementation preserved the checked model outputs and streaming state within the recorded reference tolerances. In a hardware test with speech and airflow, the integrated configuration reported a maximum processing window of about 13.6 ms against its 16 ms budget. Sampled diagnostics showed no buffer or DMA errors.

Listening remained a separate acceptance check. The stronger filter used in that hardware test was rejected because it cut speech onsets, despite meeting the timing and clipping checks. The processing deadline and speech quality both mattered.

The work established a real-time embedded processing path for the planned host-powered accessory, with numerical checks on the neural implementation and hardware timing checks on the integrated audio chain.

What would you like to achieve?

Tell us about your goal or the challenge in your way.

Let’s build a solution