On-Device Speech Enhancement

On-Device Speech Enhancement,
Same Chip, Same Engine.

We strip noise and echo from your microphone array input and hand clean audio to the speech recognition engine you already use. No cloud. No dedicated NPU.

Built for product teams putting a voice interface into robots, kiosks, vehicles, and home appliances.

64 MHz
Clock
0.571 MB
SRAM
42 dB
Noise reduction
0.852
Real-time factor

Measured in an accredited TTA test report · 65 dB speech · 60 dB noise

Diagram: six irregular waveforms from six noise sources pass through the mpWAV gate and come out as three clean waveformsOn the left are six irregular waveforms — nearby speech, TV and audio echo, road noise, motor noise, distant speech, and room reflections. They pass through the mpWAV core (AEC and ANC, on-device) in the middle and emerge on the right as three smooth waveforms feeding an ASR engine, wake word detection, and command recognition.Nearby speechTV / audio echoRoad noiseMotor noiseDistant speechRoom reflectionsASR engineWake wordCommandsmpWAVCOREAEC · ANCON-DEVICE

Research · IP

25yrs+
in speech preprocessing
14
granted patents, KR & overseas

Certification · Awards

  • CES 2024 Innovation Awards — 2 categories· Consumer Technology Association
  • NET (New Excellent Technology) certification· Ministry of Trade, Industry and Energy
  • Prime Minister's Prize, Korea Invention Patent Exhibition· 2024
  • Public Procurement Innovative Product· Public Procurement Service

Deployment — PoC

  • Robotics
  • Mobility
  • Kiosks
  • Home appliances

Why does speech recognition break down in the real world?

Because the signal reaching the microphone already carries noise, reverberation, and echo. The recognition engine receives a corrupted signal, and swapping engines does not change what arrives before them.

Nearby conversation, TV and audio playback, road noise, machinery, room reverberation — in real-world noise and reverberation, the input quality delivered to the recognition engine drops sharply.

Nearby conversation

Multiple voices arrive at the same time.

TV & audio echo

Sound played by the device loops back into the microphone.

Road & vehicle noise

Engine and driving noise interfere with voice commands.

Machine & motor noise

The product's own operating sound mixes into the input.

Far-field speech

The farther the user is from the microphone, the more noise dominates.

Room reverberation

Reflections from walls and ceilings degrade recognition.

What matters more than the recognition engine is the quality of the signal it receives.

Far-field voice environments in the smart home and the car — how noise and echo reach the microphone

How was the performance measured, and by whom?

By TTA, an accredited test laboratory. The figures below are measured in a TTA test report (2023) from the Ministry of SMEs and Startups Didimdol programme, with mpAB as the article under test.

mpWAV validates before-and-after differences against real data and real product environments.

Independently measured in accredited testing

MetricMeasured (mpWAV)Source and test conditions
SRAM footprint0.571 MBTTA test report (2023), Ministry of SMEs and Startups Didimdol · mpAB · static spec
Clock frequency64 MHzTTA test report (2023), Ministry of SMEs and Startups Didimdol · mpAB · static spec
Real-Time Factor (RTF)0.852TTA test report (2023), Ministry of SMEs and Startups Didimdol · mpAB · test conditions not stated in the report
Stereo echo and noise suppression42.088 dBTTA test report (2023), Ministry of SMEs and Startups Didimdol · mpAB · pass criterion 40 dB or higher

Measured in a TTA test report (2023, Ministry of SMEs and Startups Didimdol programme) with mpAB as the article under test.

Which products already run it?

Robots, kiosks, vehicles, factories, and care settings. These are summaries — the cases we have written consent for are published in full.

Home robot

Robot voice interface preprocessing

Validated preprocessing to improve a home robot's speech recognition amid TV audio, conversation, and household noise.

Automaker showroom robot

Speech recognition in showroom noise

Evaluated the voice interface of a showroom guide robot amid crowd conversation and room reverberation.

Care robot

Care robot voice interface

We supply the voice interface module that supports voice interaction between older adults and a care robot.

Kiosk module

Microphone-array kiosk module

Applied a multi-channel audio I/O and preprocessing module to achieve high recognition performance in store noise.

Factory motor anomaly detection

Production line acoustic inspection

Validated detecting abnormal motor sounds directly on the line in a noisy factory, without moving units to a separate soundproof room.

ClearSense Audio public pilot

Everyday conversation support in noise

Validated the effectiveness and user satisfaction of the smart listening app at the Bangbae and Nonhyeon senior welfare centers.

Why keep preprocessing as a separate stage?

1

Noise reduction that never sacrifices the target voice

Because suppressing noise also damages the speech you wanted, and recognition gets worse instead of better. Preprocessing that knows what to keep has to sit in front of the engine.

mpWAV focuses on reducing noise and echo while preserving the target voice recognition depends on.

2

Auto-optimization when microphone configurations change

Without relying on relative microphone position data, mpWAV optimizes from the actual input signals.

That reduces the repeated tuning burden when microphone count or placement changes.

3

Applies directly to your existing ASR engine

You can keep the recognition engine you use today exactly as it is — no retuning for the preprocessing.

mpWAV preprocessing sits in front of it and improves the input signal quality.

4

Integrated from software through hardware

Beyond algorithms, we support multi-channel audio I/O, DSP and FPGA porting, board design, and expansion to a dedicated SoC.

What does mpWAV's speech preprocessing do?

It removes echo, suppresses noise, and keeps only the speaker you need. Acoustic echo cancellation, noise suppression, and beamforming run before the recognition engine sees the signal.

Speech Enhancement

mpAEC

Multi-channel acoustic echo cancellation

Removes echo generated by the device itself — TV, car audio, voice prompts — that loops back into the microphone.

  • Continuous adaptation cancels echo reliably even while the user is speaking
  • No separate DTD (double-talk detection)
  • Fast convergence
  • Outstanding multi-channel echo cancellation
More about mpAEC

Speech Enhancement

mpBeamforming

Multi-microphone ambient noise reduction

Uses multiple microphone inputs to reduce ambient noise and strengthen the target voice.

  • Auto-optimization from input signals, not relative microphone positions
  • Change microphone placement without retuning
  • Minimal target-voice distortion
  • Processing designed for recognition performance
More about mpBeamforming

Speech Enhancement

mpAB

Integrated speech preprocessing

Integrates mpAEC and mpBeamforming to process speech in real time where noise and echo coexist.

  • AEC and beamforming integrated
  • Real-time FPGA implementation
  • MCU·DSP architectures
  • AP·DSP porting provided
More about mpAB

Speech Enhancement

mpNC

Single-microphone noise removal

Removes ambient noise using just one microphone, for earbuds and small smart devices where an array is not an option.

  • Built for products with limited microphone options
  • Accounts for small-device space and power limits
  • Removes everyday ambient noise
  • Runs within constrained compute budgets
More about mpNC

Speaker separation, diarization, wake-word detection, on-device recognition, on-device LLM, sound-source localization, and multi-channel hardware design all come after preprocessing.

Which products and environments is it used in?

Anywhere noise gets in the way. We have run PoCs in robotics, mobility, kiosks, and home appliances.

Robots

Helps service, home, care, and guide robots recognize user commands reliably amid ambient noise.

Kiosks

Voice interfaces for barrier-free kiosks in stores, hospitals, transit, and public institutions.

Vehicles & mobility

Delivers reliable recognition of the driver's voice where car audio, driving noise, and passenger conversation coexist.

Smart devices & appliances

Helps TVs, appliances, smart home and IoT devices work reliably in real homes.

Factory acoustic inspection

Analyzes acoustic signals so abnormal motor and equipment sounds are detected immediately, even on noisy production lines.

Defense

Delivers dependable voice interfaces for the equipment and systems used in battlefield and armored environments with extreme noise.

Meetings & voice chat

Separates speech and distinguishes speakers to support meeting minutes and clear communication across meetings, voice chat, and telecom.

Hearing assistance

ClearSense Audio raises speech clarity through ordinary earphones and a smartphone, supporting smoother everyday communication.

See ClearSense Audio

How is it delivered?

As a library, a module, a port, a license, or SoC IP. Which one fits depends on where your product is in development.

Software

Speech preprocessing algorithms delivered as a library.

Multi-channel audio I/O preprocessing module

A module integrating the preprocessing software with multi-channel audio I/O hardware.

DSP·FPGA porting

We port the technology to DSP, FPGA, or AP to meet your compute and latency requirements.

Licensing

License speech preprocessing or acoustic anomaly detection for your products and services.

SoC partnership

Dedicated chips and semiconductor IP collaboration for volume production and product expansion.

How does a project start?

With a PoC on your own data. Environment analysis, data review, PoC, then product integration — and you can start without field recordings.

  1. 1

    Environment analysis

    We review the product, the microphone array and multi-channel audio I/O hardware configuration, user distance, noise environment, and current errors.

  2. 2

    Data review

    We review your real speech data, or a jointly designed test environment.

  3. 3

    PoC & validation

    We compare recognition accuracy, latency, and compute before and after preprocessing.

  4. 4

    Product integration

    We apply the technology as software, a module, or a DSP/FPGA architecture.

  5. 5

    Field optimization

    We verify performance where the product is actually used and fine-tune.

  6. 6

    Production & expansion

    We expand to follow-up models, product lines, regions, and languages.

Smart Listening App

ClearSense Audio — a smart listening app for conversation in noise

It refines the sound your phone picks up with AI, so ordinary earphones carry conversation clearly through noise.

Built under a Seoul Business Agency assistive-technology program, it was trialed with 100 participants at the Nonhyeon and Bangbae senior welfare centers, scoring 5.8 out of 7 for satisfaction.

A smartphone running the ClearSense Audio app next to ordinary wireless earphones. The screen shows ambient listening and amplification controls

FAQ

Questions we get before a project starts

It is the stage that removes noise and echo before the signal reaches your speech recognition engine.

It raises the quality of the input signal without changing the recognition engine itself, so it applies to a speech system you have already built. mpWAV delivers echo cancellation, beamforming, and noise suppression as a single on-device module.

No. You keep the engine you have.

mpWAV is a preprocessing layer that sits between the microphone input and the recognition engine. It applies the same way to a commercial API or an in-house engine, with no retraining.

It runs in real time at 64 MHz clock, 0.571 MB of SRAM, and a 543 KB binary.

The real-time factor is 0.852, and it needs no cloud connection and no dedicated NPU. It drops into embedded environments such as robots, kiosks, and vehicles.

It was verified in an accredited third-party test.

A TTA test report from the SMBA Didimdol program recorded 42.088 dB of noise reduction and a real-time factor of 0.852 (65 dB speech, 60 dB noise). mpWAV also holds two CES 2024 Innovation Awards and NET new-technology certification from the Ministry of Trade, Industry and Energy.

You can choose a software library, a multi-channel audio I/O preprocessing module, a DSP or FPGA port, a technology license, or an SoC partnership.

Which one fits depends on your development stage and compute platform.

It starts with a PoC on your own recorded speech data.

The sequence is environment analysis, data review, PoC and performance validation, product integration, on-site tuning, and then scaling to production. We can begin with the environment analysis even if you have no field recordings yet.

Verify your product's voice problem with real data

Tell us your product type, audio I/O configuration, usage environment, and current errors — we will recommend the right validation method and integration structure.

You don't have to commit to a large project up front. Start a PoC with your most important environment and data.