Real-World AI Voice Interface Technology

From real-world voice input to recognition and dialogue

mpWAV develops the integrated voice interface stack — speech enhancement that removes ambient noise and device echo, speech separation, speaker diarization, wake-word detection, sound source localization, on-device speech recognition, and a conversational language model.

From single-microphone smart devices to multi-channel robots, kiosks, vehicles, appliances, and meeting systems, we provide the right combination of technologies for your product's audio I/O configuration and compute platform.

The mpWAV voice technology stack at a glance

One Integrated Voice Stack

What does mpWAV technology do?

mpWAV technology makes real-world voice input clearer, identifies the speaker's call and position, and converts speech into text and intent — connecting it to product functions or conversational services.

From small devices with a single microphone to products combining multiple microphones and speakers, you can select exactly the technologies you need and integrate them as one voice interface stack.

Capture → Enhancement → Refinement → Recognition → Dialogue → Product Action

Real-World Voice Challenges

Different products face different voice problems — and need different technology

A product's microphone never receives only the user's voice.

Nearby conversation, music, road and motor noise, echo from the product's own speaker, and overlapping speech can all arrive together. Some products are constrained in where microphones can be placed, and once speech is recognized, the product still has to understand intent and context.

mpWAV does not apply one technology to every product. We combine technologies based on each product's problem and hardware constraints.

Single-microphone products

Earbuds and small devices with limited microphone options need single-mic noise removal.

Multi-microphone products

Robots, kiosks, and vehicles need multi-microphone noise removal and spatial analysis.

Products with speakers

Response prompts and announcements that loop back into the microphone must be cancelled.

Multi-party conversation

Meetings and voice chat require separating overlapping voices and telling speakers apart.

Conversational services

Kiosks and robots must understand intent and context beyond recognizing the words.

The required technology stack depends on microphone count, number of speakers, product output audio, and the final service.

Find the Right Technology

Pick your product's problem

Product problemRecommended technology
Echo from the product's own speakermpAEC
Removing ambient noise with multiple microphonesmpBeamforming
Echo and ambient noise at the same timempAB
Limited microphone optionsmpNC
Detecting a wake wordmpWWD
Multiple voices mixed togethermpSeparation
Knowing the speaker's direction or positionmpLocalization
Telling who spoke whenmpDiarization
Converting speech to text or commandsmpASR
Understanding dialogue context and intentmpLLM
Multi-channel audio I/O requiredMulti-channel audio I/O HW

Real products usually face several of these at once, so two or more technologies are often applied together.

Performance Validation

We validate with real product results, not technology names

Voice technology performance varies with microphones, speakers, user distance, noise type, the connected ASR, and the execution platform.

So instead of generic benchmark figures, we measure against your own product data and field conditions and report what we find.

Speech enhancement

  • Echo removal
  • Ambient noise removal
  • Target voice preservation
  • Processed audio quality

Speech refinement

  • Speech separation
  • Per-speaker separation
  • Utterance segmentation

Speech recognition

  • Wake word detection
  • Command success rate
  • WER / CER
  • Domain terminology
  • On-device / on-premise

Dialogue processing

  • Intent understanding
  • Slot & option extraction
  • Context retention
  • Speaker direction & position estimation

Product performance

  • Processing latency
  • Memory footprint
  • FPGA·DSP·AP resources
  • On-device feasibility

Independently measured in accredited testing

MetricMeasured (mpWAV)Source and test conditions
SRAM footprint0.571 MBTTA test report (2023), Ministry of SMEs and Startups Didimdol · mpAB · static spec
Clock frequency64 MHzTTA test report (2023), Ministry of SMEs and Startups Didimdol · mpAB · static spec
Real-Time Factor (RTF)0.852TTA test report (2023), Ministry of SMEs and Startups Didimdol · mpAB · test conditions not stated in the report
Stereo echo and noise suppression42.088 dBTTA test report (2023), Ministry of SMEs and Startups Didimdol · mpAB · pass criterion 40 dB or higher

Measured in a TTA test report (2023, Ministry of SMEs and Startups Didimdol programme) with mpAB as the article under test.

Speech Enhancement

We improve the input signal before it reaches recognition

Speech recognition performance depends not only on the ASR model but on the quality of the audio arriving at the microphone.

mpWAV provides speech enhancement matched to your product's acoustic conditions — single or multiple microphones, device echo, and ambient noise.

mpAEC

Cancels echo produced by the device's own speaker

mpAEC is a multi-channel acoustic echo cancellation technology that removes sound played by the product — TV audio, car audio, robot responses, kiosk prompts — from re-entering the microphone.

It keeps adapting while the user is speaking, so cancellation stays stable, and its output can feed beamforming or your existing ASR.

Key problems

  • Speaker output re-entering the microphone
  • Degradation while the user is speaking
  • Changing echo paths
  • Real-time processing

Typical applications

  • Vehicles
  • Robots
  • Kiosks
  • Appliances
  • Smart home
  • Meeting systems

Evidence

Paper
Published in IEEE Transactions on Signal Processing
Patents
2 registered in Korea · 2 abroad · 1 PCT
In a living room, TV and speaker sound; in a car, audio and road noise — both arriving at the microphone alongside the user’s voice command. The blue dashed path is the echo mpAEC removes
How sound from the device’s own speaker returns to the microphone, at home and in a vehicle

mpBeamforming

Removes ambient noise and strengthens the target voice with multiple microphones

mpBeamforming uses the spatial information across multiple microphones to remove ambient noise and deliver the target voice crisply.

mpWAV optimizes from the actual input signals rather than relative microphone position data, so microphone placement can change between products without retuning.

Key problems

  • Far-field user speech
  • Ambient noise from multiple directions
  • Microphone array changes between products
  • Target-voice distortion during noise removal

Typical applications

  • Robots
  • Kiosks
  • Vehicles
  • Appliances
  • Smart home
  • Meeting systems

Evidence

Paper
Published in IEEE/ACM Transactions on Audio, Speech and Language Processing
Patents
3 registered in Korea · 2 abroad · 3 pending · 1 PCT
Award
Prime Minister's Award, 2024 Korea Invention Patent Competition

mpAB

Integrates echo cancellation and beamforming into one optimized pipeline

mpAB combines mpAEC and mpBeamforming to process input where product speaker echo and environmental noise are present together.

Multi-channel microphones, the speaker reference, FPGA·DSP·AP platforms, and your existing ASR can be connected into a single product architecture.

Key problems

  • Echo and ambient noise occurring together
  • Real-time multi-channel audio I/O processing
  • ASR degradation caused by signal distortion
  • Product hardware integration

Typical applications

  • Robots
  • Kiosks
  • Vehicles
  • Appliances
  • Smart devices
  • Meeting devices

Evidence

Certified
NET (New Excellent Technology), 2025 — preprocessing for speech recognition and enhancement in conversational interfaces
The red dashed box is the vacuum, speaker, and TV noise mpBeamforming removes; the blue dashed box is the device speaker echo mpAEC removes. mpAB integrates both and processes the microphone array input together before passing it to the smart device
Ambient noise handled by beamforming and speaker echo handled by mpAEC, processed in one structure

mpNC

Single-microphone noise removal for products with limited mic options

mpNC is a speech enhancement technology that removes ambient noise from a single microphone input.

It provides the technical foundation for improving voice input in earbuds, mobile accessories, small smart devices, and constrained hardware where a microphone array is not an option.

Key problems

  • Product structures that cannot add microphones
  • Space constraints of small devices
  • Everyday ambient noise
  • Limited compute resources

Typical applications

  • Earbuds
  • Wearables
  • Phone accessories
  • Small appliances
  • ClearSense Audio-linked devices

Speech Refinement

Beyond cleaning the audio — separating overlapping voices and organizing them by speaker

When several people speak at once, the product has to separate each voice and work out who spoke when.

mpWAV's speech refinement technologies provide the information needed when multiple speakers are talking.

mpSeparation

Separates multi-speaker input into per-speaker voices

mpSeparation separates a mixed speech signal into each speaker's individual voice in environments where multiple voices exist or overlap — meetings and voice chat.

Combined with mpDiarization and mpASR, it forms the input pipeline for multi-party meeting records and voice chat.

Key problems

  • Overlapping voices from multiple speakers
  • Multi-party conversation
  • Complex voice chat input
  • Signal separation before speaker diarization

Typical applications

  • Meetings
  • Remote collaboration
  • Voice chat
  • Consultation records
  • Multi-party dialogue

Evidence

Paper
NeurIPS 2024 — Separate and reconstruct: Asymmetric encoder-decoder for speech separation
mpSeparation splitting overlapping speech into per-speaker streams.

mpDiarization

Tells who spoke when

mpDiarization separates the speech segments of each speaker in multi-party audio.

It organizes per-speaker utterances in meeting minutes, consultation records, and voice chat, and attaches speaker identity to the text produced by mpASR.

Key roles

  • Per-speaker segment separation
  • Speaker-change detection
  • Speaker tagging for minutes
  • Structuring consultation and dialogue records

Typical applications

  • Meeting minutes
  • Remote meetings
  • Consultation records
  • Voice chat
  • Interview analysis

Speech Recognition and Dialogue

Turning speech into text — and understanding intent, context, and related information

Input that has passed through enhancement and refinement reaches mpWWD, which detects the wake word, and mpASR, which converts it into text.

Where a conversational interface is needed, mpLLM understands the recognition output and dialogue context, connecting questions, responses, and product functions. mpLocalization additionally provides the speaker's direction.

mpWWD

Detects the wake word and starts the voice interface

mpWWD detects a predefined wake word so the product can start listening for voice commands.

It connects to robots, smart home devices, vehicles, and edge devices that must decide the moment of activation on-device rather than streaming audio to a server at all times.

Key roles

  • Wake word detection
  • Voice interface activation
  • Always-on standby
  • Starting the on-device command flow

Typical applications

  • Robots
  • Vehicles
  • Appliances
  • Smart home
  • Kiosks
  • Wearables

mpASR

Noise-robust on-device / on-premise end-to-end ASR

mpASR is mpWAV's speech recognition technology, designed to run on servers, PCs, smart devices, and IoT/edge environments.

It supports domain fine-tuning for your product's terminology, menu names, and command set. Put mpWAV speech enhancement in front of it and recognition in noise improves along with it.

Key roles

  • Speech-to-text conversion
  • Product command recognition
  • Domain fine-tuning
  • Standalone on-device / on-premise execution

Typical applications

  • Kiosks
  • Robots
  • Vehicles
  • Appliances
  • Smart devices
  • Meeting records

mpLLM

A domain-tuned language model that understands dialogue and intent

mpLLM is a conversational language model that connects user intent, dialogue context, and product functions on top of the text produced by mpASR.

For products that need to reduce network dependence — kiosk voice ordering, for instance — it learns your menu structure and ordering rules to build a dialogue flow that fits the product.

Key roles

  • Understanding user intent
  • Maintaining dialogue context
  • Asking for missing information
  • Interpreting menu, options, and quantity

Typical applications

  • Conversational kiosks
  • Robot dialogue
  • Appliance dialogue
  • Smart devices
  • Vehicle interfaces
  • Meeting summaries

mpLocalization

Estimates the direction and position the voice comes from

mpLocalization uses multi-microphone input to estimate the direction or position of the user's voice.

It powers interfaces where a robot turns toward the user, a vehicle distinguishes speech by seat, or an omnidirectional smart device responds toward the speaker.

Key roles

  • Sound source direction estimation
  • User position information
  • Seat- and direction-based interaction
  • Robot gaze and rotation control

Typical applications

  • Robots
  • Vehicles
  • Smart home
  • Meeting devices
  • Omnidirectional smart devices

Hardware and Embedded Implementation

Turning algorithms into systems that actually run

Applying mpWAV technology to a real product requires connecting microphone input, speaker output, the AEC reference, the processing platform, and the speech recognition system together.

mpWAV supports product-specific multi-channel microphone array and audio I/O hardware design, FPGA architectures, AP·DSP porting, and system integration.

Microphone arrays

We design linear and circular arrays matched to product shape and user direction.

Multi-channel audio I/O

Captures and delivers multiple microphones and the speaker reference signal at high speed.

FPGA

Implements multi-channel preprocessing in real time and validates the pre-SoC architecture.

DSP·AP

Ports the algorithms to your product's compute environment.

Lightweight & single-mic builds

For devices that cannot fit an array, we evaluate mpNC and lightweight model structures.

Main PCB block diagram routing a MEMS microphone array through PDM2PCM, a digital amplifier, and USB FS, alongside an I/O specification comparison of mpUSB104K and mpUSB106K
Multi-channel audio I/O board block diagram, with mpUSB104K and mpUSB106K specifications

Application Technology Map

Each application needs a different stack

Robots

  1. 1

    Multi-channel audio I/O HW

    Multi-channel capture and playback

  2. 2

    mpWWD

    Wake word detection

  3. 3

    mpLocalization

    User direction estimation

  4. 4

    mpAEC

    Robot response echo cancellation

  5. 5

    mpBeamforming / mpAB

    Motor, fan, and nearby-conversation removal

  6. 6

    mpASR

    Command recognition

  7. 7

    mpLLM

    Dialogue and intent understanding

Kiosks

  1. 1

    Multi-channel audio I/O HW

    Frontal user voice capture

  2. 2

    mpAEC

    Kiosk prompt echo cancellation

  3. 3

    mpBeamforming / mpAB

    Store noise and nearby-conversation removal

  4. 4

    mpASR

    Menu and option recognition

  5. 5

    mpLLM

    Automated ordering and follow-up questions

Smart devices & earbuds

  1. 1

    Multi-channel audio I/O HW

    Voice input and output handling

  2. 2

    mpNC

    Single-microphone noise removal

  3. 3

    mpWWD

    Wake word detection

  4. 4

    mpASR

    Command recognition

  5. 5

    mpLLM

    Command understanding and dialogue generation

Meetings & voice chat

  1. 1

    Multi-channel audio I/O HW

    Multi-channel meeting capture

  2. 2

    mpSeparation

    Separating overlapping speech

  3. 3

    mpDiarization

    Per-speaker separation

  4. 4

    mpAEC

    Loudspeaker echo cancellation

  5. 5

    mpBeamforming / mpAB

    Ambient noise removal

  6. 6

    mpASR

    Meeting transcription

  7. 7

    mpLLM

    Summaries and structured notes

Vehicles & mobility

  1. 1

    Multi-channel audio I/O HW

    In-cabin multi-channel capture

  2. 2

    mpWWD

    Wake word detection

  3. 3

    mpLocalization

    Seat and speech-direction analysis

  4. 4

    mpAB

    Car audio echo and ambient noise removal

  5. 5

    mpSeparation

    Passenger voice separation

  6. 6

    mpASR

    Vehicle command recognition

  7. 7

    mpLLM

    Natural-language vehicle interface

From Algorithm to Product

Delivered in the form your development stage needs

Software

Apply mpWAV technology while keeping your existing product and ASR.

PoC · performance validation

Verify before-and-after results with your real voice data.

Mic array · HW module

Connect multi-channel audio I/O and processing to your product.

FPGA·DSP·AP porting

Port the technology to your product's compute platform.

Licensing · joint development

Integrate the technology around your product data and service goals.

On-device model optimization

Fit mpASR and mpLLM to the target platform's memory and compute budget.

SoC · semiconductor IP

Scale into dedicated silicon for volume production and miniaturization.

Research and Intellectual Property

Connecting research results to real voice interfaces

mpWAV extends speech and audio signal processing and speech recognition research into patents, software, embedded implementations, and product PoCs.

The core technology began in a Sogang University laboratory and came to the company through a transfer of eight patents. mpAEC and mpBeamforming are published in IEEE Transactions on Signal Processing and IEEE/ACM Transactions on Audio, Speech and Language Processing respectively.

Research → Patent → Algorithm → Embedded Implementation → Product

FAQ

Frequently asked questions about mpWAV technology

No.

The portfolio spans single- and multi-microphone speech enhancement, speech separation, speaker diarization, wake-word detection, localization, speech recognition, and an on-device conversational language model.

Yes. mpNC is built for exactly that case.

It removes ambient noise from a single microphone input.

It is the foundation for earbuds and small smart devices where a microphone array is not an option.

mpBeamforming for ambient noise removal, mpAEC for cancelling the product's own speaker echo, and mpAB when both problems must be handled together.

Yes. mpLocalization estimates the direction or position of the user's voice using multiple microphones.

Actual accuracy and supported arrays should be validated on your product's structure.

Yes. mpDiarization works out who spoke when in multi-party audio.

Combined with mpSeparation and mpASR, it separates meeting audio by speaker and organizes it into per-speaker text.

Yes.

mpASR converts speech to text and mpLLM interprets menu, options, quantity, and dialogue context.

A real deployment also needs the ordering API, menu data, safety policy, and validation on the target hardware.

Currently every technology is covered as a dedicated section of this overview page.

mpAEC, mpBeamforming, mpAB, mpASR, and multi-channel audio I/O HW will expand into standalone pages as material is prepared.

Yes — structurally there is no reason to replace it.

Preprocessing sits between the microphone and recognition and only cleans up the input signal. The engine's input format is unchanged, so there is nothing to modify on the engine side.

Keeping it also lets you measure the preprocessing on its own, by comparing before and after on the same engine.

No.

Results vary with microphone count and placement, speakers, the space, noise, user distance, execution platform, and the connected models. Validate with your real product data.

Build Your Voice Technology Stack

Design the voice technology your product needs as one stack

Tell us your product type, microphone count, speaker structure, noise environment, and the voice features you want to build — we will review the right mpWAV combination.

From single-mic noise removal to multi-channel echo and noise preprocessing, per-speaker separation and diarization, on-device wake-word detection and recognition, a conversational LLM, and speaker direction estimation — connected to fit your product environment.