As in digital signal processing (DSP) algorithms and chip processing power and communication network structure and other aspects of development, the increasing popularity of modern communications has.
Audio Conference is an essential feature of the communication system. There are many users participate in the audio conference, the simplest model can use the token under the control of the mutex model, so that only those with a voice that participants can speak. In this mode, each attendee a moment you can only hear all the audio signal, this kind of "half duplex" mode for audio conferences are inconvenient and impractical.Real conference call simulation more participants should be in a meeting room for dialogue.
However, due to the participating Terminal physically and not together, but each terminal only an audio output device (amplifier + speakers) should be transmitted to each terminal's audio stream can only use one channel. For each terminal at the same time receive multiple participants, must adopt a multi-channel audio synthesis programme. Teleconferencing features a Conference Hall with a microphone and speakers, this way it is easy to cause echo jamming and whistling. General meeting of the signal processing algorithm is also a major concern in this context, usually uses the echo cancellation. But this way for a meeting of the signal processing is not the most sophisticated and effective [1]. After research, the use of a silent detection, normalized calibration, Adaptive echo cancellation algorithm synthesis techniques you can implement a real session simulation results.1 Conference signals synthesis implementations
1.1 Conference signals synthesis of reasonableness and necessity
Audio stream does not like the typical video stream in the space/time domain occupies a unique position, at the same time and location of signal elements overlay is meaningless.
But the human ear can perceive in the same space/time play multiple audio streams. This is the synthesis of the Conference of signal and rationality. Synthesis by meeting, the signal of multi-channel audio stream input is processed, to provide a single output channel output audio synthesis.1.2 Conference signals synthesis of key factors
When multiple audio sources play in one space, the human ear to hear sound waves are all sound sources linear superposition of waves, which is analog audio signals synthesis.
The fact that digital voice synthesis after should also be carried out using linear superposition. Suppose there are n-channel input audio streams audio mixing, Xi (t) is the first t time I road, enter voicemail linear samples, then t time mix values are:m(t)=ΣXi(t),i=0,1,…,n-1
Voice signal is continuous, time-a streaming signal, it's time to have the characteristics of short-time smooth.
On the voice signal is processed in one of the basic concept is to create a sample speech signal, voice samples to a buffer to be processed as a unit, namely the voice sample frame. Voice processing of many concepts are based on speech frame, such as audio/silent, energy, autocorrelation, etc. Speech frame length generally use 10 ~ 20ms. Digital audio is an important parameter is the sampling rate, the road enters the audio stream is the synthesis of premise should use the same sample rate.As the need for synthesis in the increase in the number of voice channels, without taking any additional preventive measures, some not meeting effective signal (such as acoustic feedback and noise) are accumulated result quality deterioration, unacceptable.
In particular, by the local amplification system produces echoes caused by electro-acoustic feedback resulted in regeneration, the results of the reverb seriously affecting speech clarity. More deadly when acoustic feedback very seriously, since stress occurs so that the entire communications system does not work. It must be on each terminal of input audio to be silent detection and acoustic feedback.Speech synthesis to be aware of when the dynamic range of the summation samples, which leads to the normalized calibration problems.
Digital audio wave theory, calibration is to check a selected frame, locate the amplitude peak to adjust the selected frame in the overall volume, so as to allow the maximum amplitude values, without overflow. Speech synthesis is carried out on the digital waveform editing, and in particular the need to address normalized calibration problems.2 Conference signals synthesis study on key technology
2.1 Adaptive echo cancellation
Digital Echo Canceller theory based on Adaptive filter technology.
With the rapid development of DSP, digital Echo Canceller has can be applied in the DSP. In a conference call in the echo of the main reasons is the remote session signal amplification system indoors locally produced sound field caused by ECHO microphone feedback to the regeneration of the reverb.Echo Canceller must accurately estimate echo path characteristics and quickly adapt to their changing, according to the Conference call feature, use the interference cancellation model is the best way.
The model is a second input Adaptive Filter, as shown in Figure 1. It will be the local microphone output as the original signal, which will enter the local speakers as reference signal. Through Adaptive echo cancellation after treatment, be effectively inhibit local microphone output feed through the room sound field to the microphone of the electro-acoustic feedback (ECHO), thereby achieving Adaptive acoustic feedback (ECHO).
Echo cancellation of core is the Adaptive Filter algorithm.
Common algorithms including SDA algorithm and LMS algorithm. Because of the SDA algorithm calculations in gradient involves matrix and is not suitable for practical application. Through its LMS algorithm derived from the simple and practical, computing and high efficiency. TI's DSP chip TMS320C54X have special instructions for acceleration of LMS Adaptive Filter algorithm. In practice, it is also possible to LMS algorithm is modified on the basis of the filter coefficients of algorithms:
Detailed Adaptive echo cancellation algorithm calculates the steps are as follows:
(
1) sampling values;(2) in accordance with the previously calculated value and filter to modify the algorithm, a coefficient adjustment;
(3) calculation of the estimated energy; distal
δ2[k] = (1-α) δ2[k-1] +α X2[k]
(4) for FIR filter calculation, evaluated over the filter output y (n) and error signal e (n);
(5) data output;
(6) jumps to the first step.
2.2 silent energy detection
In a silent ITU-T protocol detection, voice activity detection (Voice Activity Detection).
In multipoint audio conference, silence detection, in a certain period of actual speech synthesis in terminals number significantly less than the number of participants, reducing the amount of synthetic operation, thereby reducing the burden of processing chip. At the same time also the microphone Adaptive gain control AGC Foundation.In the digital audio signal, silence detection through signal energy, zero rate parameter combinations, and out-of-the-energy threshold value.
Based on the short-time calculation of the average energy is to use a fixed width of the sliding window, enter a new sample, the sample window covering all samples of energy, on average, associates it with an energy threshold value comparison to determine if the new sample is muted or audio.As noted above, to frames on digital voice detection, if a frame in a sample is sound, the frame is sound.
The window frames, instead of taking the sample, direct with each frame of the last sample is a silent to determine the frame is a silent audio frames or frames, this simplified judgment means significant savings in operational capacity. For the purposes of the judgment does not affect the results.Use adaptive change energy threshold can more accurately to be silent.
You can sample short-term energy of first order linear low pass filter is background noise energy. And adaptive energy threshold values remain and short background noise energy a silence detection sensitivity constant ratio So. Long continuous speech elevated background noise of estimates, which correspondingly increased silence detection energy threshold, can cause immediately occur low amplitude of speech as mute and not be detected. So when detected when you change voice low pass filter cutoff frequency to estimate noise energy again.While the filter mute should pay attention to how to keep short-term energy relatively low weak audio signal, such as friction sound and a consonant.
These weak signal presence guarantees the integrity of the voice of semantic, so in the short-time average energy judgment, should have zero rate combining discriminant keep these weak audio signal. Adopt sound generator can implement retention of weak audio signal, that is the sound generator will immediately follow a voice with the first few frames. The so-called silent frames should be deemed to be the voice, so as to avoid low-level voice suppressed. ITU-T G.723.1A on sound generator algorithm for a more detailed design, this is not done detailed description.2.3 normalized calibration process
Multi-channel voice signals synthesis with linear superposition, must be solved is how to prevent stack overflow and cause distortion.
If the sample is a sample, and the sum of buffer 16bit is also that two road 16bit audio stream is easy to make the sum of overflow. Even provides high precision summation buffer so that the sum of process will not be overrun, but this does not guarantee that the sum of the results of amplitude for the requirements of the hardware device output range (the range is typically DA device 16bit).A simple method is out of range values clamped.
A better approach is to sum the result is normalized framed calibration by: a sum of all speech frame sample analysis, if the value of the sample S than the device can represent maximum range, then the S after samples are multiplied by an attenuation factor f. Where f is the ability to make S meet output devices the maximum value of the range, it is clear that the absolute value is less than 1 f. This clamp after a period of time, the voice sample size between is relatively unchanged.In experiments using a common 16bit fixed-point DSP chip for real-time simulation TMS320C549 to complete multi-channel audio synthesis.
Add the line of the sample, the sum of the values is not overrun because the sample is 16bit, accumulator is 32bit. However, and it is easy to exceed the value of output hardware devices permit (16bit).In normalized calibration process, when the attenuation factor f is initialized to 1, every time you start working on a new sample buffer, any more than a sample S to S clamp and evaluated S and the ratio of the permitted range of values, in the time series f is located on the S after the sample is divided by f.
However, in order to avoid unnecessary loss of speech, and clamping operation has let f smaller and trends, and therefore need to have let f gets bigger, this happens for each new sample buffer to begin processing. New buffer sample still needs attenuating possibilities is large, so f for each start with 1, but to some extent, the value of the inheritance of the past. That is, each new sample buffer, so long as the entrance of f is not equal to 1, they are adjusted to the slightly larger than f some value so that it becomes the new attenuation factor. If the sample does not need attenuation, after several frames after f will slowly back to 1.Fixed-point DSP is uses the Division in, so you can put all the values of f into a table, the range of f is defined as 1/16, 2/16, until 15/16, its attenuation accuracy to 1/16.
Occurs when the S clamp, used the comparative method or look-up find suitable f (one of the 15 value). The reason given is 1/16 of steps, because it has to ensure that the 16-summation of the input stream does not overflow, if you require more precision, you can get 1/32 (the n-th 2The fixed point DSP implementation seems more convenient).To sum up, the normalization of calibration of the core idea is: f must quickly become suitable attenuation factor that makes the sample does not overflow, then f is slowly getting back to 1.
S occur when the clamp is calculated immediately and f, and each time a summation frame has been processed, you tried to close to 1, f f every increase in it and the difference of 1 1/16. Namely: f ' = f + (1-f)/16. Specific calibration of the flowchart shown in Figure 2.
3 test analysis
At the same time, enter the 10-channel audio stream to the mixer module, per-channel sampling rate are 16kHz, frame, long select 10ms 160 samples.
In the electrical interference cancellation for bandwidth 3kHz (300 ~ 3 300Hz) broadband random white noise, offset level superior to 42dB.
Outdoors, the reverberation time is relatively small, on broadband noise of acoustic interference cancellation level better than 30dB. In reverberation more serious in the laboratory, the degree of acoustic interference cancellation can also be better than 15dB.After hearing test showed that after calibration and echo suppression of synthesized speech stream output can clearly make out the sound of each one.
Using Matlab comparison output simple clamping and output calibration two ways of speech waveform that can observe the former waveform has many due to overflow causes "clipping", and the latter's waveform distortion smaller.
Digital audio synthesis for multipoint audio conference system is indispensable.
First of all to the input of multiple audio streams through a silent energy detection and echo suppression processing after a valid input signal linear superposition, then gain calibration in order to reduce distortion, to meet the requirements of the output device. Fixed point DSP implementation of adopted and the experiment proves that this mode of the audio conference signals synthesis algorithm can achieve very good results of the meeting.Reference documents
1 week, DSP and communication engineering. [M]. Beijing: defense industry press, 2004: 301-315
Yangxing Jun 2. voice digital signal processing [M]-Beijing: electronic industry press, 1995: 154-157
3 ITU-T G.723.1 Annex A:Silence Compression Scheme.
ITU,1996
No comments:
Post a Comment