Wednesday, January 26, 2011

VoicePrint new technology automatically recognize judgments speaker features

Speaker recognition of research began in the 1930s.

As research tools and tools for continuous improvement, speaker recognition of gradually freed early simple ear listening mode. Bell Labs L • G • Kesta used Visual observation of the spectrogram method to be identified and put forward the "voice" concept. Our voice identification technology started relatively late, late 1980s, the Ministry of public security II (now the Ministry of public security identification centre) was introduced in the United States, conducting sound spectrometer DSP5500 VoicePrint's practice of scientific research and check-in. In 1992, the Ministry of public security identification Center completed the Ministerial key topics the 5500 spectrography in VoicePrint on application of law, 2001, the Centre of the national scientific and technical key project, 1995, the VoicePrint key technologies and speaker recognition system "passed acceptance, developed with independent intellectual property rights to the workstation, VS99 speech marks my voice identification technology matures.

"VoicePrint identification and automatic identification technology research" project by the Ministry of public security identification centres and other units, its main research achievements is the voice recognition function implantation VS99 voice workstation, the system can perform automatic speaker feature analysis, judgement and language map display and measurement, and can be combined with the expertise to identify the speaker, suitable for the practical application of forensic science.

This project developed a current VoicePrint work very practical set acoustic spectroscopy and speaker recognition system as one of the voice workstation, greatly improves the accuracy of the conclusions, as VoicePrint provides a practical system.

Innovative technology:

1. robust processing

Noise on the impact of the results is a key issue.

In this system for non-stationary noise, researchers made use of even-numbered frames section of the main components of input hidden Markov Model (HMM) is a combination of time orientation smoothing of SS methods to increase noise Chinese continuous speech recognition system robust method to achieve better recognition results.

2. speech endpoint detection

Endpoint detection can prevent the noise caused by the malfunction and the noise caused by incorrect identification, for accurate detection of speech signal of starting, such as improving the accuracy of the recognition system is significant.

Using the traditional voice activity detector SAD voice activated very easily lead to undetected. In addition, a large, interfering signal may be used at the beginning of the activation of the voice, speech virtual check is activated. To overcome this shortcoming, the researchers used a voice activated based on correlation detector, defines a valid correlation function, found the judge set the threshold, and prevent undetected and virtual testing methods.

3. recognition algorithm

This system is based on GMM model optimization algorithm.

(1) improvement of GMM's model training methods

Experiments find em algorithm exists a singular matrix of major flaws, and maximum likelihood (ML), while the recognition rate is relatively low, but not singular matrix.

Therefore researchers using maximum likelihood estimation (ML) income model for the initial model, and then use the EM algorithm for each step of the model by α value controls the ratio to be amended, as improved EM algorithm.

(2) based on genetic algorithm GMM model optimization algorithm

Researchers on traditional genetic algorithms have been improved for GMM parameter optimization, greatly improving the model optimization.

(3) GMM speaker recognition methods of optimization

Researchers presented a new optimized based on GMM speaker recognition programme, the programme through the first a pronunciation corresponds to a model of the likelihood of making a particular change and then calculate the syllable overall likelihood degrees, also is the syllable on should the total of the scoring model, as Sc, and maximum Sc belongs models speaker is the target speaker.

Social benefits:

Currently, the Ministry of public security identification Center complete the national "ninth five-year research results VS99 voice workstation is already in the country, in the actual case handling has played an important role.

The item is in adding of VS99 automatically determine the functionality to further improve the efficiency of case handling and identification of accuracy.

The project developed the VoicePrint identification system has completely independent intellectual property rights, practicality, very suitable for the practical needs of police work, in the investigation of numerous suspects for troubleshooting, you can effectively provide detection direction, narrow the scope of investigation, improve work efficiency.

At the same time the system has a language functions in Figure display real-time, applies to actions in the technology of voice signal acquisition. Since 2002, the actual inspection, identification of cases, case type 200 including criminal, economic, civil, public security cases. From closed feedback and the results of the trial, was sentenced to 100%.

No comments:

Post a Comment