The Experts below are selected from a list of 5145 Experts worldwide ranked by ideXlab platform
Eliathamby Ambikairajah - One of the best experts on this subject based on the ideXlab platform.
-
an efficient and perceptually motivated Auditory neural encoding and decoding algorithm for spiking neural networks
Frontiers in Neuroscience, 2020Co-Authors: Yansong Chua, Jibin Wu, Malu Zhang, Haizhou Li, Eliathamby AmbikairajahAbstract:: The Auditory front-end is an integral part of a spiking neural network (SNN) when performing Auditory cognitive tasks. It encodes the temporal dynamic stimulus, such as speech and audio, into an efficient, effective and reconstructable spike pattern to facilitate the subsequent processing. However, most of the Auditory front-ends in current studies have not made use of recent findings in psychoacoustics and physiology concerning human listening. In this paper, we propose a neural encoding and decoding scheme that is optimized for audio processing. The neural encoding scheme, that we call Biologically plausible Auditory Encoding (BAE), emulates the functions of the perceptual components of the human Auditory system, that include the cochlear filter bank, the inner hair cells, Auditory Masking effects from psychoacoustic models, and the spike neural encoding by the Auditory nerve. We evaluate the perceptual quality of the BAE scheme using PESQ; the performance of the BAE based on sound classification and speech recognition experiments. Finally, we also built and published two spike-version of speech datasets: the Spike-TIDIGITS and the Spike-TIMIT, for researchers to use and benchmarking of future SNN research.
-
an efficient and perceptually motivated Auditory neural encoding and decoding algorithm for spiking neural networks
arXiv: Sound, 2019Co-Authors: Yansong Chua, Jibin Wu, Malu Zhang, Haizhou Li, Eliathamby AmbikairajahAbstract:Auditory front-end is an integral part of a spiking neural network (SNN) when performing Auditory cognitive tasks. It encodes the temporal dynamic stimulus, such as speech and audio, into an efficient, effective and reconstructable spike pattern to facilitate the subsequent processing. However, most of the Auditory front-ends in current studies have not made use of recent findings in psychoacoustics and physiology concerning human listening. In this paper, we propose a neural encoding and decoding scheme that is optimized for speech processing. The neural encoding scheme, that we call Biologically plausible Auditory Encoding (BAE), emulates the functions of the perceptual components of the human Auditory system, that include the cochlear filter bank, the inner hair cells, Auditory Masking effects from psychoacoustic models, and the spike neural encoding by the Auditory nerve. We evaluate the perceptual quality of the BAE scheme using PESQ; the performance of the BAE based on speech recognition experiments. Finally, we also built and published two spike-version of speech datasets: the Spike-TIDIGITS and the Spike-TIMIT, for researchers to use and benchmarking of future SNN research.
Ashish Panda - One of the best experts on this subject based on the ideXlab platform.
-
integrating denoising autoencoder and vector taylor series with Auditory Masking for speech recognition in noisy conditions
European Signal Processing Conference, 2018Co-Authors: Biswajit A Das, Ashish PandaAbstract:We propose a new front-end feature compensation technique to improve the performance of Automatic Speech Recognition (ASR) systems in noisy environments. First, a Time Delay Neural Network (TDNN) based Denoising Autoencoder (DAE) is considered to compensate the noisy features. The DAE provides good gain in performance when it has been trained using the noise present in the test utterances (“seen” conditions). However, if the noise present in the test utterance is different to what was used in the training of the DAE (“un-seen” conditions), then the performance degrades to a great extent. To improve the ASR performance in such unseen conditions, a model compensation technique, namely the Vector Taylor Series with Auditory Masking (VTS-AM) is used. We propose a new Signal-to-Noise Ratio (SNR) based measure, which can reliably choose the type of compensation to be used for best performance gain. We show that the proposed technique improves the ASR performance significantly on noise corrupted TIMIT and Librispeech databases.
-
vector taylor series expansion with Auditory Masking for noise robust speech recognition
International Symposium on Chinese Spoken Language Processing, 2016Co-Authors: Biswajit Das, Ashish PandaAbstract:In this paper, we address the problem of speech recognition in the presence of additive noise. We investigate the applicability and efficacy of Auditory Masking in devising a robust front end for noisy features. This is achieved by introducing a Masking factor into the Vector Taylor Series (VTS) equations. The resultant first order VTS approximation is used to compensate the parameters of a clean speech model and a Minimum Mean Square Error (MMSE) estimate is used to estimate the clean speech features. The proposed algorithms are validated through experiments on a noise corrupted TIMIT speech recognition database. We show significant performance gain for the proposed method as compared to the traditional VTS algorithm.
Michael S Scordilis - One of the best experts on this subject based on the ideXlab platform.
-
psychoacoustic music analysis based on the discrete wavelet packet transform
Research Letters in Signal Processing, 2008Co-Authors: Michael S ScordilisAbstract:Psychoacoustical computational models are necessary for the perceptual processing of acoustic signals and have contributed significantly in the development of highly efficient audio analysis and coding. In this paper, we present an approach for the psychoacoustic analysis of musical signals based on the discrete wavelet packet transform. The proposed method mimics the multiresolution properties of the human ear closer than other techniques and it includes simultaneous and temporal Auditory Masking. Experimental results show that this method provides better Masking capabilities and it reduces the signal-to-Masking ratio substantially more than other approaches, without introducing audible distortion. This model can lead to greater audio compression by permitting further bit rate reduction and more secure watermarking by providing greater signal space for information hiding.
-
an enhanced psychoacoustic model based on the discrete wavelet packet transform
Journal of The Franklin Institute-engineering and Applied Mathematics, 2006Co-Authors: Michael S ScordilisAbstract:The perception of acoustic information by humans is based on the detailed temporal and spectral analysis provided by the Auditory processing of the received signal. The incorporation of this process in psychoacoustical computational models has contributed significantly both in the development of highly efficient audio compression schemes as well as in effective audio watermarking methods. In this paper, we present an approach based on the discrete wavelet packet transform, which closely mimics the multi-resolution properties of the human ear and also includes simultaneous and temporal Auditory Masking. Experimental results show that the proposed technique offers better Masking capabilities and it reduces the signal-to-Masking ratio when compared to related approaches, without introducing audible distortion. Those results have implications that are important both for audio compression by permitting further bit rate reduction, and for watermarking by providing greater signal space for information hiding.
Yansong Chua - One of the best experts on this subject based on the ideXlab platform.
-
an efficient and perceptually motivated Auditory neural encoding and decoding algorithm for spiking neural networks
Frontiers in Neuroscience, 2020Co-Authors: Yansong Chua, Jibin Wu, Malu Zhang, Haizhou Li, Eliathamby AmbikairajahAbstract:: The Auditory front-end is an integral part of a spiking neural network (SNN) when performing Auditory cognitive tasks. It encodes the temporal dynamic stimulus, such as speech and audio, into an efficient, effective and reconstructable spike pattern to facilitate the subsequent processing. However, most of the Auditory front-ends in current studies have not made use of recent findings in psychoacoustics and physiology concerning human listening. In this paper, we propose a neural encoding and decoding scheme that is optimized for audio processing. The neural encoding scheme, that we call Biologically plausible Auditory Encoding (BAE), emulates the functions of the perceptual components of the human Auditory system, that include the cochlear filter bank, the inner hair cells, Auditory Masking effects from psychoacoustic models, and the spike neural encoding by the Auditory nerve. We evaluate the perceptual quality of the BAE scheme using PESQ; the performance of the BAE based on sound classification and speech recognition experiments. Finally, we also built and published two spike-version of speech datasets: the Spike-TIDIGITS and the Spike-TIMIT, for researchers to use and benchmarking of future SNN research.
-
an efficient and perceptually motivated Auditory neural encoding and decoding algorithm for spiking neural networks
arXiv: Sound, 2019Co-Authors: Yansong Chua, Jibin Wu, Malu Zhang, Haizhou Li, Eliathamby AmbikairajahAbstract:Auditory front-end is an integral part of a spiking neural network (SNN) when performing Auditory cognitive tasks. It encodes the temporal dynamic stimulus, such as speech and audio, into an efficient, effective and reconstructable spike pattern to facilitate the subsequent processing. However, most of the Auditory front-ends in current studies have not made use of recent findings in psychoacoustics and physiology concerning human listening. In this paper, we propose a neural encoding and decoding scheme that is optimized for speech processing. The neural encoding scheme, that we call Biologically plausible Auditory Encoding (BAE), emulates the functions of the perceptual components of the human Auditory system, that include the cochlear filter bank, the inner hair cells, Auditory Masking effects from psychoacoustic models, and the spike neural encoding by the Auditory nerve. We evaluate the perceptual quality of the BAE scheme using PESQ; the performance of the BAE based on speech recognition experiments. Finally, we also built and published two spike-version of speech datasets: the Spike-TIDIGITS and the Spike-TIMIT, for researchers to use and benchmarking of future SNN research.
Haizhou Li - One of the best experts on this subject based on the ideXlab platform.
-
an efficient and perceptually motivated Auditory neural encoding and decoding algorithm for spiking neural networks
Frontiers in Neuroscience, 2020Co-Authors: Yansong Chua, Jibin Wu, Malu Zhang, Haizhou Li, Eliathamby AmbikairajahAbstract:: The Auditory front-end is an integral part of a spiking neural network (SNN) when performing Auditory cognitive tasks. It encodes the temporal dynamic stimulus, such as speech and audio, into an efficient, effective and reconstructable spike pattern to facilitate the subsequent processing. However, most of the Auditory front-ends in current studies have not made use of recent findings in psychoacoustics and physiology concerning human listening. In this paper, we propose a neural encoding and decoding scheme that is optimized for audio processing. The neural encoding scheme, that we call Biologically plausible Auditory Encoding (BAE), emulates the functions of the perceptual components of the human Auditory system, that include the cochlear filter bank, the inner hair cells, Auditory Masking effects from psychoacoustic models, and the spike neural encoding by the Auditory nerve. We evaluate the perceptual quality of the BAE scheme using PESQ; the performance of the BAE based on sound classification and speech recognition experiments. Finally, we also built and published two spike-version of speech datasets: the Spike-TIDIGITS and the Spike-TIMIT, for researchers to use and benchmarking of future SNN research.
-
an efficient and perceptually motivated Auditory neural encoding and decoding algorithm for spiking neural networks
arXiv: Sound, 2019Co-Authors: Yansong Chua, Jibin Wu, Malu Zhang, Haizhou Li, Eliathamby AmbikairajahAbstract:Auditory front-end is an integral part of a spiking neural network (SNN) when performing Auditory cognitive tasks. It encodes the temporal dynamic stimulus, such as speech and audio, into an efficient, effective and reconstructable spike pattern to facilitate the subsequent processing. However, most of the Auditory front-ends in current studies have not made use of recent findings in psychoacoustics and physiology concerning human listening. In this paper, we propose a neural encoding and decoding scheme that is optimized for speech processing. The neural encoding scheme, that we call Biologically plausible Auditory Encoding (BAE), emulates the functions of the perceptual components of the human Auditory system, that include the cochlear filter bank, the inner hair cells, Auditory Masking effects from psychoacoustic models, and the spike neural encoding by the Auditory nerve. We evaluate the perceptual quality of the BAE scheme using PESQ; the performance of the BAE based on speech recognition experiments. Finally, we also built and published two spike-version of speech datasets: the Spike-TIDIGITS and the Spike-TIMIT, for researchers to use and benchmarking of future SNN research.