The Experts below are selected from a list of 6522 Experts worldwide ranked by ideXlab platform

Arshia Cont - One of the best experts on this subject based on the ideXlab platform.

  • An online em algorithm in hidden (semi-)Markov models for Audio Segmentation and clustering
    ICASSP IEEE International Conference on Acoustics Speech and Signal Processing - Proceedings, 2015
    Co-Authors: Alberto Bietti, Francis Bach, Arshia Cont
    Abstract:

    Audio Segmentation is an essential problem in many Audio signal processing tasks, which tries to segment an Audio sig- nal into homogeneous chunks. Rather than separately find- ing change points and computing similarities between seg- ments, we focus on joint Segmentation and clustering, using the framework of hidden Markov and semi-Markov models. We introduce a new incremental EM algorithm for hidden Markov models (HMMs) and showthat it compares favorably to existing online EM algorithms for HMMs. We present re- sults for real-time Segmentation of musical notes and acoustic scenes.

  • ICASSP - An online em algorithm in hidden (semi-)Markov models for Audio Segmentation and clustering
    2015 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2015
    Co-Authors: Alberto Bietti, Francis Bach, Arshia Cont
    Abstract:

    Audio Segmentation is an essential problem in many Audio signal processing tasks, which tries to segment an Audio signal into homogeneous chunks. Rather than separately finding change points and computing similarities between segments, we focus on joint Segmentation and clustering, using the framework of hidden Markov and semi-Markov models. We introduce a new incremental EM algorithm for hidden Markov models (HMMs) and show that it compares favorably to existing online EM algorithms for HMMs. We present results for real-time Segmentation of musical notes and acoustic scenes.

  • An information-geometric approach to real-time Audio Segmentation
    IEEE Signal Processing Letters, 2013
    Co-Authors: Arnaud Dessein, Arshia Cont
    Abstract:

    We present a generic approach to real-time Audio Segmentation in the framework of information geometry for exponential families. The proposed system detects changes by monitoring the information rate of the signals as they arrive in time. We also address shortcomings of traditional cumulative sum approaches to change detection, which assume known parameters before change. This is done by considering exact generalized likelihood ratio test statistics, with a complete estimation of the unknown parameters in the respective hypotheses. We derive an efficient sequential scheme to compute these statistics through convex duality. We finally provide results for speech Segmentation in speakers, and polyphonic music Segmentation in note slices.

Alberto Bietti - One of the best experts on this subject based on the ideXlab platform.

  • An online em algorithm in hidden (semi-)Markov models for Audio Segmentation and clustering
    ICASSP IEEE International Conference on Acoustics Speech and Signal Processing - Proceedings, 2015
    Co-Authors: Alberto Bietti, Francis Bach, Arshia Cont
    Abstract:

    Audio Segmentation is an essential problem in many Audio signal processing tasks, which tries to segment an Audio sig- nal into homogeneous chunks. Rather than separately find- ing change points and computing similarities between seg- ments, we focus on joint Segmentation and clustering, using the framework of hidden Markov and semi-Markov models. We introduce a new incremental EM algorithm for hidden Markov models (HMMs) and showthat it compares favorably to existing online EM algorithms for HMMs. We present re- sults for real-time Segmentation of musical notes and acoustic scenes.

  • ICASSP - An online em algorithm in hidden (semi-)Markov models for Audio Segmentation and clustering
    2015 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2015
    Co-Authors: Alberto Bietti, Francis Bach, Arshia Cont
    Abstract:

    Audio Segmentation is an essential problem in many Audio signal processing tasks, which tries to segment an Audio signal into homogeneous chunks. Rather than separately finding change points and computing similarities between segments, we focus on joint Segmentation and clustering, using the framework of hidden Markov and semi-Markov models. We introduce a new incremental EM algorithm for hidden Markov models (HMMs) and show that it compares favorably to existing online EM algorithms for HMMs. We present results for real-time Segmentation of musical notes and acoustic scenes.

  • Online learning for Audio clustering and Segmentation
    2014
    Co-Authors: Alberto Bietti
    Abstract:

    Audio Segmentation is an essential problem in many Audio signal processing tasks which tries to segment an Audio signal into homogeneous chunks, or segments. Most current approaches rely on a change-point detection phase for finding segment boundaries, followed by a similarity matching phase which identifies similar segments. In this thesis, we focus instead on joint Segmentation and clustering algorithms which solve both tasks simultaneously, through the use of unsupervised learning techniques in sequential models. Hidden Markov and semi-Markov models are a natural choice for this modeling task, and we present their use in the context of Audio Segmentation. We then explore the use of online learning techniques in sequential models and their application to real-time Audio Segmentation tasks. We present an existing online EM algorithm for hidden Markov models and extend it to hidden semi-Markov models by introducing a different parameterization of semi-Markov chains. Finally, we develop new online learning algorithms for sequential models based on incremental optimization of surrogate functions.

Hsinmin Wang - One of the best experts on this subject based on the ideXlab platform.

  • ICASSP - BIC-based Audio Segmentation by divide-and-conquer
    2008 IEEE International Conference on Acoustics Speech and Signal Processing, 2008
    Co-Authors: Shihsian Cheng, Hsinmin Wang
    Abstract:

    Audio Segmentation has received increasing attention in recent years for its potential applications in automatic indexing and transcription of Audio data. Among existing Audio Segmentation approaches, the BIC-based approach proposed by Chen and Gopalakrishnan is most well-known for its high accuracy. However, this window-growing-based Segmentation approach suffers from the high computation cost. In this paper, we propose using the efficient divide-and-conquer strategy in Audio Segmentation. Our approaches detect acoustic changes by recursively partitioning an analysis window into two sub-windows using DeltaBIC. The results of experiments conducted on the broadcast news data demonstrate that our approaches not only have a lower computation cost but also achieve a higher Segmentation accuracy than window-growing-based Segmentation.

  • metric seqdac a hybrid approach for Audio Segmentation
    Conference of the International Speech Communication Association, 2004
    Co-Authors: Hsinmin Wang, Shihsian Cheng
    Abstract:

    This paper presents a hybrid approach for Audio Segmentation, in which the metric-based Segmentation with long sliding windows is applied first to segment an Audio stream into shorter sub-segments, and then the divide-and-conquer Segmentation is applied to a fixed-length window that slides from the beginning to the end of each sub-segment to sequentially detect the remaining acoustic changes. The experimental results on five one-hour broadcast news shows show that our approach outperforms the existing metric-based and model-selection-based approaches.

  • a sequential metric based Audio Segmentation method via the bayesian information criterion
    Conference of the International Speech Communication Association, 2003
    Co-Authors: Shihsian Cheng, Hsinmin Wang
    Abstract:

    In this paper, we propose a sequential metric-based Audio Segmentation method that has the advantage of low computation cost of metric-based methods and the advantage of high accuracy of model-selection-based methods. There are two major differences between our method and the conventional metricbased methods:(1) Each changing point has multiple chances to be detected by different pairs of windows, rather than only once by its neighboring acoustic information.(2) By introducing the Bayesian Information Criterion(BIC) into the distance computation of two windows, we can deal with the thresholding issue more easily. We used five one-hour broadcast news shows for experiments, and the experimental results show that our method performs as well as the model-selection-based methods, but with a lower computation cost.

  • INTERSPEECH - A sequential metric-based Audio Segmentation method via the Bayesian information criterion.
    2003
    Co-Authors: Shihsian Cheng, Hsinmin Wang
    Abstract:

    In this paper, we propose a sequential metric-based Audio Segmentation method that has the advantage of low computation cost of metric-based methods and the advantage of high accuracy of model-selection-based methods. There are two major differences between our method and the conventional metricbased methods:(1) Each changing point has multiple chances to be detected by different pairs of windows, rather than only once by its neighboring acoustic information.(2) By introducing the Bayesian Information Criterion(BIC) into the distance computation of two windows, we can deal with the thresholding issue more easily. We used five one-hour broadcast news shows for experiments, and the experimental results show that our method performs as well as the model-selection-based methods, but with a lower computation cost.

Meinard Muller - One of the best experts on this subject based on the ideXlab platform.

  • frame level Audio Segmentation for abridged musical works
    International Symposium Conference on Music Information Retrieval, 2014
    Co-Authors: Thomas Pratzlich, Meinard Muller
    Abstract:

    Large-scale musical works such as operas may last several hours and typically involve a huge number of musicians. For such compositions, one often finds different arrangements and abridged versions (often lasting less than an hour), which can also be performed by smaller ensembles. Abridged versions still convey the flavor of the musical work containing the most important excerpts and melodies. In this paper, we consider the task of automatically segmenting an Audio recording of a given version into semantically meaningful parts. Following previous work, the general strategy is to transfer a reference Segmentation of the original complete work to the given version. Our main contribution is to show how this can be accomplished when dealing with strongly abridged versions. To this end, opposed to previously suggested segment-level matching procedures, we adapt a frame-level matching approach for transferring the reference segment information to the unknown version. Considering the opera “Der Freischutz” as an example scenario, we discuss how to balance out flexibility and robustness properties of our proposed framelevel Segmentation procedure.

  • freischutz digital a case study for reference based Audio Segmentation for operas
    International Symposium Conference on Music Information Retrieval, 2013
    Co-Authors: Thomas Pratzlich, Meinard Muller
    Abstract:

    Music information retrieval has started to become more and more important in the humanities by providing tools for computer-assisted processing and analysis of music data. However, when applied to real-world scenarios, even established techniques, which are often developed and tested under lab conditions, reach their limits. In this paper, we illustrate some of these challenges by presenting a study on automated Audio Segmentation in the context of the interdisciplinary project “Freischutz Digital”. One basic task arising in this project is to automatically segment different recordings of the opera “Der Freischutz” according to a reference Segmentation specified by a domain expert (musicologist). As it turns out, the task is more complex as one may think at first glance due to significant acoustic and structural variations across the various recordings. As our main contribution, we reveal and discuss these variations by systematically adapting Segmentation procedures based on synchronization and matching techniques.

  • ISMIR - Freischütz Digital: A Case Study for Reference-Based Audio Segmentation for Operas.
    2013
    Co-Authors: Thomas Pratzlich, Meinard Muller
    Abstract:

    Music information retrieval has started to become more and more important in the humanities by providing tools for computer-assisted processing and analysis of music data. However, when applied to real-world scenarios, even established techniques, which are often developed and tested under lab conditions, reach their limits. In this paper, we illustrate some of these challenges by presenting a study on automated Audio Segmentation in the context of the interdisciplinary project “Freischutz Digital”. One basic task arising in this project is to automatically segment different recordings of the opera “Der Freischutz” according to a reference Segmentation specified by a domain expert (musicologist). As it turns out, the task is more complex as one may think at first glance due to significant acoustic and structural variations across the various recordings. As our main contribution, we reveal and discuss these variations by systematically adapting Segmentation procedures based on synchronization and matching techniques.

Giridharan Iyengar - One of the best experts on this subject based on the ideXlab platform.

  • unsupervised Audio Segmentation using extended baum welch transformations
    International Conference on Acoustics Speech and Signal Processing, 2007
    Co-Authors: Tara N Sainath, Dimitri Kanevsky, Giridharan Iyengar
    Abstract:

    Audio Segmentation has applications in a variety of contexts, such as Audio information retrieval, automatic sound analysis, and as a pre-processing step in speech recognition. Extended Baum-Welch (EBW) transformations are most commonly used as a discriminative technique for estimating parameters of Gaussian mixtures. In this paper, we derive an unsupervised Audio Segmentation approach using these transformations. We find that our algorithm outperforms both the Bayesian information criterion (BIC) and cumulative sum (CUSUM) Segmentation methods. In particular, our EBW Segmentation algorithm provides improvements over the baseline approaches in detecting landmarks of short duration and minimizing landmark overSegmentation. In addition, we show that the EBW approach provides faster computation compared to the baseline methods.

  • ICASSP (1) - Unsupervised Audio Segmentation using Extended Baum-Welch Transformations
    2007 IEEE International Conference on Acoustics Speech and Signal Processing - ICASSP '07, 2007
    Co-Authors: Tara N Sainath, Dimitri Kanevsky, Giridharan Iyengar
    Abstract:

    Audio Segmentation has applications in a variety of contexts, such as Audio information retrieval, automatic sound analysis, and as a pre-processing step in speech recognition. Extended Baum-Welch (EBW) transformations are most commonly used as a discriminative technique for estimating parameters of Gaussian mixtures. In this paper, we derive an unsupervised Audio Segmentation approach using these transformations. We find that our algorithm outperforms both the Bayesian information criterion (BIC) and cumulative sum (CUSUM) Segmentation methods. In particular, our EBW Segmentation algorithm provides improvements over the baseline approaches in detecting landmarks of short duration and minimizing landmark overSegmentation. In addition, we show that the EBW approach provides faster computation compared to the baseline methods.

  • INTERSPEECH - Impact of Audio Segmentation and Segment Clustering on Automated Transcription Accuracy of Large Spoken Archives
    2003
    Co-Authors: Bhuvana Ramabhadran, Giridharan Iyengar, Jing Huang, Upendra V. Chaudhari, Harriet J. Nock
    Abstract:

    This paper addresses the influence of Audio Segmentation and segment clustering on automatic transcription accuracy for large spoken archives. The work formspart of the ongoing MALACH project, which is developing advanced techniques for supporting access to the world’s largest digital archive of video oral histories collected in many languages from over 52000 survivors and witnesses of the Holocaust. We present several Audio-only and Audio-visual Segmentation schemes, including two novel schemes: the first is iterative and Audio-only, the second uses Audio-visual synchrony. Unlike most previous work, we evaluate these schemes in terms of their impact upon recognition accuracy. Results on English interviews show the automatic Segmentation schemes give performance comparable to (exhorbitantly expensive and impractically lengthy) manual Segmentation when using a single pass decoding strategy based on speaker-independent models. However, when using a multiple pass decoding strategy with adaptation, results are sensitive to both initial Audio Segmentation and the scheme for clustering segments prior to adaptation: the combination of our best automatic Segmentation and clustering scheme has an error rate 8% worse (relative) to manual Audio Segmentation and clustering due to the occurrence of “speaker-impure” segments.