The Experts below are selected from a list of 99 Experts worldwide ranked by ideXlab platform

Masataka Goto - One of the best experts on this subject based on the ideXlab platform.

  • active music listening interfaces based on signal processing
    International Conference on Acoustics Speech and Signal Processing, 2007
    Co-Authors: Masataka Goto
    Abstract:

    This paper introduces our research aimed at building "active music listening interfaces". This research approach is intended to enrich end-users' music listening experiences by applying music-understanding technologies based on signal processing. Active music listening is a way of listening to music through active interactions. We have developed seven interfaces for active music listening, such as interfaces for skipping sections of no interest within a musical piece while viewing a graphical overview of the entire song structure, for displaying virtual dancers or song lyrics synchronized with the music, for changing the timbre of instrument sounds in compact-Disc Recordings, and for browsing a large music collection to encounter interesting musical pieces or artists. These interfaces demonstrate the importance of music-understanding technologies and the benefit they offer to end users. Our hope is that this work will help change music listening into a more active, immersive experience.

  • ICASSP (1) - Integration and Adaptation of Harmonic and Inharmonic Models for Separating Polyphonic Musical Signals
    2007 IEEE International Conference on Acoustics Speech and Signal Processing - ICASSP '07, 2007
    Co-Authors: Katsutoshi Itoyama, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
    Abstract:

    This paper describes a sound source separation method for polyphonic sound mixtures of music to build an instrument equalizer for remixing multiple tracks separated from compact-Disc Recordings by changing the volume level of each track. Although such mixtures usually include both harmonic and inharmonic sounds, the difficulties in dealing with both types of sounds together have not been addressed in most previous methods that have focused on either of the two types separately. We therefore developed an integrated weighted-mixture model consisting of both harmonic-structure and inharmonic-structure tone models (generative models for the power spectrogram). On the basis of the MAP estimation using the EM algorithm, we estimated all model parameters of this integrated model under several original constraints for preventing over-training and maintaining intra-instrument consistency. Using standard MIDI files as prior information of the model parameters, we applied this model to compact-Disc Recordings and achieved the instrument equalizer.

  • ICASSP (4) - Active Music Listening Interfaces Based on Signal Processing
    2007 IEEE International Conference on Acoustics Speech and Signal Processing - ICASSP '07, 2007
    Co-Authors: Masataka Goto
    Abstract:

    This paper introduces our research aimed at building "active music listening interfaces". This research approach is intended to enrich end-users' music listening experiences by applying music-understanding technologies based on signal processing. Active music listening is a way of listening to music through active interactions. We have developed seven interfaces for active music listening, such as interfaces for skipping sections of no interest within a musical piece while viewing a graphical overview of the entire song structure, for displaying virtual dancers or song lyrics synchronized with the music, for changing the timbre of instrument sounds in compact-Disc Recordings, and for browsing a large music collection to encounter interesting musical pieces or artists. These interfaces demonstrate the importance of music-understanding technologies and the benefit they offer to end users. Our hope is that this work will help change music listening into a more active, immersive experience.

  • a chorus section detection method for musical audio signals and its application to a music listening station
    IEEE Transactions on Audio Speech and Language Processing, 2006
    Co-Authors: Masataka Goto
    Abstract:

    This paper describes a method for obtaining a list of repeated chorus ("hook") sections in compact-Disc Recordings of popular music. The detection of chorus sections is essential for the computational modeling of music understanding and is useful in various applications, such as automatic chorus-preview/search functions in music listening stations, music browsers, or music retrieval systems. Most previous methods detected as a chorus a repeated section of a given length and had difficulty identifying both ends of a chorus section and dealing with modulations (key changes). By analyzing relationships between various repeated sections, our method, called RefraiD, can detect all the chorus sections in a song and estimate both ends of each section. It can also detect modulated chorus sections by introducing a perceptually motivated acoustic feature and a similarity that enable detection of a repeated chorus section even after modulation. Experimental results with a popular music database showed that this method correctly detected the chorus sections in 80 of 100 songs. This paper also describes an application of our method, a new music-playback interface for trial listening called SmartMusicKIOSK , which enables a listener to directly jump to and listen to the chorus section while viewing a graphical overview of the entire song structure. The results of implementing this application have demonstrated its usefulness

  • a real time music scene description system predominant f0 estimation for detecting melody and bass lines in real world audio signals
    Speech Communication, 2004
    Co-Authors: Masataka Goto
    Abstract:

    Abstract In this paper, we describe the concept of music scene description and address the problem of detecting melody and bass lines in real-world audio signals containing the sounds of various instruments. Most previous pitch-estimation methods have had difficulty dealing with such complex music signals because these methods were designed to deal with mixtures of only a few sounds. To enable estimation of the fundamental frequency (F0) of the melody and bass lines, we propose a predominant-F0 estimation method called PreFEst that does not rely on the unreliable fundamental component and obtains the most predominant F0 supported by harmonics within an intentionally limited frequency range. This method estimates the relative dominance of every possible F0 (represented as a probability density function of the F0) by using MAP (maximum a posteriori probability) estimation and considers the F0’s temporal continuity by using a multiple-agent architecture. Experimental results with a set of ten music excerpts from compact-Disc Recordings showed that a real-time system implementing this method was able to detect melody and bass lines about 80% of the time these existed.

Hiroshi G. Okuno - One of the best experts on this subject based on the ideXlab platform.

  • ICASSP (1) - Integration and Adaptation of Harmonic and Inharmonic Models for Separating Polyphonic Musical Signals
    2007 IEEE International Conference on Acoustics Speech and Signal Processing - ICASSP '07, 2007
    Co-Authors: Katsutoshi Itoyama, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
    Abstract:

    This paper describes a sound source separation method for polyphonic sound mixtures of music to build an instrument equalizer for remixing multiple tracks separated from compact-Disc Recordings by changing the volume level of each track. Although such mixtures usually include both harmonic and inharmonic sounds, the difficulties in dealing with both types of sounds together have not been addressed in most previous methods that have focused on either of the two types separately. We therefore developed an integrated weighted-mixture model consisting of both harmonic-structure and inharmonic-structure tone models (generative models for the power spectrogram). On the basis of the MAP estimation using the EM algorithm, we estimated all model parameters of this integrated model under several original constraints for preventing over-training and maintaining intra-instrument consistency. Using standard MIDI files as prior information of the model parameters, we applied this model to compact-Disc Recordings and achieved the instrument equalizer.

  • ISMIR - Automatic Chord Transcription with Concurrent Recognition of Chord Symbols and Boundaries.
    2004
    Co-Authors: Takuya Yoshioka, Kazunori Komatani, Tetsuya Ogata, Tetsuro Kitahara, Hiroshi G. Okuno
    Abstract:

    This paper describes a method that recognizes musical chords from real-world audio signals in compact-Disc Recordings. The automatic recognition of musical chords is necessary for music information retrieval (MIR) systems, since the chord sequences of musical pieces capture the characteristics of their accompaniments. None of the previous methods can accurately recognize musical chords from complex audio signals that contain vocal and drum sounds. The main problem is that the chordboundary-detection and chord-symbol-identification processes are inseparable because of their mutual dependency. In order to solve this mutual dependency problem, our method generates hypotheses about tuples of chord symbols and chord boundaries, and outputs the most plausible one as the recognition result. The certainty of a hypothesis is evaluated based on three cues: acoustic features, chord progression patterns, and bass sounds. Experimental results show that our method successfully recognized chords in seven popular music songs; the average accuracy of the results was around 77%.

Katsutoshi Itoyama - One of the best experts on this subject based on the ideXlab platform.

  • ICASSP (1) - Integration and Adaptation of Harmonic and Inharmonic Models for Separating Polyphonic Musical Signals
    2007 IEEE International Conference on Acoustics Speech and Signal Processing - ICASSP '07, 2007
    Co-Authors: Katsutoshi Itoyama, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
    Abstract:

    This paper describes a sound source separation method for polyphonic sound mixtures of music to build an instrument equalizer for remixing multiple tracks separated from compact-Disc Recordings by changing the volume level of each track. Although such mixtures usually include both harmonic and inharmonic sounds, the difficulties in dealing with both types of sounds together have not been addressed in most previous methods that have focused on either of the two types separately. We therefore developed an integrated weighted-mixture model consisting of both harmonic-structure and inharmonic-structure tone models (generative models for the power spectrogram). On the basis of the MAP estimation using the EM algorithm, we estimated all model parameters of this integrated model under several original constraints for preventing over-training and maintaining intra-instrument consistency. Using standard MIDI files as prior information of the model parameters, we applied this model to compact-Disc Recordings and achieved the instrument equalizer.

  • INTEGRATIONAND ADAPTATIONOFHARMONIC AND INHARMONIC MODELS FOR SEPARATINGPOLYPHONICMUSICALSIGNALS
    2007
    Co-Authors: Katsutoshi Itoyama, Kazunori Komatani
    Abstract:

    This paper describes asound source separation method forpolyphonic soundmixtures ofmusictobuild aninstrument equalizer forremixing multiple tracks separated fromcompact-Disc Recordingsbychanging thevolume level ofeachtrack. Although such mixtures usually include bothharmonic andinharmonic sounds, the difficulties indealing withbothtypes ofsounds together havenot beenaddressed inmostprevious methods that havefocused oneither ofthetwotypes separately. Wetherefore developed anintegrated weighted-mixture modelconsisting ofbothharmonic-structure and inharmonic-structure tonemodels (generative models forthepower spectrogram). Onthebasis oftheMAPestimation using theEM algorithm, weestimated all model parameters ofthis integrated model under several original constraints forpreventing over-training and maintaining intra-instrument consistency. Using standard MIDIfiles asprior information ofthemodelparameters, weapplied this model tocompact-Disc Recordings andachieved theinstrument equalizer. IndexTerms-Music, separation, equalizers, sound source separation, music understanding.

Kazunori Komatani - One of the best experts on this subject based on the ideXlab platform.

  • ICASSP (1) - Integration and Adaptation of Harmonic and Inharmonic Models for Separating Polyphonic Musical Signals
    2007 IEEE International Conference on Acoustics Speech and Signal Processing - ICASSP '07, 2007
    Co-Authors: Katsutoshi Itoyama, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
    Abstract:

    This paper describes a sound source separation method for polyphonic sound mixtures of music to build an instrument equalizer for remixing multiple tracks separated from compact-Disc Recordings by changing the volume level of each track. Although such mixtures usually include both harmonic and inharmonic sounds, the difficulties in dealing with both types of sounds together have not been addressed in most previous methods that have focused on either of the two types separately. We therefore developed an integrated weighted-mixture model consisting of both harmonic-structure and inharmonic-structure tone models (generative models for the power spectrogram). On the basis of the MAP estimation using the EM algorithm, we estimated all model parameters of this integrated model under several original constraints for preventing over-training and maintaining intra-instrument consistency. Using standard MIDI files as prior information of the model parameters, we applied this model to compact-Disc Recordings and achieved the instrument equalizer.

  • INTEGRATIONAND ADAPTATIONOFHARMONIC AND INHARMONIC MODELS FOR SEPARATINGPOLYPHONICMUSICALSIGNALS
    2007
    Co-Authors: Katsutoshi Itoyama, Kazunori Komatani
    Abstract:

    This paper describes asound source separation method forpolyphonic soundmixtures ofmusictobuild aninstrument equalizer forremixing multiple tracks separated fromcompact-Disc Recordingsbychanging thevolume level ofeachtrack. Although such mixtures usually include bothharmonic andinharmonic sounds, the difficulties indealing withbothtypes ofsounds together havenot beenaddressed inmostprevious methods that havefocused oneither ofthetwotypes separately. Wetherefore developed anintegrated weighted-mixture modelconsisting ofbothharmonic-structure and inharmonic-structure tonemodels (generative models forthepower spectrogram). Onthebasis oftheMAPestimation using theEM algorithm, weestimated all model parameters ofthis integrated model under several original constraints forpreventing over-training and maintaining intra-instrument consistency. Using standard MIDIfiles asprior information ofthemodelparameters, weapplied this model tocompact-Disc Recordings andachieved theinstrument equalizer. IndexTerms-Music, separation, equalizers, sound source separation, music understanding.

  • ISMIR - Automatic Chord Transcription with Concurrent Recognition of Chord Symbols and Boundaries.
    2004
    Co-Authors: Takuya Yoshioka, Kazunori Komatani, Tetsuya Ogata, Tetsuro Kitahara, Hiroshi G. Okuno
    Abstract:

    This paper describes a method that recognizes musical chords from real-world audio signals in compact-Disc Recordings. The automatic recognition of musical chords is necessary for music information retrieval (MIR) systems, since the chord sequences of musical pieces capture the characteristics of their accompaniments. None of the previous methods can accurately recognize musical chords from complex audio signals that contain vocal and drum sounds. The main problem is that the chordboundary-detection and chord-symbol-identification processes are inseparable because of their mutual dependency. In order to solve this mutual dependency problem, our method generates hypotheses about tuples of chord symbols and chord boundaries, and outputs the most plausible one as the recognition result. The certainty of a hypothesis is evaluated based on three cues: acoustic features, chord progression patterns, and bass sounds. Experimental results show that our method successfully recognized chords in seven popular music songs; the average accuracy of the results was around 77%.

Tetsuya Ogata - One of the best experts on this subject based on the ideXlab platform.

  • ICASSP (1) - Integration and Adaptation of Harmonic and Inharmonic Models for Separating Polyphonic Musical Signals
    2007 IEEE International Conference on Acoustics Speech and Signal Processing - ICASSP '07, 2007
    Co-Authors: Katsutoshi Itoyama, Masataka Goto, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno
    Abstract:

    This paper describes a sound source separation method for polyphonic sound mixtures of music to build an instrument equalizer for remixing multiple tracks separated from compact-Disc Recordings by changing the volume level of each track. Although such mixtures usually include both harmonic and inharmonic sounds, the difficulties in dealing with both types of sounds together have not been addressed in most previous methods that have focused on either of the two types separately. We therefore developed an integrated weighted-mixture model consisting of both harmonic-structure and inharmonic-structure tone models (generative models for the power spectrogram). On the basis of the MAP estimation using the EM algorithm, we estimated all model parameters of this integrated model under several original constraints for preventing over-training and maintaining intra-instrument consistency. Using standard MIDI files as prior information of the model parameters, we applied this model to compact-Disc Recordings and achieved the instrument equalizer.

  • ISMIR - Automatic Chord Transcription with Concurrent Recognition of Chord Symbols and Boundaries.
    2004
    Co-Authors: Takuya Yoshioka, Kazunori Komatani, Tetsuya Ogata, Tetsuro Kitahara, Hiroshi G. Okuno
    Abstract:

    This paper describes a method that recognizes musical chords from real-world audio signals in compact-Disc Recordings. The automatic recognition of musical chords is necessary for music information retrieval (MIR) systems, since the chord sequences of musical pieces capture the characteristics of their accompaniments. None of the previous methods can accurately recognize musical chords from complex audio signals that contain vocal and drum sounds. The main problem is that the chordboundary-detection and chord-symbol-identification processes are inseparable because of their mutual dependency. In order to solve this mutual dependency problem, our method generates hypotheses about tuples of chord symbols and chord boundaries, and outputs the most plausible one as the recognition result. The certainty of a hypothesis is evaluated based on three cues: acoustic features, chord progression patterns, and bass sounds. Experimental results show that our method successfully recognized chords in seven popular music songs; the average accuracy of the results was around 77%.