The Experts below are selected from a list of 210 Experts worldwide ranked by ideXlab platform

David J C Mackay - One of the best experts on this subject based on the ideXlab platform.

  • ticker an adaptive single switch text entry method for visually impaired users
    IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019
    Co-Authors: Per Ola Kristensson, David J C Mackay
    Abstract:

    Ticker is a probabilistic stereophonic single-switch text entry method for visually-impaired users with motor disabilities who rely on single-switch scanning systems to communicate. Such scanning systems are sensitive to a variety of noise sources, which are inevitably introduced in practical use of single-switch systems. Ticker uses a novel interaction model based on stereophonic sound coupled with statistical models for robust inference of the user's intended text in the presence of noise. As a consequence of its design, Ticker is resilient to noise and therefore a practical solution for single-switch scanning systems. Ticker's performance is validated using a combination of simulations and empirical user studies.

  • Ticker: An Adaptive Single-Switch Text Entry Method for Visually Impaired Users.
    'Organisation for Economic Co-Operation and Development (OECD)', 2019
    Co-Authors: Nel Emli-mari, Kristensson, Per Ola, David J C Mackay
    Abstract:

    Ticker is a novel probabilistic stereophonic single-switch text entry method for visually-impaired users with motor disabilities who rely on single-switch scanning systems to communicate. Such scanning systems are sensitive to a variety of noise sources, which are inevitably introduced in practical use of single-switch systems. Ticker uses a novel interaction model based on stereophonic sound coupled with statistical models for robust inference of the users intended text in the presence of noise. As a consequence of its design, Ticker is resilient to noise and therefore a practical solution for single-switch scanning systems. Tickers performance is validated using a combination of simulations and empirical user studies.The work was funded by the Aegis EU project and the Gatsby Charitable Foundation

  • the statistical model for ticker an adaptive single switch text entry method for visually impaired users
    arXiv: Artificial Intelligence, 2018
    Co-Authors: Per Ola Kristensson, David J C Mackay
    Abstract:

    This paper presents the statistical model for Ticker [1], a novel probabilistic stereophonic single-switch text entry method for visually-impaired users with motor disabilities who rely on single-switch scanning systems to communicate. All terminology and notation are defined in [1].

Per Ola Kristensson - One of the best experts on this subject based on the ideXlab platform.

  • ticker an adaptive single switch text entry method for visually impaired users
    IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019
    Co-Authors: Per Ola Kristensson, David J C Mackay
    Abstract:

    Ticker is a probabilistic stereophonic single-switch text entry method for visually-impaired users with motor disabilities who rely on single-switch scanning systems to communicate. Such scanning systems are sensitive to a variety of noise sources, which are inevitably introduced in practical use of single-switch systems. Ticker uses a novel interaction model based on stereophonic sound coupled with statistical models for robust inference of the user's intended text in the presence of noise. As a consequence of its design, Ticker is resilient to noise and therefore a practical solution for single-switch scanning systems. Ticker's performance is validated using a combination of simulations and empirical user studies.

  • the statistical model for ticker an adaptive single switch text entry method for visually impaired users
    arXiv: Artificial Intelligence, 2018
    Co-Authors: Per Ola Kristensson, David J C Mackay
    Abstract:

    This paper presents the statistical model for Ticker [1], a novel probabilistic stereophonic single-switch text entry method for visually-impaired users with motor disabilities who rely on single-switch scanning systems to communicate. All terminology and notation are defined in [1].

Ziwei Liu - One of the best experts on this subject based on the ideXlab platform.

  • sep stereo visually guided stereophonic audio generation by associating source separation
    European Conference on Computer Vision, 2020
    Co-Authors: Hang Zhou, Dahua Lin, Xiaogang Wang, Ziwei Liu
    Abstract:

    Stereophonic audio is an indispensable ingredient to enhance human auditory experience. Recent research has explored the usage of visual information as guidance to generate binaural or ambisonic audio from mono ones with stereo supervision. However, this fully supervised paradigm suffers from an inherent drawback: the recording of stereophonic audio usually requires delicate devices that are expensive for wide accessibility. To overcome this challenge, we propose to leverage the vastly available mono data to facilitate the generation of stereophonic audio. Our key observation is that the task of visually indicated audio separation also maps independent audios to their corresponding visual positions, which shares a similar objective with stereophonic audio generation. We integrate both stereo generation and source separation into a unified framework, Sep-Stereo, by considering source separation as a particular type of audio spatialization. Specifically, a novel associative pyramid network architecture is carefully designed for audio-visual feature fusion. Extensive experiments demonstrate that our framework can improve the stereophonic audio generation results while performing accurate sound separation with a shared backbone (Code, models and demo video are available at https://hangz-nju-cuhk.github.io/projects/Sep-Stereo.).

  • sep stereo visually guided stereophonic audio generation by associating source separation
    European Conference on Computer Vision, 2020
    Co-Authors: Hang Zhou, Dahua Lin, Xiaogang Wang, Ziwei Liu
    Abstract:

    Stereophonic audio is an indispensable ingredient to enhance human auditory experience. Recent research has explored the usage of visual information as guidance to generate binaural or ambisonic audio from mono ones with stereo supervision. However, this fully supervised paradigm suffers from an inherent drawback: the recording of stereophonic audio usually requires delicate devices that are expensive for wide accessibility. To overcome this challenge, we propose to leverage the vastly available mono data to facilitate the generation of stereophonic audio. Our key observation is that the task of visually indicated audio separation also maps independent audios to their corresponding visual positions, which shares a similar objective with stereophonic audio generation. We integrate both stereo generation and source separation into a unified framework, Sep-Stereo, by considering source separation as a particular type of audio spatialization. Specifically, a novel associative pyramid network architecture is carefully designed for audio-visual feature fusion. Extensive experiments demonstrate that our framework can improve the stereophonic audio generation results while performing accurate sound separation with a shared backbone.

Sofiene Affes - One of the best experts on this subject based on the ideXlab platform.

  • direction of arrival estimation using the parameterized spatial correlation matrix
    IEEE Transactions on Audio Speech and Language Processing, 2007
    Co-Authors: Jacek P Dmochowski, Jacob Benesty, Sofiene Affes
    Abstract:

    The estimation of the direction-of-arrival (DOA) of one or more acoustic sources is an area that has generated much interest in recent years, with applications like automatic video camera steering and multiparty stereophonic teleconferencing entering the market. DOA estimation algorithms are hindered by the effects of background noise and reverberation. Methods based on the time-differences-of-arrival (TDOA) are commonly used to determine the azimuth angle of arrival of an acoustic source. TDOA-based methods compute each relative delay using only two microphones, even though additional microphones are usually available. This paper deals with DOA estimation based on spatial spectral estimation, and establishes the parameterized spatial correlation matrix as the framework for this class of DOA estimators. This matrix jointly takes into account all pairs of microphones, and is at the heart of several broadband spatial spectral estimators, including steered-response power (SRP) algorithms. This paper reviews and evaluates these broadband spatial spectral estimators, comparing their performance to TDOA-based locators. In addition, an eigenanalysis of the parameterized spatial correlation matrix is performed and reveals that such analysis allows one to estimate the channel attenuation from factors such as uncalibrated microphones. This estimate generalizes the broadband minimum variance spatial spectral estimator to more general signal models. A DOA estimator based on the multichannel cross correlation coefficient (MCCC) is also proposed. The performance of all proposed algorithms is included in the evaluation. It is shown that adding extra microphones helps combat the effects of background noise and reverberation. Furthermore, the link between accurate spatial spectral estimation and corresponding DOA estimation is investigated. The application of the minimum variance and MCCC methods to the spatial spectral estimation problem leads to better resolution than that of the commonly used fixed-weighted SRP spectrum. However, this increased spatial spectral resolution does not always translate to more accurate DOA estimation

  • direction of arrival estimation using eigenanalysis of the parameterized spatial correlation matrix
    International Conference on Acoustics Speech and Signal Processing, 2007
    Co-Authors: Jacek P Dmochowski, Jacob Benesty, Sofiene Affes
    Abstract:

    The estimation of the direction-of-arrival (DOA) of one or more acoustic sources is an area that has generated much interest in recent years, with applications like automatic video camera steering and multi-party stereophonic teleconferencing entering the market. Time-difference-of-arrival (TDOA) based methods compute each relative delay using only two microphones, even though additional microphones are usually available, and thus suffer from the effects of background noise and reverberation. This paper deals with DOA estimation based on spatial spectral estimation, and proposes a novel DOA estimator based on the eigenvalues of the parameterized spatial correlation matrix. Simulation results confirm the ability of the proposed method to provide reliable estimates even in heavily reverberant environments.

Hang Zhou - One of the best experts on this subject based on the ideXlab platform.

  • sep stereo visually guided stereophonic audio generation by associating source separation
    European Conference on Computer Vision, 2020
    Co-Authors: Hang Zhou, Dahua Lin, Xiaogang Wang, Ziwei Liu
    Abstract:

    Stereophonic audio is an indispensable ingredient to enhance human auditory experience. Recent research has explored the usage of visual information as guidance to generate binaural or ambisonic audio from mono ones with stereo supervision. However, this fully supervised paradigm suffers from an inherent drawback: the recording of stereophonic audio usually requires delicate devices that are expensive for wide accessibility. To overcome this challenge, we propose to leverage the vastly available mono data to facilitate the generation of stereophonic audio. Our key observation is that the task of visually indicated audio separation also maps independent audios to their corresponding visual positions, which shares a similar objective with stereophonic audio generation. We integrate both stereo generation and source separation into a unified framework, Sep-Stereo, by considering source separation as a particular type of audio spatialization. Specifically, a novel associative pyramid network architecture is carefully designed for audio-visual feature fusion. Extensive experiments demonstrate that our framework can improve the stereophonic audio generation results while performing accurate sound separation with a shared backbone (Code, models and demo video are available at https://hangz-nju-cuhk.github.io/projects/Sep-Stereo.).

  • sep stereo visually guided stereophonic audio generation by associating source separation
    European Conference on Computer Vision, 2020
    Co-Authors: Hang Zhou, Dahua Lin, Xiaogang Wang, Ziwei Liu
    Abstract:

    Stereophonic audio is an indispensable ingredient to enhance human auditory experience. Recent research has explored the usage of visual information as guidance to generate binaural or ambisonic audio from mono ones with stereo supervision. However, this fully supervised paradigm suffers from an inherent drawback: the recording of stereophonic audio usually requires delicate devices that are expensive for wide accessibility. To overcome this challenge, we propose to leverage the vastly available mono data to facilitate the generation of stereophonic audio. Our key observation is that the task of visually indicated audio separation also maps independent audios to their corresponding visual positions, which shares a similar objective with stereophonic audio generation. We integrate both stereo generation and source separation into a unified framework, Sep-Stereo, by considering source separation as a particular type of audio spatialization. Specifically, a novel associative pyramid network architecture is carefully designed for audio-visual feature fusion. Extensive experiments demonstrate that our framework can improve the stereophonic audio generation results while performing accurate sound separation with a shared backbone.