The Experts below are selected from a list of 1104 Experts worldwide ranked by ideXlab platform

Miguel A Carreiraperpinan - One of the best experts on this subject based on the ideXlab platform.

  • regularising an adaptation algorithm for Tongue Shape models
    International Conference on Acoustics Speech and Signal Processing, 2012
    Co-Authors: Mohsen Farhadloo, Miguel A Carreiraperpinan
    Abstract:

    Realistic data-driven models of the Tongue Shape can be obtained by learning a nonlinear mapping from Tongue landmarks to full contours, trained on a dataset of thousands of contours. Semiautomatic contour extraction from ultrasound takes a lot of time and effort from an expert, so practically it is preferable to adapt a reference model given just a few contours from the new speaker. However, adaptation with very few contours is unreliable and prone to overfitting. We study several forms of regularisation to constrain the adaptation, and determine the optimal amount of regularisation by leave-one-out cross-validation. Our results show that good accuracy models can be found reliably with no user intervention.

  • learning and adaptation of a Tongue Shape modelwith missing data
    International Conference on Acoustics Speech and Signal Processing, 2012
    Co-Authors: Mohsen Farhadloo, Miguel A Carreiraperpinan
    Abstract:

    Using data-driven techniques and ultrasound data, it is possible to learn models that reconstruct the Tongue Shape of a speaker with submillimetric accuracy given the location of 3–4 fleshpoints, and to adapt these models to a new speaker for which little data is available. In practice, Tongue contours extracted from ultrasound imaging are often incomplete because of shadowing, noise and other factors. We extend these models to deal with missing data during learning and adaptation, and show that submillimetric accuracy can still be achieved even with relatively large amounts of missing data.

  • adaptation of a Tongue Shape model by local feature transformations
    Conference of the International Speech Communication Association, 2010
    Co-Authors: Chao Qin, Miguel A Carreiraperpinan, Mohsen Farhadloo
    Abstract:

    Reconstructing the full contour of the Tongue from the position of 3 to 4 landmarks on it is useful in articulatory speech work. This can be done with submillimetric accuracy us- ing nonlinear predictive mappings trained on hundreds or thousands of contours extracted from ultrasound images. Collecting and segmenting this amount of data from a speaker is difficult, so a more practical solution is to adapt a well-t rained model from a reference speaker to a new speaker using a small amount of data from the latter. Previous work proposed an adaptation model with only 6 parameters and demonstrated fast, accurate re- sults using data from one speaker only. However, the estimates of this model are biased, and we show that, when adapting to a different speaker, its performance stagnates quickly with the amount of adaptation data. We then propose an unbiased adaptation approach, based on local transformations at each contour point, that achieves a significantly lower reconstruction error with a moderate amount of adaptation data.

Mohsen Farhadloo - One of the best experts on this subject based on the ideXlab platform.

  • regularising an adaptation algorithm for Tongue Shape models
    International Conference on Acoustics Speech and Signal Processing, 2012
    Co-Authors: Mohsen Farhadloo, Miguel A Carreiraperpinan
    Abstract:

    Realistic data-driven models of the Tongue Shape can be obtained by learning a nonlinear mapping from Tongue landmarks to full contours, trained on a dataset of thousands of contours. Semiautomatic contour extraction from ultrasound takes a lot of time and effort from an expert, so practically it is preferable to adapt a reference model given just a few contours from the new speaker. However, adaptation with very few contours is unreliable and prone to overfitting. We study several forms of regularisation to constrain the adaptation, and determine the optimal amount of regularisation by leave-one-out cross-validation. Our results show that good accuracy models can be found reliably with no user intervention.

  • learning and adaptation of a Tongue Shape modelwith missing data
    International Conference on Acoustics Speech and Signal Processing, 2012
    Co-Authors: Mohsen Farhadloo, Miguel A Carreiraperpinan
    Abstract:

    Using data-driven techniques and ultrasound data, it is possible to learn models that reconstruct the Tongue Shape of a speaker with submillimetric accuracy given the location of 3–4 fleshpoints, and to adapt these models to a new speaker for which little data is available. In practice, Tongue contours extracted from ultrasound imaging are often incomplete because of shadowing, noise and other factors. We extend these models to deal with missing data during learning and adaptation, and show that submillimetric accuracy can still be achieved even with relatively large amounts of missing data.

  • ICASSP - Learning and adaptation of a Tongue Shape modelwith missing data
    2012 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2012
    Co-Authors: Mohsen Farhadloo, Miguel Á. Carreira-perpiñán
    Abstract:

    Using data-driven techniques and ultrasound data, it is possible to learn models that reconstruct the Tongue Shape of a speaker with submillimetric accuracy given the location of 3–4 fleshpoints, and to adapt these models to a new speaker for which little data is available. In practice, Tongue contours extracted from ultrasound imaging are often incomplete because of shadowing, noise and other factors. We extend these models to deal with missing data during learning and adaptation, and show that submillimetric accuracy can still be achieved even with relatively large amounts of missing data.

  • adaptation of a Tongue Shape model by local feature transformations
    Conference of the International Speech Communication Association, 2010
    Co-Authors: Chao Qin, Miguel A Carreiraperpinan, Mohsen Farhadloo
    Abstract:

    Reconstructing the full contour of the Tongue from the position of 3 to 4 landmarks on it is useful in articulatory speech work. This can be done with submillimetric accuracy us- ing nonlinear predictive mappings trained on hundreds or thousands of contours extracted from ultrasound images. Collecting and segmenting this amount of data from a speaker is difficult, so a more practical solution is to adapt a well-t rained model from a reference speaker to a new speaker using a small amount of data from the latter. Previous work proposed an adaptation model with only 6 parameters and demonstrated fast, accurate re- sults using data from one speaker only. However, the estimates of this model are biased, and we show that, when adapting to a different speaker, its performance stagnates quickly with the amount of adaptation data. We then propose an unbiased adaptation approach, based on local transformations at each contour point, that achieves a significantly lower reconstruction error with a moderate amount of adaptation data.

  • INTERSPEECH - Adaptation of a Tongue Shape model by local feature transformations.
    2010
    Co-Authors: Chao Qin, Miguel Á. Carreira-perpiñán, Mohsen Farhadloo
    Abstract:

    Reconstructing the full contour of the Tongue from the position of 3 to 4 landmarks on it is useful in articulatory speech work. This can be done with submillimetric accuracy us- ing nonlinear predictive mappings trained on hundreds or thousands of contours extracted from ultrasound images. Collecting and segmenting this amount of data from a speaker is difficult, so a more practical solution is to adapt a well-t rained model from a reference speaker to a new speaker using a small amount of data from the latter. Previous work proposed an adaptation model with only 6 parameters and demonstrated fast, accurate re- sults using data from one speaker only. However, the estimates of this model are biased, and we show that, when adapting to a different speaker, its performance stagnates quickly with the amount of adaptation data. We then propose an unbiased adaptation approach, based on local transformations at each contour point, that achieves a significantly lower reconstruction error with a moderate amount of adaptation data.

Natalia Zharkova - One of the best experts on this subject based on the ideXlab platform.

  • Development of the voiceless sibilant fricative contrast in three-year-olds: an ultrasound and acoustic study.
    Journal of child language, 2020
    Co-Authors: Natalia Zharkova
    Abstract:

    The study analysed spectral and Tongue Shape dynamics of voiceless alveolar and postalveolar fricatives produced by ten children learning Scottish English. Synchronised ultrasound Tongue imaging data and acoustic data were used to characterise children's productions of the phonemic contrast. Six children had consistently accurate productions of both fricative targets, with some cross-consonant phonetic differences in the direction previously demonstrated for older children and adults, as well as some immature acoustic and articulatory dynamic patterns. Instrumental analyses made it possible to describe Tongue Shape for phonemic errors and phonetically distorted realisations. There was some evidence of articulatory contrast in production preceding contrast in perception. The observed patterns can be explained by the complex articulatory demands on the fricative production, in combination with the developing control of articulators. The paper concludes by discussing the implications of the results for phonological theory and for speech therapy practice.

  • An Ultrasound Study of the Development of Lingual Coarticulation during Childhood.
    Phonetica, 2018
    Co-Authors: Natalia Zharkova
    Abstract:

    BACKGROUND/AIMS There is growing evidence that coarticulation development is protracted and segment-specific, and yet very little information is available on the changes in the extent of coarticulation across different phonemes throughout childhood. This study describes lingual coarticulatory patterns in 6 age groups of Scottish English-speaking children between 3 and 13 years old. METHODS Vowelon-consonant anticipatory coarticulation was analysed using ultrasound imaging data on Tongue Shape from 4 consonants that differ in the degree of constraint, i.e., the extent of articulatory demand, on the Tongue. RESULTS Consonant-specific age-related patterns are reported, with consonants that have more demands on the Tongue reaching adolescent-like levels of coarticulation in older age groups. Within-speaker variability in Tongue Shape decreases with increasing age. CONCLUSION Reduced coarticulation in the youngest age group may be due to insufficient Tongue differentiation. Immature patterns for lingual consonants in 5- to 11-year-olds are explained by the goal of producing the consonant target overriding the goal of coarticulating the consonant with the following vowel.

  • The dynamics of voiceless sibilant fricative production in children between 7 and 13 years old: An ultrasound and acoustic study.
    The Journal of the Acoustical Society of America, 2018
    Co-Authors: Natalia Zharkova, William J. Hardcastle, Fiona Gibbon
    Abstract:

    This study reports on dynamic Tongue Shape and spectral characteristics of sibilant fricatives /s/ and /ʃ/ in Scottish English speaking children aged between 7 and 13 years old. The sequences /əCa/ and /əCi/ were produced by 40 children, with ten participants in each age group, and two-year intervals between successive groups. Productions of the same sequences by ten adults were used for comparison with the children's data. Quantitative dynamic analyses were carried out on spectral information and on ultrasound imaging data on Tongue Shape. All age groups differentiated between the two consonants in the fricative centroid and in Tongue Shape. Vowel-on-consonant effects showed consonant-specific patterns across age groups without a consistent increase or decrease in the extent of coarticulation with increasing age. The extent of discriminability between the two fricatives increased with age on both acoustic and articulatory measures. Younger speakers were generally more variable than older speakers. Complementary findings from the centroid and Tongue Shape measures suggest that age-related differences are due to the ongoing maturation of controlling the Tongue in coordination with other articulators, particularly the jaw, throughout childhood.

  • Voiceless alveolar stop coarticulation in typically developing 5-year-olds and 13-year-olds.
    Clinical linguistics & phonetics, 2017
    Co-Authors: Natalia Zharkova
    Abstract:

    In this study, vowel-on-consonant lingual coarticulation at [t] closure offset was compared in 5-year-old children and 13-year-old adolescents. The study aimed to establish whether, by the end of the closure, children from the younger age group adjust the Tongue Shape to the following vowels to the same extent as adolescents. Ten 5-year-olds and ten 13-year-olds, all speakers of Scottish Standard English, produced [t]-vowel syllables with the vowels [i] and [a], in a carrier phrase. Measures of Tongue Shape based on midsagittal ultrasound imaging data were used to compare anticipatory coarticulation and within-speaker variability across groups. Both age groups changed the extent of Tongue dorsum bunching in order to coarticulate the consonant with the following vowels. The 5-year-old children, unlike the adolescents, did not consistently modify the bunching location within the Tongue curve to accommodate the Tongue Shape to that of the upcoming vowel. Token-to-token variability was significantly greater in the younger age group. The results suggest that vowel-on-[t] coarticulatory patterns produced by typically developing children are affected by the development of motor control, with articulatory constraints on the Tongue limiting the extent of lingual coarticulation in 5-year-old children. The findings on typical coarticulation development are relevant for clinical practice, and they highlight the need for more detailed descriptions of how phonetic characteristics of speech sounds affect coarticulation throughout childhood.

  • Using ultrasound Tongue imaging to identify covert contrasts in children's speech.
    Clinical linguistics & phonetics, 2016
    Co-Authors: Natalia Zharkova, Fiona Gibbon, Alice Lee
    Abstract:

    Ultrasound Tongue imaging has become a promising technique for detecting covert contrasts, due to the developments in data analysis methods that allow for processing information on Tongue Shape from young children. An important feature concerning analyses of ultrasound data from children who are likely to produce covert contrasts is that the data are likely to be collected without head-to-transducer stabilisation, due to the speakers' age. This article is a review of the existing methods applicable in analysing data from non-stabilised recordings. The article describes some of the challenges of ultrasound data collection from children, and analysing these data, as well as possible ways to address those challenges. Additionally, there are examples from typical and disordered productions featuring covert contrasts, with illustrations of quantifying differences in Tongue Shape between target speech sounds.

Miguel Á. Carreira-perpiñán - One of the best experts on this subject based on the ideXlab platform.

  • ICASSP - Learning and adaptation of a Tongue Shape modelwith missing data
    2012 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2012
    Co-Authors: Mohsen Farhadloo, Miguel Á. Carreira-perpiñán
    Abstract:

    Using data-driven techniques and ultrasound data, it is possible to learn models that reconstruct the Tongue Shape of a speaker with submillimetric accuracy given the location of 3–4 fleshpoints, and to adapt these models to a new speaker for which little data is available. In practice, Tongue contours extracted from ultrasound imaging are often incomplete because of shadowing, noise and other factors. We extend these models to deal with missing data during learning and adaptation, and show that submillimetric accuracy can still be achieved even with relatively large amounts of missing data.

  • INTERSPEECH - Adaptation of a Tongue Shape model by local feature transformations.
    2010
    Co-Authors: Chao Qin, Miguel Á. Carreira-perpiñán, Mohsen Farhadloo
    Abstract:

    Reconstructing the full contour of the Tongue from the position of 3 to 4 landmarks on it is useful in articulatory speech work. This can be done with submillimetric accuracy us- ing nonlinear predictive mappings trained on hundreds or thousands of contours extracted from ultrasound images. Collecting and segmenting this amount of data from a speaker is difficult, so a more practical solution is to adapt a well-t rained model from a reference speaker to a new speaker using a small amount of data from the latter. Previous work proposed an adaptation model with only 6 parameters and demonstrated fast, accurate re- sults using data from one speaker only. However, the estimates of this model are biased, and we show that, when adapting to a different speaker, its performance stagnates quickly with the amount of adaptation data. We then propose an unbiased adaptation approach, based on local transformations at each contour point, that achieves a significantly lower reconstruction error with a moderate amount of adaptation data.

James M. Scobbie - One of the best experts on this subject based on the ideXlab platform.

  • The Effects of Syllable and Utterance Position on Tongue Shape and Gestural Magnitude in /l/ and /r/
    2019
    Co-Authors: Eleanor Lawson, Gregory Leplatre, Jane Stuart-smith, James M. Scobbie
    Abstract:

    This paper is an ultrasound-based articulatory study of the impact of syllable-position and utterance position on Tongue Shape and Tongue-gesture magnitude in liquid consonants in American, Irish and Scottish English. Mixed effects modelling was used to analyse variation in normalised Tonguegesture magnitude for /r/ and /l/ in syllable-onset and coda position and in utterance-initial, medial and final position. Variation between onset and coda mean midsagittal Tongue surfaces was also quantified using normalised root-mean-square distances, and patterns of articulatory onset-coda allophony were identified. Despite the fact that some speakers in all varieties used tip-up /r/ in syllable-onset position and bunched /r/ in coda position, RMS distance results show greater degrees of similarity between onset and coda /r/ than between onset and coda /l/. Gesture magnitude was significantly reduced for both /l/ and /r/ in coda position. Utterance position had a significant effect on /l/ only.

  • Tongue Shape dynamics in swallowing using sagittal ultrasound
    Dysphagia, 2019
    Co-Authors: Mai Ohkubo, James M. Scobbie
    Abstract:

    Ultrasound imaging is simple, repeatable, gives real-time feedback, and its dynamic soft tissue imaging may make it superior to other modalities for swallowing research. We tested this hypothesis and measured certain spatial and dynamic aspects of the swallowing to investigate its efficacy. Eleven healthy adults wearing a headset to stabilize the probe participated in the study. Both thickened and thin liquids were used, and liquid bolus volumes of 10 and 25 ml were administered to the subjects by using a cup. The Tongue’s surface was traced as a spline superimposed on a fan-Shaped measurement space for every image from the time at which the Tongue blade started moving up toward the palate at the start of swallowing to the time when the entire Tongue was in contact with the palate. To measure depression depth, the distance (in mm) was measured along each radial fan line from the location at which the Tongue’s surface spline intersected the fan line to the point where the hard palate intersected the fan line at each timepoint. There were differences between individual participants in the imageability of the swallow, and so we defined quantitatively “measureable” and “unmeasurable” types. The most common type was measureable, in which we could find a clear bolus depression in the cupped Tongue’s surface. Indeed, with 10 ml of thin liquids, we were able to find and measure the depression depth for all participants. The average maximum radial distance from the palate to the Tongue’s surface was 20.9 mm (median) (IQR: 4.3 mm) for swallowing 10 ml of thin liquid compared to 24.6 mm (IQR: 3.3 mm) for 25 ml of thin liquid swallow (p < 0.001). We conclude that it is possible to use ultrasound imaging of the Tongue to capture spatial aspects of swallowing.

  • Tongue Shape Dynamics in Swallowing Using Sagittal Ultrasound
    Dysphagia, 2018
    Co-Authors: Mai Ohkubo, James M. Scobbie
    Abstract:

    Ultrasound imaging is simple, repeatable, gives real-time feedback, and its dynamic soft tissue imaging may make it superior to other modalities for swallowing research. We tested this hypothesis and measured certain spatial and dynamic aspects of the swallowing to investigate its efficacy. Eleven healthy adults wearing a headset to stabilize the probe participated in the study. Both thickened and thin liquids were used, and liquid bolus volumes of 10 and 25 ml were administered to the subjects by using a cup. The Tongue's surface was traced as a spline superimposed on a fan-Shaped measurement space for every image from the time at which the Tongue blade started moving up toward the palate at the start of swallowing to the time when the entire Tongue was in contact with the palate. To measure depression depth, the distance (in mm) was measured along each radial fan line from the location at which the Tongue's surface spline intersected the fan line to the point where the hard palate intersected the fan line at each timepoint. There were differences between individual participants in the imageability of the swallow, and so we defined quantitatively "measureable" and "unmeasurable" types. The most common type was measureable, in which we could find a clear bolus depression in the cupped Tongue's surface. Indeed, with 10 ml of thin liquids, we were able to find and measure the depression depth for all participants. The average maximum radial distance from the palate to the Tongue's surface was 20.9 mm (median) (IQR: 4.3 mm) for swallowing 10 ml of thin liquid compared to 24.6 mm (IQR: 3.3 mm) for 25 ml of thin liquid swallow (p 

  • A mimicry study of adaptation towards socially-salient Tongue Shape variants
    2014
    Co-Authors: Eleanor Lawson, Jane Stuart-smith, James M. Scobbie
    Abstract:

    We know that fine phonetic variation is exploited by speakers to construct and index social identity (Hay and Drager 2007). Sociophonetic work to date has tended to focus on acoustic analysis, e.g. Docherty and Foulkes (1999); however, some aspects of speech production are not readily recoverable from an acoustic analysis. New articulatory analysis techniques, such as ultrasound Tongue imaging (UTI), have helped to identify seemingly covert aspects of speech articulation, which pattern consistently with indexical factors, e.g. underlyingly, Scottish English middle-class and working-class coda /r/, have radically different Tongue Shapes and Tongue gesture timings (Lawson, Scobbie and Stuart-Smith, 2011). This articulatory variation has gone unidentified, despite decades of auditory and acoustic analysis (Romaine, 1979; Speitel and Johnston, 1983; Stuart-Smith, 2007). UTI revealed that middle-class speakers tend to produce bunched variants of postvocalic /r/, while working-class speakers tend to produce Tongue-tip raised variants (Lawson, Stuart-Smith and Scobbie 2011). We present the results of an ultrasound-based, speech-mimicry pilot study, which investigates how subtle articulatory information might be passed from speaker to speaker. Baseline articulatory information on /r/ was gathered for three Central-Scottish female, middle-class, pilot participants, who all used bunched /r/ variants in baseline. Participants mimicked audio-only stimuli, extracted from a socially-stratified audio-ultrasound corpus collected in Glasgow, Scotland. Analysis showed a range of mimicking behaviours from pilot participants including no modification from baseline, accurate discrimination between middle-class and working-class stimuli and adaptation from baseline, but failure to discriminate between middle-class and working class stimuli. This small-scale study provides new insights into the transmission of phonetic variation.

  • The social stratification of Tongue Shape for postvocalic /r/ in Scottish English
    Journal of Sociolinguistics, 2011
    Co-Authors: Eleanor Lawson, James M. Scobbie, Jane Stuart-smith
    Abstract:

    The sociolinguistic modelling of phonological variation and change is almost exclusively based on auditory and acoustic analyses of speech. One phenomenon which has proved elusive when considered in these ways is the variation in postvocalic /r/ in Scottish English. This study therefore shifts to speech production: we present a socioarticulatory study of variation of postvocalic /r/ in CVr words, using a socially-stratified ultrasound Tongue imaging corpus of speech collected in eastern central Scotland in 2008. Our results show social stratification of /r/ at the articulatory level, with middle-class speakers using bunched articulations, while working-class speakers use greater proportions of Tongue-tip and Tongue-front raised variants. Unlike articulatory variation of /r/ in American English, the articulatory variants in our Scottish English corpus are both auditorily distinct from one another, and correlate with strong and weak ends of an auditory rhotic continuum, which also shows clear social stratification