The Experts below are selected from a list of 7299 Experts worldwide ranked by ideXlab platform

Junsong Yuan - One of the best experts on this subject based on the ideXlab platform.

  • sibnet sibling Convolutional Encoder for video captioning
    IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021
    Co-Authors: Sheng Liu, Zhou Ren, Junsong Yuan
    Abstract:

    Visual captioning, the task of describing an image or a video using one or few sentences, is a challenging task owing to the complexity of understanding the copious visual information and describing it using natural language. Motivated by the success of applying neural networks for machine translation, previous work applies sequence to sequence learning to translate videos into sentences. In this work, different from previous work that encodes visual information using a single flow, we introduce a novel Sibling Convolutional Encoder (SibNet) for visual captioning, which employs a dual-branch architecture to collaboratively encode videos. The first content branch encodes visual content information of the video with an autoEncoder, capturing the visual appearance information of the video as other networks often do. While the second semantic branch encodes semantic information of the video via visual-semantic joint embedding, which brings complementary representation by considering the semantics when extracting features from videos. Then both branches are effectively combined with soft-attention mechanism and finally fed into a RNN decoder to generate captions. With our SibNet explicitly capturing both content and semantic information, the proposed model can better represent the rich information in videos. To validate the advantages of the proposed model, we conduct experiments on two benchmarks for video captioning, YouTube2Text and MSR-VTT. Our results demonstrate that the proposed SibNet consistently outperforms existing methods across different evaluation metrics.

  • sibnet sibling Convolutional Encoder for video captioning
    IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020
    Co-Authors: Sheng Liu, Zhou Ren, Junsong Yuan
    Abstract:

    Visual captioning, the task of describing an image or a video using one or few sentences, is challenging owing to the complexity of understanding copious visual information and describing it using natural language. Motivated by the success neural machine translation, previous work applies sequence to sequence learning to translate videos into sentences. In this work, different from previous work that encodes visual information using a single flow, we introduce a novel Sibling Convolutional Encoder (SibNet) for visual captioning, which employs a two-branch architecture to collaboratively encode videos. The first content branch encodes visual content information of the video with an autoEncoder, capturing visual appearance information of the video as other networks often do. While the second semantic branch encodes semantic information of the video via visual-semantic joint embedding, which brings complementary representation by considering the semantics when extracting features from videos. Then both branches are effectively combined with soft-attention mechanism and finally fed into a RNN decoder to generate captions. With our SibNet explicitly capturing both content and semantic information, the proposed model can better represent rich information in videos. To validate the advantages of SibNet, we conduct experiments on two video captioning benchmarks, YouTube2Text and MSR-VTT. Our results demonstrates that SibNet outperforms existing methods across different evaluation metrics.

  • sibnet sibling Convolutional Encoder for video captioning
    ACM Multimedia, 2018
    Co-Authors: Sheng Liu, Zhou Ren, Junsong Yuan
    Abstract:

    Video captioning is a challenging task owing to the complexity of understanding the copious visual information in videos and describing it using natural language. Different from previous work that encodes video information using a single flow, in this work, we introduce a novel Sibling Convolutional Encoder (SibNet) for video captioning, which utilizes a two-branch architecture to collaboratively encode videos. The first content branch encodes the visual content information of the video via autoEncoder, and the second semantic branch encodes the semantic information by visual-semantic joint embedding. Then both branches are effectively combined with soft-attention mechanism and finally fed into a RNN decoder to generate captions. With our SibNet explicitly capturing both content and semantic information, the proposed method can better represent the rich information in videos. Extensive experiments on YouTube2Text and MSR-VTT datasets validate that the proposed architecture outperforms existing methods by a large margin across different evaluation metrics.

Arman Rahmim - One of the best experts on this subject based on the ideXlab platform.

  • direct attenuation correction of brain pet images using only emission data via a deep Convolutional Encoder decoder deep dac
    European Radiology, 2019
    Co-Authors: Isaac Shiri, Kevin Leung, Pardis Ghafarian, Parham Geramifar, Mehrdad Oveisi, Arman Rahmim, Mostafa Ghelichoghli
    Abstract:

    To obtain attenuation-corrected PET images directly from non-attenuation-corrected images using a Convolutional Encoder-decoder network. Brain PET images from 129 patients were evaluated. The network was designed to map non-attenuation-corrected (NAC) images to pixel-wise continuously valued measured attenuation-corrected (MAC) PET images via an Encoder-decoder architecture. Image quality was evaluated using various evaluation metrics. Image quantification was assessed for 19 radiomic features in 83 brain regions as delineated using the Hammersmith atlas (n30r83). Reliability of measurements was determined using pixel-wise relative errors (RE; %) for radiomic feature values in reference MAC PET images. Peak signal-to-noise ratio (PSNR) and structural similarity index metric (SSIM) values were 39.2 ± 3.65 and 0.989 ± 0.006 for the external validation set, respectively. RE (%) of SUVmean was − 0.10 ± 2.14 for all regions, and only 3 of 83 regions depicted significant differences. However, the mean RE (%) of this region was 0.02 (range, − 0.83 to 1.18). SUVmax had mean RE (%) of − 3.87 ± 2.84 for all brain regions, and 17 regions in the brain depicted significant differences with respect to MAC images with a mean RE of − 3.99 ± 2.11 (range, − 8.46 to 0.76). Homogeneity amongst Haralick-based radiomic features had the highest number (20) of regions with significant differences with a mean RE (%) of 7.22 ± 2.99. Direct AC of PET images using deep Convolutional Encoder-decoder networks is a promising technique for brain PET images. The proposed deep learning method shows significant potential for emission-based AC in PET images with applications in PET/MRI and dedicated brain PET scanners. • We demonstrate direct emission-based attenuation correction of PET images without using anatomical information. • We performed radiomics analysis of 83 brain regions to show robustness of direct attenuation correction of PET images. • Deep learning methods have significant promise for emission-based attenuation correction in PET images with potential applications in PET/MRI and dedicated brain PET scanners.

  • simultaneous attenuation correction and reconstruction of pet images using deep Convolutional Encoder decoder networks from emission data
    The Journal of Nuclear Medicine, 2019
    Co-Authors: Isaac Shiri, Kevin Leung, Pardis Ghafarian, Parham Geramifar, Mehrdad Oveisi, Arman Rahmim
    Abstract:

    1370 Aim: Quantitative positron emission tomography (PET) image reconstruction is challenging due to the attenuation correction needed during the reconstruction process. In this study, we propose and investigate methods that perform attenuation correction and reconstruction of PET images via deep Convolutional Encoder-decoder networks, without using anatomical information for attenuation correction. Methods: Brain PET images from 120 patients were included in our study. Patient data were divided into training, test, and external validation sets with 80, 20 and 20 patients, respectively. We proposed and investigated 3 main methods. First, we directly mapped the non-attenuation-corrected sinogram to original reconstructed image. The second method consisted of two Convolutional Encoder and decoder networks where the first part tries to reconstruct non-attention-corrected images and the second network tries to perform attenuation correction of the generated image in the image space. Third method’s architecture consists of two Convolutional Encoder and decoder networks for end-to-end learning: the first part tries to correct attenuation in sinogram space from non-attenuated corrected sinogram passed as input through the Encoder, and the decoder tries to reconstruct the attenuation-corrected sinogram. This latter second part mapped the sinogram to the pixel-wise continuously-valued measured attenuation-corrected PET images as obtained from reference CT images. Quality of the synthesized images were quantitatively assessed by mean squared error (MSE), peak signal-to-noise ratio (PSNR), and structural Similarity index metrics (SSIM). Image quantification was assessed using SUV bias map and joint histogram of pixel-wise SUV correlation between generated images and reference (attenuated corrected by CT, and reconstruction with iterative reconstruction). Results: With respect to reference PET images, MSE, PSNR and SSIM values were 0.0023±0.0011, 26.85±2.44, 0.85±0.09 and 0.0029±0.0011, 25.66±2.07, 0.82±0.01 and 0.0020±0.0013, 28.45±1.03 ,0.87±0.01 for the first, second and third methods, respectively. Relative error (%) of SUV was -34.12±6.01, 36.12±7.03 and 18.12±8.12 for the first, second and third method respectively. Pixel wise SUV correlation (Pearson correlation, R2) between generated images and reference was 0.71± 0.01, 0.69± 0.05 and 0.81± 0.09 for the first, second and third method, respectively. Conclusions: In this present study, we developed new approaches to attenuation correction and reconstruction of PET images from emission data without using anatomical information for attenuation correction. The highest performance was obtained by the deep neural network architecture that consisted of two Convolutional Encoder and decoder networks where first part performed attenuation correction in sinogram space and the second network reconstructed the image. The present study showed that attenuation correction and reconstruction of PET images using deep Convolutional Encoder-decoder networks from emission data is a promising technique for PET images.

Roberto Cipolla - One of the best experts on this subject based on the ideXlab platform.

  • segnet a deep Convolutional Encoder decoder architecture for image segmentation
    IEEE Transactions on Pattern Analysis and Machine Intelligence, 2017
    Co-Authors: Vijay Badrinarayanan, Alex Kendall, Roberto Cipolla
    Abstract:

    We present a novel and practical deep fully Convolutional neural network architecture for semantic pixel-wise segmentation termed SegNet. This core trainable segmentation engine consists of an Encoder network, a corresponding decoder network followed by a pixel-wise classification layer. The architecture of the Encoder network is topologically identical to the 13 Convolutional layers in the VGG16 network [1] . The role of the decoder network is to map the low resolution Encoder feature maps to full input resolution feature maps for pixel-wise classification. The novelty of SegNet lies is in the manner in which the decoder upsamples its lower resolution input feature map(s). Specifically, the decoder uses pooling indices computed in the max-pooling step of the corresponding Encoder to perform non-linear upsampling. This eliminates the need for learning to upsample. The upsampled maps are sparse and are then convolved with trainable filters to produce dense feature maps. We compare our proposed architecture with the widely adopted FCN [2] and also with the well known DeepLab-LargeFOV [3] , DeconvNet [4] architectures. This comparison reveals the memory versus accuracy trade-off involved in achieving good segmentation performance. SegNet was primarily motivated by scene understanding applications. Hence, it is designed to be efficient both in terms of memory and computational time during inference. It is also significantly smaller in the number of trainable parameters than other competing architectures and can be trained end-to-end using stochastic gradient descent. We also performed a controlled benchmark of SegNet and other architectures on both road scenes and SUN RGB-D indoor scene segmentation tasks. These quantitative assessments show that SegNet provides good performance with competitive inference time and most efficient inference memory-wise as compared to other architectures. We also provide a Caffe implementation of SegNet and a web demo at http://mi.eng.cam.ac.uk/projects/segnet/ .

  • bayesian segnet model uncertainty in deep Convolutional Encoder decoder architectures for scene understanding
    British Machine Vision Conference, 2017
    Co-Authors: Alex Kendall, Vijay Badrinarayanan, Roberto Cipolla
    Abstract:

    © 2017. The copyright of this document resides with its authors. We present a deep learning framework for probabilistic pixel-wise semantic segmentation, which we term Bayesian SegNet. Semantic segmentation is an important tool for visual scene understanding and a meaningful measure of uncertainty is essential for decision making. Our contribution is a practical system which is able to predict pixel-wise class labels with a measure of model uncertainty using Bayesian deep learning. We achieve this by Monte Carlo sampling with dropout at test time to generate a posterior distribution of pixel class labels. In addition, we show that modelling uncertainty improves segmentation performance by 2-3% across a number of datasets and architectures such as SegNet, FCN, Dilation Network and DenseNet.

  • bayesian segnet model uncertainty in deep Convolutional Encoder decoder architectures for scene understanding
    arXiv: Computer Vision and Pattern Recognition, 2015
    Co-Authors: Alex Kendall, Vijay Badrinarayanan, Roberto Cipolla
    Abstract:

    We present a deep learning framework for probabilistic pixel-wise semantic segmentation, which we term Bayesian SegNet. Semantic segmentation is an important tool for visual scene understanding and a meaningful measure of uncertainty is essential for decision making. Our contribution is a practical system which is able to predict pixel-wise class labels with a measure of model uncertainty. We achieve this by Monte Carlo sampling with dropout at test time to generate a posterior distribution of pixel class labels. In addition, we show that modelling uncertainty improves segmentation performance by 2-3% across a number of state of the art architectures such as SegNet, FCN and Dilation Network, with no additional parametrisation. We also observe a significant improvement in performance for smaller datasets where modelling uncertainty is more effective. We benchmark Bayesian SegNet on the indoor SUN Scene Understanding and outdoor CamVid driving scenes datasets.

  • segnet a deep Convolutional Encoder decoder architecture for image segmentation
    arXiv: Computer Vision and Pattern Recognition, 2015
    Co-Authors: Vijay Badrinarayanan, Alex Kendall, Roberto Cipolla
    Abstract:

    We present a novel and practical deep fully Convolutional neural network architecture for semantic pixel-wise segmentation termed SegNet. This core trainable segmentation engine consists of an Encoder network, a corresponding decoder network followed by a pixel-wise classification layer. The architecture of the Encoder network is topologically identical to the 13 Convolutional layers in the VGG16 network. The role of the decoder network is to map the low resolution Encoder feature maps to full input resolution feature maps for pixel-wise classification. The novelty of SegNet lies is in the manner in which the decoder upsamples its lower resolution input feature map(s). Specifically, the decoder uses pooling indices computed in the max-pooling step of the corresponding Encoder to perform non-linear upsampling. This eliminates the need for learning to upsample. The upsampled maps are sparse and are then convolved with trainable filters to produce dense feature maps. We compare our proposed architecture with the widely adopted FCN and also with the well known DeepLab-LargeFOV, DeconvNet architectures. This comparison reveals the memory versus accuracy trade-off involved in achieving good segmentation performance. SegNet was primarily motivated by scene understanding applications. Hence, it is designed to be efficient both in terms of memory and computational time during inference. It is also significantly smaller in the number of trainable parameters than other competing architectures. We also performed a controlled benchmark of SegNet and other architectures on both road scenes and SUN RGB-D indoor scene segmentation tasks. We show that SegNet provides good performance with competitive inference time and more efficient inference memory-wise as compared to other architectures. We also provide a Caffe implementation of SegNet and a web demo at this http URL

  • segnet a deep Convolutional Encoder decoder architecture for robust semantic pixel wise labelling
    Computer Vision and Pattern Recognition, 2015
    Co-Authors: Vijay Badrinarayanan, Ankur Handa, Roberto Cipolla
    Abstract:

    We propose a novel deep architecture, SegNet, for semantic pixel wise image labelling. SegNet has several attractive properties; (i) it only requires forward evaluation of a fully learnt function to obtain smooth label predictions, (ii) with increasing depth, a larger context is considered for pixel labelling which improves accuracy, and (iii) it is easy to visualise the effect of feature activation(s) in the pixel label space at any depth. SegNet is composed of a stack of Encoders followed by a corresponding decoder stack which feeds into a soft-max classification layer. The decoders help map low resolution feature maps at the output of the Encoder stack to full input image size feature maps. This addresses an important drawback of recent deep learning approaches which have adopted networks designed for object categorization for pixel wise labelling. These methods lack a mechanism to map deep layer feature maps to input dimensions. They resort to ad hoc methods to upsample features, e.g. by replication. This results in noisy predictions and also restricts the number of pooling layers in order to avoid too much upsampling and thus reduces spatial context. SegNet overcomes these problems by learning to map Encoder outputs to image pixel labels. We test the performance of SegNet on outdoor RGB scenes from CamVid, KITTI and indoor scenes from the NYU dataset. Our results show that SegNet achieves state-of-the-art performance even without use of additional cues such as depth, video frames or post-processing with CRF models.

Sheng Liu - One of the best experts on this subject based on the ideXlab platform.

  • sibnet sibling Convolutional Encoder for video captioning
    IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021
    Co-Authors: Sheng Liu, Zhou Ren, Junsong Yuan
    Abstract:

    Visual captioning, the task of describing an image or a video using one or few sentences, is a challenging task owing to the complexity of understanding the copious visual information and describing it using natural language. Motivated by the success of applying neural networks for machine translation, previous work applies sequence to sequence learning to translate videos into sentences. In this work, different from previous work that encodes visual information using a single flow, we introduce a novel Sibling Convolutional Encoder (SibNet) for visual captioning, which employs a dual-branch architecture to collaboratively encode videos. The first content branch encodes visual content information of the video with an autoEncoder, capturing the visual appearance information of the video as other networks often do. While the second semantic branch encodes semantic information of the video via visual-semantic joint embedding, which brings complementary representation by considering the semantics when extracting features from videos. Then both branches are effectively combined with soft-attention mechanism and finally fed into a RNN decoder to generate captions. With our SibNet explicitly capturing both content and semantic information, the proposed model can better represent the rich information in videos. To validate the advantages of the proposed model, we conduct experiments on two benchmarks for video captioning, YouTube2Text and MSR-VTT. Our results demonstrate that the proposed SibNet consistently outperforms existing methods across different evaluation metrics.

  • sibnet sibling Convolutional Encoder for video captioning
    IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020
    Co-Authors: Sheng Liu, Zhou Ren, Junsong Yuan
    Abstract:

    Visual captioning, the task of describing an image or a video using one or few sentences, is challenging owing to the complexity of understanding copious visual information and describing it using natural language. Motivated by the success neural machine translation, previous work applies sequence to sequence learning to translate videos into sentences. In this work, different from previous work that encodes visual information using a single flow, we introduce a novel Sibling Convolutional Encoder (SibNet) for visual captioning, which employs a two-branch architecture to collaboratively encode videos. The first content branch encodes visual content information of the video with an autoEncoder, capturing visual appearance information of the video as other networks often do. While the second semantic branch encodes semantic information of the video via visual-semantic joint embedding, which brings complementary representation by considering the semantics when extracting features from videos. Then both branches are effectively combined with soft-attention mechanism and finally fed into a RNN decoder to generate captions. With our SibNet explicitly capturing both content and semantic information, the proposed model can better represent rich information in videos. To validate the advantages of SibNet, we conduct experiments on two video captioning benchmarks, YouTube2Text and MSR-VTT. Our results demonstrates that SibNet outperforms existing methods across different evaluation metrics.

  • sibnet sibling Convolutional Encoder for video captioning
    ACM Multimedia, 2018
    Co-Authors: Sheng Liu, Zhou Ren, Junsong Yuan
    Abstract:

    Video captioning is a challenging task owing to the complexity of understanding the copious visual information in videos and describing it using natural language. Different from previous work that encodes video information using a single flow, in this work, we introduce a novel Sibling Convolutional Encoder (SibNet) for video captioning, which utilizes a two-branch architecture to collaboratively encode videos. The first content branch encodes the visual content information of the video via autoEncoder, and the second semantic branch encodes the semantic information by visual-semantic joint embedding. Then both branches are effectively combined with soft-attention mechanism and finally fed into a RNN decoder to generate captions. With our SibNet explicitly capturing both content and semantic information, the proposed method can better represent the rich information in videos. Extensive experiments on YouTube2Text and MSR-VTT datasets validate that the proposed architecture outperforms existing methods by a large margin across different evaluation metrics.

Roger Tam - One of the best experts on this subject based on the ideXlab platform.

  • grey matter segmentation in spinal cord mris via 3d Convolutional Encoder networks with shortcut connections
    DLMIA ML-CDS@MICCAI, 2017
    Co-Authors: Adam Porisky, Tom Brosch, Youngjin Yoo, Lisa Tang, Anthony Traboulsee, Emil Ljungberg, Benjamin De Leener, Julien Cohenadad, Roger Tam
    Abstract:

    Segmentation of grey matter in magnetic resonance images of the spinal cord is an important step in assessing disease state in neurological disorders such as multiple sclerosis. However, manual delineation of spinal cord tissue is time-consuming and susceptible to variability introduced by the rater. We present a novel segmentation method for spinal cord tissue that uses fully Convolutional Encoder networks (CENs) for direct end-to-end training and includes shortcut connections to combine multi-scale features, similar to a u-net. While CENs with shortcuts have been used successfully for brain tissue segmentation, spinal cord images have very different features, and therefore deserve their own investigation. In particular, we develop the methodology by evaluating the impact of the number of layers, filter sizes, and shortcuts on segmentation accuracy in standard-resolution cord MRIs. This deep learning-based method is trained on data from a recent public challenge, consisting of 40 MRIs from 4 unique scan sites, with each MRI having 4 manual segmentations from 4 expert raters, resulting in a total of 160 image-label pairs. Performance of the method is evaluated using an independent test set of 40 scans and compared against the challenge results. Using a comprehensive suite of performance metrics, including the Dice similarity coefficient (DSC) and Jaccard index, we found shortcuts to have the strongest impact (0.60 to 0.80 in DSC), while filter size (0.76 to 0.80) and the number of layers (0.77 to 0.80) are also important considerations. Overall, the method is highly competitive with other state-of-the-art methods.

  • deep 3d Convolutional Encoder networks with shortcuts for multiscale feature integration applied to multiple sclerosis lesion segmentation
    IEEE Transactions on Medical Imaging, 2016
    Co-Authors: Tom Brosch, Youngjin Yoo, Lisa Tang, Anthony Traboulsee, Roger Tam
    Abstract:

    We propose a novel segmentation approach based on deep 3D Convolutional Encoder networks with shortcut connections and apply it to the segmentation of multiple sclerosis (MS) lesions in magnetic resonance images. Our model is a neural network that consists of two interconnected pathways, a Convolutional pathway, which learns increasingly more abstract and higher-level image features, and a deConvolutional pathway, which predicts the final segmentation at the voxel level. The joint training of the feature extraction and prediction pathways allows for the automatic learning of features at different scales that are optimized for accuracy for any given combination of image types and segmentation task. In addition, shortcut connections between the two pathways allow high- and low-level features to be integrated, which enables the segmentation of lesions across a wide range of sizes. We have evaluated our method on two publicly available data sets (MICCAI 2008 and ISBI 2015 challenges) with the results showing that our method performs comparably to the top-ranked state-of-the-art methods, even when only relatively small data sets are available for training. In addition, we have compared our method with five freely available and widely used MS lesion segmentation methods (EMS, LST-LPA, LST-LGA, Lesion-TOADS, and SLS) on a large data set from an MS clinical trial. The results show that our method consistently outperforms these other methods across a wide range of lesion sizes.

  • deep Convolutional Encoder networks for multiple sclerosis lesion segmentation
    Medical Image Computing and Computer-Assisted Intervention, 2015
    Co-Authors: Tom Brosch, Youngjin Yoo, Lisa Tang, Anthony Traboulsee, Roger Tam
    Abstract:

    We propose a novel segmentation approach based on deep Convolutional Encoder networks and apply it to the segmentation of multiple sclerosis (MS) lesions in magnetic resonance images. Our model is a neural network that has both Convolutional and deConvolutional layers, and combines feature extraction and segmentation prediction in a single model. The joint training of the feature extraction and prediction layers allows the model to automatically learn features that are optimized for accuracy for any given combination of image types. In contrast to existing automatic feature learning approaches, which are typically patch-based, our model learns features from entire images, which eliminates patch selection and redundant calculations at the overlap of neighboring patches and thereby speeds up the training. Our network also uses a novel objective function that works well for segmenting underrepresented classes, such as MS lesions. We have evaluated our method on the publicly available labeled cases from the MS lesion segmentation challenge 2008 data set, showing that our method performs comparably to the state-of-theart. In addition, we have evaluated our method on the images of 500 subjects from an MS clinical trial and varied the number of training samples from 5 to 250 to show that the segmentation performance can be greatly improved by having a representative data set.