The Experts below are selected from a list of 12591 Experts worldwide ranked by ideXlab platform
Yingli Tian - One of the best experts on this subject based on the ideXlab platform.
-
self guiding multimodal lstm when we do not have a perfect training dataset for image captioning
IEEE Transactions on Image Processing, 2019Co-Authors: Yang Xian, Yingli TianAbstract:In this paper, a self-guiding multimodal LSTM (sgLSTM) image captioning model is proposed to handle an uncontrolled imbalanced real-world image-sentence dataset. We collect a FlickrNYC dataset from Flickr as our testbed with 306,165 images and the original text descriptions uploaded by the users are utilized as the ground truth for training. Descriptions in the FlickrNYC dataset vary dramatically ranging from short term-descriptions to long paragraph-descriptions and can describe any visual aspects, or even refer to objects that are not depicted. To deal with the imbalanced and noisy situation and to fully explore the dataset itself, we propose a novel guiding Textual Feature extracted utilizing a multimodal LSTM (mLSTM) model. Training of mLSTM is based on the portion of data in which the image content and the corresponding descriptions are strongly bonded. Afterward, during the training of sgLSTM on the rest training data, this guiding information serves as additional input to the network along with the image representations and the ground-truth descriptions. By integrating these input components into a multimodal block, we aim to form a training scheme with the Textual information tightly coupled with the image content. The experimental results demonstrate that the proposed sgLSTM model outperforms the traditional state-of-the-art multimodal RNN captioning framework in successfully describing the key components of the input images.
-
self guiding multimodal lstm when we do not have a perfect training dataset for image captioning
arXiv: Computer Vision and Pattern Recognition, 2017Co-Authors: Yang Xian, Yingli TianAbstract:In this paper, a self-guiding multimodal LSTM (sg-LSTM) image captioning model is proposed to handle uncontrolled imbalanced real-world image-sentence dataset. We collect FlickrNYC dataset from Flickr as our testbed with 306,165 images and the original text descriptions uploaded by the users are utilized as the ground truth for training. Descriptions in FlickrNYC dataset vary dramatically ranging from short term-descriptions to long paragraph-descriptions and can describe any visual aspects, or even refer to objects that are not depicted. To deal with the imbalanced and noisy situation and to fully explore the dataset itself, we propose a novel guiding Textual Feature extracted utilizing a multimodal LSTM (m-LSTM) model. Training of m-LSTM is based on the portion of data in which the image content and the corresponding descriptions are strongly bonded. Afterwards, during the training of sg-LSTM on the rest training data, this guiding information serves as additional input to the network along with the image representations and the ground-truth descriptions. By integrating these input components into a multimodal block, we aim to form a training scheme with the Textual information tightly coupled with the image content. The experimental results demonstrate that the proposed sg-LSTM model outperforms the traditional state-of-the-art multimodal RNN captioning framework in successfully describing the key components of the input images.
Yang Xian - One of the best experts on this subject based on the ideXlab platform.
-
self guiding multimodal lstm when we do not have a perfect training dataset for image captioning
IEEE Transactions on Image Processing, 2019Co-Authors: Yang Xian, Yingli TianAbstract:In this paper, a self-guiding multimodal LSTM (sgLSTM) image captioning model is proposed to handle an uncontrolled imbalanced real-world image-sentence dataset. We collect a FlickrNYC dataset from Flickr as our testbed with 306,165 images and the original text descriptions uploaded by the users are utilized as the ground truth for training. Descriptions in the FlickrNYC dataset vary dramatically ranging from short term-descriptions to long paragraph-descriptions and can describe any visual aspects, or even refer to objects that are not depicted. To deal with the imbalanced and noisy situation and to fully explore the dataset itself, we propose a novel guiding Textual Feature extracted utilizing a multimodal LSTM (mLSTM) model. Training of mLSTM is based on the portion of data in which the image content and the corresponding descriptions are strongly bonded. Afterward, during the training of sgLSTM on the rest training data, this guiding information serves as additional input to the network along with the image representations and the ground-truth descriptions. By integrating these input components into a multimodal block, we aim to form a training scheme with the Textual information tightly coupled with the image content. The experimental results demonstrate that the proposed sgLSTM model outperforms the traditional state-of-the-art multimodal RNN captioning framework in successfully describing the key components of the input images.
-
self guiding multimodal lstm when we do not have a perfect training dataset for image captioning
arXiv: Computer Vision and Pattern Recognition, 2017Co-Authors: Yang Xian, Yingli TianAbstract:In this paper, a self-guiding multimodal LSTM (sg-LSTM) image captioning model is proposed to handle uncontrolled imbalanced real-world image-sentence dataset. We collect FlickrNYC dataset from Flickr as our testbed with 306,165 images and the original text descriptions uploaded by the users are utilized as the ground truth for training. Descriptions in FlickrNYC dataset vary dramatically ranging from short term-descriptions to long paragraph-descriptions and can describe any visual aspects, or even refer to objects that are not depicted. To deal with the imbalanced and noisy situation and to fully explore the dataset itself, we propose a novel guiding Textual Feature extracted utilizing a multimodal LSTM (m-LSTM) model. Training of m-LSTM is based on the portion of data in which the image content and the corresponding descriptions are strongly bonded. Afterwards, during the training of sg-LSTM on the rest training data, this guiding information serves as additional input to the network along with the image representations and the ground-truth descriptions. By integrating these input components into a multimodal block, we aim to form a training scheme with the Textual information tightly coupled with the image content. The experimental results demonstrate that the proposed sg-LSTM model outperforms the traditional state-of-the-art multimodal RNN captioning framework in successfully describing the key components of the input images.
Reda Alhajj - One of the best experts on this subject based on the ideXlab platform.
-
employing structural and Textual Feature extraction for semistructured document classification
Systems Man and Cybernetics, 2012Co-Authors: Mohammad Khabbaz, Keivan Kianmehr, Reda AlhajjAbstract:This paper addresses XML document classification by considering both structural and content-based Features of the documents. This approach leads to better constructing a set of informative Feature vectors that represents both structural and Textual aspects of XML documents. For this purpose, we integrate soft clustering of words and Feature reduction into the process. To extract structural information, we employ an existing frequent tree-mining algorithm combined with an information gain filter to retrieve the most informative substructures from XML documents. However, for extracting content information, we propose soft clustering of words using each cluster as a Textual Feature. We have conducted extensive experiments on a benchmark dataset, namely 20NewsGroups, and an XML documents dataset given in LOGML that describes the web-server logs of user sessions. With regards to the classifier built only using our Textual Features, the results show that it outperforms a naive support-vector-machine (SVM)-based classifier, as well as an information retrieval classifier (IRC). We further demonstrate the effectiveness of incorporating both structural and content information into the process of learning, by comparing our classifier model and several XML document classifiers. In particular, by applying SVM and decision tree algorithms using our Feature vector representation of XML documents dataset, we have achieved 85.79% and 87.04% classification accuracy, respectively, which are higher than accuracy achieved by XRules, a well-known structural-based XML document classifier.
Christoph Meinel - One of the best experts on this subject based on the ideXlab platform.
-
visual Textual late semantic fusion using deep neural network for document categorization
International Conference on Neural Information Processing, 2015Co-Authors: Cheng Wang, Haojin Yang, Christoph MeinelAbstract:Multi-modality fusion has recently drawn much attention due to the fast increasing of multimedia data. Document that consists of multiple modalities i.e. image, text and video, can be better understood by machines if information from different modalities semantically combined. In this paper, we propose to fuse image and text information with deep neural network (DNN) based approach. By jointly fusing visual-Textual Feature and taking the correlation between image and text into account, fusion Features can be learned for representing document. We investigated the fusion Features on document categorization, found that DNN-based fusion outperforms mainstream algorithms include K-Nearest Neighbor(KNN), Support Vector Machine (SVM) and Naive Bayes (NB) and 3-layer Neural Network (3L-NN) in both early and late fusion strategies.
Mohammad Khabbaz - One of the best experts on this subject based on the ideXlab platform.
-
employing structural and Textual Feature extraction for semistructured document classification
Systems Man and Cybernetics, 2012Co-Authors: Mohammad Khabbaz, Keivan Kianmehr, Reda AlhajjAbstract:This paper addresses XML document classification by considering both structural and content-based Features of the documents. This approach leads to better constructing a set of informative Feature vectors that represents both structural and Textual aspects of XML documents. For this purpose, we integrate soft clustering of words and Feature reduction into the process. To extract structural information, we employ an existing frequent tree-mining algorithm combined with an information gain filter to retrieve the most informative substructures from XML documents. However, for extracting content information, we propose soft clustering of words using each cluster as a Textual Feature. We have conducted extensive experiments on a benchmark dataset, namely 20NewsGroups, and an XML documents dataset given in LOGML that describes the web-server logs of user sessions. With regards to the classifier built only using our Textual Features, the results show that it outperforms a naive support-vector-machine (SVM)-based classifier, as well as an information retrieval classifier (IRC). We further demonstrate the effectiveness of incorporating both structural and content information into the process of learning, by comparing our classifier model and several XML document classifiers. In particular, by applying SVM and decision tree algorithms using our Feature vector representation of XML documents dataset, we have achieved 85.79% and 87.04% classification accuracy, respectively, which are higher than accuracy achieved by XRules, a well-known structural-based XML document classifier.