The Experts below are selected from a list of 36687 Experts worldwide ranked by ideXlab platform
Lingyu Duan - One of the best experts on this subject based on the ideXlab platform.
-
live Sports event detection based on broadcast video and web casting text
ACM Multimedia, 2006Co-Authors: Jinjun Wang, Kongwah Wan, Lingyu DuanAbstract:Event detection is essential for Sports video summarization, indexing and retrieval and extensive research efforts have been devoted to this area. However, the previous approaches are heavily relying on video content itself and require the whole video content for event detection. Due to the semantic gap between low-level features and high-level events, it is difficult to come up with a generic framework to achieve a high accuracy of event detection. In addition, the dynamic structures from different Sports domains further complicate the analysis and impede the implementation of live event detection systems. In this paper, we present a novel approach for event detection from the live Sports Game using web-casting text and broadcast video. Web-casting text is a text broadcast source for Sports Game and can be live captured from the web. Incorporating web-casting text into Sports video analysis significantly improves the event detection accuracy. Compared with previous approaches, the proposed approach is able to: (1) detect live event only based on the partial content captured from the web and TV; (2) extract detailed event semantics and detect exact event boundary, which are very difficult or impossible to be handled by previous approaches; and (3) create personalized summary related to certain event, player or team according to user's preference. We present the framework of our approach and details of text analysis, video analysis and text/video alignment. We conducted experiments on both live Games and recorded Games. The results are encouraging and comparable to the manually detected events. We also give scenarios to illustrate how to apply the proposed solution to professional and consumer services.
-
a mid level representation framework for semantic Sports video analysis
ACM Multimedia, 2003Co-Authors: Lingyu Duan, Min Xu, Tatseng Chua, Qi Tian, Changsheng XuAbstract:Sports video has been widely studied due to its tremendous commercial potentials. Despite encouraging results from various specific Sports Games, it is almost impossible to extend a system for a new Sports Game because they usually employ different sets of low-level features appropriate for the specific Games and closely coupled with the use of Game specific rules to detect events or highlights. There is a lack of internal representation and structure to be generic and applicable for many different Sports. In this paper, we present a generic mid-level representation framework for semantic Sports video analysis. The mid-level representation layer is introduced between the low-level audio-visual processing and high-level semantic analysis. It allows us to separate Sports specific knowledge and rules from the low-level and mid-level feature extraction. This makes Sports video analysis more efficient, effective, and less ad-hoc for various types of Sports. To achieve robustness of the low-level feature analysis, a non-parametric clustering, mean shift procedure, has been successfully applied to both color and motion analysis. The proposed framework has been tested for five field-ball type Sports covering duration of about 8 hours. Experiments have shown its robust performance in semantic analysis and event detection. We believe that the proposed mid-level representation framework can be used for event detection, highlight extraction, summarization and personalization of many types of Sports video.
-
a mid level representation framework for semantic Sports video analysis
ACM Multimedia, 2003Co-Authors: Lingyu Duan, Tatseng Chua, Qi TianAbstract:Sports video has been widely studied due to its tremendous commercial potentials. Despite encouraging results from various specific Sports Games, it is almost impossible to extend a system for a new Sports Game because they usually employ different sets of low-level features appropriate for the specific Games and closely coupled with the use of Game specific rules to detect events or highlights. There is a lack of internal representation and structure to be generic and applicable for many different Sports. In this paper, we present a generic mid-level representation framework for semantic Sports video analysis. The mid-level representation layer is introduced between the low-level audio-visual processing and high-level semantic analysis. It allows us to separate Sports specific knowledge and rules from the low-level and mid-level feature extraction. This makes Sports video analysis more efficient, effective, and less ad-hoc for various types of Sports. To achieve robustness of the low-level feature analysis, a non-parametric clustering, mean shift procedure, has been successfully applied to both color and motion analysis. The proposed framework has been tested for five field-ball type Sports covering duration of about 8 hours. Experiments have shown its robust performance in semantic analysis and event detection. We believe that the proposed mid-level representation framework can be used for event detection, highlight extraction, summarization and personalization of many types of Sports video.
-
a fusion scheme of visual and auditory modalities for event detection in Sports video
International Conference on Acoustics Speech and Signal Processing, 2003Co-Authors: Lingyu Duan, Qi TianAbstract:In this paper, we propose an effective fusion scheme of visual and auditory modalities to detect events in Sports video. The proposed scheme is built upon semantic shot classification, where we classify video shots into several major or interesting classes, each of which has clear semantic meanings. Among major shot classes we perform classification of the different auditory signal segments (i.e. silence, hitting ball, applause, commentator speech) with the goal of detecting events with strong semantic meaning. For instance, for tennis video, we have identified five interesting events: serve, reserve, ace, return, and score. Since we have developed a unified framework for semantic shot classification in Sports videos and a set of audio mid-level representation with supervised learning methods, the proposed fusion scheme can be easily adapted to a new Sports Game. We are extending this fusion scheme to three additional typical Sports videos: basketball, volleyball and soccer. Correctly detected Sports video events will greatly facilitate further structural and temporal analysis, such as Sports video skimming, table of content, etc.
Qi Tian - One of the best experts on this subject based on the ideXlab platform.
-
a mid level representation framework for semantic Sports video analysis
ACM Multimedia, 2003Co-Authors: Lingyu Duan, Min Xu, Tatseng Chua, Qi Tian, Changsheng XuAbstract:Sports video has been widely studied due to its tremendous commercial potentials. Despite encouraging results from various specific Sports Games, it is almost impossible to extend a system for a new Sports Game because they usually employ different sets of low-level features appropriate for the specific Games and closely coupled with the use of Game specific rules to detect events or highlights. There is a lack of internal representation and structure to be generic and applicable for many different Sports. In this paper, we present a generic mid-level representation framework for semantic Sports video analysis. The mid-level representation layer is introduced between the low-level audio-visual processing and high-level semantic analysis. It allows us to separate Sports specific knowledge and rules from the low-level and mid-level feature extraction. This makes Sports video analysis more efficient, effective, and less ad-hoc for various types of Sports. To achieve robustness of the low-level feature analysis, a non-parametric clustering, mean shift procedure, has been successfully applied to both color and motion analysis. The proposed framework has been tested for five field-ball type Sports covering duration of about 8 hours. Experiments have shown its robust performance in semantic analysis and event detection. We believe that the proposed mid-level representation framework can be used for event detection, highlight extraction, summarization and personalization of many types of Sports video.
-
a mid level representation framework for semantic Sports video analysis
ACM Multimedia, 2003Co-Authors: Lingyu Duan, Tatseng Chua, Qi TianAbstract:Sports video has been widely studied due to its tremendous commercial potentials. Despite encouraging results from various specific Sports Games, it is almost impossible to extend a system for a new Sports Game because they usually employ different sets of low-level features appropriate for the specific Games and closely coupled with the use of Game specific rules to detect events or highlights. There is a lack of internal representation and structure to be generic and applicable for many different Sports. In this paper, we present a generic mid-level representation framework for semantic Sports video analysis. The mid-level representation layer is introduced between the low-level audio-visual processing and high-level semantic analysis. It allows us to separate Sports specific knowledge and rules from the low-level and mid-level feature extraction. This makes Sports video analysis more efficient, effective, and less ad-hoc for various types of Sports. To achieve robustness of the low-level feature analysis, a non-parametric clustering, mean shift procedure, has been successfully applied to both color and motion analysis. The proposed framework has been tested for five field-ball type Sports covering duration of about 8 hours. Experiments have shown its robust performance in semantic analysis and event detection. We believe that the proposed mid-level representation framework can be used for event detection, highlight extraction, summarization and personalization of many types of Sports video.
-
a fusion scheme of visual and auditory modalities for event detection in Sports video
International Conference on Acoustics Speech and Signal Processing, 2003Co-Authors: Lingyu Duan, Qi TianAbstract:In this paper, we propose an effective fusion scheme of visual and auditory modalities to detect events in Sports video. The proposed scheme is built upon semantic shot classification, where we classify video shots into several major or interesting classes, each of which has clear semantic meanings. Among major shot classes we perform classification of the different auditory signal segments (i.e. silence, hitting ball, applause, commentator speech) with the goal of detecting events with strong semantic meaning. For instance, for tennis video, we have identified five interesting events: serve, reserve, ace, return, and score. Since we have developed a unified framework for semantic shot classification in Sports videos and a set of audio mid-level representation with supervised learning methods, the proposed fusion scheme can be easily adapted to a new Sports Game. We are extending this fusion scheme to three additional typical Sports videos: basketball, volleyball and soccer. Correctly detected Sports video events will greatly facilitate further structural and temporal analysis, such as Sports video skimming, table of content, etc.
Changsheng Xu - One of the best experts on this subject based on the ideXlab platform.
-
Using webcast text for semantic event detection in broadcast Sports video
IEEE Transactions on Multimedia, 2008Co-Authors: Changsheng Xu, Yong Rui, Hanqing Lu, Yi-fan Zhang, Guangyu Zhu, Qingming HuangAbstract:Sports video semantic event detection is essential for Sports video summarization and retrieval. Extensive research efforts have been devoted to this area in recent years. However, the existing Sports video event detection approaches heavily rely on either video content itself, which face the difficulty of high-level semantic information extraction from video content using computer vision and image processing techniques, or manually generated video ontology, which is domain specific and difficult to be automatically aligned with the video content. In this paper, we present a novel approach for Sports video semantic event detection based on analysis and alignment of Webcast text and broadcast video. Webcast text is a text broadcast channel for Sports Game which is co-produced with the broadcast video and is easily obtained from the Web. We first analyze Webcast text to cluster and detect text events in an unsupervised way using probabilistic latent semantic analysis (pLSA). Based on the detected text event and video structure analysis, we employ a conditional random field model (CRFM) to align text event and video event by detecting event moment and event boundary in the video. Incorporation of Webcast text into Sports video analysis significantly facilitates Sports video semantic event detection. We conducted experiments on 33 hours of soccer and basketball Games for Webcast analysis, broadcast video analysis and text/video semantic alignment. The results are encouraging and compared with the manually labeled ground truth.
-
a mid level representation framework for semantic Sports video analysis
ACM Multimedia, 2003Co-Authors: Lingyu Duan, Min Xu, Tatseng Chua, Qi Tian, Changsheng XuAbstract:Sports video has been widely studied due to its tremendous commercial potentials. Despite encouraging results from various specific Sports Games, it is almost impossible to extend a system for a new Sports Game because they usually employ different sets of low-level features appropriate for the specific Games and closely coupled with the use of Game specific rules to detect events or highlights. There is a lack of internal representation and structure to be generic and applicable for many different Sports. In this paper, we present a generic mid-level representation framework for semantic Sports video analysis. The mid-level representation layer is introduced between the low-level audio-visual processing and high-level semantic analysis. It allows us to separate Sports specific knowledge and rules from the low-level and mid-level feature extraction. This makes Sports video analysis more efficient, effective, and less ad-hoc for various types of Sports. To achieve robustness of the low-level feature analysis, a non-parametric clustering, mean shift procedure, has been successfully applied to both color and motion analysis. The proposed framework has been tested for five field-ball type Sports covering duration of about 8 hours. Experiments have shown its robust performance in semantic analysis and event detection. We believe that the proposed mid-level representation framework can be used for event detection, highlight extraction, summarization and personalization of many types of Sports video.
Tatseng Chua - One of the best experts on this subject based on the ideXlab platform.
-
a mid level representation framework for semantic Sports video analysis
ACM Multimedia, 2003Co-Authors: Lingyu Duan, Min Xu, Tatseng Chua, Qi Tian, Changsheng XuAbstract:Sports video has been widely studied due to its tremendous commercial potentials. Despite encouraging results from various specific Sports Games, it is almost impossible to extend a system for a new Sports Game because they usually employ different sets of low-level features appropriate for the specific Games and closely coupled with the use of Game specific rules to detect events or highlights. There is a lack of internal representation and structure to be generic and applicable for many different Sports. In this paper, we present a generic mid-level representation framework for semantic Sports video analysis. The mid-level representation layer is introduced between the low-level audio-visual processing and high-level semantic analysis. It allows us to separate Sports specific knowledge and rules from the low-level and mid-level feature extraction. This makes Sports video analysis more efficient, effective, and less ad-hoc for various types of Sports. To achieve robustness of the low-level feature analysis, a non-parametric clustering, mean shift procedure, has been successfully applied to both color and motion analysis. The proposed framework has been tested for five field-ball type Sports covering duration of about 8 hours. Experiments have shown its robust performance in semantic analysis and event detection. We believe that the proposed mid-level representation framework can be used for event detection, highlight extraction, summarization and personalization of many types of Sports video.
-
a mid level representation framework for semantic Sports video analysis
ACM Multimedia, 2003Co-Authors: Lingyu Duan, Tatseng Chua, Qi TianAbstract:Sports video has been widely studied due to its tremendous commercial potentials. Despite encouraging results from various specific Sports Games, it is almost impossible to extend a system for a new Sports Game because they usually employ different sets of low-level features appropriate for the specific Games and closely coupled with the use of Game specific rules to detect events or highlights. There is a lack of internal representation and structure to be generic and applicable for many different Sports. In this paper, we present a generic mid-level representation framework for semantic Sports video analysis. The mid-level representation layer is introduced between the low-level audio-visual processing and high-level semantic analysis. It allows us to separate Sports specific knowledge and rules from the low-level and mid-level feature extraction. This makes Sports video analysis more efficient, effective, and less ad-hoc for various types of Sports. To achieve robustness of the low-level feature analysis, a non-parametric clustering, mean shift procedure, has been successfully applied to both color and motion analysis. The proposed framework has been tested for five field-ball type Sports covering duration of about 8 hours. Experiments have shown its robust performance in semantic analysis and event detection. We believe that the proposed mid-level representation framework can be used for event detection, highlight extraction, summarization and personalization of many types of Sports video.
Qingming Huang - One of the best experts on this subject based on the ideXlab platform.
-
Using webcast text for semantic event detection in broadcast Sports video
IEEE Transactions on Multimedia, 2008Co-Authors: Changsheng Xu, Yong Rui, Hanqing Lu, Yi-fan Zhang, Guangyu Zhu, Qingming HuangAbstract:Sports video semantic event detection is essential for Sports video summarization and retrieval. Extensive research efforts have been devoted to this area in recent years. However, the existing Sports video event detection approaches heavily rely on either video content itself, which face the difficulty of high-level semantic information extraction from video content using computer vision and image processing techniques, or manually generated video ontology, which is domain specific and difficult to be automatically aligned with the video content. In this paper, we present a novel approach for Sports video semantic event detection based on analysis and alignment of Webcast text and broadcast video. Webcast text is a text broadcast channel for Sports Game which is co-produced with the broadcast video and is easily obtained from the Web. We first analyze Webcast text to cluster and detect text events in an unsupervised way using probabilistic latent semantic analysis (pLSA). Based on the detected text event and video structure analysis, we employ a conditional random field model (CRFM) to align text event and video event by detecting event moment and event boundary in the video. Incorporation of Webcast text into Sports video analysis significantly facilitates Sports video semantic event detection. We conducted experiments on 33 hours of soccer and basketball Games for Webcast analysis, broadcast video analysis and text/video semantic alignment. The results are encouraging and compared with the manually labeled ground truth.
-
player action recognition in broadcast tennis video with applications to semantic analysis of Sports Game
ACM Multimedia, 2006Co-Authors: Guangyu Zhu, Qingming Huang, Wen Gao, Liyuan XingAbstract:Recognition of player actions in broadcast Sports video is a challenging task due to low resolution of the players in video frames. In this paper, we present a novel method to recognize the basic player actions in broadcast tennis video. Different from the existing appearance-based approaches, our method is based on motion analysis and considers the relationship between the movements of different body parts and the regions in the image plane. A novel motion descriptor is proposed and supervised learning is employed to train the action classifier. We also propose a novel framework by combining the player action recognition with other multimodal features for semantic and tactic analysis of the broadcast tennis video. Incorporating action recognition into the framework not only improves the semantic indexing and retrieval performance of the video content, but also conducts highlights ranking and tactics analysis in tennis matches, which is the first solution to our knowledge for tennis Game. The experimental results demonstrate that our player action recognition method outperforms existing appearance-based approaches and the multimodal framework is effective for broadcast tennis video analysis.