The Experts below are selected from a list of 239244 Experts worldwide ranked by ideXlab platform

Qi Tian - One of the best experts on this subject based on the ideXlab platform.

  • a mid Level Representation framework for semantic sports video analysis
    ACM Multimedia, 2003
    Co-Authors: Lingyu Duan, Min Xu, Tatseng Chua, Qi Tian, Changsheng Xu
    Abstract:

    Sports video has been widely studied due to its tremendous commercial potentials. Despite encouraging results from various specific sports games, it is almost impossible to extend a system for a new sports game because they usually employ different sets of low-Level features appropriate for the specific games and closely coupled with the use of game specific rules to detect events or highlights. There is a lack of internal Representation and structure to be generic and applicable for many different sports. In this paper, we present a generic mid-Level Representation framework for semantic sports video analysis. The mid-Level Representation layer is introduced between the low-Level audio-visual processing and high-Level semantic analysis. It allows us to separate sports specific knowledge and rules from the low-Level and mid-Level feature extraction. This makes sports video analysis more efficient, effective, and less ad-hoc for various types of sports. To achieve robustness of the low-Level feature analysis, a non-parametric clustering, mean shift procedure, has been successfully applied to both color and motion analysis. The proposed framework has been tested for five field-ball type sports covering duration of about 8 hours. Experiments have shown its robust performance in semantic analysis and event detection. We believe that the proposed mid-Level Representation framework can be used for event detection, highlight extraction, summarization and personalization of many types of sports video.

  • a mid Level Representation framework for semantic sports video analysis
    ACM Multimedia, 2003
    Co-Authors: Lingyu Duan, Tatseng Chua, Qi Tian
    Abstract:

    Sports video has been widely studied due to its tremendous commercial potentials. Despite encouraging results from various specific sports games, it is almost impossible to extend a system for a new sports game because they usually employ different sets of low-Level features appropriate for the specific games and closely coupled with the use of game specific rules to detect events or highlights. There is a lack of internal Representation and structure to be generic and applicable for many different sports. In this paper, we present a generic mid-Level Representation framework for semantic sports video analysis. The mid-Level Representation layer is introduced between the low-Level audio-visual processing and high-Level semantic analysis. It allows us to separate sports specific knowledge and rules from the low-Level and mid-Level feature extraction. This makes sports video analysis more efficient, effective, and less ad-hoc for various types of sports. To achieve robustness of the low-Level feature analysis, a non-parametric clustering, mean shift procedure, has been successfully applied to both color and motion analysis. The proposed framework has been tested for five field-ball type sports covering duration of about 8 hours. Experiments have shown its robust performance in semantic analysis and event detection. We believe that the proposed mid-Level Representation framework can be used for event detection, highlight extraction, summarization and personalization of many types of sports video.

  • ACM Multimedia - A mid-Level Representation framework for semantic sports video analysis
    Proceedings of the eleventh ACM international conference on Multimedia - MULTIMEDIA '03, 2003
    Co-Authors: Lingyu Duan, Tatseng Chua, Qi Tian
    Abstract:

    Sports video has been widely studied due to its tremendous commercial potentials. Despite encouraging results from various specific sports games, it is almost impossible to extend a system for a new sports game because they usually employ different sets of low-Level features appropriate for the specific games and closely coupled with the use of game specific rules to detect events or highlights. There is a lack of internal Representation and structure to be generic and applicable for many different sports. In this paper, we present a generic mid-Level Representation framework for semantic sports video analysis. The mid-Level Representation layer is introduced between the low-Level audio-visual processing and high-Level semantic analysis. It allows us to separate sports specific knowledge and rules from the low-Level and mid-Level feature extraction. This makes sports video analysis more efficient, effective, and less ad-hoc for various types of sports. To achieve robustness of the low-Level feature analysis, a non-parametric clustering, mean shift procedure, has been successfully applied to both color and motion analysis. The proposed framework has been tested for five field-ball type sports covering duration of about 8 hours. Experiments have shown its robust performance in semantic analysis and event detection. We believe that the proposed mid-Level Representation framework can be used for event detection, highlight extraction, summarization and personalization of many types of sports video.

  • ICIP - A generic mid-Level Representation for semantic video analysis
    2004 International Conference on Image Processing 2004. ICIP '04., 1
    Co-Authors: Qing Tang, Joo-hwee Lim, Jesse S. Jin, Haiping Sun, Qi Tian
    Abstract:

    The paper presents a generic, mid-Level Representation for efficient semantic video analysis, which adopts a frame-by-frame scheme using P-frames rather than shot-based schemes. Each P-frame is partitioned into an m/spl times/n grid (row by column), and each cell is called a 'block'. The Representation can bridge the semantic gap and build an intermediate description of video features across frames and blocks. Soccer video is used to showcase the potential of the framework for real video processing. Experiments with tennis video and news video have also been conducted. Results demonstrate the excellent performance of the framework in semantic analysis and also indicate its further potential for automatic video analysis.

Song-chun Zhu - One of the best experts on this subject based on the ideXlab platform.

  • Video Primal Sketch: A Unified Middle-Level Representation for Video
    arXiv: Computer Vision and Pattern Recognition, 2015
    Co-Authors: Zhi Han, Song-chun Zhu
    Abstract:

    This paper presents a middle-Level video Representation named Video Primal Sketch (VPS), which integrates two regimes of models: i) sparse coding model using static or moving primitives to explicitly represent moving corners, lines, feature points, etc., ii) FRAME /MRF model reproducing feature statistics extracted from input video to implicitly represent textured motion, such as water and fire. The feature statistics include histograms of spatio-temporal filters and velocity distributions. This paper makes three contributions to the literature: i) Learning a dictionary of video primitives using parametric generative models; ii) Proposing the Spatio-Temporal FRAME (ST-FRAME) and Motion-Appearance FRAME (MA-FRAME) models for modeling and synthesizing textured motion; and iii) Developing a parsimonious hybrid model for generic video Representation. Given an input video, VPS selects the proper models automatically for different motion patterns and is compatible with high-Level action Representations. In the experiments, we synthesize a number of textured motion; reconstruct real videos using the VPS; report a series of human perception experiments to verify the quality of reconstructed videos; demonstrate how the VPS changes over the scale transition in videos; and present the close connection between VPS and high-Level action models.

  • Video Primal Sketch: A Unified Middle-Level Representation for Video
    Journal of Mathematical Imaging and Vision, 2015
    Co-Authors: Zhi Han, Song-chun Zhu
    Abstract:

    This paper presents a middle-Level video Representation named video primal sketch (VPS), which integrates two regimes of models: (i) sparse coding model using static or moving primitives to explicitly represent moving corners, lines, feature points, etc., (ii) FRAME /MRF model reproducing feature statistics extracted from input video to implicitly represent textured motion, such as water and fire. The feature statistics include histograms of spatio-temporal filters and velocity distributions. This paper makes three contributions to the literature: (i) Learning a dictionary of video primitives using parametric generative models; (ii) Proposing the spatio-temporal FRAME and motion-appearance FRAME models for modeling and synthesizing textured motion; and (iii) Developing a parsimonious hybrid model for generic video Representation. Given an input video, VPS selects the proper models automatically for different motion patterns and is compatible with high-Level action Representations. In the experiments, we synthesize a number of textured motion; reconstruct real videos using the VPS; report a series of human perception experiments to verify the quality of reconstructed videos; demonstrate how the VPS changes over the scale transition in videos; and present the close connection between VPS and high-Level action models.

  • ICCV - Video Primal Sketch: A generic middle-Level Representation of video
    2011 International Conference on Computer Vision, 2011
    Co-Authors: Zhi Han, Song-chun Zhu
    Abstract:

    This paper presents a middle-Level video Representation named Video Primal Sketch (VPS), which integrates two regimes of models: i) sparse coding model using static or moving primitives to explicitly represent moving corners, lines, feature points, etc., ii) FRAME/MRF model with spatio-temporal filters to implicitly represent textured motion, such as water and fire, by matching feature statistics, i.e. histograms. This paper makes three contributions: i) learning a dictionary of video primitives as parametric generative model; ii) studying the Spatio-Temporal FRAME (ST-FRAME) model for modeling and synthesizing textured motion; and iii) developing a parsimonious hybrid model for generic video Representation. VPS selects the proper Representation automatically and is compatible with high-Level action Representations. In the experiments, we synthesize a series of dynamic textures, reconstruct real videos and show varying VPS over the change of densities causing by the scale transition in videos.

Stefano Soatto - One of the best experts on this subject based on the ideXlab platform.

  • a mid Level Representation of visual structures for video compression
    Workshop on Applications of Computer Vision, 2016
    Co-Authors: Georgios Georgiadis, Stefano Soatto
    Abstract:

    A video coding system is presented that partitions the scene into "visual structures" anda residual "background" layer. A low-Level Representation ("track-template") of visual structures is proposed that exploits their temporal redundancy. A dictionary of track-templates is constructed that is used to encode video frames. We make optimal use of the dictionary in terms of rate-distortion by choosing a subset of the dictionary's elements for encoding using a Markov Random Field (MRF) formulation that places the track-templates in "depth" layers. The selected "track-templates" form the mid-Level Representation of the "visual structure" regions of the video. Our video coding system offers improvements over H.265/H.264 and other methods in a rate-distortion comparison.

  • WACV - A mid-Level Representation of visual structures for video compression
    2016 IEEE Winter Conference on Applications of Computer Vision (WACV), 2016
    Co-Authors: Georgios Georgiadis, Stefano Soatto
    Abstract:

    A video coding system is presented that partitions the scene into "visual structures" anda residual "background" layer. A low-Level Representation ("track-template") of visual structures is proposed that exploits their temporal redundancy. A dictionary of track-templates is constructed that is used to encode video frames. We make optimal use of the dictionary in terms of rate-distortion by choosing a subset of the dictionary's elements for encoding using a Markov Random Field (MRF) formulation that places the track-templates in "depth" layers. The selected "track-templates" form the mid-Level Representation of the "visual structure" regions of the video. Our video coding system offers improvements over H.265/H.264 and other methods in a rate-distortion comparison.

  • ECCV Workshops (3) - SuperFloxels: a mid-Level Representation for video sequences
    Computer Vision – ECCV 2012. Workshops and Demonstrations, 2012
    Co-Authors: Avinash Ravichandran, Chaohui Wang, Michalis Raptis, Stefano Soatto
    Abstract:

    We describe an approach for grouping trajectories extracted from a video that preserves motion discontinuities due, for instance, to occlusions, but not color or intensity boundaries. Our method takes as input trajectories with variable length and onset time, and outputs a membership function as well as an indicator function denoting the exemplar trajectory of each group. This can be used for several applications such as compression, segmentation, and background removal.

Lingyu Duan - One of the best experts on this subject based on the ideXlab platform.

  • a mid Level Representation framework for semantic sports video analysis
    ACM Multimedia, 2003
    Co-Authors: Lingyu Duan, Min Xu, Tatseng Chua, Qi Tian, Changsheng Xu
    Abstract:

    Sports video has been widely studied due to its tremendous commercial potentials. Despite encouraging results from various specific sports games, it is almost impossible to extend a system for a new sports game because they usually employ different sets of low-Level features appropriate for the specific games and closely coupled with the use of game specific rules to detect events or highlights. There is a lack of internal Representation and structure to be generic and applicable for many different sports. In this paper, we present a generic mid-Level Representation framework for semantic sports video analysis. The mid-Level Representation layer is introduced between the low-Level audio-visual processing and high-Level semantic analysis. It allows us to separate sports specific knowledge and rules from the low-Level and mid-Level feature extraction. This makes sports video analysis more efficient, effective, and less ad-hoc for various types of sports. To achieve robustness of the low-Level feature analysis, a non-parametric clustering, mean shift procedure, has been successfully applied to both color and motion analysis. The proposed framework has been tested for five field-ball type sports covering duration of about 8 hours. Experiments have shown its robust performance in semantic analysis and event detection. We believe that the proposed mid-Level Representation framework can be used for event detection, highlight extraction, summarization and personalization of many types of sports video.

  • a mid Level Representation framework for semantic sports video analysis
    ACM Multimedia, 2003
    Co-Authors: Lingyu Duan, Tatseng Chua, Qi Tian
    Abstract:

    Sports video has been widely studied due to its tremendous commercial potentials. Despite encouraging results from various specific sports games, it is almost impossible to extend a system for a new sports game because they usually employ different sets of low-Level features appropriate for the specific games and closely coupled with the use of game specific rules to detect events or highlights. There is a lack of internal Representation and structure to be generic and applicable for many different sports. In this paper, we present a generic mid-Level Representation framework for semantic sports video analysis. The mid-Level Representation layer is introduced between the low-Level audio-visual processing and high-Level semantic analysis. It allows us to separate sports specific knowledge and rules from the low-Level and mid-Level feature extraction. This makes sports video analysis more efficient, effective, and less ad-hoc for various types of sports. To achieve robustness of the low-Level feature analysis, a non-parametric clustering, mean shift procedure, has been successfully applied to both color and motion analysis. The proposed framework has been tested for five field-ball type sports covering duration of about 8 hours. Experiments have shown its robust performance in semantic analysis and event detection. We believe that the proposed mid-Level Representation framework can be used for event detection, highlight extraction, summarization and personalization of many types of sports video.

  • ACM Multimedia - A mid-Level Representation framework for semantic sports video analysis
    Proceedings of the eleventh ACM international conference on Multimedia - MULTIMEDIA '03, 2003
    Co-Authors: Lingyu Duan, Tatseng Chua, Qi Tian
    Abstract:

    Sports video has been widely studied due to its tremendous commercial potentials. Despite encouraging results from various specific sports games, it is almost impossible to extend a system for a new sports game because they usually employ different sets of low-Level features appropriate for the specific games and closely coupled with the use of game specific rules to detect events or highlights. There is a lack of internal Representation and structure to be generic and applicable for many different sports. In this paper, we present a generic mid-Level Representation framework for semantic sports video analysis. The mid-Level Representation layer is introduced between the low-Level audio-visual processing and high-Level semantic analysis. It allows us to separate sports specific knowledge and rules from the low-Level and mid-Level feature extraction. This makes sports video analysis more efficient, effective, and less ad-hoc for various types of sports. To achieve robustness of the low-Level feature analysis, a non-parametric clustering, mean shift procedure, has been successfully applied to both color and motion analysis. The proposed framework has been tested for five field-ball type sports covering duration of about 8 hours. Experiments have shown its robust performance in semantic analysis and event detection. We believe that the proposed mid-Level Representation framework can be used for event detection, highlight extraction, summarization and personalization of many types of sports video.

Zhi Han - One of the best experts on this subject based on the ideXlab platform.

  • Video Primal Sketch: A Unified Middle-Level Representation for Video
    arXiv: Computer Vision and Pattern Recognition, 2015
    Co-Authors: Zhi Han, Song-chun Zhu
    Abstract:

    This paper presents a middle-Level video Representation named Video Primal Sketch (VPS), which integrates two regimes of models: i) sparse coding model using static or moving primitives to explicitly represent moving corners, lines, feature points, etc., ii) FRAME /MRF model reproducing feature statistics extracted from input video to implicitly represent textured motion, such as water and fire. The feature statistics include histograms of spatio-temporal filters and velocity distributions. This paper makes three contributions to the literature: i) Learning a dictionary of video primitives using parametric generative models; ii) Proposing the Spatio-Temporal FRAME (ST-FRAME) and Motion-Appearance FRAME (MA-FRAME) models for modeling and synthesizing textured motion; and iii) Developing a parsimonious hybrid model for generic video Representation. Given an input video, VPS selects the proper models automatically for different motion patterns and is compatible with high-Level action Representations. In the experiments, we synthesize a number of textured motion; reconstruct real videos using the VPS; report a series of human perception experiments to verify the quality of reconstructed videos; demonstrate how the VPS changes over the scale transition in videos; and present the close connection between VPS and high-Level action models.

  • Video Primal Sketch: A Unified Middle-Level Representation for Video
    Journal of Mathematical Imaging and Vision, 2015
    Co-Authors: Zhi Han, Song-chun Zhu
    Abstract:

    This paper presents a middle-Level video Representation named video primal sketch (VPS), which integrates two regimes of models: (i) sparse coding model using static or moving primitives to explicitly represent moving corners, lines, feature points, etc., (ii) FRAME /MRF model reproducing feature statistics extracted from input video to implicitly represent textured motion, such as water and fire. The feature statistics include histograms of spatio-temporal filters and velocity distributions. This paper makes three contributions to the literature: (i) Learning a dictionary of video primitives using parametric generative models; (ii) Proposing the spatio-temporal FRAME and motion-appearance FRAME models for modeling and synthesizing textured motion; and (iii) Developing a parsimonious hybrid model for generic video Representation. Given an input video, VPS selects the proper models automatically for different motion patterns and is compatible with high-Level action Representations. In the experiments, we synthesize a number of textured motion; reconstruct real videos using the VPS; report a series of human perception experiments to verify the quality of reconstructed videos; demonstrate how the VPS changes over the scale transition in videos; and present the close connection between VPS and high-Level action models.

  • ICCV - Video Primal Sketch: A generic middle-Level Representation of video
    2011 International Conference on Computer Vision, 2011
    Co-Authors: Zhi Han, Song-chun Zhu
    Abstract:

    This paper presents a middle-Level video Representation named Video Primal Sketch (VPS), which integrates two regimes of models: i) sparse coding model using static or moving primitives to explicitly represent moving corners, lines, feature points, etc., ii) FRAME/MRF model with spatio-temporal filters to implicitly represent textured motion, such as water and fire, by matching feature statistics, i.e. histograms. This paper makes three contributions: i) learning a dictionary of video primitives as parametric generative model; ii) studying the Spatio-Temporal FRAME (ST-FRAME) model for modeling and synthesizing textured motion; and iii) developing a parsimonious hybrid model for generic video Representation. VPS selects the proper Representation automatically and is compatible with high-Level action Representations. In the experiments, we synthesize a series of dynamic textures, reconstruct real videos and show varying VPS over the change of densities causing by the scale transition in videos.