The Experts below are selected from a list of 4602 Experts worldwide ranked by ideXlab platform

Philip S Yu - One of the best experts on this subject based on the ideXlab platform.

  • Spatiotemporal Pyramid Network for Video Action Recognition
    2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017
    Co-Authors: Yunbo Wang, Mingsheng Long, Jianmin Wang, Philip S Yu
    Abstract:

    Two-stream convolutional networks have shown strong performance in video action recognition tasks. The key idea is to learn spatiotemporal features by fusing convolutional networks spatially and temporally. However, it remains unclear how to model the correlations between the spatial and temporal structures at multiple abstraction levels. First, the spatial stream tends to fail if two videos share similar backgrounds. Second, the temporal stream may be fooled if two actions resemble in short snippets, though appear to be distinct in the long term. We propose a novel spatiotemporal pyramid network to fuse the spatial and temporal features in a pyramid structure such that they can reinforce each other. From the architecture perspective, our network constitutes hierarchical fusion strategies which can be trained as a whole using a unified spatiotemporal loss. A series of ablation experiments support the importance of each fusion strategy. From the technical perspective, we introduce the spatiotemporal compact Bilinear Operator into video analysis tasks. This Operator enables efficient training of Bilinear fusion operations which can capture full interactions between the spatial and temporal features. Our final network achieves state-of-the-art results on standard video datasets.

  • CVPR - Spatiotemporal Pyramid Network for Video Action Recognition
    2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017
    Co-Authors: Yunbo Wang, Mingsheng Long, Jianmin Wang, Philip S Yu
    Abstract:

    Two-stream convolutional networks have shown strong performance in video action recognition tasks. The key idea is to learn spatiotemporal features by fusing convolutional networks spatially and temporally. However, it remains unclear how to model the correlations between the spatial and temporal structures at multiple abstraction levels. First, the spatial stream tends to fail if two videos share similar backgrounds. Second, the temporal stream may be fooled if two actions resemble in short snippets, though appear to be distinct in the long term. We propose a novel spatiotemporal pyramid network to fuse the spatial and temporal features in a pyramid structure such that they can reinforce each other. From the architecture perspective, our network constitutes hierarchical fusion strategies which can be trained as a whole using a unified spatiotemporal loss. A series of ablation experiments support the importance of each fusion strategy. From the technical perspective, we introduce the spatiotemporal compact Bilinear Operator into video analysis tasks. This Operator enables efficient training of Bilinear fusion operations which can capture full interactions between the spatial and temporal features. Our final network achieves state-of-the-art results on standard video datasets.

Yunbo Wang - One of the best experts on this subject based on the ideXlab platform.

  • Spatiotemporal Pyramid Network for Video Action Recognition
    2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017
    Co-Authors: Yunbo Wang, Mingsheng Long, Jianmin Wang, Philip S Yu
    Abstract:

    Two-stream convolutional networks have shown strong performance in video action recognition tasks. The key idea is to learn spatiotemporal features by fusing convolutional networks spatially and temporally. However, it remains unclear how to model the correlations between the spatial and temporal structures at multiple abstraction levels. First, the spatial stream tends to fail if two videos share similar backgrounds. Second, the temporal stream may be fooled if two actions resemble in short snippets, though appear to be distinct in the long term. We propose a novel spatiotemporal pyramid network to fuse the spatial and temporal features in a pyramid structure such that they can reinforce each other. From the architecture perspective, our network constitutes hierarchical fusion strategies which can be trained as a whole using a unified spatiotemporal loss. A series of ablation experiments support the importance of each fusion strategy. From the technical perspective, we introduce the spatiotemporal compact Bilinear Operator into video analysis tasks. This Operator enables efficient training of Bilinear fusion operations which can capture full interactions between the spatial and temporal features. Our final network achieves state-of-the-art results on standard video datasets.

  • CVPR - Spatiotemporal Pyramid Network for Video Action Recognition
    2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017
    Co-Authors: Yunbo Wang, Mingsheng Long, Jianmin Wang, Philip S Yu
    Abstract:

    Two-stream convolutional networks have shown strong performance in video action recognition tasks. The key idea is to learn spatiotemporal features by fusing convolutional networks spatially and temporally. However, it remains unclear how to model the correlations between the spatial and temporal structures at multiple abstraction levels. First, the spatial stream tends to fail if two videos share similar backgrounds. Second, the temporal stream may be fooled if two actions resemble in short snippets, though appear to be distinct in the long term. We propose a novel spatiotemporal pyramid network to fuse the spatial and temporal features in a pyramid structure such that they can reinforce each other. From the architecture perspective, our network constitutes hierarchical fusion strategies which can be trained as a whole using a unified spatiotemporal loss. A series of ablation experiments support the importance of each fusion strategy. From the technical perspective, we introduce the spatiotemporal compact Bilinear Operator into video analysis tasks. This Operator enables efficient training of Bilinear fusion operations which can capture full interactions between the spatial and temporal features. Our final network achieves state-of-the-art results on standard video datasets.

John Gracey - One of the best experts on this subject based on the ideXlab platform.

  • Off-shell quark Bilinear Operator Green's functions at two loops
    Physical Review D, 2019
    Co-Authors: John Gracey
    Abstract:

    We construct the two loop Green's functions for a quark Bilinear Operator inserted at non-zero momentum in a quark 2-point function for the most general off-shell configuration. In particular we consider the quark mass Operator, vector and tensor currents as well as the second moment of the flavour non-singlet Wilson Operator.

  • Fermion Bilinear Operator critical exponents at O (1 /N 2 ) in the QED-Gross-Neveu universality class
    Physical Review D, 2018
    Co-Authors: John Gracey
    Abstract:

    We use the critical point large $N$ formalism to calculate the critical exponents corresponding to the fermion mass Operator and flavour non-singlet fermion Bilinear Operator in the universality class of Quantum Electrodynamics (QED) coupled to the Gross-Neveu model for an $SU(N)$ flavour symmetry in $d$-dimensions. The $\epsilon$ expansion of the exponents in $d$ $=$ $4$ $-$ $2\epsilon$ dimensions are in agreement with recent three and four loop perturbative evaluations of both renormalization group functions of these Operators. Estimates of the value of the non-singlet Operator exponent in three dimensions are provided.

Jianmin Wang - One of the best experts on this subject based on the ideXlab platform.

  • Spatiotemporal Pyramid Network for Video Action Recognition
    2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017
    Co-Authors: Yunbo Wang, Mingsheng Long, Jianmin Wang, Philip S Yu
    Abstract:

    Two-stream convolutional networks have shown strong performance in video action recognition tasks. The key idea is to learn spatiotemporal features by fusing convolutional networks spatially and temporally. However, it remains unclear how to model the correlations between the spatial and temporal structures at multiple abstraction levels. First, the spatial stream tends to fail if two videos share similar backgrounds. Second, the temporal stream may be fooled if two actions resemble in short snippets, though appear to be distinct in the long term. We propose a novel spatiotemporal pyramid network to fuse the spatial and temporal features in a pyramid structure such that they can reinforce each other. From the architecture perspective, our network constitutes hierarchical fusion strategies which can be trained as a whole using a unified spatiotemporal loss. A series of ablation experiments support the importance of each fusion strategy. From the technical perspective, we introduce the spatiotemporal compact Bilinear Operator into video analysis tasks. This Operator enables efficient training of Bilinear fusion operations which can capture full interactions between the spatial and temporal features. Our final network achieves state-of-the-art results on standard video datasets.

  • CVPR - Spatiotemporal Pyramid Network for Video Action Recognition
    2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017
    Co-Authors: Yunbo Wang, Mingsheng Long, Jianmin Wang, Philip S Yu
    Abstract:

    Two-stream convolutional networks have shown strong performance in video action recognition tasks. The key idea is to learn spatiotemporal features by fusing convolutional networks spatially and temporally. However, it remains unclear how to model the correlations between the spatial and temporal structures at multiple abstraction levels. First, the spatial stream tends to fail if two videos share similar backgrounds. Second, the temporal stream may be fooled if two actions resemble in short snippets, though appear to be distinct in the long term. We propose a novel spatiotemporal pyramid network to fuse the spatial and temporal features in a pyramid structure such that they can reinforce each other. From the architecture perspective, our network constitutes hierarchical fusion strategies which can be trained as a whole using a unified spatiotemporal loss. A series of ablation experiments support the importance of each fusion strategy. From the technical perspective, we introduce the spatiotemporal compact Bilinear Operator into video analysis tasks. This Operator enables efficient training of Bilinear fusion operations which can capture full interactions between the spatial and temporal features. Our final network achieves state-of-the-art results on standard video datasets.

Mingsheng Long - One of the best experts on this subject based on the ideXlab platform.

  • Spatiotemporal Pyramid Network for Video Action Recognition
    2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017
    Co-Authors: Yunbo Wang, Mingsheng Long, Jianmin Wang, Philip S Yu
    Abstract:

    Two-stream convolutional networks have shown strong performance in video action recognition tasks. The key idea is to learn spatiotemporal features by fusing convolutional networks spatially and temporally. However, it remains unclear how to model the correlations between the spatial and temporal structures at multiple abstraction levels. First, the spatial stream tends to fail if two videos share similar backgrounds. Second, the temporal stream may be fooled if two actions resemble in short snippets, though appear to be distinct in the long term. We propose a novel spatiotemporal pyramid network to fuse the spatial and temporal features in a pyramid structure such that they can reinforce each other. From the architecture perspective, our network constitutes hierarchical fusion strategies which can be trained as a whole using a unified spatiotemporal loss. A series of ablation experiments support the importance of each fusion strategy. From the technical perspective, we introduce the spatiotemporal compact Bilinear Operator into video analysis tasks. This Operator enables efficient training of Bilinear fusion operations which can capture full interactions between the spatial and temporal features. Our final network achieves state-of-the-art results on standard video datasets.

  • CVPR - Spatiotemporal Pyramid Network for Video Action Recognition
    2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017
    Co-Authors: Yunbo Wang, Mingsheng Long, Jianmin Wang, Philip S Yu
    Abstract:

    Two-stream convolutional networks have shown strong performance in video action recognition tasks. The key idea is to learn spatiotemporal features by fusing convolutional networks spatially and temporally. However, it remains unclear how to model the correlations between the spatial and temporal structures at multiple abstraction levels. First, the spatial stream tends to fail if two videos share similar backgrounds. Second, the temporal stream may be fooled if two actions resemble in short snippets, though appear to be distinct in the long term. We propose a novel spatiotemporal pyramid network to fuse the spatial and temporal features in a pyramid structure such that they can reinforce each other. From the architecture perspective, our network constitutes hierarchical fusion strategies which can be trained as a whole using a unified spatiotemporal loss. A series of ablation experiments support the importance of each fusion strategy. From the technical perspective, we introduce the spatiotemporal compact Bilinear Operator into video analysis tasks. This Operator enables efficient training of Bilinear fusion operations which can capture full interactions between the spatial and temporal features. Our final network achieves state-of-the-art results on standard video datasets.