The Experts below are selected from a list of 20766 Experts worldwide ranked by ideXlab platform
Thomas Brox - One of the best experts on this subject based on the ideXlab platform.
-
coot cooperative hierarchical transformer for video text representation learning
Neural Information Processing Systems, 2020Co-Authors: Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, Thomas BroxAbstract:Many real-world video-text tasks involve different levels of granularity, such as frames and words, clip and sentences or videos and paragraphs, each with distinct semantics. In this paper, we propose a Cooperative hierarchical Transformer (COOT) to leverage this hierarchy information and model the interactions between different levels of granularity and different modalities. The method consists of three major components: an attention-aware feature Aggregation Layer, which leverages the local temporal context (intra-level, e.g., within a clip), a contextual transformer to learn the interactions between low-level and high-level semantics (inter-level, e.g. clip-video, sentence-paragraph), and a cross-modal cycle-consistency loss to connect video and text. The resulting method compares favorably to the state of the art on several benchmarks while having few parameters. All code is available open-source at this https URL
Simon Ging - One of the best experts on this subject based on the ideXlab platform.
-
coot cooperative hierarchical transformer for video text representation learning
Neural Information Processing Systems, 2020Co-Authors: Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, Thomas BroxAbstract:Many real-world video-text tasks involve different levels of granularity, such as frames and words, clip and sentences or videos and paragraphs, each with distinct semantics. In this paper, we propose a Cooperative hierarchical Transformer (COOT) to leverage this hierarchy information and model the interactions between different levels of granularity and different modalities. The method consists of three major components: an attention-aware feature Aggregation Layer, which leverages the local temporal context (intra-level, e.g., within a clip), a contextual transformer to learn the interactions between low-level and high-level semantics (inter-level, e.g. clip-video, sentence-paragraph), and a cross-modal cycle-consistency loss to connect video and text. The resulting method compares favorably to the state of the art on several benchmarks while having few parameters. All code is available open-source at this https URL
Hamed Pirsiavash - One of the best experts on this subject based on the ideXlab platform.
-
coot cooperative hierarchical transformer for video text representation learning
Neural Information Processing Systems, 2020Co-Authors: Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, Thomas BroxAbstract:Many real-world video-text tasks involve different levels of granularity, such as frames and words, clip and sentences or videos and paragraphs, each with distinct semantics. In this paper, we propose a Cooperative hierarchical Transformer (COOT) to leverage this hierarchy information and model the interactions between different levels of granularity and different modalities. The method consists of three major components: an attention-aware feature Aggregation Layer, which leverages the local temporal context (intra-level, e.g., within a clip), a contextual transformer to learn the interactions between low-level and high-level semantics (inter-level, e.g. clip-video, sentence-paragraph), and a cross-modal cycle-consistency loss to connect video and text. The resulting method compares favorably to the state of the art on several benchmarks while having few parameters. All code is available open-source at this https URL
Mohammadreza Zolfaghari - One of the best experts on this subject based on the ideXlab platform.
-
coot cooperative hierarchical transformer for video text representation learning
Neural Information Processing Systems, 2020Co-Authors: Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash, Thomas BroxAbstract:Many real-world video-text tasks involve different levels of granularity, such as frames and words, clip and sentences or videos and paragraphs, each with distinct semantics. In this paper, we propose a Cooperative hierarchical Transformer (COOT) to leverage this hierarchy information and model the interactions between different levels of granularity and different modalities. The method consists of three major components: an attention-aware feature Aggregation Layer, which leverages the local temporal context (intra-level, e.g., within a clip), a contextual transformer to learn the interactions between low-level and high-level semantics (inter-level, e.g. clip-video, sentence-paragraph), and a cross-modal cycle-consistency loss to connect video and text. The resulting method compares favorably to the state of the art on several benchmarks while having few parameters. All code is available open-source at this https URL
Krzysztof Janowicz - One of the best experts on this subject based on the ideXlab platform.
-
contextual graph attention for answering logical queries over incomplete knowledge graphs
International Conference on Knowledge Capture, 2019Co-Authors: Krzysztof JanowiczAbstract:Recently, several studies have explored methods for using KG embedding to answer logical queries. These approaches either treat embedding learning and query answering as two separated learning tasks, or fail to deal with the variability of contributions from different query paths. We proposed to leverage a graph attention mechanism to handle the unequal contribution of different query paths. However, commonly used graph attention assumes that the center node embedding is provided, which is unavailable in this task since the center node is to be predicted. To solve this problem we propose a multi-head attention-based end-to-end logical query answering model, called Contextual Graph Attention model (CGA), which uses an initial neighborhood Aggregation Layer to generate the center embedding, and the whole model is trained jointly on the original KG structure as well as the sampled query-answer pairs. We also introduce two new datasets, DB18 and WikiGeo19, which are rather large in size compared to the existing datasets and contain many more relation types, and use them to evaluate the performance of the proposed model. Our result shows that the proposed CGA with fewer learnable parameters consistently outperforms the baseline models on both datasets as well as Bio dataset.