The Experts below are selected from a list of 137136 Experts worldwide ranked by ideXlab platform

Taras Holotyak - One of the best experts on this subject based on the ideXlab platform.

  • fast Content Identification in high dimensional feature spaces using sparse ternary codes
    International Workshop on Information Forensics and Security, 2016
    Co-Authors: Sohrab Ferdowsi, Dimche Kostadinov, Slava Voloshynovskiy, Taras Holotyak
    Abstract:

    We consider the problem of fast Content Identification in high-dimensional feature spaces where a sub-linear search complexity is required. By formulating the problem as sparse approximation of projected coefficients, a closed-form solution can be found which we approximate as a ternary representation. Hence, as opposed to dense binary codes, a framework of Sparse Ternary Codes (STC) is proposed resulting in sparse, but robust representation and sub-linear complexity of search. The proposed method is compared with the Locality Sensitive Hashing (LSH) and the memory vectors on several large-scale synthetic and public image databases, showing its superiority.

  • WIFS - Fast Content Identification in high-dimensional feature spaces using Sparse Ternary Codes
    2016 IEEE International Workshop on Information Forensics and Security (WIFS), 2016
    Co-Authors: Sohrab Ferdowsi, Dimche Kostadinov, Slava Voloshynovskiy, Taras Holotyak
    Abstract:

    We consider the problem of fast Content Identification in high-dimensional feature spaces where a sub-linear search complexity is required. By formulating the problem as sparse approximation of projected coefficients, a closed-form solution can be found which we approximate as a ternary representation. Hence, as opposed to dense binary codes, a framework of Sparse Ternary Codes (STC) is proposed resulting in sparse, but robust representation and sub-linear complexity of search. The proposed method is compared with the Locality Sensitive Hashing (LSH) and the memory vectors on several large-scale synthetic and public image databases, showing its superiority.

  • ICME Workshops - Privacy preserving multimedia Content Identification for cloud based bag-of-feature architectures
    2015 IEEE International Conference on Multimedia & Expo Workshops (ICMEW), 2015
    Co-Authors: Sviatoslav Voloshynovskiy, Maurits Diephuis, Taras Holotyak
    Abstract:

    In this paper we consider privacy-preserving multimedia Content Identification for a cloud based Bag-of-Feature (BoF) framework. We analytically model how geometric information can be used as a shared secret and derive the tradeoff between Identification capability, privacy and computational load. In addition we suggest a descriptor ambiguization method that introduces uncertainty to the server with respect to the true interest of the data users.

  • ICASSP - Performance analysis of Bag-of-Features based Content Identification systems
    2014 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2014
    Co-Authors: Sviatoslav Voloshynovskiy, Maurits Diephuis, Taras Holotyak
    Abstract:

    Many state-of-the-art methods in image retrieval, classification and copy detection are based on the Bag-of-Features (BOF) framework. However, the performance of these systems is mostly experimentally evaluated and little results are reported on theoretical performance. In this paper, we present a statistical framework that makes it possible to analyse the performance of a simple BOF-system and to better understand the impact of different design elements such as the robustness of descriptors, the accuracy of encoding/assignment, information preserving pooling and finally decision making. The proposed framework can be also of interest for a security and privacy analysis of BOF systems.

  • Information-theoretic analysis of privacy-preserving Identification
    2012
    Co-Authors: Taras Holotyak
    Abstract:

    Digital Content fingerprinting has emerged as a possible technique for fast, robust and privacy-preserving Identification, which is a highly demanded application due to the increased interaction with humans and physical objects as well as explosive amount of multimedia data. Besides of the high attractiveness, the existing methods of Content Identification based on Content fingerprinting still lack both deep theoretical understanding of achievable performance limits and practical methods capable to achieve these limits. Additionally, yet little is known about privacy protection technologies, which can provide zero-privacy leakage together with sufficient Identification rate. In this thesis, the information-theoretic fundamentals of privacy-preserving Content Identification are introduced and analyzed. To cover the majority of existing practical techniques, the generalized model of Content Identification system is proposed which consists of decorrelation transform, dimensionality reduction mapper, privacy protection, indexed database and decoder. The state-of-the-art methods are analyzed according to the defined performance measures based on the achievable Identification rate and privacy leak. More particularly, the trade-off between them is investigated with respect to each element of Content Identification system. The thesis provides four major contributions in part of generalized analysis of dimensionality reduction mapping, a new technique of channel decomposition according to the signs and magnitudes of signal components, a concept of channel splitting and polarization and, finally, a new method of zero-leak privacy protection.

Farzad Farhadzadeh - One of the best experts on this subject based on the ideXlab platform.

  • ICASSP - Efficient two stage decoding scheme to achieve Content Identification capacity
    2014 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP), 2014
    Co-Authors: Farzad Farhadzadeh, Sohrab Fredowsi
    Abstract:

    We introduce a scheme to address the trade-off between the Identification rate, search and memory complexities in large-scale Identification systems. We use a special database organization by assigning database entries to a set of possibly overlapping clusters. The clusters are generated based on statistics of both database entries and queries. The decoding procedure is accomplished in two stages. First, a list of clusters related to the query is detected. Then, refinement checks are performed on members of the detected clusters to produce a unique index. We investigate the minimum achievable search complexity for binary symmetric sources.

  • Content authentication and Identification under informed attacks
    International Workshop on Information Forensics and Security, 2012
    Co-Authors: Fokko Beekhof, Sviatoslav Voloshynovskiy, Farzad Farhadzadeh
    Abstract:

    We consider the problem of Content Identification and authentication based on digital Content fingerprinting. Contrary to existing work in which the performance of these systems under blind attacks is analysed, we investigate the information-theoretic performance under informed attacks. In the case of binary Content fingerprinting, in a blind attack, a probe is produced at random independently from the fingerprints of the original Contents. Contrarily, informed attacks assume that the attacker might have some information about the original Content and is thus able to produce a counterfeit probe that is related to an authentic fingerprint corresponding to an original item, thus leading to an increased probability of false acceptance. We demonstrate the impact of the ability of an attacker to create counterfeit items whose fingerprints are related to fingerprints of authentic items, and consider the influence of the length of the fingerprint on the performance of finite-length systems. Finally, the information-theoretic achieveble rate of Content Identification systems sustaining informed attacks is derived under asymptotic assumptions about the fingerprint length.

  • Media Forensics and Security - Private Content Identification based on soft fingerprinting
    Media Watermarking Security and Forensics III, 2011
    Co-Authors: Sviatoslav Voloshynovskiy, Oleksiy Koval, Fokko Beekhof, Taras Holotyak, Farzad Farhadzadeh
    Abstract:

    In many problems such as biometrics, multimedia search, retrieval, recommendation systems requiring privacypreserving similarity computations and Identification, some binary features are stored in the public domain or outsourced to third parties that might raise certain privacy concerns about the original data. To avoid this privacy leak, privacy protection is used. In most cases, privacy protection is uniformly applied to all binary features resulting in data degradation and corresponding loss of performance. To avoid this undesirable effect we propose a new privacy amplification technique that is based on data hiding principles and benefits from side information about bit reliability a.k.a. soft fingerprinting. In this paper, we investigate the Identification-rate vs privacy-leak trade-off. The analysis is performed for the case of a perfect match between side information shared between the encoder and decoder as well as for the case of partial side information.

  • ITW - Sign-magnitude decomposition of mutual information with polarization effect in digital Identification
    2011 IEEE Information Theory Workshop, 2011
    Co-Authors: Svyatoslav Voloshynovskiy, Fokko Beekhof, Oleksiy Koval, Taras Holotyak, Farzad Farhadzadeh
    Abstract:

    Content Identification based on digital fingerprinting attracts a lot of attention in different emerging applications. In this paper, we consider digital Identification based on the sign-magnitude decomposition of fingerprint codewords and analyze the achievable rates for each component. We introduce a channel splitting approach and reveal certain interesting phenomena related to channel polarization. It is demonstrated that under certain conditions almost all rate in the sign channel is concentrated in reliable components, this can be of interest for complexity and security in various Content Identification applications. The envisioned extensions cover applications where the input and output alphabets of the channel are different at the encoding and decoding stages. Additionally, the reduction of the input data dimensionality at the encoding/enrollment stage can increase the cryptographic protection in terms of privacy leakage and simplify the decoding algorithms in biometric applications.

  • information theoretical analysis of private Content Identification
    Information Theory Workshop, 2010
    Co-Authors: Sviatoslav Voloshynovskiy, Farzad Farhadzadeh, Oleksiy Koval, Fokko Beekhof, Taras Holotyak
    Abstract:

    In recent years, Content Identification based on digital fingerprinting attracts a lot of attention in different emerging applications. At the same time, the theoretical analysis of digital fingerprinting systems for finite length case remains an open issue. Additionally, privacy leaks caused by fingerprint storage, distribution and sharing in a public domain via third party outsourced services cause certain concerns in the cryptographic community. In this paper, we perform an information-theoretic analysis of finite length digital fingerprinting systems in a private Content Identification setup and reveal certain connections between fingerprint based Content Identification and Forney's erasure/list decoding [1]. Along this analysis, we also consider complexity issues of fast Content Identification in large databases on remote untrusted servers.

Avinash L. Varna - One of the best experts on this subject based on the ideXlab platform.

  • Modeling and Analysis of Correlated Binary Fingerprints for Content Identification
    IEEE Transactions on Information Forensics and Security, 2011
    Co-Authors: Avinash L. Varna
    Abstract:

    Multimedia Identification via Content fingerprints is used in many applications, such as Content filtering on user-generated Content websites, and automatic multimedia Identification and tagging. A compact “fingerprint” is computed for each multimedia signal that captures robust and unique properties of the perceptual Content, which is later used for identifying the multimedia. Several different multimedia fingerprinting schemes have been proposed in the literature and have been evaluated through experiments. To complement these experimental evaluations and provide guidelines for choosing system parameters and designing better schemes, this paper develops models for Content fingerprinting and provides an analysis of the Identification performance under these models. As a first step, bounds on the Identification accuracy and the required fingerprint length for the simplest case when the fingerprint bits are modeled as i.i.d. are summarized. Markov Random Fields are then used to address more realistic settings of fingerprints with correlated components. The optimal likelihood ratio detector is derived and a statistical physics inspired approach for computing the probability of detection and probability of false alarm is described. The analysis shows that the commonly used Hamming distance detection criterion is susceptible to correlations among fingerprint bits, whereas the optimal log-likelihood ratio decision rule yields 5-20% improvement in the accuracy over a range of correlations. Simulation results demonstrate the validity of the theoretical predictions.

  • ICME - Modeling and analysis of Content Identification
    2009 IEEE International Conference on Multimedia and Expo, 2009
    Co-Authors: Avinash L. Varna
    Abstract:

    Content fingerprinting provides a compact Content-based representation of a multimedia document. An important application of fingerprinting is the Identification of modified copies of the original media Content. These modifications may be incidental changes that occur during the usage of multimedia, or intentional modifications made by an adversary to avoid detection. Currently, the effectiveness of Content Identification techniques is often assessed through benchmark databases. To complement these experimental performance evaluations, this paper develops a theoretical framework for analyzing Content Identification techniques. Beneficial aspects from decision theory and game theory are exploited to gain insights toward optimal system design and parameter selection.

  • a decision theoretic framework for analyzing binary hash based Content Identification systems
    Digital Rights Management, 2008
    Co-Authors: Avinash L. Varna, Ashwin Swaminathan
    Abstract:

    Content Identification has many applications, ranging from preventing illegal sharing of copyrighted Content on video sharing websites, to automatic Identification and tagging of Content. Several Content Identification techniques based on watermarking or robust hashes have been proposed in the literature, but they have mostly been evaluated through experiments. This paper analyzes binary hash-based Content Identification schemes under a decision theoretic framework and presents a lower bound on the length of the hash required to correctly identify multimedia Content that may have undergone modifications. A practical scheme for Content Identification is evaluated under the proposed framework. The results obtained through experiments agree very well with the performance suggested by the theoretical analysis.

  • Digital Rights Management Workshop - A decision theoretic framework for analyzing binary hash-based Content Identification systems
    Proceedings of the 8th ACM workshop on Digital rights management - DRM '08, 2008
    Co-Authors: Avinash L. Varna, Ashwin Swaminathan
    Abstract:

    Content Identification has many applications, ranging from preventing illegal sharing of copyrighted Content on video sharing websites, to automatic Identification and tagging of Content. Several Content Identification techniques based on watermarking or robust hashes have been proposed in the literature, but they have mostly been evaluated through experiments. This paper analyzes binary hash-based Content Identification schemes under a decision theoretic framework and presents a lower bound on the length of the hash required to correctly identify multimedia Content that may have undergone modifications. A practical scheme for Content Identification is evaluated under the proposed framework. The results obtained through experiments agree very well with the performance suggested by the theoretical analysis.

Svyatoslav Voloshynovskiy - One of the best experts on this subject based on the ideXlab platform.

  • Content Identification binary Content fingerprinting versus binary Content encoding
    Proceedings of SPIE, 2014
    Co-Authors: Sohrab Ferdowsi, Svyatoslav Voloshynovskiy, Dimche Kostadinov
    Abstract:

    In this work, we address the problem of Content Identification. We consider Content Identification as a special case of multiclass classification. The conventional approach towards Identification is based on Content fingerprinting where a short binary Content description known as a fingerprint is extracted from the Content. We propose an alternative solution based on elements of machine learning theory and digital communications. Similar to binary Content fingerprinting, binary Content representation is generated based on a set of trained binary classifiers. We consider several training/encoding strategies and demonstrate that the proposed system can achieve the upper theoretical performance limits of Content Identification. The experimental results were carried out both on a synthetic dataset with different parameters and the FAMOS dataset of microstructures from consumer packages.

  • Media Watermarking, Security, and Forensics - Content Identification: binary Content fingerprinting versus binary Content encoding
    Media Watermarking Security and Forensics 2014, 2014
    Co-Authors: Sohrab Ferdowsi, Svyatoslav Voloshynovskiy, Dimche Kostadinov
    Abstract:

    In this work, we address the problem of Content Identification. We consider Content Identification as a special case of multiclass classification. The conventional approach towards Identification is based on Content fingerprinting where a short binary Content description known as a fingerprint is extracted from the Content. We propose an alternative solution based on elements of machine learning theory and digital communications. Similar to binary Content fingerprinting, binary Content representation is generated based on a set of trained binary classifiers. We consider several training/encoding strategies and demonstrate that the proposed system can achieve the upper theoretical performance limits of Content Identification. The experimental results were carried out both on a synthetic dataset with different parameters and the FAMOS dataset of microstructures from consumer packages.

  • CBMI - DCT sign based robust privacy preserving image copy detection for cloud-based systems
    2012 10th International Workshop on Content-Based Multimedia Indexing (CBMI), 2012
    Co-Authors: Maurits Diephuis, Oleksiy Koval, Svyatoslav Voloshynovskiy, Fokko Beekhof
    Abstract:

    In this paper we propose an architecture for message-privacy preserving copy detection and Content Identification for images based on the signs of the Discrete Cosine Transform (DCT) coefficients. The architecture allows for searching in encrypted data and places the computational burden on the server. Sign components of the low frequency DCT coefficients of an image are used to generate a dual set of keys that in turn are used to encrypt the source image and serve as a robust hash that can be queried for Content Identification. The statistical properties of these DCT sign vectors are modelled and we analyse their robustness against real world image distortions. Finally, the trade-off between the discriminative power of such vectors, the offered security and the resilience against errors is demonstrated.

  • ITW - Sign-magnitude decomposition of mutual information with polarization effect in digital Identification
    2011 IEEE Information Theory Workshop, 2011
    Co-Authors: Svyatoslav Voloshynovskiy, Fokko Beekhof, Oleksiy Koval, Taras Holotyak, Farzad Farhadzadeh
    Abstract:

    Content Identification based on digital fingerprinting attracts a lot of attention in different emerging applications. In this paper, we consider digital Identification based on the sign-magnitude decomposition of fingerprint codewords and analyze the achievable rates for each component. We introduce a channel splitting approach and reveal certain interesting phenomena related to channel polarization. It is demonstrated that under certain conditions almost all rate in the sign channel is concentrated in reliable components, this can be of interest for complexity and security in various Content Identification applications. The envisioned extensions cover applications where the input and output alphabets of the channel are different at the encoding and decoding stages. Additionally, the reduction of the input data dimensionality at the encoding/enrollment stage can increase the cryptographic protection in terms of privacy leakage and simplify the decoding algorithms in biometric applications.

  • 2010 IEEE Information Theory Workshop (ITW) - Information-theoretical analysis of private Content Identification
    2010 IEEE Information Theory Workshop, 2010
    Co-Authors: Svyatoslav Voloshynovskiy, Farzad Farhadzadeh, Oleksiy Koval, Fokko Beekhof, Taras Holotyak
    Abstract:

    In recent years, Content Identification based on digital fingerprinting attracts a lot of attention in different emerging applications. At the same time, the theoretical analysis of digital fingerprinting systems for finite length case remains an open issue. Additionally, privacy leaks caused by fingerprint storage, distribution and sharing in a public domain via third party outsourced services cause certain concerns in the cryptographic community. In this paper, we perform an information-theoretic analysis of finite length digital fingerprinting systems in a private Content Identification setup and reveal certain connections between fingerprint based Content Identification and Forney's erasure/list decoding [1]. Along this analysis, we also consider complexity issues of fast Content Identification in large databases on remote untrusted servers.

Mehul Motani - One of the best experts on this subject based on the ideXlab platform.

  • Exponential Strong Converse for Content Identification With Lossy Recovery
    IEEE Transactions on Information Theory, 2018
    Co-Authors: Lin Zhou, Vincent Y. F. Tan, Mehul Motani
    Abstract:

    We revisit the high-dimensional Content Identification with lossy recovery problem (Tuncel and Gunduz, 2014) and establish an exponential strong converse theorem. As a corollary of the exponential strong converse theorem, we derive an upper bound on the joint Identification-error and excess-distortion exponent for the problem. Our main results can be specialized to the biometrical Identification problem (Willems, 2003) and the Content Identification problem (Tuncel, 2009) since these two problems are both special cases of the Content Identification with lossy recovery problem. We leverage the information spectrum method introduced by Oohama and adapt the strong converse techniques therein to be applicable to the problem at hand.

  • ISIT - Strong converse for Content Identification with lossy recovery
    2017 IEEE International Symposium on Information Theory (ISIT), 2017
    Co-Authors: Lin Zhou, Vincent Y. F. Tan, Mehul Motani
    Abstract:

    In this paper, we revisit the Content Identification problem with lossy recovery (Tuncel and Gunduz, 2014) and establish the exponential strong converse theorem for the problem. Further, we derive an upper bound on the joint excess-distortion and error exponent for the problem.