The Experts below are selected from a list of 3966 Experts worldwide ranked by ideXlab platform
Mark Gerstein - One of the best experts on this subject based on the ideXlab platform.
-
Tiling Array data analysis a multiscale approach using wavelets
BMC Bioinformatics, 2011Co-Authors: Alexander Karpikov, Joel Rozowsky, Mark GersteinAbstract:Tiling Array data is hard to interpret due to noise. The wavelet transformation is a widely used technique in signal processing for elucidating the true signal from noisy data. Consequently, we attempted to denoise representative Tiling Array datasets for ChIP-chip experiments using wavelets. In doing this, we used specific wavelet basis functions, Coiflets, since their triangular shape closely resembles the expected profiles of true ChIP-chip peaks. In our wavelet-transformed data, we observed that noise tends to be confined to small scales while the useful signal-of-interest spans multiple large scales. We were also able to show that wavelet coefficients due to non-specific cross-hybridization follow a log-normal distribution, and we used this fact in developing a thresholding procedure. In particular, wavelets allow one to set an unambiguous, absolute threshold, which has been hard to define in ChIP-chip experiments. One can set this threshold by requiring a similar confidence level at different length-scales of the transformed signal. We applied our algorithm to a number of representative ChIP-chip data sets, including those of Pol II and histone modifications, which have a diverse distribution of length-scales of biochemical activity, including some broad peaks. Finally, we benchmarked our method in comparison to other approaches for scoring ChIP-chip data using spike-ins on the ENCODE Nimblegen Tiling Array. This comparison demonstrated excellent performance, with wavelets getting the best overall score.
-
an approach to comparing Tiling Array and high throughput sequencing technologies for genomic transcript mapping
BMC Research Notes, 2009Co-Authors: Joel Rozowsky, Rajkumar Sasidharan, Ashish Agarwal, Mark GersteinAbstract:Background There are two main technologies for transcriptome profiling, namely, Tiling microArrays and high-throughput sequencing. Recently there has been a tremendous amount of excitement about the latter because of the advent of next-generation sequencing technologies and its promises. Consequently, the question of the moment is how these two technologies compare. Here we attempt to develop an approach to do a fair comparison of transcripts identified from Tiling microArray and MPSS sequencing data.
-
peakseq enables systematic scoring of chip seq experiments relative to controls
Nature Biotechnology, 2009Co-Authors: Joel Rozowsky, Zhengdong D Zhang, Michael Snyder, Ghia Euskirchen, Raymond K Auerbach, Theodore Gibson, Robert D Bjornson, Nicholas Carriero, Mark GersteinAbstract:Chromatin immunoprecipitation (ChIP) followed by tag sequencing (ChIP-seq) using high-throughput next-generation instrumentation is fast, replacing chromatin immunoprecipitation followed by genome Tiling Array analysis (ChIP-chip) as the preferred approach for mapping of sites of transcription-factor binding and chromatin modification. Using two deeply sequenced data sets for human RNA polymerase II and STAT1, each with matching input-DNA controls, we describe a general scoring approach to address unique challenges in ChIP-seq data analysis. Our approach is based on the observation that sites of potential binding are strongly correlated with signal peaks in the control, likely revealing features of open chromatin. We develop a two-pass strategy called PeakSeq to compensate for this. A two-pass strategy compensates for signal caused by open chromatin, as revealed by inclusion of the controls. The first pass identifies putative binding sites and compensates for genomic variation in the 'mappability' of sequences. The second pass filters out sites not significantly enriched compared to the normalized control, computing precise enrichments and significances. Our scoring procedure enables us to optimize experimental design by estimating the depth of sequencing required for a desired level of coverage and demonstrating that more than two replicates provides only a marginal gain in information.
-
tilescope online analysis pipeline for high density Tiling microArray data
Genome Biology, 2007Co-Authors: Zhengdong D Zhang, Joel Rozowsky, Hugo Y K Lam, Michael Snyder, Mark GersteinAbstract:We developed Tilescope, a fully integrated data processing pipeline for analyzing high-density Tiling-Array data http://tilescope.gersteinlab.org. In a completely automated fashion, Tilescope will normalize signals between channels and across Arrays, combine replicate experiments, score each Array element, and identify genomic features. The program is designed with a modular, three-tiered architecture, facilitating parallelism, and a graphic user-friendly interface, presenting results in an organized web page, downloadable for further analysis.
-
comparative analysis of genome Tiling Array data reveals many novel primate specific functional rnas in human
BMC Evolutionary Biology, 2007Co-Authors: Zhaolei Zhang, Andy Wing Chun Pang, Mark GersteinAbstract:Widespread transcription activities in the human genome were recently observed in high-resolution Tiling Array experiments, which revealed many novel transcripts that are outside of the boundaries of known protein or RNA genes. Termed as "TARs" (Transcriptionally Active Regions), these novel transcribed regions represent "dark matter" in the genome, and their origin and functionality need to be explained. Many of these transcripts are thought to code for novel proteins or non-protein-coding RNAs. We have applied an integrated bioinformatics approach to investigate the properties of these TARs, including cross-species conservation, and the ability to form stable secondary structures. The goal of this study is to identify a list of potential candidate sequences that are likely to code for functional non-protein-coding RNAs. We are particularly interested in the discovery of those functional RNA candidates that are primate-specific, i.e. those that do not have homologs in the mouse or dog genomes but in rhesus. Using sequence conservation and the probability of forming stable secondary structures, we have identified ~300 possible candidates for primate-specific noncoding RNAs. We are currently in the process of sequencing the orthologous regions of these candidate sequences in several other primate species. We will then be able to apply a "phylogenetic shadowing" approach to analyze the functionality of these ncRNA candidates. The existence of potential primate-specific functional transcripts has demonstrated the limitation of previous genome comparison studies, which put too much emphasis on conservation between human and rodents. It also argues for the necessity of sequencing additional primate species to gain a better and more comprehensive understanding of the human genome.
Joel Rozowsky - One of the best experts on this subject based on the ideXlab platform.
-
Tiling Array data analysis a multiscale approach using wavelets
BMC Bioinformatics, 2011Co-Authors: Alexander Karpikov, Joel Rozowsky, Mark GersteinAbstract:Tiling Array data is hard to interpret due to noise. The wavelet transformation is a widely used technique in signal processing for elucidating the true signal from noisy data. Consequently, we attempted to denoise representative Tiling Array datasets for ChIP-chip experiments using wavelets. In doing this, we used specific wavelet basis functions, Coiflets, since their triangular shape closely resembles the expected profiles of true ChIP-chip peaks. In our wavelet-transformed data, we observed that noise tends to be confined to small scales while the useful signal-of-interest spans multiple large scales. We were also able to show that wavelet coefficients due to non-specific cross-hybridization follow a log-normal distribution, and we used this fact in developing a thresholding procedure. In particular, wavelets allow one to set an unambiguous, absolute threshold, which has been hard to define in ChIP-chip experiments. One can set this threshold by requiring a similar confidence level at different length-scales of the transformed signal. We applied our algorithm to a number of representative ChIP-chip data sets, including those of Pol II and histone modifications, which have a diverse distribution of length-scales of biochemical activity, including some broad peaks. Finally, we benchmarked our method in comparison to other approaches for scoring ChIP-chip data using spike-ins on the ENCODE Nimblegen Tiling Array. This comparison demonstrated excellent performance, with wavelets getting the best overall score.
-
an approach to comparing Tiling Array and high throughput sequencing technologies for genomic transcript mapping
BMC Research Notes, 2009Co-Authors: Joel Rozowsky, Rajkumar Sasidharan, Ashish Agarwal, Mark GersteinAbstract:Background There are two main technologies for transcriptome profiling, namely, Tiling microArrays and high-throughput sequencing. Recently there has been a tremendous amount of excitement about the latter because of the advent of next-generation sequencing technologies and its promises. Consequently, the question of the moment is how these two technologies compare. Here we attempt to develop an approach to do a fair comparison of transcripts identified from Tiling microArray and MPSS sequencing data.
-
peakseq enables systematic scoring of chip seq experiments relative to controls
Nature Biotechnology, 2009Co-Authors: Joel Rozowsky, Zhengdong D Zhang, Michael Snyder, Ghia Euskirchen, Raymond K Auerbach, Theodore Gibson, Robert D Bjornson, Nicholas Carriero, Mark GersteinAbstract:Chromatin immunoprecipitation (ChIP) followed by tag sequencing (ChIP-seq) using high-throughput next-generation instrumentation is fast, replacing chromatin immunoprecipitation followed by genome Tiling Array analysis (ChIP-chip) as the preferred approach for mapping of sites of transcription-factor binding and chromatin modification. Using two deeply sequenced data sets for human RNA polymerase II and STAT1, each with matching input-DNA controls, we describe a general scoring approach to address unique challenges in ChIP-seq data analysis. Our approach is based on the observation that sites of potential binding are strongly correlated with signal peaks in the control, likely revealing features of open chromatin. We develop a two-pass strategy called PeakSeq to compensate for this. A two-pass strategy compensates for signal caused by open chromatin, as revealed by inclusion of the controls. The first pass identifies putative binding sites and compensates for genomic variation in the 'mappability' of sequences. The second pass filters out sites not significantly enriched compared to the normalized control, computing precise enrichments and significances. Our scoring procedure enables us to optimize experimental design by estimating the depth of sequencing required for a desired level of coverage and demonstrating that more than two replicates provides only a marginal gain in information.
-
tilescope online analysis pipeline for high density Tiling microArray data
Genome Biology, 2007Co-Authors: Zhengdong D Zhang, Joel Rozowsky, Hugo Y K Lam, Michael Snyder, Mark GersteinAbstract:We developed Tilescope, a fully integrated data processing pipeline for analyzing high-density Tiling-Array data http://tilescope.gersteinlab.org. In a completely automated fashion, Tilescope will normalize signals between channels and across Arrays, combine replicate experiments, score each Array element, and identify genomic features. The program is designed with a modular, three-tiered architecture, facilitating parallelism, and a graphic user-friendly interface, presenting results in an organized web page, downloadable for further analysis.
-
a supervised hidden markov model framework for efficiently segmenting Tiling Array data in transcriptional and chip chip experiments systematically incorporating validated biological knowledge
Bioinformatics, 2006Co-Authors: Joel Rozowsky, Zhengdong D Zhang, Michael Snyder, Jan O Korbel, Thomas Royce, Martin H Schultz, Mark GersteinAbstract:Motivation: Large-scale Tiling Array experiments are becoming increasingly common in genomics. In particular, the ENCODE project requires the consistent segmentation of many different Tiling Array datasets into 'active regions' (e.g. finding transfrags from transcriptional data and putative binding sites from ChIP-chip experiments). Previously, such segmentation was done in an unsupervised fashion mainly based on characteristics of the signal distribution in the Tiling Array data itself. Here we propose a supervised framework for doing this. It has the advantage of explicitly incorporating validated biological knowledge into the model and allowing for formal training and testing. Methodology: In particular, we use a hidden Markov model (HMM) framework, which is capable of explicitly modeling the dependency between neighboring probes and whose extended version (the generalized HMM) also allows explicit description of state duration density. We introduce a formal definition of the Tiling-Array analysis problem, and explain how we can use this to describe sampling small genomic regions for experimental validation to build up a gold-standard set for training and testing. We then describe various ideal and practical sampling strategies (e.g. maximizing signal entropy within a selected region versus using gene annotation or known promoters as positives for transcription or ChIP-chip data, respectively). Results: For the practical sampling and training strategies, we show how the size and noise in the validated training data affects the performance of an HMM applied to the ENCODE transcriptional and ChIP-chip experiments. In particular, we show that the HMM framework is able to efficiently process Tiling Array data as well as or better than previous approaches. For the idealized sampling strategies, we show how we can assess their performance in a simulation framework and how a maximum entropy approach, which samples sub-regions with very different signal intensities, gives the maximally performing gold-standard. This latter result has strong implications for the optimum way medium-scale validation experiments should be carried out to verify the results of the genome-scale Tiling Array experiments. Supplementary information: The supplementary data are available at http://Tiling.gersteinlab.org/hmm/ Contact: mark.gerstein@yale.edu
Krishna Karuturi - One of the best experts on this subject based on the ideXlab platform.
-
deciphering transcription factor binding patterns from genome wide high density chip chip Tiling Array data
BMC Proceedings, 2011Co-Authors: Lei Zhu, Majid Eshaghi, Jianhua Liu, Krishna KaruturiAbstract:The binding events of DNA-interacting proteins and their patterns can be extensively characterized by high density ChIP-chip Tiling Array data. The characteristics of the binding events could be different for different transcription factors. They may even vary for a given transcription factor among different interaction loci. The knowledge of binding sites and binding occupancy patterns are all very useful to understand the DNA-protein interaction and its role in the transcriptional regulation of genes. In the view of the complexity of the DNA-protein interaction and the opportunity offered by high density tiled ChIP-chip data, we present a statistical procedure which focuses on identifying the interaction signal regions instead of signal peaks using moving window binomial testing method and deconvolving the patterns of interaction using peakedness and skewness scores. We analyzed ChIP-chip data of 4 different DNA interacting proteins including transcription factors and RNA polymerase in fission yeast using our procedure. Our analysis revealed the variation of binding patterns within and across different DNA interacting proteins. We present their utility in understanding transcriptional regulation from ChIP-chip data. Our method can successfully detect the signal regions and characterize the binding patterns in ChIP-chip data which help appropriate analysis of the ChIP-chip data.
Achim Tresch - One of the best experts on this subject based on the ideXlab platform.
-
starr simple Tiling Array analysis of affymetrix chip chip data
BMC Bioinformatics, 2010Co-Authors: Benedikt Zacher, Pei Fen Kuan, Achim TreschAbstract:Chromatin immunoprecipitation combined with DNA microArrays (ChIP-chip) is an assay used for investigating DNA-protein-binding or post-translational chromatin/histone modifications. As with all high-throughput technologies, it requires thorough bioinformatic processing of the data for which there is no standard yet. The primary goal is to reliably identify and localize genomic regions that bind a specific protein. Further investigation compares binding profiles of functionally related proteins, or binding profiles of the same proteins in different genetic backgrounds or experimental conditions. Ultimately, the goal is to gain a mechanistic understanding of the effects of DNA binding events on gene expression. We present a free, open-source R/Bioconductor package Starr that facilitates comparative analysis of ChIP-chip data across experiments and across different microArray platforms. The package provides functions for data import, quality assessment, data visualization and exploration. Starr includes high-level analysis tools such as the alignment of ChIP signals along annotated features, correlation analysis of ChIP signals with complementary genomic data, peak-finding and comparative display of multiple clusters of binding profiles. It uses standard Bioconductor classes for maximum compatibility with other software. Moreover, Starr automatically updates microArray probe annotation files by a highly efficient remapping of microArray probe sequences to an arbitrary genome. Starr is an R package that covers the complete ChIP-chip workflow from data processing to binding pattern detection. It focuses on the high-level data analysis, e.g., it provides methods for the integration and combined statistical analysis of binding profiles and complementary functional genomics data. Starr enables systematic assessment of binding behaviour for groups of genes that are alingned along arbitrary genomic features.
-
softwaresimple Tiling Array analysis of affymetrix chip chip data
2010Co-Authors: Benedikt Zacher, Pei Fen Kuan, Achim TreschAbstract:Background: Chromatin immunoprecipitation combined with DNA microArrays (ChIP-chip) is an assay used for investigating DNA-protein-binding or post-translational chromatin/histone modifications. As with all high-throughput technologies, it requires thorough bioinformatic processing of the data for which there is no standard yet. The primary goal is to reliably identify and localize genomic regions that bind a specific protein. Further investigation compares binding profiles of functionally related proteins, or binding profiles of the same proteins in different genetic backgrounds or experimental conditions. Ultimately, the goal is to gain a mechanistic understanding of the effects of DNA binding events on gene expression. Results: We present a free, open-source R/Bioconductor package Starr that facilitates comparative analysis of ChIPchip data across experiments and across different microArray platforms. The package provides functions for data import, quality assessment, data visualization and exploration. Starr includes high-level analysis tools such as the alignment of ChIP signals along annotated features, correlation analysis of ChIP signals with complementary genomic data, peak-finding and comparative display of multiple clusters of binding profiles. It uses standard Bioconductor classes for maximum compatibility with other software. Moreover, Starr automatically updates microArray probe annotation files by a highly efficient remapping of microArray probe sequences to an arbitrary genome. Conclusion: Starr is an R package that covers the complete ChIP-chip workflow from data processing to binding pattern detection. It focuses on the high-level data analysis, e.g., it provides methods for the integration and combined statistical analysis of binding profiles and complementary functional genomics data. Starr enables systematic assessment of binding behaviour for groups of genes that are alingned along arbitrary genomic features.
-
starr simple Tiling Array analysis of affymetrix chip chip data
arXiv: Quantitative Methods, 2009Co-Authors: Benedikt Zacher, Achim TreschAbstract:Chromatin immunoprecipitation combined with DNA microArrays (ChIP-chip) is an assay for DNA-protein-binding or post-translational chromatin/histone modifications. As with all high-throughput technologies, it requires a thorough bioinformatic processing of the data for which there is no standard yet. The primary goal is the reliable identification and localization of genomic regions that bind a specific protein. The second step comprises comparison of binding profiles of functionally related proteins, or of binding profiles of the same protein in different genetic backgrounds or environmental conditions. Ultimately, one would like to gain a mechanistic understanding of the effects of DNA binding events on gene expression. We present a free, open-source R package Starr that, in combination with the package Ringo, facilitates the comparative analysis of ChIP-chip data across experiments and across different microArray platforms. Core features are data import, quality assessment, normalization and visualization of the data, and the detection of ChIP-enriched genomic regions. The use of common Bioconductor classes ensures the compatibility with other R packages. Most importantly, Starr provides methods for integration of complementary genomics data, e.g., it enables systematic investigation of the relation between gene expression and dna binding.
Lei Zhu - One of the best experts on this subject based on the ideXlab platform.
-
deciphering transcription factor binding patterns from genome wide high density chip chip Tiling Array data
BMC Proceedings, 2011Co-Authors: Lei Zhu, Majid Eshaghi, Jianhua Liu, Krishna KaruturiAbstract:The binding events of DNA-interacting proteins and their patterns can be extensively characterized by high density ChIP-chip Tiling Array data. The characteristics of the binding events could be different for different transcription factors. They may even vary for a given transcription factor among different interaction loci. The knowledge of binding sites and binding occupancy patterns are all very useful to understand the DNA-protein interaction and its role in the transcriptional regulation of genes. In the view of the complexity of the DNA-protein interaction and the opportunity offered by high density tiled ChIP-chip data, we present a statistical procedure which focuses on identifying the interaction signal regions instead of signal peaks using moving window binomial testing method and deconvolving the patterns of interaction using peakedness and skewness scores. We analyzed ChIP-chip data of 4 different DNA interacting proteins including transcription factors and RNA polymerase in fission yeast using our procedure. Our analysis revealed the variation of binding patterns within and across different DNA interacting proteins. We present their utility in understanding transcriptional regulation from ChIP-chip data. Our method can successfully detect the signal regions and characterize the binding patterns in ChIP-chip data which help appropriate analysis of the ChIP-chip data.
-
deciphering transcription factor binding patterns from genome wide high density chip chip Tiling Array data
International Symposium on Bioinformatics Research and Applications, 2010Co-Authors: Lei Zhu, Majid Eshaghi, Jianhua Liu, Radha Krishna Murthy KaruturiAbstract:The binding events of DNA-binding proteins can be extensively characterized by high density ChIP-chip Tiling Array data. The binding sites and binding occupancy patterns are all very useful to understand the DNA-protein interaction. We propose a statistical procedure which focuses on identifying the interaction signal regions and the patterns of interaction using peakedness and skewness tests. Its utility to annotate the binding signals by analyzing the Tbp1 and Rpb1 ChIP-chip datasets in fission yeast is demonstrated.