The Experts below are selected from a list of 174 Experts worldwide ranked by ideXlab platform
Pedro Miramontes - One of the best experts on this subject based on the ideXlab platform.
-
Diminishing Return for increased mappability with longer sequencing reads implications of the k mer distributions in the human genome
BMC Bioinformatics, 2014Co-Authors: Jan Freudenberg, Pedro MiramontesAbstract:The amount of non-unique sequence (non-singletons) in a genome directly affects the difficulty of read alignment to a reference assembly for high throughput-sequencing data. Although a longer read is more likely to be uniquely mapped to the reference genome, a quantitative analysis of the influence of read lengths on mappability has been lacking. To address this question, we evaluate the k-mer distribution of the human reference genome. The k-mer frequency is determined for k ranging from 20 bp to 1000 bp. We observe that the proportion of non-singletons k-mers decreases slowly with increasing k, and can be fitted by piecewise power-law functions with different exponents at different ranges of k. A slower decay at greater values for k indicates more limited gains in mappability for read lengths between 200 bp and 1000 bp. The frequency distributions of k-mers exhibit long tails with a power-law-like trend, and rank frequency plots exhibit a concave Zipf’s curve. The most frequent 1000-mers comprise 172 regions, which include four large stretches on chromosomes 1 and X, containing genes of biomedical relevance. Comparison with other databases indicates that the 172 regions can be broadly classified into two types: those containing LINE transposable elements and those containing segmental duplications. Read mappability as measured by the proportion of singletons increases steadily up to the length scale around 200 bp. When read length increases above 200 bp, smaller gains in mappability are expected. Moreover, the proportion of non-singletons decreases with read lengths much slower than linear. Even a read length of 1000 bp would not allow the unique alignment of reads for many coding regions of human genes. A mix of techniques will be needed for efficiently producing high-quality data that cover the complete human genome.
-
Diminishing Return for increased mappability with longer sequencing reads implications of the k mer distributions in the human genome
arXiv: Genomics, 2013Co-Authors: Jan Freudenberg, Pedro MiramontesAbstract:The amount of non-unique sequence (non-singletons) in a genome directly affects the difficulty of read alignment to a reference assembly for high throughput-sequencing data. Although a greater length increases the chance for reads being uniquely mapped to the reference genome, a quantitative analysis of the influence of read lengths on mappability has been lacking. To address this question, we evaluate the k-mer distribution of the human reference genome. The k-mer frequency is determined for k ranging from 20 to 1000 basepairs. We use the proportion of non-singleton k-mers to evaluate the mappability of reads for a corresponding read length. We observe that the proportion of non-singletons decreases slowly with increasing k, and can be fitted by piecewise power-law functions with different exponents at different k ranges. A faster decay at smaller values for k indicates more limited gains for read lengths > 200 basepairs. The frequency distributions of k-mers exhibit long tails in a power-law-like trend, and rank frequency plots exhibit a concave Zipf's curve. The location of the most frequent 1000-mers comprises 172 kilobase-ranged regions, including four large stretches on chromosomes 1 and X, containing genes with biomedical implications. Even the read length 1000 would be insufficient to reliably sequence these specific regions.
Yuichi Yoshida - One of the best experts on this subject based on the ideXlab platform.
-
a generalization of submodular cover via the Diminishing Return property on the integer lattice
Neural Information Processing Systems, 2015Co-Authors: Tasuku Soma, Yuichi YoshidaAbstract:We consider a generalization of the submodular cover problem based on the concept of Diminishing Return property on the integer lattice. We are motivated by real scenarios in machine learning that cannot be captured by (traditional) sub-modular set functions. We show that the generalized submodular cover problem can be applied to various problems and devise a bicriteria approximation algorithm. Our algorithm is guaranteed to output a log-factor approximate solution that satisfies the constraints with the desired accuracy. The running time of our algorithm is roughly O(n log(nr) log r), where n is the size of the ground set and r is the maximum value of a coordinate. The dependency on r is exponentially better than the naive reduction algorithms. Several experiments on real and artificial datasets demonstrate that the solution quality of our algorithm is comparable to naive algorithms, while the running time is several orders of magnitude faster.
-
maximizing submodular functions with the Diminishing Return property over the integer lattice
arXiv: Data Structures and Algorithms, 2015Co-Authors: Tasuku Soma, Yuichi YoshidaAbstract:The problem of maximizing non-negative monotone submodular functions under a certain constraint has been intensively studied in the last decade. In this paper, we address the problem for functions defined over the integer lattice. Suppose that a non-negative monotone submodular function $f:\mathbb{Z}_+^n \to \mathbb{R}_+$ is given via an evaluation oracle. Furthermore, we assume that $f$ satisfies the Diminishing Return property, which is not an immediate consequence of the submodularity when the domain is the integer lattice. Then, we show (i) a $(1-1/e-\epsilon)$-approximation algorithm for a cardinality constraint with $\widetilde{O}(\frac{n}{\epsilon}\log \frac{r}{\epsilon})$ queries, where $r$ is the maximum cardinality of feasible solutions, (ii) a $(1-1/e-\epsilon)$-approximation algorithm for a polymatroid constraint with $\widetilde{O}(\frac{nr}{\epsilon^4}+n^6)$ queries, where $r$ is the rank of the polymatroid, and (iii) a $(1-1/e-\epsilon)$-approximation algorithm for a knapsack constraint with $\widetilde{O}(\frac{n^2}{\epsilon^{18}}\log \frac{1}{w})(\frac{1}{\epsilon})^{O(1/\epsilon^8)}$ queries, where $w$ is the minumum weight of elements. Our algorithms for polymatroid constraints and knapsack constraints first extend the domain of the objective function to the Euclidean space and then run the continuous greedy algorithm. We give two different kinds of continuous extensions, one is for knapsack constraints and the other is for polymatroid constraints, which might be of independent interest.
Jan Freudenberg - One of the best experts on this subject based on the ideXlab platform.
-
Diminishing Return for increased mappability with longer sequencing reads implications of the k mer distributions in the human genome
BMC Bioinformatics, 2014Co-Authors: Jan Freudenberg, Pedro MiramontesAbstract:The amount of non-unique sequence (non-singletons) in a genome directly affects the difficulty of read alignment to a reference assembly for high throughput-sequencing data. Although a longer read is more likely to be uniquely mapped to the reference genome, a quantitative analysis of the influence of read lengths on mappability has been lacking. To address this question, we evaluate the k-mer distribution of the human reference genome. The k-mer frequency is determined for k ranging from 20 bp to 1000 bp. We observe that the proportion of non-singletons k-mers decreases slowly with increasing k, and can be fitted by piecewise power-law functions with different exponents at different ranges of k. A slower decay at greater values for k indicates more limited gains in mappability for read lengths between 200 bp and 1000 bp. The frequency distributions of k-mers exhibit long tails with a power-law-like trend, and rank frequency plots exhibit a concave Zipf’s curve. The most frequent 1000-mers comprise 172 regions, which include four large stretches on chromosomes 1 and X, containing genes of biomedical relevance. Comparison with other databases indicates that the 172 regions can be broadly classified into two types: those containing LINE transposable elements and those containing segmental duplications. Read mappability as measured by the proportion of singletons increases steadily up to the length scale around 200 bp. When read length increases above 200 bp, smaller gains in mappability are expected. Moreover, the proportion of non-singletons decreases with read lengths much slower than linear. Even a read length of 1000 bp would not allow the unique alignment of reads for many coding regions of human genes. A mix of techniques will be needed for efficiently producing high-quality data that cover the complete human genome.
-
Diminishing Return for increased mappability with longer sequencing reads implications of the k mer distributions in the human genome
arXiv: Genomics, 2013Co-Authors: Jan Freudenberg, Pedro MiramontesAbstract:The amount of non-unique sequence (non-singletons) in a genome directly affects the difficulty of read alignment to a reference assembly for high throughput-sequencing data. Although a greater length increases the chance for reads being uniquely mapped to the reference genome, a quantitative analysis of the influence of read lengths on mappability has been lacking. To address this question, we evaluate the k-mer distribution of the human reference genome. The k-mer frequency is determined for k ranging from 20 to 1000 basepairs. We use the proportion of non-singleton k-mers to evaluate the mappability of reads for a corresponding read length. We observe that the proportion of non-singletons decreases slowly with increasing k, and can be fitted by piecewise power-law functions with different exponents at different k ranges. A faster decay at smaller values for k indicates more limited gains for read lengths > 200 basepairs. The frequency distributions of k-mers exhibit long tails in a power-law-like trend, and rank frequency plots exhibit a concave Zipf's curve. The location of the most frequent 1000-mers comprises 172 kilobase-ranged regions, including four large stretches on chromosomes 1 and X, containing genes with biomedical implications. Even the read length 1000 would be insufficient to reliably sequence these specific regions.
Tasuku Soma - One of the best experts on this subject based on the ideXlab platform.
-
a generalization of submodular cover via the Diminishing Return property on the integer lattice
Neural Information Processing Systems, 2015Co-Authors: Tasuku Soma, Yuichi YoshidaAbstract:We consider a generalization of the submodular cover problem based on the concept of Diminishing Return property on the integer lattice. We are motivated by real scenarios in machine learning that cannot be captured by (traditional) sub-modular set functions. We show that the generalized submodular cover problem can be applied to various problems and devise a bicriteria approximation algorithm. Our algorithm is guaranteed to output a log-factor approximate solution that satisfies the constraints with the desired accuracy. The running time of our algorithm is roughly O(n log(nr) log r), where n is the size of the ground set and r is the maximum value of a coordinate. The dependency on r is exponentially better than the naive reduction algorithms. Several experiments on real and artificial datasets demonstrate that the solution quality of our algorithm is comparable to naive algorithms, while the running time is several orders of magnitude faster.
-
maximizing submodular functions with the Diminishing Return property over the integer lattice
arXiv: Data Structures and Algorithms, 2015Co-Authors: Tasuku Soma, Yuichi YoshidaAbstract:The problem of maximizing non-negative monotone submodular functions under a certain constraint has been intensively studied in the last decade. In this paper, we address the problem for functions defined over the integer lattice. Suppose that a non-negative monotone submodular function $f:\mathbb{Z}_+^n \to \mathbb{R}_+$ is given via an evaluation oracle. Furthermore, we assume that $f$ satisfies the Diminishing Return property, which is not an immediate consequence of the submodularity when the domain is the integer lattice. Then, we show (i) a $(1-1/e-\epsilon)$-approximation algorithm for a cardinality constraint with $\widetilde{O}(\frac{n}{\epsilon}\log \frac{r}{\epsilon})$ queries, where $r$ is the maximum cardinality of feasible solutions, (ii) a $(1-1/e-\epsilon)$-approximation algorithm for a polymatroid constraint with $\widetilde{O}(\frac{nr}{\epsilon^4}+n^6)$ queries, where $r$ is the rank of the polymatroid, and (iii) a $(1-1/e-\epsilon)$-approximation algorithm for a knapsack constraint with $\widetilde{O}(\frac{n^2}{\epsilon^{18}}\log \frac{1}{w})(\frac{1}{\epsilon})^{O(1/\epsilon^8)}$ queries, where $w$ is the minumum weight of elements. Our algorithms for polymatroid constraints and knapsack constraints first extend the domain of the objective function to the Euclidean space and then run the continuous greedy algorithm. We give two different kinds of continuous extensions, one is for knapsack constraints and the other is for polymatroid constraints, which might be of independent interest.
Frank A Anania - One of the best experts on this subject based on the ideXlab platform.
-
universities and sponsored research indirect cost recovery and the law of Diminishing Return
Hepatology, 2015Co-Authors: Frank A AnaniaAbstract:Two recent reports, one from the Government Accountability Office (GAO) of the US government1 and an editorial in the Proceedings of the National Academy of Sciences by Ronald J. Daniels, president of Johns Hopkins University,2 underscore another major issue facing research universities. This issue concerns the recovery of indirect costs (IDC), or facilities and administration (F&A) costs, which sustain biomedical research in major research institutions. Because schools of medicine (SOM)—and other health science–related schools—are pivotal players in garnering research funds, SOM are caught between respective parent universities’ priorities, faculty who invariably will have lapses in (direct cost) research funding in this environment, and a “new normal” among federal agencies which traditionally could be counted on for IDC support. Traditionally, SOM have received the bulk of direct and indirect dollars from the National Institutes of Health (NIH). In this financial climate, however, SOM—whether public or private—are facing simultaneous downward economic pressures from long-standing, traditional revenue streams. Dollars derived from clinical revenues at partner teaching hospitals, tuition costs for education, as well as direct costs for sponsored research can no longer be counted on to sustain the cost of doing research or supplement universities for other non-SOM-related missions. In the case of public universities, many states are sharply reducing financial support. This leaves SOM at major universities scrambling to seek additional funds, for example, by canvasing for donors so that endowments can generate higher interest income to make up for other financial shortfalls. While some SOM have been more successful in building their endowments, others have not. Most parent universities—seeking general revenues for necessary non-SOM academic functions—used to look to SOM for general funds derived from IDC recovery and for good reason. As an example, at the close of the fiscal year at Emory University (August 31, 2014), 93% of all university-sponsored research support was derived from the schools of the health sciences and nearly 67% of the health sciences–sponsored research was awarded to Emory University School of Medicine faculty (data from Woodruff Health Sciences Center of Emory University). I am not suggesting Emory or any other university redirect IDC from the health sciences, but because most stakeholders—including faculty and the NIH—are not exactly aware of how most F&A dollars flow, clear accounting practices available for review might help universities and SOM go a long way in making their case. The case that universities are making was detailed by Daniels2 of Hopkins. He contends the cost of sponsoring research in the United States is far outpacing the shared financial responsibility between the NIH and sponsoring institutions on which the nation’s biomedical research platform was initially established.3 He asserts that universities are paying more than their fair share. Is he correct? There are two financial components to F&A: a facilities component and an administrative component. The administrative component was set as a fixed rate of 26% for administrative reimbursement back in the 1990s. The facilities component is not fixed but is a negotiated rate between an individual university and the Department of Health and Human Services every 2 years. Because total direct costs awarded are ebbing downward, it would follow that there is also an overall decline in F&A to universities. However, university officials contend that even if direct funding were keeping pace with the rate of inflation, indirect costs are not keeping pace and are stilling falling behind. Many sponsored federal and nonfederal awards for junior faculty and postdoctoral research training have little, if any, F&A costs associated with them. Anyone who conducts research recognizes that administrative compliance with institutional review boards, animal care and use committees, etc. is very cumbersome. To accommodate the increasing regulatory burden, universities have hired additional administrators with specialized training and expertise to assure compliance with federal and state oversight. Finally, the rising cost of fringe benefits and gap funding for tenured faculty, who invariably in this environment will have lapses in extramural support, along with establishing and maintain state-of-the-art core facilities are also adding to research institutions’ F&A concerns. Data from the National Science Foundation further support universities’ concerns that their costs for sponsoring research on campus have significantly risen from 8.7% in 1962 to 19.4% in 2012.4 The GAO report released in 2013 was requested by Jeff Sessions, R-AL, then the ranking member of the US Senate Budget Committee.1 The analysis was designed to specifically address the potential changes in reimbursement by the NIH to universities for IDCs, the key factors affecting NIH reimbursements to universities for such costs, and whether IDCs were having a negative impact on direct costs associated with the NIH’s extramural research mission. The GAO recommended that “the NIH Director should assess the impact of growth in indirect costs on its research mission … planning for how to deal with potential future increases in indirect costs that could limit the amount of funding available for direct costs of research projects.” The GAO’s report “Biomedical Research: NIH Should Assess the Impact on Its Mission of Growth in Indirect Costs” was refuted by the NIH. The NIH disagreed with a key statement of the report: “According to … available data, indirect costs as a proportion of the NIH budget have remained below thirty percent over the past 25 years” and, “from FY 2002 through FY 2012”—the period for which the GAO conducted its independent study—“indirect costs remained at approximately 27%.” The NIH indicates that larger grant awards drove a proportionate, and not a disproportionate, increase in IDC expenses (a 28.1% increase for F&A over this period compared to 27% in direct costs). The NIH’s statements would seem to support university presidents and deans that F&A have not kept pace with inflation. These findings suggest that administrative costs associated with sponsored research are not sustainable. However, universities made a contract with society through the federal government. By pursuing federal contracts, universities are financial stakeholders of public resources, making sponsored research a shared responsibility.3 Financial transparency, not just with trustees but also with faculty and society at large, would go a long way toward convincing many that a financial “bailout” may be needed for research sustainability. Arguably, administrative costs should be more flexible and at the very least could be increased from the current rate. However, universities should examine more critically how they account for facility fees which are negotiable. Given the shrinking research workforce, it may no longer be prudent to pass on to taxpayers the cost of libraries and interest on debts incurred from building construction—costs associated with the facility fees. As universities become leaner and adopt fiscal restraint, they will need to revisit non–health science faculty policies as well as the economics of general operations. Universities should reveal to the government (federal and states), their faculty, and the American public a clear assessment of how IDCs are spent and apportioned within research-intensive universities. Finally, universities should be more forthcoming about how their endowments are structured because most people are not aware that the vast majority of endowed funds are restricted by donors. Because a major goal of a university is to generate knowledge, is it reasonable for universities to continue to pursue donations that are restricted when public funding of sponsored research is in a prolonged period of deflationary growth? A recent quote from a related article in the New York Times succinctly sums up the quandary with restricted philanthropy: “For better or worse, the practice of science in the 21st century is becoming shaped less by national priorities or by peer-review groups and more by the particular preferences of individuals with huge amounts of money.”5 Finally, major universities should reexamine the cost of their own administrative structures and associated perks. While some may argue these costs are unlikely to provide additional funds directed at administrating research, these measures would go a long way toward stemming the silent crisis of morale among faculty and students, who are also research stakeholders in a very uncertain future.