The Experts below are selected from a list of 555 Experts worldwide ranked by ideXlab platform
Antony Williams - One of the best experts on this subject based on the ideXlab platform.
-
Twenty Five Years in Cheminformatics – A Career Path Through a Diverse Series of Roles and Responsibilities
2017Co-Authors: Antony WilliamsAbstract:Antony Williams is a Computational Chemist at the US Environmental Protection Agency in the National Center for Computational Toxicology. He has been involved in cheminformatics and the dissemination of chemical information for over twenty-five years. He has worked for a Fortune 500 company (Eastman Kodak), in two successful start-ups (ACD/Labs and ChemSpider), for the Royal Society of Chemistry (in publishing) and, now, at the EPA. Throughout his diverse career path he has experienced multiple work cultures and focused his efforts on understanding the needs of his employers and the often unrecognized needs of a larger community. Antony will provide a short overview of his career path and discuss the various decisions that helped motivate his change in career from professional spectroscopist to website host and innovator, to working for one of the world's foremost scientific societies and now for one of the most impactful government organizations in the world.
-
Serving the medicinal chemistry community with Royal Society of Chemistry cheminformatics platforms
2015Co-Authors: Antony WilliamsAbstract:The Royal Society of Chemistry (RSC) is a major participant in providing access to chemistry related data via the web. As an internationally renowned society for the chemical sciences, a scientific publisher and the host of the ChemSpider database for the community, RSC continues to make dramatic strides in providing online access to data. ChemSpider provides access to over 30 million chemicals sourced from over 500 data suppliers and linked out to related information on the web. The platform is a crowdsourcing environment whereby members of the community can participate in validating and expanding the content of the database. With a set of application programming interfaces ChemSpider is used by various organizations and projects to serve up data for various purposes. These include structure identification for mass spectrometry instrument vendors, RSC databases such as the Marinlit natural products database and a European grant-based project from the Innovative Medicines Initiative fund. This presentation will provide an overview of various cheminformatics activities and projects that RSC is involved with to serve the medicinal chemistry community. This will include the Open PHACTS semantic web project, the PharmaSea project to identify new pharmaceutical leads from the ocean and the UK National Compound Collection to identify new lead compounds contained within PhD theses.
-
The marriage of ACD/Labs technologies to eScience Projects at the Royal Society of Chemistry
2014Co-Authors: Antony WilliamsAbstract:The Royal Society of Chemistry is one of the worlds foremost scientific societies, a primary publisher for the chemical sciences and an innovator in the domain of eScience. In order to deliver on a number of our eScience projects we utilize a number of components of Advanced Chemistry Development software including nomenclature, physchem prediction, spectroscopy tools and the ACD/Ilab web-based system. This presentation will provide an overview of a number of RSC projects where ACS/Labs software has played an important role in the delivery of the systems including ChemSpider and the National Chemical Database Service for the United Kingdom. We will also provide an overview of our vision to deliver a repository for various types of experimental chemistry data and how we foresee utilizing various prediction and validation software approaches to characterize the data as well as the potential to generate predictive models from the data. This couples directly with our intention to data enable our publication archive of over 300,000 articles extracting chemicals, reactions and analytical data from the historical records.
-
Jean-Claude Bradley Open Melting Point Dataset
2014Co-Authors: Jean-claude Bradley, Antony Williams, Andrew LangAbstract:Jean-Claude Bradley's Legacy Dataset of Open Melting Points. 28,645 measurements including those found to be incorrect (marked as 'do not use'). csid corresponds to ChemSpider ID.
-
Jean-Claude Bradley Double Plus Good (Highly Curated and Validated) Melting Point Dataset
2014Co-Authors: Jean-claude Bradley, Andrew Lang, Antony WilliamsAbstract:3041 highly curated and validated melting point measurements taken from the Jean-Claude Bradley Open Melting Point Dataset. Values were only kept if there were multiple measurements and the range of values was between 0.01 C and 5 C inclusive. csid corresponds to ChemSpider ID.
Valery Tkachenko - One of the best experts on this subject based on the ideXlab platform.
-
METHODOLOGY Open Access The ChEMBL database as linked open data
2014Co-Authors: Egon L. Willighagen, Valery Tkachenko, Peter Ansell, Andra Waagmeester, Antony J Williams, Ola Spjuth, Janna Hastings, David J WildAbstract:Background: Making data available as Linked Data using Resource Description Framework (RDF) promotes integration with other web resources. RDF documents can natively link to related data, and others can link back using Uniform Resource Identifiers (URIs). RDF makes the data machine-readable and uses extensible vocabularies for additional information, making it easier to scale up inference and data analysis. Results: This paper describes recent developments in an ongoing project converting data from the ChEMBL database into RDF triples. Relative to earlier versions, this updated version of ChEMBL-RDF uses recently introduced ontologies, including CHEMINF and CiTO; exposes more information from the database; and is now available as dereferencable, linked data. To demonstrate these new features, we present novel use cases showing further integration with other web resources, including Bio2RDF, Chem2Bio2RDF, and ChemSpider, and showing the use of standard ontologies for querying. Conclusions: We have illustrated the advantages of using open standards and ontologies to link the ChEMBL database to other databases. Using those links and the knowledge encoded in standards and ontologies, the ChEMBL-RDF resource creates a foundation for integrated semantic web cheminformatics applications, such as the presented decision support
-
Royal Society of Chemistry developments to support open drug discovery
2014Co-Authors: Antony Williams, Valery Tkachenko, Ken Karapetyan, Alexey Pshenichnov, Colin Batchelor, Jon SteeleAbstract:In recent years the Royal Society of Chemistry has become known for our development of freely accessible data platforms including ChemSpider, ChemSpider Reactions and our new chemistry data repository. In order to support drug discovery RSC participates in a number of projects including the Open PHACTS semantic web project, the PharmaSea natural products discovery project and the Open Source Drug Discovery project in collaboration with a team in India. Our most recent developments include extending our efforts to support neglected diseases by the provision of high quality datasets resulting from our curation efforts to support modeling, the delivery of enhanced application programming interfaces to allow open source drug discovery teams to both source and deposit data from our chemistry databases and the provision of a micropublishing platform to report on various aspects of work supporting neglected disease drug discovery. This presentation will review our existing efforts and our plans for extended development.
-
Data enhancing the Royal Society of Chemistry publication archive
2014Co-Authors: Antony Williams, Ken Karapetyan, Colin Batchelor, Peter Corbett, Valery TkachenkoAbstract:The Royal Society of Chemistry has an archive of hundreds of thousands of published articles containing various types of chemistry related data – compounds, reactions, property data, spectral data etc. RSC has a vision of extracting as much of these data as possible and providing access via ChemSpider and its related projects. To this end we have applied a combination of text-mining extraction, image conversion and chemical validation and standardization approaches. The outcome of this project will result in new chemistry related data being added to our chemical and reaction databases and in the ability to more tightly couple web-based versions of the articles with these extracted data. The ability to search across the archive will be enhanced as a result. This presentation will report on our progress in this data extraction project and discuss how we will ultimately use similar approaches in our publishing pipeline to enhance article markup for new publications.
-
Mining public domain data as a basis for drug repurposing
2013Co-Authors: Antony Williams, Valery Tkachenko, Sean EkinsAbstract:Online databases containing high throughput screening and other property data continue to proliferate in number. Many pharmaceutical chemists will have used databases such as PubChem, ChemSpider, DrugBank, BindingDB and many others. This work will report on the potential value of these databases for providing data to be used to repurpose drugs using cheminformatics-based approaches (e.g. docking, ligand-based machine learning methods). This work will also discuss the potentially related applications of the Open PHACTS project, a European Union Innovative Medicines Initiative project, that is utilizing semantic web based approaches to integrate large scale chemical and biological data in new ways. We will report on how compound and data quality should be taken into account when utilizing data from online databases and how their careful curation can provide high quality data that can be used to underpin the delivery of molecular models that can in turn identify new uses for old drugs.
-
ChemValidator – an online service for validating and standardizing chemical structure files
2013Co-Authors: Antony Williams, Valery Tkachenko, Ken Karapetyan, David Sharpe, Colin BatchelorAbstract:The production of valid and appropriate chemical structure representations which are appropriate for deposition into chemical structure databases and for inclusion into scientific publications requires adoption of a set of pre-processing filters and standardization procedures. As part of our ongoing effort to improve the quality of data for deposition into the RSC ChemSpider database, to provide a manner by which to validate and prepare data for publication and to provide a valuable service to the chemistry community, we have delivered the ChemValidator online service. This website provides access to an intuitive user interface for the upload of chemical compounds in various formats, pre-processing and standardization relative to a defined set of standards and validation checking of the chemicals according to a number of rules including hypervalency, absence of stereochemistry and charge balance. This presentation will report on the development of ChemValidator. This presentation was given by David Sharpe at the ACS Fall Meeting in 2012
Antony J Williams - One of the best experts on this subject based on the ideXlab platform.
-
Parallel Worlds of Public and Commercial Bioactive Chemistry Data
2016Co-Authors: Christopher A. Lipinski, Antony J Williams, Nadia K. Litterman, Christopher Southan, Alex M. Clark, Sean EkinsAbstract:ABSTRACT: The availability of structures and linked bioactivity data in databases is powerfully enabling for drug discovery and chemical biology. However, we now review some confounding issues with the divergent expansions of public and commercial sources of chemical structures. These are associated with not only expanding patent extraction but also increasingly large vendor collections amassed via different selection criteria between SciFinder from Chemical Abstracts Service (CAS) and major public sources such as PubChem, ChemSpider, UniChem, and others. These increasingly massive collections may include both real and virtual compounds, as well as so-called prophetic compounds from patents. We address a range of issues raised by the challenges faced resolving the NIH probe compounds. In addition we highlight the confounding of prior-art searching by virtual compounds that could impact the composition of matter patentability of a new medicinal chemistry lead. Finally, we propose some potential solutions
-
ChemTrove: Enabling a Generic ELN To Support Chemistry through the Use of Transferable Plug-ins and Online Data Sources
2015Co-Authors: Aileen E. Day, Simon J. Coles, Colin L. Bird, Jeremy G. Frey, Richard J. Whitby, Valery E. Tkachenko, Antony J WilliamsAbstract:In designing an Electronic Lab Notebook (ELN), there is a balance to be struck between keeping it as general and multidisciplinary as possible for simplicity of use and maintenance and introducing more domain-specific functionality to increase its appeal to target research areas. Here, we describe the results of a collaboration between the Royal Society of Chemistry (RSC) and the University of Southampton, guided by the aims of the Dial-a-Molecule Grand Challenge, intended to achieve the best of both worlds and augment a discipline-agnostic ELN, LabTrove, with chemistry-specific functionality and using data provided by the ChemSpider platform. This has been done using plug-in technology to ensure maximum transferability with minimal effort of the chemistry functionality to other ELNs and equally other subject-specific functionality to LabTrove. The resulting product, ChemTrove, has undergone a usability trial by selected academics, and the resulting feedback will guide the future development of the underlying ELN technology
-
METHODOLOGY Open Access The ChEMBL database as linked open data
2014Co-Authors: Egon L. Willighagen, Valery Tkachenko, Peter Ansell, Andra Waagmeester, Antony J Williams, Ola Spjuth, Janna Hastings, David J WildAbstract:Background: Making data available as Linked Data using Resource Description Framework (RDF) promotes integration with other web resources. RDF documents can natively link to related data, and others can link back using Uniform Resource Identifiers (URIs). RDF makes the data machine-readable and uses extensible vocabularies for additional information, making it easier to scale up inference and data analysis. Results: This paper describes recent developments in an ongoing project converting data from the ChEMBL database into RDF triples. Relative to earlier versions, this updated version of ChEMBL-RDF uses recently introduced ontologies, including CHEMINF and CiTO; exposes more information from the database; and is now available as dereferencable, linked data. To demonstrate these new features, we present novel use cases showing further integration with other web resources, including Bio2RDF, Chem2Bio2RDF, and ChemSpider, and showing the use of standard ontologies for querying. Conclusions: We have illustrated the advantages of using open standards and ontologies to link the ChEMBL database to other databases. Using those links and the knowledge encoded in standards and ontologies, the ChEMBL-RDF resource creates a foundation for integrated semantic web cheminformatics applications, such as the presented decision support
-
The ChEMBL database as linked open data
Journal of Cheminformatics, 2013Co-Authors: Egon L. Willighagen, Valery Tkachenko, Peter Ansell, John Hastings, Andra Waagmeester, Antony J Williams, Ola Spjuth, Bin Chen, David J WildAbstract:BACKGROUND: Making data available as Linked Data using Resource Description Framework (RDF) promotes integration with other web resources. RDF documents can natively link to related data, and others can link back using Uniform Resource Identifiers (URIs). RDF makes the data machine-readable and uses extensible vocabularies for additional information, making it easier to scale up inference and data analysis.\n\nRESULTS: This paper describes recent developments in an ongoing project converting data from the ChEMBL database into RDF triples. Relative to earlier versions, this updated version of ChEMBL-RDF uses recently introduced ontologies, including CHEMINF and CiTO; exposes more information from the database; and is now available as dereferencable, linked data. To demonstrate these new features, we present novel use cases showing further integration with other web resources, including Bio2RDF, Chem2Bio2RDF, and ChemSpider, and showing the use of standard ontologies for querying.\n\nCONCLUSIONS: We have illustrated the advantages of using open standards and ontologies to link the ChEMBL database to other databases. Using those links and the knowledge encoded in standards and ontologies, the ChEMBL-RDF resource creates a foundation for integrated semantic web cheminformatics applications, such as the presented decision support.
-
Identification of “Known Unknowns” Utilizing Accurate Mass Data and ChemSpider
Journal of The American Society for Mass Spectrometry, 2012Co-Authors: James L. Little, Alexey Pshenichnov, Antony J Williams, Valery TkachenkoAbstract:In many cases, an unknown to an investigator is actually known in the chemical literature, a reference database, or an internet resource. We refer to these types of compounds as “known unknowns.” ChemSpider is a very valuable internet database of known compounds useful in the identification of these types of compounds in commercial, environmental, forensic, and natural product samples. The database contains over 26 million entries from hundreds of data sources and is provided as a free resource to the community. Accurate mass mass spectrometry data is used to query the database by either elemental composition or a monoisotopic mass. Searching by elemental composition is the preferred approach. However, it is often difficult to determine a unique elemental composition for compounds with molecular weights greater than 600 Da. In these cases, searching by the monoisotopic mass is advantageous. In either case, the search results are refined by sorting the number of references associated with each compound in descending order. This raises the most useful candidates to the top of the list for further evaluation. These approaches were shown to be successful in identifying “known unknowns” noted in our laboratory and for compounds of interest to others.
Jan A Kors - One of the best experts on this subject based on the ideXlab platform.
-
Automatic vs. manual curation of a multi-source chemical dictionary: the impact on text mining
2013Co-Authors: Antony Williams, Valery Tkachenko, Kristina M Hettne, Erik M Van Mulligen, Jos C S Kleinjans, Jan A KorsAbstract:Previously, we developed a combined dictionary dubbed Chemlist for the identification of small molecules and drugs in text based on a number of publicly available databases and tested it on an annotated corpus. To achieve an acceptable recall and precision we used a number of automatic and semi-automatic processing steps together with disambiguation rules. However, it remained to be investigated which impact an extensive manual curation of a multi-source chemical dictionary would have on chemical term identification in text. ChemSpider is a chemical database that has undergone extensive manual curation aimed at establishing valid chemical name-to-structure relationships.
-
automatic vs manual curation of a multi source chemical dictionary the impact on text mining
Journal of Cheminformatics, 2010Co-Authors: Valery Tkachenko, Antony J Williams, Kristina M Hettne, Erik M Van Mulligen, Jos C S Kleinjans, Jan A KorsAbstract:Previously, we developed a combined dictionary dubbed Chemlist for the identification of small molecules and drugs in text based on a number of publicly available databases and tested it on an annotated corpus. To achieve an acceptable recall and precision we used a number of automatic and semi-automatic processing steps together with disambiguation rules. However, it remained to be investigated which impact an extensive manual curation of a multi-source chemical dictionary would have on chemical term identification in text. ChemSpider is a chemical database that has undergone extensive manual curation aimed at establishing valid chemical name-to-structure relationships. We acquired the component of ChemSpider containing only manually curated names and synonyms. Rule-based term filtering, semi-automatic manual curation, and disambiguation rules were applied. We tested the dictionary from ChemSpider on an annotated corpus and compared the results with those for the Chemlist dictionary. The ChemSpider dictionary of ca. 80 k names was only a 1/3 to a 1/4 the size of Chemlist at around 300 k. The ChemSpider dictionary had a precision of 0.43 and a recall of 0.19 before the application of filtering and disambiguation and a precision of 0.87 and a recall of 0.19 after filtering and disambiguation. The Chemlist dictionary had a precision of 0.20 and a recall of 0.47 before the application of filtering and disambiguation and a precision of 0.67 and a recall of 0.40 after filtering and disambiguation. We conclude the following: (1) The ChemSpider dictionary achieved the best precision but the Chemlist dictionary had a higher recall and the best F-score; (2) Rule-based filtering and disambiguation is necessary to achieve a high precision for both the automatically generated and the manually curated dictionary. ChemSpider is available as a web service at http://www.ChemSpider.com/ and the Chemlist dictionary is freely available as an XML file in Simple Knowledge Organization System format on the web at http://www.biosemantics.org/chemlist .
-
Automatic vs. manual curation of a multi-source chemical dictionary: the impact on text mining
Journal of Cheminformatics, 2010Co-Authors: Kristina M Hettne, Valery Tkachenko, Antony J Williams, Jos C S Kleinjans, Erik M Van Mulligen, Jan A KorsAbstract:Background Previously, we developed a combined dictionary dubbed Chemlist for the identification of small molecules and drugs in text based on a number of publicly available databases and tested it on an annotated corpus. To achieve an acceptable recall and precision we used a number of automatic and semi-automatic processing steps together with disambiguation rules. However, it remained to be investigated which impact an extensive manual curation of a multi-source chemical dictionary would have on chemical term identification in text. ChemSpider is a chemical database that has undergone extensive manual curation aimed at establishing valid chemical name-to-structure relationships. Results We acquired the component of ChemSpider containing only manually curated names and synonyms. Rule-based term filtering, semi-automatic manual curation, and disambiguation rules were applied. We tested the dictionary from ChemSpider on an annotated corpus and compared the results with those for the Chemlist dictionary. The ChemSpider dictionary of ca. 80 k names was only a 1/3 to a 1/4 the size of Chemlist at around 300 k. The ChemSpider dictionary had a precision of 0.43 and a recall of 0.19 before the application of filtering and disambiguation and a precision of 0.87 and a recall of 0.19 after filtering and disambiguation. The Chemlist dictionary had a precision of 0.20 and a recall of 0.47 before the application of filtering and disambiguation and a precision of 0.67 and a recall of 0.40 after filtering and disambiguation. Conclusions We conclude the following: (1) The ChemSpider dictionary achieved the best precision but the Chemlist dictionary had a higher recall and the best F-score; (2) Rule-based filtering and disambiguation is necessary to achieve a high precision for both the automatically generated and the manually curated dictionary. ChemSpider is available as a web service at http://www.ChemSpider.com/ and the Chemlist dictionary is freely available as an XML file in Simple Knowledge Organization System format on the web at http://www.biosemantics.org/chemlist .
Kristina M Hettne - One of the best experts on this subject based on the ideXlab platform.
-
Automatic vs. manual curation of a multi-source chemical dictionary: the impact on text mining
2013Co-Authors: Antony Williams, Valery Tkachenko, Kristina M Hettne, Erik M Van Mulligen, Jos C S Kleinjans, Jan A KorsAbstract:Previously, we developed a combined dictionary dubbed Chemlist for the identification of small molecules and drugs in text based on a number of publicly available databases and tested it on an annotated corpus. To achieve an acceptable recall and precision we used a number of automatic and semi-automatic processing steps together with disambiguation rules. However, it remained to be investigated which impact an extensive manual curation of a multi-source chemical dictionary would have on chemical term identification in text. ChemSpider is a chemical database that has undergone extensive manual curation aimed at establishing valid chemical name-to-structure relationships.
-
automatic vs manual curation of a multi source chemical dictionary the impact on text mining
Journal of Cheminformatics, 2010Co-Authors: Valery Tkachenko, Antony J Williams, Kristina M Hettne, Erik M Van Mulligen, Jos C S Kleinjans, Jan A KorsAbstract:Previously, we developed a combined dictionary dubbed Chemlist for the identification of small molecules and drugs in text based on a number of publicly available databases and tested it on an annotated corpus. To achieve an acceptable recall and precision we used a number of automatic and semi-automatic processing steps together with disambiguation rules. However, it remained to be investigated which impact an extensive manual curation of a multi-source chemical dictionary would have on chemical term identification in text. ChemSpider is a chemical database that has undergone extensive manual curation aimed at establishing valid chemical name-to-structure relationships. We acquired the component of ChemSpider containing only manually curated names and synonyms. Rule-based term filtering, semi-automatic manual curation, and disambiguation rules were applied. We tested the dictionary from ChemSpider on an annotated corpus and compared the results with those for the Chemlist dictionary. The ChemSpider dictionary of ca. 80 k names was only a 1/3 to a 1/4 the size of Chemlist at around 300 k. The ChemSpider dictionary had a precision of 0.43 and a recall of 0.19 before the application of filtering and disambiguation and a precision of 0.87 and a recall of 0.19 after filtering and disambiguation. The Chemlist dictionary had a precision of 0.20 and a recall of 0.47 before the application of filtering and disambiguation and a precision of 0.67 and a recall of 0.40 after filtering and disambiguation. We conclude the following: (1) The ChemSpider dictionary achieved the best precision but the Chemlist dictionary had a higher recall and the best F-score; (2) Rule-based filtering and disambiguation is necessary to achieve a high precision for both the automatically generated and the manually curated dictionary. ChemSpider is available as a web service at http://www.ChemSpider.com/ and the Chemlist dictionary is freely available as an XML file in Simple Knowledge Organization System format on the web at http://www.biosemantics.org/chemlist .
-
Automatic vs. manual curation of a multi-source chemical dictionary: the impact on text mining
Journal of Cheminformatics, 2010Co-Authors: Kristina M Hettne, Valery Tkachenko, Antony J Williams, Jos C S Kleinjans, Erik M Van Mulligen, Jan A KorsAbstract:Background Previously, we developed a combined dictionary dubbed Chemlist for the identification of small molecules and drugs in text based on a number of publicly available databases and tested it on an annotated corpus. To achieve an acceptable recall and precision we used a number of automatic and semi-automatic processing steps together with disambiguation rules. However, it remained to be investigated which impact an extensive manual curation of a multi-source chemical dictionary would have on chemical term identification in text. ChemSpider is a chemical database that has undergone extensive manual curation aimed at establishing valid chemical name-to-structure relationships. Results We acquired the component of ChemSpider containing only manually curated names and synonyms. Rule-based term filtering, semi-automatic manual curation, and disambiguation rules were applied. We tested the dictionary from ChemSpider on an annotated corpus and compared the results with those for the Chemlist dictionary. The ChemSpider dictionary of ca. 80 k names was only a 1/3 to a 1/4 the size of Chemlist at around 300 k. The ChemSpider dictionary had a precision of 0.43 and a recall of 0.19 before the application of filtering and disambiguation and a precision of 0.87 and a recall of 0.19 after filtering and disambiguation. The Chemlist dictionary had a precision of 0.20 and a recall of 0.47 before the application of filtering and disambiguation and a precision of 0.67 and a recall of 0.40 after filtering and disambiguation. Conclusions We conclude the following: (1) The ChemSpider dictionary achieved the best precision but the Chemlist dictionary had a higher recall and the best F-score; (2) Rule-based filtering and disambiguation is necessary to achieve a high precision for both the automatically generated and the manually curated dictionary. ChemSpider is available as a web service at http://www.ChemSpider.com/ and the Chemlist dictionary is freely available as an XML file in Simple Knowledge Organization System format on the web at http://www.biosemantics.org/chemlist .