Why bother with new technology?
Creators & Contributors
Kinasepro has blogged about discussions of new chemoinformatics technology (specifically CML (Chemical Markup Language) and InChI (chemical identifier)). Here's the post and some correspondence. It's basically about the introduction of new technology. Obviously I'm not neutral but I will try to discuss it in a neutral manner. For that reason I have copied it more or less in full.
There's been a fair amount of talk [ChemBark] over the last little while on the topic of chemoinformatics and chemblogs. Here's my two cents.smilesinchiAldrich
smilesinchiChemExper
smilesinchiThe PDB
smilesinchiChemdraw (until v10)
smilesinchiThe entire pharmaceutical industry.
smilesinchi Peter Murray Rust
smilesinchi IUPACSo somehow a couple librarians have convinced Google that inchi > smiles. Result? Google may well do Inchi, but noone but the librarians are currently using it, and meanwhile google doesn't index smiles very well. I'm reminded of a day when it was thought to be a good idea to put the CAS#s of new entities at the bottom of ACS journal articles. Don't worry, we survived those librarians too.
Lookit, we don't need a string of XML code that you need an advanced degree to use. We don't need people telling us to tag our blog posts, we need an integrated solution. We need something that can draw structures and present them attractively in an index friendly HTML format. Near term: Get google to index picture descriptions, and code a firefox plugin that can insert smiles into said descriptions.
Till google has a smiles substructure search, I'm not going to bother.
- client (i.e. your browser)
- Google (we have discussed this with Google and it's not impossible)
- third party (who may or may not charge for it).
Given that Openbabel can search millions of structures quite rapidly there are some encouraging opportunities.
- 1 totallymedicinal Dec 5th, 2006 at 3:09 pm
Couldn't agree more with the sentiment – not only does my ancient version of ChemDraw not support this exotic format, but I have enuff hassle in my life without learning some obscure new coding system.
PMR: Again this is a perfectly valid response. Any approach to chemoinformatics requires tools. And I suspect or your institution would have to pay for an upgrade to Chemdraw. Obviously there is the opportunity of some Open Source free tools but they are not yet widely deployed and are effectively for early adopters.
- 3 Dec 7th, 2006 at 4:06 am
I could not agree more about the need for an integrated solution! I got a really thoughtful response from Peter Murray-Rust and friends, and I feel kind of bad about not acting on it, but putting random InChI designations at the bottom of all our blog posts doesn't seem worth it to me. I think that CML is indeed the future, and I look forward to the day of being able to download a CML plugin for WordPress that will take care of everything for us lazy bloggers.
- 4 Dec 9th, 2006 at 12:23 pm
The argument against SMILES seems to be they are not an Open Format and it is possible to represent a single molecule with multiple SMILES strings. For my part I can read and write SMILES, (and SMARTS and SMIRKS). I find InChi impenetrable and I don't think there is syntax for substructure or similarity queries, in addition I don't think there is a system for describing reactions.I've started to add SMILES to my web pages in the hope that someone will build an index at some point, I guess it would help if there was a SMILES tag.
-
5 Dec 9th, 2006 at 7:03 pm
InchI and CML may well be the future, and no-one will embrace it more then me, but SMILES is the present. For people working in the field not to understand that boggles the mind!
PMR: I'm not sure who "people working in the field" are. If it includes me, then I fully understand it. I am simply trying to bring the future to the present a bit quicker and a bit more predictably. 🙂
-
I've experimented on this site a little with smiles. For instance a google search of the following string brings you here:O=C(C2=CN=C(NC3=NC(C)=NC(N4CCN(CCO)CC4)=C3)S2)NC1=C(C)C=CC=C1ClOf note I'm not the only one with that string on the web! Maybe thats an important compound? Sadly google indexed that page under my SRC tag rather then as a standalone page. Put that together with the fact that smiles strings are not substructure searchable via google and its clear to me that google is not ready to be a chemistry informatics platform. It's sad really, because it doesn't seem to me that it would be that difficult for them to make SMILES strings substructure searchable via the same algorithm the PDB, relibase, aldrich and everybody else is using.
PMR: This is a very important point and at the heart of the problem. Google works by indexing text. It's good at it and can distinguish different roles for text and can look for substrings. This is a simple, powerful model. But at present it doesn't index other objects (faces, maps, etc.) These are both harder and require specialist software. By contrast PDB, Relibase, Aldrich do index chemical structures. That means that they have to have specialist software running on their servers. Which means a business model. And that someone has to pay somewhere. PDB gets a grant, Relibase is commercial, Aldrich will see this as the basis for selling more compounds. All completely valid. But there is no business reason for Google to invest in chemistry-specific software – as I said chemistry is too small for Google to bother with. It's not helped by the fact that all the information is proprietary and that one of the major chemical information suppliers (CAS/ACS) sued Google. So unless you convince them differently – and I have gently tried – it won't happen.
So this is all about the introduction of new technology. The primary messages from the chemistry community are something like:
- We're happy with what we've got – it's worked for the last 20 years and will go on doing so. Yes, for a little while.
- When it's necessary CambridgeSoft, Chemical Abstracts, Elsevier will develop a new technology and we'll pay them to use it. Unfortunately I don't see any movement from any of these to embrace the new Web metaphors. Biology, geoscience, etc. are working hard to develope the semantic web in the subjects – apart from a few of us noone in chemistry is.
- Well, it's a bit of a mess, but it's not at the top of my priorities. I'll come back in a few years.
- computational chemistry. We are having a visit of COST D37 (EU) to Cambridge tomorrow to create an interoperable infrastructure for computational chemistry. It will be based on communal agreements and use XML/CML as the infrastructure.
- chemoinformatics. The Open Source community (e.g. Blue Obelisk) supports both current (legacy) formats (SMILES, Mol) etc. and also CML/InChI. This can provide a smooth path towards the wider adoption of these newer approaches, including toolkits. The toolkits are free, which some see as a disadvantage, in which case you will have to convince the commercial suppliers to create them.
- publishing. Commercial publishing is universally based on XML (and variants) so it is easy for them to include CML and related systems. I won't give details but I'd be surprised if there weren't major changes in the next 2-3 years here which I hope will answer some of the obejctions raised here.
There are also general major drivers elsewhere for the abandonment of legacy formats. They include the semantic web, RSS, institutional repositories, archival, etc. These efforts require interoperability and freely available tools – you can't archive – say – a binary chemistry file and expect it to be readable in 5 years time. There are a lot of people to whom that matters.
So I'm not telling anyone to do anything – I'm putting ideas, protocols and tools where they may wish to pick them up. If 5% of a community is enthusiastic that's a good beginning. It worries me that the pharma industry has no concept of interoperability. But I've said that already.
Additional details
Description
Kinasepro has blogged about discussions of new chemoinformatics technology (specifically CML (Chemical Markup Language) and InChI (chemical identifier)). Here's the post and some correspondence. It's basically about the introduction of new technology. Obviously I'm not neutral but I will try to discuss it in a neutral manner. For that reason I have copied it more or less in full. PMR: I'm not sure who the librarians are.
Identifiers
- GUID
- https://blogs.ch.cam.ac.uk/pmr/?p=214
- URL
- https://blogs.ch.cam.ac.uk/pmr/2006/12/10/why-bother-with-new-technology/
Dates
- Issued
-
2006-12-10T14:26:00Z
- Updated
-
2006-12-10T14:26:00Z