Open Notebook NMR – motivations and confusions
Creators & Contributors
I have been pleased by the interest in Open Notebook NMR but the current discussions have widened far too useful to be useful, so I want to be absolutely clear what the project and its limits are.
This is a part of Nick Day's PhD thesis at Cambridge.
The only motivation is for Nick to be able to do good science, with advice and help which is publishable in his thesis. I repeat:
This is a part of Nick Day's PhD thesis at Cambridge.
I made that absolutely clear at the beginning. Any broadening of the project is a distraction and could be detrimental to his work. For example I follow what Alicia is doing with Jean-Claude but I would never dream of suggesting she does other than what is agreed between them.
It is extremely unusual for a PhD student to be exposing his work as Open Notebook Science. I only suggested it because he is a good student and I believe his technical competence and commitment is such that he will do careful and valuable work. Remember that if anything is wrong it is extremely public.
We also made clear what the limits of the project were and I will repeat them as our hypothetical report:
We adapted Rychnovksy's method of calculating 13C NMR shifts by adding (XXX) basis set and functionals (Henry has done this). We extracted 1234 spectra with predicted 3D geometries for rigid molecules in NMRShiftDB (no acyclic-acyclic bond nodes for heavy atoms). Molecules had < = 21 heavy atoms (<= Cl). These were optimised using Gaussian XXX and the isotropic magnetic tensors calculated using correction for the known solvent. The shift was subtracted from the calculated TMS shift (in the same solvent) and the predicted shift compared with the observed.
Initially the RMS deviation was xxx. This was due to a small number of structures where there appeared to be gross errors of assignment. These were exposed to the community who agreed that these should be removed. The RMS dropped to yyy. The largest deviations were then due to Y-C-X systems, where a correction was applied (with theoretical backing). The RMS then dropped to zzz. The main outliers then appeared to be from laboratory AAA to whom we wrote and they agreed that their output format introduced systematic errors. They have now corrected this. The RMS was now zzz. The deviations were analysed by standard chemoinformatics methods and were found to correlate with the XC(=Z)Y group which probably has two conformations. A conformational analysis of the system was undertaken for any system with this group and the contributions from different conformers averaged. The RMS now dropped to vvv.
This established a protocol for predicting NMR spectra to 99.3% confidence. We then applied this to spectra published in 2007 in major chemical journals. We found that aa% of spectra appeared to be misassigned, and that bb% of suggested structures were "wrong" – i.e. the reported chemical shifts did not fit the reported spectra values.
This is precisely what we have been doing and we are sticking to it. It would be irresponsible for a supervisor and student to do elsewise until unforeseen difficulties arose. It has gone according to plan:
- Initially the RMS deviation was ca. 3 ppm. This was due to a small number of structures where there appeared to be gross errors of assignment. These were exposed to the community who agreed that these should be removed. We have had comments from Christoph, Henry, Egon and Jean-Claude which have allowed us to remove 3 entries.
- The largest deviations were then due to C-Hal systems, where a correction was applied (with theoretical backing from Henry). We are now applying this correction and will report as soon as the new code has been written (because we use XML-CML this has a rapid turnround).
- The main outliers then appeared to be from laboratory AAA to whom we wrote and they agreed that their output format introduced systematic errors. A very serious error occurred in an entry from DKFZ spektren which was due to human editing of a apectrum (identified by Christoph). It is therefore not unreasonable to remove all entries from this source (note we have to remove all without inspecting them – we cannot just choose which we like).
- A conformational analysis of the system was undertaken for any system with this group and the contributions from different conformers averaged. It is clear from inspection of the data that some compounds have equipopulated conformers (e.g. C6H5-CH(=O)) and that is legitimate to identify this framework symmetry group and average over equivalent atoms if they have equal shifts.
-
This established a protocol for predicting NMR spectra to xxx% confidence. This is the primary aim of the work – to be able to show that machines can make decisions without humans within a given confidence level. It depends on being able to find a set of data which are accepted as accurately assigned. That is what we are asking the community for.
The immediate Open task is to help annotate outliers, which are being released as soon as they are identified and we repeat – any help with this is very much appreciated.
The project that Chemspider has identified is completely distinct from Nick Day's thesis:
"The results from the GIAO calculations were compared with three other prediction approaches provided by Advanced Chemistry Development. These algorithms were not limited in the number of heavy atoms that could be handled by the algorithm, The algorithms were a HOSE-code based approach, a neural network approach and an "increment approach". A distinct advantage of these approaches is the time for prediction relative to the quantum-mechanical calculations. The QM calculation took a number of weeks to perform on the dataset of 23475 structures on a cluster of computers. However, a standard PC enabled the HOSE code based predictions to be performed in a few hours, the Neural Net predictions in about 4 minutes and the Increment based predictions in less than 3 minutes.
A comparison of the approaches gave statistics for the non-QM approaches superior to those of the QM approach.
This has no bearing in any way on Nick's work – it does not help certify entries as "correct". The last sentence actually suggests that Chemspider believes their work is superior to what Nick is doing. Without a probably annotated data set such claims are meaningless.
Finally I should clarify what is meant by Open since there is confusion:
Ryan Sasaki Says:
October 24th, 2007 at 12:16 pm e
Hi Peter,
While I cannot speak for Brent Lefebvre, I have a couple of comments in relation to Open Notebook NMR and the potential involvement of commercial software companies in this study.
First of all, you mention that you will probably share your insights but not your data. If that is true, then I do not understand how this project can be referred as OPEN. If the dataset being used is not shared publicly, how can this be considered an open project?
We have so far shared every piece of data and metadata that we feel is fit to publish. Open does not mean "immediate". The data is taken from NMRShiftDB which is Open. We have published our protocol as it is refined. We have said we are going to publish our outliers and we are doing so. When we get input from the community we shall publish more. The product will be a data set which will have community approval – along with the protocol this will be our major Open deliverables. Nick will have enough breathing space to make scientific discoveries – if any are to be made.
The biological community already operates in this way. When a lab does a protein structure they do not publish their photographs immediately. We do not expect them to, any more than we expect Rosalind Franklin to do so. But when the time comes for publication we shall publish all that is necessary to replicate the experiment – that is the key. And, along the way, we are asking the community for input. If they do so the result could be a data set that is relatively small but of very high quality and therefore useful for testing computational approaches.
Additional details
Description
I have been pleased by the interest in Open Notebook NMR but the current discussions have widened far too useful to be useful, so I want to be absolutely clear what the project and its limits are. This is a part of Nick Day's PhD thesis at Cambridge. The only motivation is for Nick to be able to do good science, with advice and help which is publishable in his thesis. I repeat: This is a part of Nick Day's PhD thesis at Cambridge.
Identifiers
- GUID
- https://blogs.ch.cam.ac.uk/pmr/?p=734
- URL
- https://blogs.ch.cam.ac.uk/pmr/2007/10/24/open-notebook-nmr-motivations-and-confusions/
Dates
- Issued
-
2007-10-24T23:55:00Z
- Updated
-
2007-10-24T23:55:00Z