| Abstract
| - Nowadays millions of different compounds are known, their structures stored in electronic databases. Analysisof these data could yield valuable insights into the laws of chemistry and the habits of chemists. We havetherefore explored the public database of the National Cancer Institute (>250 000 compounds) by patternsearching. We split the molecules of this database into fragments to find out which fragments exist, howfrequent they are, and whether the occurrence of one fragment in a molecule is related to the occurrence ofanother, nonoverlapping fragment. It turns out that some fragments and combinations of fragments are sofrequent that they can be called “chemical clichés”. We believe that the fragment data can give insight intothe chemical space explored so far by synthesis. The lists of fragments and their (co-)occurrences can helpcreate novel chemical compounds by (i) systematically listing the most popular and therefore most easilyused substituents and ring systems for synthesizing new compounds, (ii) being an easily accessible repositoryfor rarer fragments suitable for lead compound optimization, and (iii) pointing out some of the yet unexploredparts of chemical space.
|