High-dimensional property spaces for compound opti- mization or data set analysis are generally difficult to represent and navigate. While the potency-centric AL concept has substantially contributed to graphical SAR exploration, especially for larger and structurally hetero- geneous data sets, little efforts have thus far been made to visualize multi-dimensional property landscapes that combine activity with other optimization-relevant proper- ties. Typically, dimension reduction techniques such as PCA are applied to evaluate feature contributions in multi- dimensional space. Different types of graphical analysis are expected to aid in the rationalization of multi-dimensional property spaces. Therefore, a visualization methodology for multi-dimensional property spaces has been developed, as reported herein. Our analysis was based upon the gen- eration of drug-like subspaces in chemical space, which takes molecular similarity relationships implicitly into account. However, it would also be feasible to focus an analysis explicitly on selected distance relationships in chemical space (or generate subspaces for compound ref- erence sets with other characteristic properties).
Our study introduces the STC and PAC concepts, adapted from computer graphics, to the medicinal chem- istry community. STC/PAC visualization of compound data is designed to complement multi-objective optimiza- tion, provide access to multi-dimensional data distribu- tions, and aid in compound selection. For a given
Fig. 4 Numerical comparison of projections. A projection was created for each weight value setting of the multi-objective function containing 14 descriptors and the number of drugs within the 20 top ranked compounds was determined. The graph reveals the number of weight combinations yielding largest numbers of highly-ranked drugs across the different target sets (colored by target IDs given in Table1)
Descriptor Settings -1 -0.33 0.33 1 a_acc a_aroR a_don a_ringR b_rotR chiral_u FCharge logP(o/w) logS PEOE_VSA_FHYD PEOE_VSA_FPNEG PEOE_VSA_FPPOS Pot Weight Setting 1 Setting 2 Projection 1 Projection 2
Top ranked drug compounds Top ranked bioactive compounds Other drug compounds Other bioactive compounds
Descriptor Settings -1 -0.33 0.33 1 a_acc a_aroR a_don a_ringR b_rotR chiral_u FCharge logP(o/w) logS PEOE_VSA_FHYD PEOE_VSA_FPNEG PEOE_VSA_FPPOS Pot Weight Setting 1 Setting 2 Projection 1 Projection 2 (a) (b) (c) (d)
Fig. 5 Visualization of projections. Exemplary projections are visu- alized and compared. In (a) and (b), two projections generated for beta-2 adrenergsic receptors (ChEMBL target ID 210) are shown. The corresponding top 20 rankings contained 13 drugs each (11 of which were the same). a Compares the weight combinations (settings) for these projections and b their STC visualizations. Points represent individual compounds and are color-coded according to Fig.3a. In (c) and (d), two projections generated for alpha-2a adrenergic receptor ligands (ID 1867) are shown. The corresponding top 20
rankings contained eight drugs each (seven of which were the same). cCompares the weight combinations (settings) for these projections and d their STC visualizations. In (b) and (d), STC visualizations were scaled to the same value ranges. e PCA-based data set projections (using the first two PCs) with unweighted descriptors (top, drugs colored cyan and bioactive compounds gray) and weighted descriptors from projection 1 (middle) and 2 (bottom) taken from (c). PCA plots of projections are color-coded as in (d) J Comput Aided Mol Des
projection and compound ranking, the STC visualization provides a 2D representation of a compound distribution in multi-dimensional property space and views highly ranked compound subsets in the data set context. In addition, the PAC representation compares individual property contri- butions and identifies property settings that distinguish highly ranked compounds from others. We have demon- strated that STC visualizations help to differentiate numerically equivalent optimization solutions with similar or distinct property settings. The data sets used herein are made freely available [30].
References
1. Stumpfe D, Bajorath J (2012) Methods for SAR visualization. RSC Adv 2:369–378
2. Wassermann AM, Wawer M, Bajorath J (2010) Activity land- scape representations for structure-activity relationship analysis. J Med Chem 53:8209–8223
3. Shanmugasundaram V, Maggiora GM (2001) Characterizing property and activity landscapes using an information-theoretic approach. In: Proceedings of 222nd American chemical society national meeting, division of chemical information, Chicago, IL, August 26–30, 2001; American Chemical Society: Washington, D.C., 2001; abstract no. 77
4. Wawer M, Peltason L, Weskamp N, Teckentrup A, Bajorath J (2008) Structure-activity relationship anatomy by network-like similarity graphs and local structure-activity relationship indices. J Med Chem 51:6075–6084
5. Wollenhaupt S, Baumann K (2014) inSARa: Intuitive and inter- active SAR interpretation by reduced graphs and hierarchical MCS-based network navigation. J Chem Inf Model 54:1395–1409 6. Agrafiotis DK, Shemanarev M, Connolly PJ, Farnum M, Lobanov VS (2007) SAR maps: a new SAR visualization technique for medicinal chemists. J Med Chem 50:5926–5937
7. Wassermann AM, Bajorath J (2012) Directed R-group combi- nation graph: a methodology to uncover structure-activity rela- tionship patterns in a series of analogues. J Med Chem 55:1215–1226
8. Peltason L, Weskamp N, Teckentrup A, Bajorath J (2009) Exploration of structure-activity relationship determinants in analogue series. J Med Chem 52:3212–3224
9. Wawer M, Bajorath J (2010) Similarity-potency trees: a method to search for SAR information in compound data sets and derive SAR rules. J Chem Inf Model 50:1395–1409
10. Peltason L, Iyer P, Bajorath J (2010) Rationalizing three-di- mensional activity landscapes and the influence of molecular representations on landscape topology and the formation of activity cliffs. J Chem Inf Model 50:1021–1033
11. Reutlinger M, Guba W, Martin RE, Alanine AI, Hoffmann T, Klenner A, Hiss JA, Schneider P, Schneider G (2011) Neigh- borhood-preserving visualization of adaptive structure-activity landscapes: application to drug discovery. Angew Chem Int Ed 50:11633–11636
12. Zwierzyna M, Vogt M, Maggiora GM, Bajorath J (2015) Design and characterization of chemical space networks for different compound data sets. J Comput-Aided Mol Des 29:113–125 13. Ertl P, Rohde B (2012) The molecule cloud-compact visualiza-
tion of large collections of molecules. J Cheminf 4:12
-15 -10 -5 0 5 -4 -2 0 2 4 6 8 PC1 PC2 -15 -10 -5 0 5 -4 -2 02 46 8 PC1 PC2 -15 -10 -5 0 5 -4 -2 0 2 4 6 8 PC1 PC2
(e)
Fig. 5 continued14. Awale M, van Deursen R, Reymond J-L (2010) MQN-mapplet: visualization of chemical space with interactive maps of Drug- Bank, ChEMBL, PubChem, GDB-11, and GDB-13. J Chem Inf Model 50:1395–1409
15. Reymond J-L (2015) The chemical space project. Acc Chem Res 48:722–730
16. Kireeva N, Baskin II, Gaspar HA, Horvath D, Marcou G, Varnek A (2012) Generative topographic mapping (GTM): universal tool for data visualization, structure-activity modeling, and dataset comparison. Mol Inf 3(4):301–312
17. Wermuth CG (2008) The practice of medicinal chemistry, 3rd edn. Academic Press-Elsevier, Burlington, London
18. Gillet VJ, Khatib W, Willett P, Fleming P, Green DVS (2002) Combinatorial library design using multiobjective genetic algo- rithm. J Chem Inf Comput Sci 42:375–385
19. Gillet VJ (2004) Applications of evolutionary computation in drug design. Struct Bond 110:133–152
20. Nicolaou CA, Brown N, Pattichis CS (2007) Molecular opti- mization using computational multi-objective methods. Curr Opin Drug Discov Develop 10:316–324
21. Gaulton A, Bellis LJ, Bento AP, Chambers J, Davies M, Hersey A, Light Y, McGlinchey S, Michalovich D, Al-Lazi- kani B, Overington JP (2012) ChEMBL: a large-scale bioac- tivity database for drug discovery. Nucleic Acids Res 40:D1100–D1107
22. Law V, Knox C, Djoumbou Y, Jewison T, Guo AC, Liu Y, Maciejewski A, Arndt D, Wilson M, Neveu V, Tang A, Gabriel G, Ly C, Adamjee S, Dame ZT, Han B, Zhou Y, Wishart DS (2014) DrugBank 4.0: shedding new light on drug metabolism. Nucleic Acids Res 42:D1091–D1097
23. OEChem TK (2012) OpenEye scientific software Inc, Santa Fe, NM, USA
24. Molecular Operating Environment (2012) Chemical computing group Inc.: Montreal, Quebec, Canada
25. Cook D, Buja A, Lee EK, Wickham H (2008) Grand tours, projection pursuit guided tours and manual controls. In: Chen C, Ha¨rdle W, Unwin A (eds) Handbook of data visualization. Springer, Heidelberg, pp 295–314
26. Kandogan E (2000) Star coordinates: a multi-dimensional visu- alization technique with uniform treatment of dimensions. In: LBHT Proc IEEE information visualization symposium, pp 9–12 27. Java universal network/graph framework. http://jung.source
fourge.net/. Accessed May 1, 2014
28. Inselberg A (1985) The plane with parallel coordinates. Visual Comput 1:69–91
29. R: a language and environment for statistical computing. R foundation for statistical computing, Vienna, Austria, 2012 30. de la Vega de Leo´n A, Kayastha S, Dimova D, Schultz T,
Bajorath J (2015) ChEMBL20 data sets for multi-property land- scape analysis. ZENODO. doi:10.5281/zenodo.21782