QSAR Basics
On this page
Direct answer
Quantitative structure-activity relationship (QSAR) methods replace trial-and-error synthesis with equations linking structural descriptors to biological activity. The classical Hansch analysis is a multiple linear regression of potency, expressed as log(1/C), on three descriptor families: hydrophobic (π, the substituent's contribution to partitioning, or log P for the whole molecule), electronic (Hammett σ) and steric (Taft Es), with molar refractivity often added — log(1/C) = a·π + b·σ + c·Es + d. Its intellectual ancestor is the Meyer-Overton correlation between anaesthetic potency and lipid solubility, the oldest linear free-energy relationship in pharmacology. Practical tools include the Craig plot (σ against π), which flags substituent sets that decorrelate the variables so regression can separate their effects; Free-Wilson analysis, which treats group contributions additively; and modern descendants — topological indices and 3D-QSAR such as CoMFA, where steric and electrostatic fields sampled around aligned molecules feed partial least squares. Statistics decide credibility: r² for fit, cross-validated q² (leave-one-out) for predictive power on held-out compounds, and an applicability domain outside which predictions are not claimed.
What you must remember
- The Hansch equation: log(1/C) = a·π + b·σ + c·Es + d, with C the molar concentration for a standard effect — the form exam papers ask you to expand term by term.
- Descriptor families: hydrophobic (π, log P), electronic (Hammett σ, resonance and inductive), steric (Taft Es, molar refractivity, Verloop parameters) — the classic triad.
- Meyer-Overton origin: anaesthetic potency tracking lipid solubility — the founding observation that potency could be predicted from a physical property.
- Craig plot: σ versus π scatter that lets medicinal chemists pick substituent sets spanning both axes orthogonally, avoiding collinearity in the regression.
- Free-Wilson model: purely additive contributions of each substituent position, no physical descriptor needed — the simpler sibling of Hansch.
- CoMFA and 3D-QSAR: aligned molecules in a lattice; steric and electrostatic field values at grid points become thousands of predictors handled by partial least squares, producing contour maps for design.
- Statistical honesty: r² measures fit, q² from leave-one-out cross-validation measures prediction; a model is judged by compounds it has never seen and by its applicability domain.
- Companion filter: Lipinski's rule of five — molecular weight ≤ 500, log P ≤ 5, ≤ 5 hydrogen-bond donors, ≤ 10 acceptors — a peroral solubility-permeability screen, not a QSAR itself but always asked alongside.
Reading a Hansch equation like a designer
Consider, as a teaching example, substituted aromatics whose Hansch analysis returns log(1/C) = 0.8·π − 1.2·σ + 0.3·Es + constant — numbers chosen for illustration, but read them the way a chemist would. The positive π coefficient says potency climbs with lipophilicity, plausibly better passage to a hydrophobic pocket; the negative σ coefficient says electron-donating groups (negative σ) help, pointing to ring electron density as part of the binding interaction; the small Es term says size matters little at this position. The design move follows: pick strong electron donors with decent lipophilicity — methoxy and methyl, not nitro — and check the Craig plot so the chosen set decorrelates σ and π for the next regression.
Contrast the failure mode every examiner loves: add a bulkier, more lipophilic analogue and potency collapses despite a "better" π — activity versus log P is usually parabolic, an optimum near 2 for many centrally acting drugs, past which solubility and desolvation costs dominate. The linear equation was honest only inside its domain, and modern practice formalises the caution: split the data into training and test sets, report q² on held-out compounds, and state the applicability domain — a QSAR is a prediction engine, not a law.
Where students slip
Sign conventions sink more answers than the mathematics: negative σ means electron donation (Hammett defined σ against benzoic acid ionisation), so a negative σ coefficient favours donors — careless readers design the opposite substituent. Second, π versus log P: π is the substituent's incremental constant, log P the whole molecule's. Third, r² worship: heavily intercorrelated descriptors can return an impressive training r² and fail every new compound, which is why q² and test-set validation exist. Fourth, CoMFA answers that skip alignment miss the method's known sensitivity to superposition. Finally, Lipinski's rule is quoted as an absolute bar — it only flags probable poor oral absorption, and many successful drugs violate it.
Frequently asked questions
What are the terms of the Hansch equation?
log(1/C) = a·π + b·σ + c·Es + d, where π is hydrophobic, σ the Hammett electronic and Es the Taft steric substituent constant; the coefficients guide next-round synthesis.
What is a Craig plot used for?
It plots σ against π for candidate substituents so a series can be chosen spanning both axes without collinearity, letting the regression separate hydrophobic from electronic effects.
Why is cross-validated q² preferred over r²?
r² only describes fit to the training data and inflates with intercorrelated descriptors; q² from leave-one-out or test-set validation estimates true predictive power on new compounds.
How does CoMFA differ from classical QSAR?
It uses three-dimensional steric and electrostatic fields sampled at lattice points around aligned molecules, analysed by partial least squares, and outputs contour maps showing where bulk or charge helps or hurts.
What does Lipinski's rule of five state?
Molecular weight ≤ 500, log P ≤ 5, hydrogen-bond donors ≤ 5 and acceptors ≤ 10; exceeding two thresholds flags probable poor oral absorption — a guideline, not a law.