Statistical Foundations - ScreenAssist
Classical vs. Robust Statistics
Understanding the distinction between classical and robust statistical approaches is essential for interpreting uHTS assay metrics correctly.
Classical Statistics
Classical statistics (sometimes called parametric statistics) use the arithmetic mean (μ) and standard deviation (σ) to describe data distributions. These metrics assume that the data follow a roughly normal (Gaussian) distribution and are highly sensitive to outliers. A single extreme data point, such as a contaminated well, a precipitated compound, or an autofluorescent compound, can dramatically inflate the standard deviation and artificially deflate quality metrics such as Z Prime (Z’).
Classical statistics are most appropriate when the dataset is clean and outliers have been removed prior to analysis.
Robust Statistics
Robust statistics use measures that are resistant to the influence of outliers. Instead of the mean and standard deviation, robust approaches substitute:
Median in place of the arithmetic mean, the middle value of a sorted dataset, unaffected by extreme values at either end.
Median Absolute Deviation (MAD) in place of the standard deviation. the median of the absolute differences between each data point and the dataset median.
Where xᵢ are the individual data points, and is the median of the dataset.
The robust standard deviation σrobust is derived from the MAD by applying a consistency factor:
This scaling factor (1.4826) ensures that for a perfectly normal distribution, equals the classical standard deviation. For datasets with outliers, will be substantially smaller than the classical σ, providing a more realistic picture of true assay variability.
Note: In uHTS on Dianthus, it is recommended to calculate both classical and robust versions of key metrics (Z’ and Robust Z’) as a cross-check. Large differences between the two suggest significant outlier contamination in the data. We recommend to use Robust Statistics for assay validation.
Z Prime (Z’)
Z’ is the gold-standard metric for assessing the suitability of a high-throughput screening assay. Originally described by Zhang et al. (1999) , it quantifies the separation between the positive control and the neutral reference (negative control) distributions, taking into account the variability of both.
The formula captures the ratio of the combined 3σ “noise” zone (the 99.7% confidence windows of both control populations) to the total signal separation between them.
Z’ Value |
Interpretation |
> 0.5 |
Excellent |
0.4 – 0.5 |
Marginal — review carefully |
< 0.4 |
Poor — not suitable for screening |
In the Dianthus uHTS context, Z’ is calculated over the Spectral Shift ratio (650 nm / 670 nm) values of positive control and neutral reference wells.
Note: Z’ is sensitive to outliers because it uses the classical mean and standard deviation. A single autofluorescent compound in the control population can dramatically reduce the apparent Z’. As such, for Dianthus uHTS, we recommend the use of Robust Z Prime (RZ’) for validating assays.
Robust Z Prime (Robust Z’)
Robust Z’ replaces the mean (μ) with the median (x̃) and the standard deviation (σ) with the robust standard deviation equivalent (derived from the MAD) in the Z’ formula:
where x̃ denotes the median.
Robust Z’ is the preferred metric for real-world screening datasets. It provides a more stable and realistic reflection of assay quality.
Diagnostic tip: When Robust Z’ and classical Z’ diverge significantly — for example, Robust Z’ is excellent (≥ 0.5) but classical Z’ is marginal (< 0.5) — this is a strong indicator that a small number of outlier wells are inflating the classical standard deviation. Outlier investigation (are they spatially co-located, is there patterns of outliers) should be conducted before making final quality decisions.
Z Score
The Z score is a per-compound normalization metric that expresses how many standard deviations a compound’s measurement lies from the neutral reference mean. It is applied at the data normalization stage to enable comparison of measurements across plates and across a full screen.
Where is the ratio value for an individual compound well, is the mean of the neutral reference wells on the same plate, and is the standard deviation of the neutral reference wells.
A compound with a Z score ≥ 3 (i.e., 3 standard deviations from the reference) would typically be flagged as a potential hit, though the appropriate cut-off depends on the assay and the desired hit rate.
Robust Z Score
The Robust Z score substitutes the median and robust standard deviation of the neutral reference in place of the classical mean and SD:
Robust Z scores are preferred for Dianthus uHTS. They reduce the rate of false positives introduced by statistical noise.
Strictly Standardized Mean Difference (SSMD)
SSMD is a statistical measure of effect size that quantifies the separation between the positive control and neutral reference populations relative to their combined variability. Unlike Z’, SSMD is derived from probability theory and has a direct relationship to false positive and false negative rates.
SSMD accounts for the variance of both control populations independently, making it statistically more rigorous than Z’ in certain scenarios. In the HTS field, |SSMD| ≥ 6 is considered an excellent assay, equivalent to well-separated control populations with a very low false positive/negative rate.
SSMD complements Z’ as a validation metric and is particularly useful when:
Control population sizes are unequal.
The two populations have markedly different variances.
A direct probabilistic interpretation of assay quality is required.
Note: Robust SSMD variants using medians and MAD-derived standard deviations can also be calculated for datasets with outliers.
Fluorescence Coefficient of Variation (CV)
The Coefficient of Variation (CV) is a normalized measure of signal dispersion. It is calculated as the ratio of the standard deviation to the mean, expressed as a percentage:
In the Dianthus (uHTS) context, the Fluorescence CV is measured at 650 nm for the positive control wells and neutral reference wells. It reports how reproducibly the labelled target has been dispensed into the wells and how consistent the fluorescence signal is across the plate.
A high CV indicates variability in:
Target loading volume
Target stability within the dispensing system
Target stability within the assay plates
Sudden changes in Fluorescence CV between plates could be:
Changes in target stability
Changes in the dispense of control compounds and/or DMSO
All of these will reduce assay sensitivity and inflate false positive or negative rates. Keeping the 650 nm CV within acceptable limits (≤ 5%) ensures that signal variability is dominated by biology (binding events) rather than technical noise.
Neutral RSDE (Robust Standard Deviation Equivalent)
The Neutral RSDE is a plate-level quality metric that describes the intrinsic variability of the neutral reference wells using robust statistics. It provides information on signal stability, and therefore behavior and stability of the labelled target molecule. It is calculated as the robust standard deviation of the Spectral Shift ratio values across all neutral reference wells on a plate:
RSDE is reported in ratio units (dimensionless).
A low Neutral RSDE confirms that the neutral reference population is tight and reproducible, a prerequisite for reliable hit identification.
A high RSDE indicates excessive noise in the reference population, which will broaden the hit-calling distribution and increase the false positive rate. When the RSDE is high, Z score-based hit thresholds become less reliable because the statistical noise floor is elevated.