On this page
- What standard deviation measures
- Worked example 1: step by step
- Relative standard deviation (RSD)
- Standard error of the mean
- Confidence intervals
- Comparing a result with a known value
- Testing a suspect value: the Q-test
- Precision in a calibration
- Repeatability vs reproducibility
- Using a spreadsheet or calculator
- Common mistakes
- Key takeaways
The range tells you roughly how spread out a set of results is, but it only uses two numbers: the highest and the lowest. Standard deviation uses every result, and it’s the measure of precision used throughout analytical chemistry, from university practicals to pharmaceutical quality control. This article works through the calculations step by step, then shows how standard deviation leads to confidence intervals and to deciding whether results genuinely differ.
For simpler measures, start with averages and anomalies in lab data.
What standard deviation measures
Standard deviation (s) is, roughly, the typical distance of individual results from the mean. A small s means results are tightly clustered (high precision); a large s means they’re scattered.
For a set of n repeat measurements (a sample of all possible measurements), the sample standard deviation is:
s = √[ Σ(xᵢ − x̄)² ÷ (n − 1) ]
where x̄ is the mean and xᵢ are the individual results.
Why n − 1 rather than n? Because the mean was calculated from the same data, one “degree of freedom” is used up; dividing by n − 1 corrects for the tendency of small samples to underestimate the true spread. Calculators usually label this sₙ₋₁ or s; the version dividing by n is labelled σₙ.
Worked example 1: step by step
Five titres (cm³): 24.50, 24.55, 24.45, 24.60, 24.50.
Step 1: mean x̄ = (24.50 + 24.55 + 24.45 + 24.60 + 24.50) ÷ 5 = 122.60 ÷ 5 = 24.52 cm³
Step 2: deviations and squares
| xᵢ | xᵢ − x̄ | (xᵢ − x̄)² |
|---|---|---|
| 24.50 | −0.02 | 0.0004 |
| 24.55 | +0.03 | 0.0009 |
| 24.45 | −0.07 | 0.0049 |
| 24.60 | +0.08 | 0.0064 |
| 24.50 | −0.02 | 0.0004 |
| Sum | 0.0130 |
Step 3: divide by n − 1 0.0130 ÷ 4 = 0.00325
Step 4: square root s = √0.00325 = 0.057 cm³
Result: 24.52 ± 0.06 cm³ (mean ± one standard deviation), n = 5.
Relative standard deviation (RSD)
To compare precision between measurements of different sizes, express s as a percentage of the mean:
RSD (%) = (s ÷ x̄) × 100
For example 1: RSD = 0.057 ÷ 24.52 × 100 = 0.23%. Analytical methods in quality control often have RSD requirements, such as below 1% or 2%, depending on the analysis.
Worked example 2: comparing precision
Method A measures iron in tablets: mean 65.2 mg, s = 0.8 mg. Method B measures lead in water: mean 4.1 μg/L, s = 0.3 μg/L. Which is more precise?
- A: RSD = 0.8 ÷ 65.2 × 100 = 1.2%
- B: RSD = 0.3 ÷ 4.1 × 100 = 7.3%
Method A is more precise relative to its measurement, even though the units and sizes differ. Precision is usually poorer at trace levels.
Standard error of the mean
The mean of several results is more reliable than any single result. Its uncertainty is the standard error:
Standard error = s ÷ √n
For example 1: 0.057 ÷ √5 = 0.057 ÷ 2.236 = 0.025 cm³
Taking four times as many measurements halves the standard error, which is why repeats improve the reliability of a mean, but with diminishing returns.
Confidence intervals
A 95% confidence interval gives a range in which we’re 95% confident the true mean lies (assuming no systematic error):
x̄ ± t × s ÷ √n
where t comes from the t-distribution for n − 1 degrees of freedom:
| n − 1 | 2 | 3 | 4 | 5 | 9 | large |
|---|---|---|---|---|---|---|
| t (95%) | 4.30 | 3.18 | 2.78 | 2.57 | 2.26 | 1.96 |
Worked example 3
For example 1 (n = 5, degrees of freedom = 4, t = 2.78):
24.52 ± 2.78 × 0.057 ÷ √5 = 24.52 ± 2.78 × 0.0255 = 24.52 ± 0.071 cm³
95% confidence interval: 24.45 to 24.59 cm³
Notice how t is large for small n: with only three results, the interval is much wider. This is the statistical reason for doing enough repeats.
Comparing a result with a known value
Worked example 4: is there a systematic error?
A student analyses a certified reference material containing 50.0 mg/g of calcium. Six results give a mean of 48.9 mg/g, s = 0.9 mg/g.
95% confidence interval: t (5 degrees of freedom) = 2.57 48.9 ± 2.57 × 0.9 ÷ √6 = 48.9 ± 2.57 × 0.367 = 48.9 ± 0.94 Interval: 47.96 to 49.84 mg/g
The certified value, 50.0 mg/g, lies outside the interval. So the difference is unlikely to be due to random error alone: there is probably a systematic error (such as incomplete dissolution or a calibration problem). This logic, formally a t-test, is how labs validate methods. See calibration curves and experimental errors.
Testing a suspect value: the Q-test
When one result looks out of line, Dixon’s Q-test gives a simple, objective check for small data sets:
Q = |suspect value − nearest value| ÷ range
If Q is greater than the critical value, the suspect value may be rejected.
| n | 3 | 4 | 5 | 6 | 7 |
|---|---|---|---|---|---|
| Q crit (95%) | 0.970 | 0.829 | 0.710 | 0.625 | 0.568 |
Worked example 5
Results (%): 12.47, 12.51, 12.49, 12.72, 12.50.
- Suspect value: 12.72. Nearest value: 12.51. Range: 12.72 − 12.47 = 0.25.
- Q = (12.72 − 12.51) ÷ 0.25 = 0.21 ÷ 0.25 = 0.84
- Critical Q for n = 5 is 0.710. Since 0.84 > 0.710, 12.72 can be rejected at 95% confidence.
- New mean of the remaining four: (12.47 + 12.51 + 12.49 + 12.50) ÷ 4 = 12.49%
Grubbs’ test, which compares the suspect value’s distance from the mean with s, is now often recommended instead of Q, particularly by international standards bodies. Either way, only one outlier should be tested at a time, and rejected data should still be recorded.
Precision in a calibration
For calibration graphs, the scatter of points around the best-fit line gives a standard deviation of the residuals, which is used to calculate the uncertainty in concentrations read from the graph and to estimate a method’s limit of detection (often taken as about three times the standard deviation of the blank signal).
Repeatability vs reproducibility
Analysts distinguish two kinds of precision. Repeatability is the spread when one analyst repeats a measurement on the same equipment over a short time. Reproducibility is the spread when different analysts, instruments or labs measure the same sample. Reproducibility standard deviations are almost always larger, because they include extra sources of variation, and they are what matters when results from different labs are compared.
Using a spreadsheet or calculator
- Spreadsheets:
=AVERAGE(range),=STDEV.S(range)(sample standard deviation),=STDEV.S(range)/SQRT(COUNT(range))for standard error, and=T.INV.2T(0.05, n-1)for t. - Scientific calculators: use statistics mode, enter data, and read x̄ and sₓ (or σₙ₋₁).
Always check whether the function divides by n or n − 1.
Common mistakes
- Using σ (divide by n) instead of s (divide by n − 1) for a small set of repeats.
- Confusing standard deviation (spread of individual results) with standard error (uncertainty of the mean).
- Using 1.96 instead of the proper t-value for small n.
- Rejecting several “outliers” until the data look nice.
- Assuming a small s means an accurate result; precision says nothing about systematic error.
Key takeaways
- Sample standard deviation s = √[Σ(xᵢ − x̄)² ÷ (n − 1)] measures precision using every result.
- RSD (%) = s ÷ x̄ × 100 lets you compare precision across different measurements.
- Standard error s ÷ √n gives the uncertainty of the mean; a 95% confidence interval is x̄ ± t·s/√n.
- If a known value lies outside the confidence interval, suspect a systematic error.
- Use Q or Grubbs’ tests to judge single outliers objectively, and always record rejected data. See propagating uncertainty.
Advertisement