Explainer

Sampling: Why the First Step of Analysis Matters Most

Lab Techniques & AnalysisAdvanced6 min read
On this page
  1. The goal: a representative sample
  2. Sampling strategies
  3. How much sample?
  4. From field sample to lab portion: subsampling
  5. Preserving the sample
  6. Documentation and chain of custody
  7. Worked example: why sampling error dominates
  8. Examples across fields
  9. Key takeaways

Imagine an analyst with a state-of-the-art ICP-MS, perfect calibration and flawless technique. The instrument reports that a field contains 42 mg of lead per kilogram of soil, to three significant figures. But the 0.5 g that went into the instrument was scooped from one spot next to an old painted fence. The number is precise and completely misleading.

This is the uncomfortable truth of analytical chemistry: the measurement can never be better than the sample. In many real analyses, most of the total uncertainty comes not from the instrument but from how the sample was collected. That’s why sampling deserves as much thought as any instrument.

The goal: a representative sample

The population (or lot) is the whole thing you want to know about: a lake, a field, a batch of 10,000 tablets, a lorry-load of grain. You can’t analyse all of it. A sample is the part you actually take, and it must be representative: its composition should match the population’s, within an acceptable uncertainty.

The difficulty is heterogeneity. Real materials are rarely uniform:

  • A lake may be warmer, more oxygenated and more polluted near the surface or near an outflow.
  • A field has hot spots of contamination.
  • Powders separate by particle size when shaken or poured (fine particles sink, large ones rise, the so-called “Brazil nut effect”).
  • A batch of tablets may drift in composition from the start to the end of a production run.

Sampling strategies

Random sampling

Every part of the population has an equal chance of being chosen, for example by using random numbers to pick grid squares in a field or tablets from a batch. This avoids bias from the sampler’s choices (such as always sampling the most accessible spots).

Systematic sampling

Samples are taken at regular intervals: every 10 m on a grid, every hour from a production line. It’s practical and gives even coverage, but it can miss patterns that happen to repeat at the same interval as the sampling.

Stratified sampling

If the population has known distinct zones (strata), such as the top and bottom of a tank, or different soil types in a field, each stratum is sampled separately. This gives more information and lower uncertainty for the same effort.

Judgemental (targeted) sampling

The analyst deliberately samples suspected problem areas, such as next to a leaking drum. It’s useful for finding contamination but can’t be used to estimate an average for the whole site.

Composite samples

Several increments from different places or times are combined and mixed into one sample. This gives a good estimate of the average at lower cost, but information about variation between locations is lost. A 24-hour composite of wastewater, for example, averages out changes through the day.

How much sample?

For particulate materials, the key issue is the size of the particles relative to the mass taken. If a contaminant occurs as a few large grains, a small sample may contain none of them, or one, which would make the result wildly high.

Two general principles:

  • Larger particles need larger samples. Coarse ore may need kilograms; a fine, well-mixed powder may need only grams.
  • Grinding helps. Reducing particle size makes the material more homogeneous, so smaller subsamples become representative.

The number of increments matters too. Taking many small increments reduces the random sampling error roughly in proportion to 1 ÷ √n, where n is the number of increments. So four times as many increments halves the sampling error. See calculating uncertainty.

From field sample to lab portion: subsampling

A field sample of a few kilograms must be reduced to perhaps 0.2 g for analysis without losing representativeness.

Coning and quartering is the classic method for powders and soils:

  1. Pour the (dried, ground) material into a cone on a clean surface.
  2. Flatten the cone into a circular cake.
  3. Divide it into four quarters.
  4. Keep two opposite quarters and discard the other two.
  5. Mix and repeat until the mass is small enough.

Riffle splitters, boxes with alternating chutes that divide a poured sample into two equal halves, do the same job more reproducibly.

Preserving the sample

A sample can change between collection and analysis, giving a result that doesn’t reflect the original state.

Problem Example Prevention
biological activity bacteria consume nitrate or organic matter cool to about 4 °C, analyse quickly
adsorption onto container walls metals sticking to glass acidify with nitric acid; use suitable plastic
loss of volatile compounds solvents escape fill bottles with no air gap; seal tightly
oxidation Fe²⁺ → Fe³⁺ exclude air; acidify
contamination from the container plasticisers leaching into water use glass for organic analysis
reaction with disinfectant chlorine continuing to react add a neutralising agent

See how drinking water is tested for how these rules apply in practice.

Documentation and chain of custody

Every sample needs a record: location (often with GPS coordinates), date and time, who took it, the method, the container, the preservation used and conditions at the time. For regulatory or legal work, a chain of custody record tracks every transfer of the sample, just as in forensic chemistry.

Blank samples help too. A field blank (pure water taken to the site, opened and handled like a real sample) reveals contamination picked up during sampling and transport.

Worked example: why sampling error dominates

An analyst measures lead in soil from a garden. Five separate samples from across the garden give 38, 55, 29, 71 and 47 mg/kg. Repeat measurements on a single, well-ground sample agree within ±1 mg/kg.

  • The analytical variation is about ±1 mg/kg.
  • The sampling variation, shown by the spread across the garden (29 to 71), is roughly ±15 mg/kg.
  • The mean of the five samples is (38 + 55 + 29 + 71 + 47) ÷ 5 = 48 mg/kg.

Improving the instrument would barely change the uncertainty of the garden’s average lead content. Taking more samples, or combining many increments into composites, would. This is typical: sampling often contributes most of the uncertainty. See experimental errors for how random and systematic errors differ.

Examples across fields

  • Pharmaceuticals: tablets are sampled from the beginning, middle and end of a batch, and content uniformity is tested on individual tablets.
  • Food: a shipment of nuts may be tested for aflatoxins. Because contamination is concentrated in a few nuts, large samples (several kilograms) are ground and subsampled. See food analysis.
  • Air: pumps draw a measured volume of air through filters or sorbent tubes over a set time.
  • Blood: the time since a drug dose matters as much as the method.

Key takeaways

  • An analysis is only as good as its sample; sampling often contributes more uncertainty than the measurement.
  • A representative sample must account for heterogeneity in space and time.
  • Random, systematic, stratified, targeted and composite strategies each suit different questions.
  • Grinding, adequate sample size and many increments reduce sampling error; coning and quartering or riffling reduce mass without bias.
  • Preservation, documentation, blanks and chain of custody protect the sample’s integrity. For the analysis that follows, see qualitative vs quantitative analysis.

Advertisement

More from this topic: Lab Techniques & Analysis