The Test Uncertainty Ratio: What It Is, Why It Beat TAR, and How to Use It

Your calibration certificate says Pass. Before you file it, ask a question almost nobody asks. What decision rule did the lab use to make the PASS decision, and how much risk did it share to you? The test uncertainty ratio (TUR) is one the tools to help determine this. It compares your equipment’s tolerance against the real uncertainty of the calibration that judged it. That turns a simple acceptance passing result into something that could carry more confidence.

This guide covers what the TUR test uncertainty ratio is and how it can be used. It defines the term, walks the calculation, and explains why the industry moved away from the older test accuracy ratio. It then shows how the ratio feeds a decision rule, and then we demonstrate an example using a pipette calibration case with 2 different balances (gravimetric measurement calibration).

What the Test Uncertainty Ratio Actually Measures

ANSI/NCSL Z540.3, Clause 3.11 defines it. Paraphrased, you divide the width of the tolerance band on the item being calibrated by 2 times the expanded uncertainty of the calibration process at 95 percent confidence. The clause applies that form to two sided tolerances.

TUR  =  ΔUUT2×UU  =  k×u12+u22++un2TUR \;=\; \frac{\Delta_{UUT}}{2 \times U} \qquad\qquad U \;=\; k \times \sqrt{u_{1}^{2} + u_{2}^{2} + \cdots + u_{n}^{2}}

The key here is that we are comparing the entire Measurement process uncertainty to the tolerance band. Not just the reference standard. The denominator covers the whole calibration. That means the reference, the resolution of your device, the repeatability of the readings, and any environmental or operator effect that moves the result.

So the ratio answers a practical question. How much of your tolerance band does the measurement doubt consume? A high ratio means the calibration resolves your tolerance comfortably. A low ratio means uncertainty crowds the limits, and a reading near the edge becomes close to a coin flip. A low ration means that if you have a tolerance band of 1 unit and the measurement uncertainty is 0.5 units x 2, that means that your uncertainty is about equal to your tolerance band. So, how are you going to know if your equipment passes or not if the tolerance window is the same as your measurement uncertainty?

One note on wording. ILAC G8:09/2019 defines the same idea as the tolerance limit divided by the expanded uncertainty. For a symmetric two sided tolerance both forms give the identical number. Cite whichever you use, and do not blend them.

How to Calculate the Test Uncertainty Ratio

Six steps, about ten minutes per test point.

  1. Find the tolerance span. Take the tolerance at your test point and subtract the lower limit from the upper limit. A tolerance of ± 0.25 psi gives a span of 0.50 psi.
  2. Convert the laboratory’s method uncertainty. Take the CMC from the scope of accreditation, published as an expanded value at k = 2, and divide by that coverage factor. Two cautions here. A CMC describes capability on a best existing device, so it can understate the real process uncertainty and flatter your ratio. And the uncertainty printed on your certificate usually already covers your device’s resolution and repeatability, so adding Steps 3 and 4 on top counts them twice.
  3. Convert your device’s resolution. The displayed increment is the full width of a rectangular distribution, so divide it by the square root of twelve.
  4. Measure repeatability. Take at least five readings at the test point without changing anything, then calculate the sample standard deviation.
  5. Combine and expand. Square each standard uncertainty, sum the squares, take the square root, then multiply by k = 2.
  6. Divide. Tolerance span divided by twice the expanded uncertainty.

Watch the two factors of two. The k stretches the result to roughly 95 percent confidence. The 2 in the ratio equation converts the full span back to a half width. Drop either and you overstate the ratio twofold.

Need help with Calculating TUR? Schedule a free consultation with Precision ISO.
Schedule a Free Consult

Why TUR Replaced TAR

The test accuracy ratio compares two published specifications. You divide the tolerance of the unit under test by the tolerance of the reference standard, and you stop there.

Start with a fact that surprises most people. TAR carries no definition in the Z540 documents. Clause 3 of Z540.3 defines the test uncertainty ratio and says nothing about TAR, and neither document is current in any case. The definition in general use comes from ASQ. NASA metrologist Scott Mimbs has argued in NCSLI conference work that this definition complies with neither Z540 document. It addresses tolerance ratios without accounting for the actual uncertainty components.

Look at what TAR leaves out.

What moves your readingCounted by TARCounted by TUR
Reference standard specificationYesYes
Repeatability of the readingsNoYes
Resolution of the deviceNoYes
Environmental effectsNoYes
Operator techniqueNoYes
Other devices in the calibration chainNoYes
Demonstrated performance rather than a datasheet claimNoYes

Two consequences follow. First, TAR overstates capability, sometimes dramatically. A pressure gauge calibration can return a TAR of 10:1 and a test uncertainty ratio of 3.4:1. Same job, same day, same instrument. Second, TAR stops scaling up the traceability chain, because no tier stays four times better than the one below it forever.

TAR still appears in older procedures and legacy contracts. Honor it where a contract binds you. Just recognize that it compares two claims, while TUR measures a process.

Simple Acceptance: The Decision Rule Nobody Agreed To

A decision rule describes how measurement uncertainty is accounted for when stating conformity. Simple Acceptance is the rule where the acceptance limit equals the tolerance limit, so the guard band width is zero. ILAC G8:09/2019 calls it exactly that, and it notes the other name the industry uses: shared risk.

Why would it be called shared risk? Who is sharing the risk exactly?

Shared Risk – Is this Industry Standard?

I could find no published survey quantifying how many laboratories default to simple acceptance, so I will not invent a percentage. What I can point to is stronger, because it comes from the providers themselves. One major calibration lab publishes simple acceptance as its default decision rule, gated at 4:1, with guard banded methods available when a customer selects one. Another major lab lists in and out of tolerance determination with no guard band as its commercial default, with guard banded options alongside as alternatives. In Transcat’s comparison of Z540.3 against ISO/IEC 17025 goes further. It describes simple acceptance as the industry standard practice, and notes that it permits false accept probability up to 50 percent.

So the risk transfer does not happen because anyone acted carelessly. It happens through a default. The customer receives a certificate that says Pass. The provider applied the rule it publishes. Very often nobody discussed it.

What does ISO 17025 say about it?

ISO/IEC 17025:2017, Clause 7.1.3 sets the expectation. Where a customer requests a statement of conformity, both the specification and the decision rule must be clearly defined. Unless the rule is inherent in the requested specification, the laboratory must communicate it to the customer and agree it with them. A published default that a customer actively selects at order placement can satisfy that duty. ILAC G8 explicitly contemplates laboratories offering service tiers with differing guard bands, including zero. The problem is not the tier. The problem is accepting the default without knowing what it costs you.

Specific Risk and Global Risk Are Not the Same Number

You will read that simple acceptance carries a 50 percent false accept risk. That statement is true in a narrow sense and badly misleading in a broad one. Understanding the difference puts you ahead of most people in this conversation.

ILAC G8 defines two quantities.

  • Specific risk is the probability that one particular accepted item is nonconforming. It rests on the measurement of a single item.
  • Global risk is the average probability that an accepted item is nonconforming across the whole population.

The 50 percent figure is specific risk, and it applies to a result sitting exactly on the tolerance limit. G8 puts it plainly. Under simple acceptance the probability of falling outside the tolerance limit may reach 50 percent when a result lands right on the limit, assuming a symmetric normal distribution. That is the coin flip. It is not the average across everything you calibrate.

Now the part worth carrying into an audit. Z540.3 Clause 5.3 (b) caps the probability of false accept at 2 percent. The clause itself does not say global or average, so read that carefully. The settled interpretation treats it as global risk, and ILAC G8 states it that way when it describes a 2 percent false acceptance criterion as a global risk criterion. G8 then spells out the consequence. An instrument passing a 2 percent global criterion may still carry a specific risk approaching 50 percent on an individual result.

Global risk protects the population. It does not protect the one instrument whose reading sat on the line.

What a Test Uncertainty Ratio Actually Buys You

Now connect the ratio to the risk. Under simple acceptance the global false accept probability depends on two inputs. One is the test uncertainty ratio itself. The other is how reliable the submitted population is when it arrives. That second input is the in tolerance probability (ITR), sometimes called end of period reliability (EOPR).

The table below gives global false accept probability under simple acceptance. I computed it from the JCGM 106 model at k = 2 for a two sided tolerance. It assumes a zero centered normal population, normal measurement error, and no measurement bias. ITP is defined as the Population In-Tolerance Probability, or you can think of it as the probability that your devices coming back from calibration will be in tolerance. It may even be called: End-of-Period Reliability (EOPR).

ITP is calculated using this equation

In-Tolerance Probability (ITP) is the probability that an item from the population is actually within its specified tolerance before the conformity measurement is made. For calibration programs, it is often estimated from end-of-period reliability data by dividing the number of as-found in-tolerance calibrations by the total number of calibrations. For example, if 95 of 100 comparable instruments are found in tolerance at their scheduled calibration, the estimated ITP is 95%.

For example, if 95 of 100 comparable instruments are found in tolerance at their scheduled calibration, the estimated ITP is 95%

Test uncertainty ratioITP 99 %ITP 95 %ITP 90 %ITP 85 %ITP 80 %Worst ITP
1 : 10.41 %1.80 %3.25 %4.45 %5.42 %7.37 %
2 : 10.33 %1.34 %2.26 %2.95 %3.46 %4.18 %
3 : 10.28 %1.05 %1.71 %2.18 %2.51 %2.91 %
4 : 10.23 %0.86 %1.37 %1.73 %1.97 %2.24 %
10 : 10.12 %0.41 %0.62 %0.76 %0.85 %0.94 %

Reading the Risk Table

Read across the 4:1 row and notice that ITP can still be more than 2%. The same ratio delivers anywhere from 0.23 percent to 2.24 percent false accept risk across this range of fleet reliability. That is nearly a tenfold spread, and it turns on a number that never appears on a calibration certificate. Widen the range and the spread widens with it, which is the point.

The last column deserves its label. It gives the reliability that maximizes risk within this normal population model, not an absolute ceiling. Read down it and a second lesson lands. Even a 1:1 ratio holds global risk near 7 percent under these assumptions, far below the 50 percent figure people quote. Change the population shape, or push most of the fleet just outside the limit, and global risk climbs well beyond that. So treat 50 percent as what G8 says it is, a specific risk at the limit, rather than an average you can cite anytime.

The 4:1 Test Uncertainty Ratio Rule Does Not Deliver What People Think

Look again at that last column. At 4:1 it reads 2.24 percent, which sits above the 2 percent Z540.3 names as its primary requirement.

Be precise about what that does and does not mean. Clause 5.3 (b) offers 4:1 as an alternative compliance route for when estimating the probability is impractical, so a laboratory on the 4:1 route is compliant. What the number shows is that the two routes are not equivalent in the protection they deliver. Choosing the ratio route can leave you above the risk the other route would have held you to.

That is not a rounding quibble. Solve for the test uncertainty ratio where worst case global risk finally drops to 2 percent. You get 4.5:1 at k = 2, or 4.6:1 at k = 1.96. NASA metrologists Harben and Reese reached the same conclusion in NCSLI Measure. They report that under worst case reliability, false accept risk stays below 2 percent only once the ratio reaches 4.6:1 or greater.

History of the 4:1 Rule

The history explains the gap. The 4:1 rule traces to Jerry Hayes at the Naval Ordnance Laboratory in the mid 1950s, building on papers by Eagle and by Grubbs and Coon. Hayes started from a 1 percent consumer risk target, which pointed to roughly 3:1. He then added a margin with Stan Crandon for uncertainty in the reliability of tolerances. That is how 3:1 became 4:1. Nobody derived it to satisfy a 2 percent requirement, because that requirement did not exist yet.

Read Z540.3 Clause 5.3 (b) in the right order and this makes sense. The clause leads with the risk requirement and offers 4:1 only where estimating that probability proves impractical. The ratio is a proxy. The risk analysis outranks it.

None of this makes 4:1 useless. At a well managed fleet with 95 percent in tolerance probability, a 4:1 ratio delivers 0.86 percent risk and everybody sleeps. The catch is that the laboratory performing your calibration usually has no idea which fleet you are.

Turning the Test Uncertainty Ratio Into a Decision Rule

ISO/IEC 17025:2017 gives you three obligations, and the ratio informs all three.

Clause 7.1.3 comes first, at the point of taking the work on. Define the specification and the decision rule, and agree the rule with the customer unless the specification already contains it. Clause 7.8.6.1 then requires three things whenever you provide a statement of conformity. Document the decision rule. Take account of the level of risk, including false accept, false reject, and the statistical assumptions behind it. Then apply the rule. A note follows: where the customer, a regulation, or a normative document prescribes the rule, further consideration of risk is not necessary. Clause 7.8.6.2 governs the report. It must identify which results the statement covers, which specifications are met or not met, and which decision rule you applied.

Guard Banding

Guard banding is how you act on a thin test uncertainty ratio. You pull the acceptance limit inside the tolerance limit by some multiple of the expanded uncertainty. ILAC G8 tabulates the common choices against specific risk.

Decision ruleGuard bandSpecific risk
Six sigmaw = 3 UUnder 1 ppm false accept
Three sigmaw = 1.5 UUnder 0.16 % false accept
ILAC G8:2009 rulew = 1 UUnder 2.5 % false accept
ISO 14253-1:2017w = 0.83 UUnder 5 % false accept
Simple acceptancew = 0Under 50 % false accept
Uncritical acceptancew = −UUnder 2.5 % false reject
Customer definedw = r UWhatever the multiplier r delivers

G8 states those figures for a single sided specification with normally distributed results, so treat them as indicative rather than transferable to the two sided case without rework.

A practical policy falls out of the two tables together. Where the test uncertainty ratio sits comfortably above 4:1 and fleet reliability is good, simple acceptance stays defensible. Document that you chose it anyway. Where the ratio runs thin, apply a guard band and state the multiplier. Where it falls below 1:1, a guard band still pulls risk down, but only by rejecting good equipment in bulk. At that point you need a better method rather than better paperwork.

Questions to ask your Calibration Provider

Ask your provider three questions.

  • Which decision rule did you apply?
  • What test uncertainty ratio did you achieve at my test points on my models?
  • What guard band options do you offer, and what do they cost?

The answer you get to these questions will say a lot about the calibration lab you are dealing with. If they have someone who can answer these questions, chances are they are thinking about shared risk and trying to help their customers get the best quality services.

Need help with Decision Rules? Schedule a free consultation with Precision ISO.
Schedule a Free Consult

Case Study: Calibrating a 1000 µL Pipette

Theory settles once you watch it work. This case study calibrates a BRAND Transferpette S, catalogue number 705880. It is a single channel adjustable volume mechanical air displacement pipette covering 100 to 1000 µL. The method is gravimetric, following ISO 8655-6:2022. You dispense water into a weighing vessel, read the mass, and convert.

The Z factor converts milligrams of water to microlitres. It bundles water density, air buoyancy, and the density of the reference weights. So it depends on water temperature, air temperature, and barometric pressure, and at 22 °C and 101.3 kPa it comes to about 1.0033 µL/mg.

ISO 8655-2:2022 sets three test volumes: nominal, roughly half of nominal, and either a tenth of nominal or the smallest settable volume. For this pipette that means 1000, 500, and 100 µL, with ten dispenses at each.

The tolerance. Every budget below uses BRAND’s own published systematic error limits, which run tighter than the ISO values.

Test volumeBRAND systematic error limitTolerance spanISO 8655-2 maximum permissible systematic error
1000 µL± 6 µL (± 0.6 %)12 µL± 8 µL
500 µL± 4 µL (± 0.8 %)8 µL± 8 µL
100 µL± 3 µL (± 3.0 %)6 µL± 8 µL

Note what ISO does there. Its maximum permissible error references the nominal volume and stays constant across all three test points. The same ± 8 µL applies whether you dispense 1000 µL or 100.

Types of Tolerance Limits

Two caveats before the budgets. ISO 8655-2 also sets a maximum permissible random error of 3 µL, and a unit can pass the systematic limit while failing that one. Everything below addresses the systematic error, because that is the tolerance the test uncertainty ratio speaks to. Second, the repeatability values are stated assumptions: standard deviations of 1.2, 0.9, and 0.35 µL across ten dispenses, near 60 percent of BRAND’s published CV limits.

The two references.

XPR56 micro analyticalLA204E analytical
Capacity52 g220 g
Readability0.001 mg0.1 mg
Typical repeatability0.0007 mg0.08 mg
Weighing pan40 by 40 mm80 mm diameter
USP minimum weight, typical1.4 mg160 mg
Assumed certificate U, k = 20.010 mg0.25 mg

The certificate values are stated assumptions consistent with each published specification, not quoted certificates. The instrument figures come from Mettler-Toledo documentation.

Both balances satisfy ISO 8655-6. Its Table 1 keys the requirement to the nominal volume of the apparatus, not to each test volume, so a 1000 µL pipette sits in the 200 µL to 10 mL row throughout: 0.1 mg resolution, 0.2 mg repeatability, and expanded uncertainty at or below 0.4 mg. The LA204E meets resolution exactly on the limit and clears the other two.

The Assumption That Decides This Whole Case Study

Before any arithmetic, one contributor has to be settled, because it turns out to own the budget.

Gravimetric pipette calibration carries a handling term covering operator technique, tip seal, pre wetting, hand warmth, and dispensing rhythm. Two current guides quantify it, and they disagree on the basis.

GuideMinimum handling contributionReferenced to
DKD-R 81, Edition 12/2011, Section 8.90.1 % for single channel variable volume pipettesNominal volume
EURAMET cg-19, version 4.1, Section 6.3.7.20.1 %Measured volume

EURAMET cites DKD as its source, then restates the figure on a different basis. At the 100 µL test point the two readings give 1.0 µL and 0.1 µL. A factor of ten, on the term that dominates everything.

Neither guide is wrong to cite, so the budgets below run both ways. The difference turns out to be more instructive than the balance comparison that prompted the exercise.

Case One and Case Two at Nominal Volume

Start at 1000 µL, where the dispensed water weighs about 996.7 mg. At nominal volume the two handling bases coincide, so there is a single answer.

XPR56 micro analytical balance

Uncertainty Contribution SourceRationaleStandard Uncertainty% Contribution
Pipette handling0.1 % minimum contribution1.0000 µL86.6 %
Pipette repeatabilitystd dev of 10 dispenses ÷ √100.3795 µL12.5 %
Z factor0.01 % of volume, temperature and pressure0.1000 µL0.9 %
Residual evaporation after correction± 0.05 mg, rectangular0.0290 µL0.1 %
Balance, from calibration certificate0.010 ÷ 20.0050 µL0.0 %
Balance resolution, both readings√2 × 0.001 ÷ √120.0004 µL0.0 %

uc=1.075 µLU=2.149 µLTUR=12.02×2.149=2.8:1u_c = 1.075 \text{ µL} \qquad U = 2.149 \text{ µL} \qquad TUR = \frac{12.0}{2 \times 2.149} = 2.8{:}1

LA204E analytical balance

Uncertainty Contribution SourceRationaleStandard Uncertainty% Contribution
Pipette handling0.1 % minimum contribution1.0000 µL85.3 %
Pipette repeatabilitystd dev of 10 dispenses ÷ √100.3795 µL12.3 %
Balance, from calibration certificate0.25 ÷ 20.1254 µL1.3 %
Z factor0.01 % of volume, temperature and pressure0.1000 µL0.9 %
Balance resolution, both readings√2 × 0.1 ÷ √120.0410 µL0.1 %
Residual evaporation after correction± 0.05 mg, rectangular0.0290 µL0.1 %

uc=1.083 µLU=2.165 µLTUR=12.02×2.165=2.8:1u_c = 1.083 \text{ µL} \qquad U = 2.165 \text{ µL} \qquad TUR = \frac{12.0}{2 \times 2.165} = 2.8{:}1

Read those two results side by side, because this is what the exercise was built to test. A balance whose certificate uncertainty is twenty five times larger produced 2.77:1 against 2.79:1. The risk moves from 3.11 percent to 3.13 percent. At nominal volume the choice between these two instruments changes nothing you could measure.

Stacked bar chart comparing the uncertainty budget for a 1000 µL pipette calibration against two reference balances. Pipette handling holds 86.6 percent of the budget with the XPR56 and 85.3 percent with the LA204E, pipette repeatability holds 12.5 and 12.3 percent, and the balance calibration certificate holds 0.002 and 1.34 percent, so the test uncertainty ratio moves only from 2.79 to 1 down to 2.77 to 1.
Figure 1. Where the uncertainty actually comes from at nominal volume. The 25 times larger balance certificate is the thin teal sliver near the right edge of the second bar.

Note how the balance certificate enters the budget: once, not twice. A tare and a gross reading seconds apart share their sensitivity error, reference weight uncertainty, and eccentricity, so those largely cancel in the difference. Only the resolution genuinely enters twice, which is why that line carries the √2.

Where the Balance Choice Does and Does Not Matter

Now run all three test volumes under both handling bases.

Handling basisTest volumeXPR56 TURRiskLA204E TURRisk
DKD, 0.1 % of nominal1000 µL2.8 : 13.11 %2.8 : 13.13 %
DKD, 0.1 % of nominal500 µL1.9 : 14.32 %1.9 : 14.35 %
DKD, 0.1 % of nominal100 µL1.5 : 15.36 %1.5 : 15.40 %
EURAMET, 0.1 % of measured1000 µL2.8 : 13.11 %2.8 : 13.13 %
EURAMET, 0.1 % of measured500 µL3.5 : 12.56 %3.4 : 12.62 %
EURAMET, 0.1 % of measured100 µL9.8 : 10.95 %7.4 : 11.24 %
Two line charts showing the test uncertainty ratio at 1000, 500 and 100 µL against a 4 to 1 reference line. Under the DKD-R 8-1 handling basis both balances trace the same falling line, from 2.8 to 1 down to 1.5 to 1. Under the EURAMET cg-19 basis the ratio rises instead, reaching 9.8 to 1 with the XPR56 and 7.4 to 1 with the LA204E at the 100 µL test point.
Figure 2. The same pipette, the same two balances, and one line of the budget read two different ways. The handling basis decides whether the 100 µL point clears a 4:1 decision rule or falls well short of it.

The handling basis matters far more than the balance

Switching from DKD’s nominal basis to EURAMET’s measured basis moves the 100 µL test uncertainty ratio from 1.5:1 to 9.8:1 and the risk from 5.36 percent to 0.95 percent. Switching from a 0.25 mg balance certificate to a 0.010 mg one moves the same point by a quarter of that, and only on the EURAMET basis. You are choosing between a factor of six and a factor of one and a third.

On the DKD basis the balance never matters at all

The handling term is a constant 1.0 µL at every test volume, and it holds 86, 92, and 99 percent of the budget as you step down. Under that reading, nothing about the reference instrument is visible anywhere in the working range.

On the EURAMET basis the balance finally appears at the low end

At 100 µL the handling term shrinks to 0.1 µL, the LA204E certificate becomes 38.7 percent of its budget, and the test uncertainty ratio separates, 9.8:1 against 7.4:1. That is the only place in this entire exercise where the two instruments give materially different answers.

Notice too that the DKD ratios fall as the volume drops while the EURAMET ratios rise. Neither pattern is about the method. Under DKD the uncertainty stays fixed while BRAND’s tolerance shrinks. Under EURAMET both shrink, and the tolerance shrinks more slowly.

On Minimum Weight, and What It Does Not Mean

A 100 µL dispense weighs 99.7 mg, which sits below the LA204E’s published USP minimum weight of 160 mg. That comparison looks alarming and is widely misused, so it is worth getting right.

USP General Chapter 41 governs balances used for materials that must be accurately weighed in compendial procedures. Chapter 1251 draws the boundary explicitly, noting that its own guidance applies more broadly than Chapter 41 does. Weighing dispensed water to characterize an instrument is not a compendial weighing of a material, and the documents that actually govern this activity agree: ISO 8655-6 sets its own balance requirements, and EDQM guidance on qualifying piston pipettes does the same without invoking Chapter 41 at all.

Three further points sharpen it. Minimum weight derives from a required weighing tolerance, and the 160 mg figure applies Chapter 41’s 0.10 percent criterion, which has nothing to do with the permissible error for a pipette at a tenth of nominal. In addition, Chapter 41 also requires the minimum weight to come from your own repeatability test on the installed balance, not a catalogue typical value. And its own floor rule, 820 times the scale interval, gives 82 mg here, which the 99.7 mg dispense clears.

So the honest statement is narrower than the alarming one. Nothing forbids the measurement. What the minimum weight figure does tell you is that this balance works near the bottom of its useful range at that test point. The test uncertainty ratio already showed you that by putting the certificate at 38.7 percent of the budget. Wrong standard, right instinct.

What the Tolerance Choice Does to the Ratio and the Risk

Swap BRAND’s limits for the ISO 8655-2:2022 maximum permissible error of ± 8 µL, which applies at every test volume because it references nominal.

Handling basisTest volumeXPR56 TURRiskLA204E TURRisk
DKD, 0.1 % of nominal1000 µL3.7 : 12.39 %3.7 : 12.41 %
DKD, 0.1 % of nominal500 µL3.8 : 12.32 %3.8 : 12.34 %
DKD, 0.1 % of nominal100 µL4.0 : 12.25 %3.9 : 12.27 %
EURAMET, 0.1 % of measured1000 µL3.7 : 12.39 %3.7 : 12.41 %
EURAMET, 0.1 % of measured500 µL6.9 : 11.34 %6.7 : 11.37 %
EURAMET, 0.1 % of measured100 µL26.3 : 10.36 %19.9 : 10.48 %

The DKD rows produce something elegant. Both the tolerance and the dominant uncertainty reference the nominal volume, so they scale together and the ratio sits almost flat at 3.7 to 4.0 across the range. That is what happens when your acceptance criterion and your largest uncertainty contributor share a basis.

Two lessons follow.

First, a test uncertainty ratio means nothing until you name the tolerance behind it. The same pipette, the same balance, and the same ten dispenses return 9.8:1 or 26.3:1 at 100 µL. Only the acceptance criterion changed. Anyone quoting you a ratio without stating the tolerance has told you almost nothing.

Second, holding your pipettes to the manufacturer’s specification rather than the ISO limit is a real choice with a real cost, and at nominal volume it is the difference between 2.8:1 and 3.7:1. It is defensible and often right in a GMP environment. Just make it deliberately, and document that you chose it.

The Contributor That Controls Your Test Uncertainty Ratio

Since the handling term owns the budget, look at what happens when you change it at nominal volume.

Handling contributionU at 1000 µLTURWorst case risk
0.10 % of volume, the stated minimum2.149 µL2.8 : 13.11 %
0.05 % of volume1.273 µL4.7 : 11.92 %
Excluded entirely0.787 µL7.6 : 11.22 %

That single assumption moves the test uncertainty ratio from 2.8:1 to 7.6:1. It also takes the risk from above the Z540.3 threshold to well below it.

One caution on reading published scopes. A 0.1 percent handling term at 1000 µL forces the expanded uncertainty above 2 µL on its own, so any accredited scope reporting materially less is either evaluating handling differently or not applying the minimum. Ask a prospective provider two questions. Do you include an operator handling contribution, and on what basis, nominal or measured? Those answers tell you more than any equipment list.

One more caution, aimed at the reader rather than the provider. This budget uses the standard deviation of ten dispenses divided by the square root of ten, because the measurand is the systematic error, which ISO 8655-6 derives from the mean. That is correct for the uncertainty of the calibration result. It is not the uncertainty of your next single pipetting step in routine work. Use the full standard deviation for that, and the expanded uncertainty at nominal rises from 2.15 µL to roughly 3.13 µL.

Making the Decision Rule Affordable

At 2.8:1 with simple acceptance you carry up to 3.11 percent false accept risk, which misses the 2 percent that Z540.3 names. A guard band fixes it, and here the fix is cheap.

Pull the acceptance limit in by 0.36 µL, from ± 6 µL to ± 5.64 µL, and worst case risk falls to 2 percent. Total false rejects then run near 4 percent at 95 percent in tolerance probability. Measurement error was already rejecting about 2.5 percent of conforming pipettes before you touched the limit, so the guard band itself costs roughly 1.4 percent more, about one pipette in seventy.

Compare the full one-U guard band from the ILAC G8 table. That moves the acceptance limit to ± 3.85 µL, drops worst case risk to 0.07 percent, and rejects nearly one pipette in five. Both are defensible. They buy different things, and the test uncertainty ratio is what lets you price the choice before you commit.

Putting the Test Uncertainty Ratio to Work

Work it in this order. Name your acceptance tolerance and record where it came from, because every ratio here moved when that number moved. Build the budget, combine in quadrature, expand at k = 2, then divide the tolerance span by twice that value. Look up the risk your test uncertainty ratio implies, and be honest about your fleet’s in tolerance probability rather than assuming the best case. Then choose a decision rule deliberately and document it, applying a guard band wherever the ratio runs thin. Finally, read the percent contribution column, because it names the one thing worth fixing.

The pipette case makes that last instruction uncomfortable and useful at the same time. A laboratory that reads its budget invests in technique, training, and a documented dispensing rhythm. A laboratory that does not read it buys a better balance and changes the answer in the third significant figure.

Interpreting Results

The broader point is simple. A calibration certificate reports a measurement, but the word Pass reports a decision, and decisions carry risk whether or not anyone quantifies it. The test uncertainty ratio is how you find out how much.

Want a starting point that already carries this structure? The Precision ISO measurement uncertainty budget template and reporting SOP lay out the contributor table, the combination step, and the decision rule section. Tailor them to your instruments and your test points, and this analysis becomes a controlled procedure your assessor can follow.


A note on accuracy: This article teaches; it does not give legal, regulatory, or accreditation advice, and it guarantees no audit or inspection outcome. Risk figures were computed from the JCGM 106 model under stated assumptions and reproduce published values from Dobbert (2008 NCSLI), Harben and Reese (NCSLI Measure), and Deaver (Fluke). They describe the model, not your equipment. Instrument specifications come from Mettler-Toledo and BRAND published documentation. References cover ISO/IEC 17025:2017, ILAC G8:09/2019, ILAC P-14:09/2020, ANSI/NCSL Z540.3-2006 (R2013), ISO 8655-2:2022, ISO 8655-6:2022, ISO/TR 20461:2023, DKD-R 8-1, EURAMET cg-19, and USP General Chapters 41 and 1251. Verify every clause number, specification, and requirement against your organization’s controlled copies, and build your budgets and reliability estimates from your own data.

Tags
Share this post:
Contact Info
Let our experts simplify your compliance journey and guide you from assessment to successful certification.

Ready to Get ISO Certified Without the Stress?

Let our experts simplify your compliance journey and guide you from assessment to successful certification.

Scroll to Top