Worked example
The number is not the justification.
You get the arithmetic in full: where confidence and reliability come from, what attribute data costs against variables data, why 59 keeps appearing, and what five units actually support. A sample size with no reasoning behind it is the finding, whatever the number happens to be. Every figure follows from a rule you can see, and the one input that has to come from a published table rather than from arithmetic is flagged where it appears.
In short
How a sample size is justified
A sample size is justified by stating what claim it supports, not by naming a number. Write down the risk being protected against, whether the data is attribute or variables, the confidence and reliability being claimed, and the arithmetic connecting them. For attribute data with zero failures allowed, 95 percent confidence in 95 percent reliability takes 59 units. Five units, with zero failures, supports about 55 percent.
The problem, stated once
A number can sit in a protocol with no record of where it came from.
Take a protocol that says n = 5. It may have come from a laboratory quote, from a previous submission, from a template, or from the last person who held the job. The protocol records the five. If it does not also record what claim the five is supposed to support, nobody reading the file later can tell whether five was generous, adequate or nowhere near enough.
That missing sentence is the finding. Not the five. A small number with a written basis and a large number with none are two different records, and the one without the basis is the weaker of the two however much it cost to produce. Write it, and the number, whatever it is, has a basis someone can check.
The tool
Put your numbers in. Get the sentence that justifies them.
Ask what a claim costs in units, or what the number already in your protocol supports, and the answer comes back as a plain sentence you can lift straight into the rationale. You pick the claim from a list rather than typing one, because which claim your file is entitled to is a risk decision this page will not make for you. Every answer is a sentence first and a number second, so it is obvious at a glance whether the number is enough.
What does a claim cost in units?
Pick the claim your file needs to support. Which one you are entitled to is a risk decision, so the reasoning for it belongs in your rationale rather than in any calculator.
29 units, all passing, are what this claim costs with zero failures allowed: 95 percent confidence that at least 90 percent conform.
- With your numbers
- 28.433
- The result
- 29
- The check
- 0.047101 (95.29) / 0.052335 (94.77)
Zero failures allowed is the whole calculation rather than a footnote. What happens if a unit does fail is written before any unit is tested, along with what that allowance did to the sample size.
What does the number in my protocol support?
The number already written in the protocol, as a whole number of units, with zero failures allowed.
Confidence is held at 95 percent here, because the ladder above is published at 95 percent and at no other level.
5 units, all passing, support the statement that at least 54.9 percent of units conform, with 95 percent confidence.
- The result
- 0.5493
- The check
- 0.0500
Zero failures allowed is the whole calculation rather than a footnote. What happens if a unit does fail is written before any unit is tested, along with what that allowance did to the sample size.
What if one unit is allowed to fail?
93 units, with one unit allowed to fail, hold the same claim at 95.002 percent confidence.
One unit fewer, 92, reaches 94.79 percent and misses.
- The check
- 0.049976 / 0.052136
Zero failures allowed is the whole calculation rather than a footnote. What happens if a unit does fail is written before any unit is tested, along with what that allowance did to the sample size.
What does variables data change?
The one-sided lower limit the measurements have to clear. This is the direction the page works, and the only one it demonstrates.
The arithmetic mean of the measured values.
The spread of the measured values. Variables data buys the reduction in units by using this number, so a process with more spread gives the reduction back.
How many units the mean and the standard deviation were computed from.
From your own table, for your n, confidence and reliability. This site has not read a tolerance factor table at source, so it publishes no value, no default and no suggestion here.
Free text. It is printed beside every figure below and in the rationale.
Required. A factor with no source behind it is the same defect as a sample size with no source behind it, and this line goes into the rationale beside the factor.
The assumptions in that number
n belongs to a stated claim, about a stated attribute, at a stated condition, rather than to a program. Stability and performance are separate branches asking different questions of the same package, so those arms carry their own sampling reasoning.
What still has to be written
The number is one line of six. Everything the arithmetic above can fill in is filled in for you here; the rest is the part only your file can answer. Work down it in order and you have a rationale.
- 1The claim
- Yours to write
- 2The risk
- Yours to write
- 3The data type and the method
- Attribute data
- 4The confidence and the reliability
- 5 units, all passing, support the statement that at least 54.9 percent of units conform, with 95 percent confidence.
- 5The arithmetic
- 0.5493 (0.0500)
- 6The failure rule
- Yours to write
Why the search does not terminate
There is no table to look the number up in.
Two different questions are easy to run together, and running them together is reasonable, because both are called sampling.
Lot acceptance sampling asks whether to accept a batch that has already been made, at a stated quality level. Those schemes are indexed by lot size, because the lot is the population being sampled from and the decision applies to that lot alone.
Validation sampling asks something else: how confident you are that the packaging system your process makes will hold. The population there is what the process produces under the conditions you sampled, not a batch counted out in advance, so batch size does not appear in the arithmetic at all. Searching a lot acceptance table for a validation sample size is a search that cannot terminate, because the number is not in there.
Both halves of that are our reading and are stated as ours: that the two kinds of sampling answer different questions, and that there is therefore no validation sample size table to look a number up in. Either way the job is the same. You derive the number from the claim you are making, and you write the claim down first. The qualifier about conditions is not decoration, and it carries real weight when the units all come from one run.
The choice that sets the price
Attribute data and variables data are two different purchases.
Before anyone can compute a sample size, the test method has to be settled, because the method decides what each unit tells you. This is worth knowing before the protocol is signed rather than after the quote arrives. Variables data is also written variable data and continuous data; the three phrases mean the same thing here.
- Attribute data: each unit returns one bit
Dye penetration on the seal, bubble emission under internal pressurization, a visual inspection against a defect definition. The unit passes or it does not, and how close it came to failing stays invisible, so nothing about the margin is recoverable. With the claim held fixed, the only lever left is the number of units, which is why attribute plans are expensive in samples. The other lever is to claim less, and claiming less is a decision that belongs in the risk file rather than in the sample size line.
- Variables data: each unit returns a measurement
Seal strength in newtons per 15 mm from a peel test, burst pressure in kilopascals. A measurement carries distance from the limit as well as pass or fail, so a smaller set of units can support the same statement. The price is an assumption about how the measurements are distributed, which then has to be stated, checked and written into the rationale rather than assumed silently.
Neither is the correct answer in general. The correct answer is the one your risk assessment justifies, and the file has to say which you chose and why. Note also that n is not one number for the whole program: it belongs to a stated claim, about a stated attribute, at a stated condition. Stability and performance are separate branches asking different questions of the same package, so those arms carry their own sampling reasoning rather than sharing one figure.
FIG 01
The arithmetic behind 59, in full.
An attribute plan with zero failures allowed is the simplest case there is: one formula, one assumption, one number at the end. It is also where 59 comes from, so it is worth doing slowly.
Start with the claim, in words: with 95 percent confidence, at least 95 percent of units conform. Those are two different ideas and the order matters. Reliability, the second 95, is the proportion of the population you are claiming conforms. Confidence, the first 95, is how sure you are of that statement given that you only tested a sample. A file that states one and not the other has stated half a claim.
- If reliability is R, the chance that one unit passes when the true conformance rate is exactly R is R.
- The chance that n independent units all pass is R to the power n.
- Choose n so that R^n is at or below 5 percent. Then, if the true rate really were as low as R, seeing every unit pass would be an outcome with no more than a 5 percent chance of happening.
- Every unit did pass. So either something with a 5 percent chance happened, or the true rate is higher than R. Taking the second is what the 95 percent confidence means.
- So solve R^n = 1 - C for n: n = ln(1 - C) / ln(R).
- Put the numbers in: ln(0.05) = -2.995732, ln(0.95) = -0.0512933, and the quotient is 58.404.
- Sample sizes are whole units, and rounding down would fail the claim, so n = 59.
Check it the other way and the rounding stops being a matter of taste. 0.95^59 = 0.048495, which is below 0.05, so 59 units with zero failures reach 95.15 percent confidence. 0.95^58 = 0.051047, which is above 0.05, so 58 units reach 94.90 percent and miss. Fifty-nine is the smallest whole number that supports the sentence.
FIG 02
What each claim costs, with zero failures allowed.
Same formula, different claims. Every row is n = ln(1 - C) / ln(R) rounded up, and every row can be checked in a spreadsheet in about fifteen seconds.
| The claim, stated in words | Units required, zero failures |
|---|---|
| The claim, stated in words95 percent confidence that at least 90 percent conform | Units required, zero failures29 |
| The claim, stated in words90 percent confidence that at least 95 percent conform | Units required, zero failures45 |
| The claim, stated in words95 percent confidence that at least 95 percent conform | Units required, zero failures59 |
| The claim, stated in words95 percent confidence that at least 99 percent conform | Units required, zero failures299 |
| The claim, stated in words99 percent confidence that at least 99 percent conform | Units required, zero failures459 |
Two things fall out of the table. The first is that the cost is driven by the claim, not by the device. Holding confidence at 95 percent, moving the reliability claim from 90 percent to 95 doubles the units, 29 to 59. Moving it from 90 percent to 99 multiplies them by ten, 29 to 299. The second is that nothing in the table mentions batch size, and that is not an omission.
Which row you are entitled to sit on is a risk question, and it is the part a website cannot answer for you. The reasoning that has to appear in the file is why this level of reliability is appropriate for this failure mode on this device, with the link into your ISO 14971 file made explicit rather than implied.
FIG 03
What five units support, worked the same way.
The formula runs backwards, and that is the move worth having. Take the n already written in your protocol and compute the claim it actually supports, instead of choosing a claim and computing n.
- Rearrange R^n = 1 - C to get R = (1 - C)^(1/n).
- At 95 percent confidence with five units and zero failures: R = 0.05^(1/5) = 0.5493.
- Check it: 0.5493^5 = 0.0500.
Five units, all passing, support the statement that at least 54.9 percent of units conform, with 95 percent confidence. That is what five buys as an attribute plan. Written that way, in a sentence, it is immediately clear whether it is enough for the failure mode in question, which is exactly the clarity the file is missing when the number sits there alone.
| Units tested, zero failures | Reliability supported at 95 percent confidence |
|---|---|
| Units tested, zero failures3 | Reliability supported at 95 percent confidence36.8 percent |
| Units tested, zero failures5 | Reliability supported at 95 percent confidence54.9 percent |
| Units tested, zero failures10 | Reliability supported at 95 percent confidence74.1 percent |
| Units tested, zero failures20 | Reliability supported at 95 percent confidence86.1 percent |
| Units tested, zero failures29 | Reliability supported at 95 percent confidence90.2 percent |
| Units tested, zero failures30 | Reliability supported at 95 percent confidence90.5 percent |
| Units tested, zero failures59 | Reliability supported at 95 percent confidence95.0 percent |
None of this makes five wrong. Five may be entirely appropriate for a low risk attribute where the file says so, and the arithmetic above is the way to say so. What five cannot do is support a 95/95 statement, and a protocol that runs five units under a heading claiming 95/95 contains a contradiction that is visible to anyone who does the multiplication.
FIG 04
The same question with variables data, and what it costs instead.
Measurements let you say more with fewer units, because each unit reports its distance from the limit rather than just which side of it fell. The structure of the calculation changes: instead of counting passes, you build a one-sided tolerance bound from the mean and the standard deviation and compare that bound with the specification limit.
- Specification: seal strength not less than 1.0 N per 15 mm.
- Thirty units measured. Mean 2.6 N per 15 mm, standard deviation 0.45.
- Suppose the one-sided tolerance factor table your rationale cites gives k = 2.220 for n = 30 at 95 percent confidence and 95 percent reliability. This is the one number on this page that comes from a table rather than from a formula, and it is a supposition here rather than a value we are publishing. Which table and which edition is part of your rationale, and a factor with no source behind it is the same defect as a sample size with no source behind it.
- Lower bound = mean minus k times the standard deviation = 2.6 - (2.220 x 0.45) = 2.6 - 0.999 = 1.601 N per 15 mm.
- 1.601 is above the 1.0 limit, so the claim holds on thirty units rather than the fifty-nine an attribute plan would have taken.
Now change one input and nothing else. With the same mean of 2.6 and the same thirty units, a standard deviation of 0.72 gives 2.6 - (2.220 x 0.72) = 2.6 - 1.5984 = 1.0016, which clears the limit by 0.0016. A standard deviation of 0.80 gives 2.6 - (2.220 x 0.80) = 2.6 - 1.776 = 0.824, which is below the limit and fails, even though the limit still sits a full two standard deviations under the mean and every one of the thirty measured values could have been above 1.0.
That is the trade, stated plainly. Variables data buys the reduction in units by using the spread of your process, so a process with more spread gives the reduction back. It also buys it with an assumption: the tolerance factor assumes the measurements follow a known distribution, usually the normal one, and that assumption is part of the rationale and has to be checked rather than asserted. If you would rather assume nothing about the distribution and take the lowest measured value as the bound, the arithmetic returns to the attribute case and 95/95 costs 59 units again.
The rule nobody writes first
Zero failures allowed is the whole calculation, not a footnote.
Every number above assumes no unit fails. That assumption is doing all the work. Allow one failure and the single power term becomes a binomial sum, and the rule changes to this: find the smallest n for which the chance of seeing zero or one failure, when the true conformance rate is 95 percent, is at or below 5 percent. That n is 93, against 59 for the zero-failure plan, and 92 misses at 94.79 percent. The margin does not come free just because the failure was unexpected.
Which is why the protocol has to say, before any unit is tested, what happens if one fails. Investigation and root cause, a restated claim supported by fresh evidence, or a design change: any of those can be the right answer, and all of them have to be chosen in advance. A plan that is silent on the point leaves the acceptance criteria to be written after the data has arrived. Our view, stated as ours: that is the most serious integrity failure a validation file can carry, because a study whose criteria are written after its results cannot fail, and the document trail shows which of the two was written first.
What the sample spans
A sample drawn from one run answers a question about one run.
The arithmetic is correct whatever the units have in common. What it cannot do is widen the claim beyond the conditions the units actually came from. Fifty-nine units built in a single run, on one machine, from one material lot, with one sealer profile, support a statement about the process as it was configured for that run.
Whether that is the statement you need depends on what varies in your process. Material lot, sealer setup and profile, tooling, operator, line and site are all candidates, and your risk file either identifies them as sources of variation or does not. Where a source of variation is real and no unit in the sample spans it, the arithmetic remains correct and the claim quietly narrows to something smaller than the one being made in the report.
In place of a number of runs, you get the reasoning that produces one, because a figure you saw elsewhere describes what somebody else did rather than what you are doing. What you can write down, and what a reviewer can follow, is the list of variation sources, which ones your sampling spans, and why the ones it does not span were reasoned to be immaterial.
The deliverable
Six sentences, and the rationale is written.
A sampling rationale is not a long document. It is six sentences. If your file already contains them, the work is done and you do not need anyone.
- 1The claim
What is being demonstrated, about which attribute, at which condition, and at what point in the packaging system's life. Written as a sentence, not as a heading.
- 2The risk
Which failure this protects against, what it would do to the patient or the user, and where that failure sits in your ISO 14971 file. This is the sentence that justifies the confidence and reliability levels you are about to pick.
- 3The data type and the method
Attribute or variables, which test method produces it, and why that method answers the claim in sentence one rather than an adjacent question.
- 4The confidence and the reliability
Both numbers, stated explicitly, with the reason those levels were chosen and not higher or lower ones. The reason is the risk from sentence two, made specific.
- 5The arithmetic
The formula, the numbers put into it, and the result, in a form the reader can recompute without asking you anything. Where a table factor is used, name the table and the edition it came from.
- 6The failure rule
What happens if a unit fails, written before any unit is tested, along with how many failures the plan allows and what that allowance did to the sample size.
Six sentences. Their absence is the finding, and their presence is the thing that makes the number, whatever it is, defensible in writing.
The answer that is not one
Common practice is a description, not a justification.
"It is what everyone uses" contains no statement about your device, your risk file or your claim. It is also not checkable: there is no source to read, no arithmetic to reproduce, and no way for the person reviewing your file to agree or disagree with it on the evidence in front of them. It moves the question somewhere it cannot be answered.
A poster on the Elsmar Cove quality forum, asking for the statistical basis of a sample size a laboratory had recommended, reported that the whole answer was: "They only say it is common use." That is a forum post rather than guidance, and it is quoted here as one practitioner's account of one exchange. The shape of it is what to notice: the number may well have been the right number, and the file still could not show why.
None of this says laboratories are careless. A laboratory is engaged to run a method and report data, and it does not hold your risk assessment. The rationale is not a laboratory deliverable, which is precisely why it goes missing.
A number that circulates this way can still encode a claim, which is often how it came to circulate. Twenty-nine units with zero failures is the 95 percent confidence, 90 percent reliability plan, whichever route it arrived by. Fifty-nine is 95/95. The problem is not that such a number is wrong. The problem is that the file does not say what it means, so the reader has to reconstruct the reasoning, and the reconstruction was your job.
There is a related trap in the other direction. A rationale copied from a stranger's website, this one included, is the same finding as no rationale at all, because it is a claim about somebody else's device and somebody else's risk file. The arithmetic here transfers. The reasoning does not.
Failure modes
What makes a written rationale fall over anyway.
- Acceptance criteria that appear in the record after the results. The document trail shows the order, and no amount of correct arithmetic recovers from it.
- A stated number with no stated claim, which is the original problem in a longer document.
- A claim and a test that do not match: a statement about whole package integrity supported only by seal strength data, or the reverse.
- Arithmetic that cannot be reproduced: a tolerance factor with no table named, a confidence level with no reliability level beside it, a number that does not follow from the inputs given.
- A sample that spans none of the variation the risk file itself identifies, so the claim in the report is wider than the claim the units support.
- A distributional assumption used and never mentioned, which turns a variables plan into an unstated bet.
- A rationale written for the protocol and never updated when the method, the claim or the packaging system changed.
Limits of this page
Where this page stops, and what your own file has to carry.
No number to copy. The arithmetic is general and the choice is not: the confidence and reliability you are entitled to claim come from your risk assessment, and nobody can supply that from outside your file. A page that told you which row of the table to use would be doing the thing it spends its length arguing against.
No tolerance factor table, and no statement about what any reviewer or regulator will do with a rationale. What a written basis gives you is a file someone can follow and disagree with on the evidence. That is worth having and it is not a guarantee of anything.
This is also general information rather than a review of anything you hold. We manage accredited testing and we do not operate a laboratory, so no part of this is an offer to generate data. The manufacturer writes, approves and owns the rationale, and our name is never in your approval block.
Follow-ons
Questions this arithmetic raises.
- How many samples does ISO 11607 require?
There is no sample size table to look a number up in, which is our reading and is stated as ours. The number follows from the claim you are making and the risk that claim addresses, and what has to exist in the file is the reasoning that connects them.
- Why does 59 keep coming up?
Because it is what a 95/95 attribute plan with zero failures costs. n = ln(0.05)/ln(0.95) = 58.404, rounded up to 59, and 0.95^59 = 0.0485. It is not a rule and it is not a threshold anyone set. It is the output of one particular claim, and a number can travel a long way after the record of choosing that claim has gone.
- Is there a sample size based on batch size?
Not for this question. Batch size drives lot acceptance sampling, which decides whether to accept a batch that already exists. Validation sampling makes a statement about what the process produces under the conditions you sampled, so batch size does not enter the calculation.
- Can five samples ever be defensible?
Yes. With variables data, a stated distributional assumption and a modest claim, five measurements can carry real information. With attribute data, five supports about 55 percent reliability at 95 percent confidence, which is a defensible claim only if that is genuinely all the failure mode requires and the file says so. What is never defensible is five with no statement of what it supports.
- Do we need the same sample size for every test?
No. The number belongs to a claim, not to a program. Different attributes, different conditions and different arms of the study each carry their own reasoning, and stability and performance are separate branches rather than two readings of one number.
- Whose job is the rationale, ours or the laboratory's?
The manufacturer owns it. A laboratory runs the method and reports the data; the statement about what the data demonstrates for your device goes in your file under your approval. A laboratory can be entirely competent and still have no basis on which to write your rationale, because it does not hold your risk assessment.
- What if one unit fails?
Then the claim as written is not demonstrated, because the arithmetic assumed zero failures. What happens next has to have been decided before the test ran: investigation and root cause, a restated claim with fresh evidence, or a change to the packaging system. Deciding afterwards is the defect, not the failure itself.
- Does a larger sample size make the file safer?
Not on its own. A large number with no stated basis is still a number with no stated basis, and it costs more. The record is improved by the sentence, not by the units.
Next
Bring the protocol and the number that is in it.
Thirty minutes, no slides. Tell us the claim the number is meant to support and the method that produces the data, and we will talk through what the number in your protocol currently supports and what it would take to state it in writing. If the reasoning already holds, we will say so and there is nothing to buy.
Or email directly. A personal reply the same business day, not a queue.
Book a 30-minute call