Statistical Significance: Is the Comparables Set Big Enough?
Statistical significance in benchmarking: the set-size adequacy question — the minimum set, the TPO’s first ask, and the documentation that answers it.
Definition
Statistical significance in the benchmarking context is the set-size adequacy question: is the comparables set large enough for the range computed on it (the IQR, the mid-point) to be a reliable measure of the arm’s length position — or is the set too small, and the range an artifact of the few members it happens to contain? It is the TPO’s first ask on the pool (“how many comparables, and is that enough?”), and the documentation’s answer is the set size, the screens that produced it, and the adequacy read — not a statistical test, but the record that the size is the comparability result, not the compromise.
| The element | The content |
|---|---|
| The set size | The number of members in the final pool (after the quantitative screens, the qualitative screen, the outlier and the loss-making treatments) — the number the range is computed on |
| The practice threshold | The 10+ members as the working adequacy (the set below 10 is the thin set — the range’s reliability questioned, the TPO’s challenge invited); the 12–20 the comfortable band; the larger the better, within the comparability |
| The screens’ record | The candidate count at each screen (the search’s candidates → the quantitative screen’s survivors → the qualitative screen’s accepted → the final pool) — the size is the result of the recorded screens, not the target the screens were bent to |
| The thin-set treatment | Where the final pool is thin (below the working threshold): the search design revisited (the code set widened, the size band relaxed — the documented relaxation, the comparability cost stated), or the thin-set reliance stated (the range on the thin set, the reliability note, the supplement — the extra events, the trend) |
The working read (the benchmarking mistakes self-audit): the significance question is answered by the screen sequence’s record — the candidate population at each stage, the filter and its threshold, the survivors — so the final size is explained (the comparability produced the size) rather than defended (the size is what it is, help it). The thin set is not automatically the failed set — the documented thin set (the search’s genuine result on the comparability facts, the reliability noted, the supplement applied) is a different file from the bent set (the screens relaxed to hit the count, the relaxation unstated). The IQR’s construction is robust to the small n (the p25–p75 on the few values, the interpolation stated) — but the reliability of the range on the thin set is the significance question, and it is the documentation’s, not the statistic’s, to carry.
Example
The search returns 140 candidates on the NIC code set; the quantitative screens (the size band, the sector, the financial ratios) narrow to 48; the qualitative screen (the products, the mix, the assets) accepts 12; the outlier treatment excludes 1 (the product mix, the reason stated); the final pool is 11 — above the working 10+ threshold, the comfortable band’s edge. The significance read in the file: the screen sequence’s record (140 → 48 → 12 → 11, each screen’s filter and threshold), the size as the comparability result, the IQR on the 11 (the bounds, the median, the interpolation stated). The TPO’s “is 11 enough?” is answered by the sequence — the 11 is what the comparability produced, on the recorded screens, not the count the screens were bent to.
See also
- Benchmarking Mistakes (the self-audit)
- Interquartile Range (IQR)
- Quantitative Screening · Search Design
FAQ
Is there a minimum set size the law or the OECD prescribes? No fixed number is prescribed — the OECD’s position is the reliability of the range (the set’s adequacy to the comparability), not a count. The 10+ working threshold is the practice convention (the set below 10 is the thin set the TPO challenges) — the documentation carries the size, the screens that produced it, and the adequacy read, on the reliability standard rather than the count standard.
What is the thin-set documentation? The search design revisited (the code set, the size band, the relaxation — stated, with the comparability cost of each relaxation), the final size as the result (the screen sequence’s record, the genuine result on the comparability facts), the reliability note (the thin set’s range, the reliability caveat stated), and the supplement (the extra events, the trend read, the regional view where the supplement is the wider search). The thin set, documented, is a defensible file; the thin set, unstated, is the mistake list row.
Does the significance question change with the multi-year data? The set size is the same question (the members, not the years) — but the multi-year averaging (the averaging methods) changes the values the significance runs on (the averaged PLIs, the one-year noise smoothed) — and the per-period availability (the member with the missing year, the one-time year) is the data-quality read the significance note carries. The set size is the members; the years’ availability is the data quality; the documentation carries both.
Run the screens as a study, not a spreadsheet
Quartyl applies the method, PLI and screening steps above as a pipeline — and keeps a documented reason for every exclusion.
Related docs
Interquartile Range (IQR): The OECD Arm's Length Range
The interquartile range defined: the band between the first and third quartiles of the comparable pool — the OECD\'s preferred construction of the arm's length range.
Read doc12 Common Benchmarking Mistakes (Self-Audit Checklist)
Twelve mistakes that fail benchmarking studies in examination — the wrong code, the mixed PLI, the silent loss-maker, the range shop — each with its consequence and the fix.
Read doc