Reading the Statistical Results: Range, IQR and Percentiles
How to read the statistical results: accepted comparables distribution, median, the applied arm’s length band, percentiles, CV, and tested party position inside the range.
The last pipeline step, Final Analysis, is where a run becomes a result: the statistical record of the benchmark. This page is how to read that record — what each number is, which band is the arm’s-length range, and how the tested party sits inside it. Every value is computed on the accepted comparable set only; rejected and flagged companies never enter the pool that produces the range, and the range is downstream of every disposition in the comparables grid.
Where the numbers come from
When Final Analysis completes, the engine computes descriptive statistics on the accepted set:
- Shape of the set — count, mean, median, standard deviation, minimum, maximum.
- Quartiles — first quartile (Q1), third quartile (Q3) and the interquartile range (IQR = Q3 − Q1).
- Interpolated percentiles — p10, p25, p35, p50, p65, p75, p90 of the accepted PLI distribution, linear-interpolated.
- Coefficient of variation — the CV as a percentage: the spread of the set relative to its mean.
- The applied arm’s-length band — the lower and upper bounds the engine resolved for the study’s jurisdiction, plus the statutory basis text and whether the prescribed band was in fact applied.
- Sensitivity ranges — two fixed rows, the 25th–75th and the 35th–65th band, each with its own tested-party-within flag. They are reference bands, not the conclusion.
- Risk and reliability metrics — the benchmark risk score and the benchmark reliability score, explained in Risk and Reliability Scores, and the acceptance rate, which is reported alongside them but feeds neither score.
Three things are worth knowing before you read a number:
- A thin pool produces no statistics at all: the pipeline stops reporting
below two accepted companies, and the statistics function needs at least
three accepted companies carrying a PLI value. Below that,
statisticsandconclusioncome back empty and the deterministic insights flag the set against the OECD’s 3–5 reliable comparables guidance (backend/app/services/statistics/statistics_bridge.py:46-52). - Multi-year studies apply the Golden Rule consistently to the tested party and the pool: raw financials are summed first, then divided. Per-year margins are never averaged.
Which band is the arm’s length range
The band Quartyl judges the tested party against is resolved at Final
Analysis from the study’s jurisdiction range rule, not from a global
default: the engine looks the jurisdiction up, takes that rule’s prescribed
percentile pair, and reports the statutory basis with the result
(backend/app/services/statistics/statistics.py:100-118). Explicit rules
exist for six jurisdictions; everything else falls to the default
OECD-style interquartile rule.
| Jurisdiction | Applied band | Basis | Adjustment target |
|---|---|---|---|
| India | 35th–65th percentile where the accepted set has six or more entries; below six, the arithmetic mean as a point estimate | Income-tax Rules, 1962 — Rule 10CA | Median |
| UAE | 25th–75th percentile, no sample-size gate | Ministerial Decision No. 97 of 2023 | Nearest bound |
| USA | 25th–75th percentile | Treasury Regulations §1.482-1(e)(2)(iii) | Nearest bound |
| UK | 25th–75th percentile | TIOPA 2010, Part 4 | Nearest bound |
| Singapore | 25th–75th percentile | Income Tax (Transfer Pricing) Rules 2018 | Nearest bound |
| Germany | 25th–75th percentile | Section 1 AStG; Verwaltungsgrundsätze 2020 | Nearest bound |
| Every other jurisdiction | 25th–75th percentile (the default rule) | OECD Transfer Pricing Guidelines, Chapter III | Nearest bound |
The interquartile range is therefore the default construction — it uses the central half of the distribution and keeps single outsized companies from stretching the range. The interquartile range glossary entry has the quartile math; IQR vs full range covers when the full min–max range is defensible and how to document that choice.
Alongside the applied band, the ranges table always shows two fixed
sensitivity rows — 25th–75th and 35th–65th — each with its own within flag,
so a reviewer can see how the tested party would sit under the other
construction without re-running anything. Where a jurisdiction prescribes a
band other than the study’s requested percentile pair, the prescribed one
wins and the result records that in its basis text and its
prescribed_band_applied flag.
Where the tested party stands
The engine compares the tested party’s PLI against the applied band and records one of three conclusions:
| Conclusion | Condition | What it means |
|---|---|---|
| Within Arm’s Length | lower bound ≤ tested party ≤ upper bound | No transfer-pricing adjustment is required |
| Below Arm’s Length | tested party < lower bound | The tested party earns less than the comparable central range — an upward adjustment may be required |
| Above Arm’s Length | tested party > upper bound | The tested party earns more than the comparable central range — a downward adjustment may be considered |
In the mean-point-estimate framework the band has already collapsed onto the
mean, so the test is against that single value with a small tolerance — a
point estimate can never be met exactly in floating point
(backend/app/services/statistics/statistics.py:316-333).
Inside the band, the analysis is done: the controlled transaction is documented as at arm’s length. Outside it, the study carries adjustment exposure — the target of the adjustment is what the jurisdiction’s rule prescribes (the median under Rule 10CA, the nearest bound by default), and the adjustment computation (formula-driven, not narrative) is produced for the report. The first question after an outside conclusion is never “is the study wrong” but “which input does this incriminate” — the tested party profile, a screen that over- or under-selected, the PLI denominator, or the data years.
The results head carries three items: the Conclusion with a check icon
when the tested party is inside and a warning icon when it is not, the
Tested Party Margin, and the Arm’s Length Range rendered as
lower – upper (for example, 2.50% – 7.50%).
The range stats card
The Charts & Analysis section opens with the Arm’s Length Range card — the distribution of the accepted set drawn as a single horizontal track:
- The grey track spans the observed range (minimum to maximum of the accepted companies, with a little padding at each end).
- The green Arm’s Length Zone band spans the applied lower and upper bounds — the 25th to 75th percentile by default, Rule 10CA’s 35th–65th for an Indian set of six or more.
- A dot marks the median with its value.
- A dashed vertical marker pins the tested party on the same scale, so inside/below/above is visible at a glance.
- Below the track, five chips read the scale off numerically under the fixed labels Min, P25, Median, P75, Max. Min, Median and Max print the observed values; the two band chips print the applied bounds, so on a Rule 10CA study of six or more comparables they show the 35th and 65th percentiles under the P25 / P75 labels.
This card is the fastest read in the study: the shape of the pool, the defensible zone, and where the tested party stands, in one glance.
The statistics tables
The Statistics section gives the same data in table form — what you cite in a review, and what the reports take their numbers from:
- Summary Statistics — Count, Mean, Median, Std Dev, Min, Max, IQR and CV in one block.
- Percentile Distribution — the interpolated percentile table, p10 through p90.
- Arm’s Length Ranges — one row per fixed sensitivity band (25th–75th and 35th–65th) with lower bound, upper bound and a Yes/No badge for whether the tested party’s margin falls inside. The band that decides the conclusion is the applied one, shown in the head and on the range card; these two rows are the sensitivity check.
Reading a result the way a reviewer does
- Tight set, tested party inside the band. The textbook defensible result. Note the CV and the count — they are what the range’s reliability indicators score (see risk and reliability).
- Tight set, tested party near the band edge. Inside, but a close call: one or two comparables of difference would move the bounds. Review the grid to confirm the pool is genuinely representative.
- Wide set (high CV), tested party inside. The conclusion holds, but the range is doing less work. The re-benchmark question — narrower industry scope, size filters, geography — is on the table.
- Tested party outside the range. Work the grid before the parameters: the dispositions are the record of why the pool is what it is. Outside with a thin pool is a re-benchmark signal; outside with a rich, tight pool is a finding about the tested party.
FAQ
Why don’t rejected companies appear in the range? Only the accepted set produces the statistics. Rejections are the record of why the pool is the pool — they live in the comparables grid and the Accept-Reject matrix, not in the distribution.
Which range drives the conclusion — 25–75 or 35–65? Neither, as a rule of thumb: the conclusion follows the band the study’s jurisdiction prescribes. India applies Rule 10CA’s 35th–65th percentile range once the accepted set reaches six entries (and an arithmetic-mean point estimate below that); the other five jurisdictions with curated rules and every other country apply the 25th–75th interquartile default. The two rows in the ranges table are fixed sensitivity bands, so the sensitivity of the conclusion to the construction is visible without a re-run.
The tested party is outside the range — is the study wrong? No, it is a finding. The review question is which input is responsible: the profile, a screen, the PLI, or the years. The grid and the parameter record are where the answer lives.
See it working in your workspace
Sign in to run the steps above on a real study — or book a demo and we will walk the workflow end to end.
Related docs
Risk and Reliability Scores: What the Numbers Mean
What the benchmark risk and reliability scores actually are: fixed additive point tables over the accepted set’s dispersion, size, loss share, outliers and range tightness.
Read docAccess and Security Logs
The firm-level security event log: logins, failed logins, MFA events, role changes, invitations, ledger exports and API key activity, with IP and user-agent context.
Read doc