Skip to main content
Quartyl
TP Software & AIprofessional

AI in Transfer Pricing: Use Cases, Controls and Governance

Where AI genuinely helps in TP work — screening, enrichment, drafting, risk — and the controls that keep it defensible: evidence capture, human decisions, override governance and the audit trail.

Quartyl Team

AI in transfer pricing has a narrow, high-value centre — reading a lot of structured and semi-structured company data quickly and consistently — and a wide, dangerous periphery, where a confident-sounding output replaces a judgement that was the point of the exercise. The difference between the two is not the model; it is the control architecture around it: what the AI reads, what it must show, who decides, and where the decision is recorded. This guide maps the use cases that have earned their place, the controls each one requires, and the governance layer that makes the whole thing defensible in front of a TPO, an IRS examiner, or the firm’s own quality review.

The use cases, ranked by maturity

Use case What the AI does Maturity The control that makes it defensible
Quantitative pre-screening Applies the size, sector and financial-ratio filters to the database, flags the borderline cases High — deterministic, fully auditable The filter sequence and thresholds are recorded; the AI reports the pass/fail, it does not invent the threshold
Qualitative screening assistance Reads each comparable’s business description (filings, website, database text) and drafts a per-company comparability assessment with the source passages cited High — with evidence Every assessment carries the source excerpt; a human accepts, rejects or overrides with a reason
Data enrichment GLEIF/LEI lookups, corporate linkage, business-description capture from websites, ratio computation High — lookup + extraction The enrichment is data with a source, not a conclusion; the source URL/record is stored
Drafting support Drafts the functional analysis narrative, the benchmarking rationale, the documentation sections from the study data Medium Drafts only, against the study’s own data; every factual claim maps to a study field or a cited source; a professional rewrites, not approves
Risk analytics Scores audit probability, litigation exposure, defensibility from the study’s own signals (range position, comparables quality, override density) Medium — transparent scoring The factors and weights are documented and inspectable; the score is a triage input, not a substitute for the professional’s view
Search string / classification Suggests NIC/NACE codes and search strings from the tested party description Medium The suggestion is a starting point; the search design judgement (scope, over/under-inclusiveness) stays human
Autonomous comparables selection Picks the final comparables set end to end, without a human decision per company Not defensible No control architecture cures a missing human decision — the accept/reject judgement is the professional’s, and it must exist as a decision, not an absence of one

The pattern: AI is defensible where it produces data or drafts with a source, and a human produces the decision. The moment the AI’s output is the decision — the comparables set, the method choice, the range — without a recorded human act on top of it, the file has a gap the audit will find.

The control architecture (the four layers)

Layer 1 — Input control: what the AI may read

  • The AI reads the study’s own data (the tested party file, the parameters, the comparable records) and the declared sources (filings, the GLEIF record, the website extract). It does not read “the internet” and cite whatever it finds.
  • Every input is versioned: if the company’s filing changes between the screening run and the review, the run’s inputs are the recorded ones, not the current ones.
  • Tenant isolation: the firm’s data — the tested party financials, the pricing positions — is scoped to the firm. This is a security requirement as much as a control, but it is the precondition for every other layer.

Layer 2 — Output control: what the AI must show

  • Per-company evidence. Every AI-assisted screening output carries the source passages the assessment was based on — the filing line, the website paragraph, the database field. An assessment without its source is an opinion, and opinions do not defend comparability.
  • Structured reasons, not prose only. The accept/reject rationale is captured against defined dimensions (products, revenue mix, assets, functions) so the review can be checked row by row — this is the qualitative screening standard applied with the AI’s draft as the first pass.
  • Confidence, honestly labeled. Where the source data is thin (a one-paragraph website, a stale filing), the output should say so. A screening tool that presents a thin-source assessment with the same confidence as a deep-source one is generating false assurance.

Layer 3 — Decision control: the human act

  • The accept/reject/flag decision is a human decision recorded against the AI’s assessment — the reviewer can agree (accept the recommendation), disagree (override, with the reason), or hold (flag for data). The override is not an exception; it is the normal record of professional judgement, and the override rationale is part of the defense, not a mark of shame.
  • The decision is role-gated: the reviewer is a named user with the role to review, and the decision carries that identity and a timestamp.
  • Nothing in the downstream statistics (the range, the percentiles) changes silently: every comparables-set change — AI-suggested or human-made — goes through the same recorded decision.

Layer 4 — Record control: the audit trail

  • An immutable ledger of the AI’s runs: what model/configuration, what inputs, what outputs, what the human decided, in sequence.
  • The ledger is exportable — the audit packet for the proceeding includes the AI layer’s record exactly like any other working paper.
  • The ground-truth linkage: where the AI’s assessment was overruled and the overrule was later validated (the TPO accepted the reviewer’s exclusion), that outcome is recorded against the original assessment — the dataset that tells the firm whether its AI layer is earning its keep.

What AI cannot do (the boundary list)

The boundary is not technical — the models are capable of more — it is defensibility:

  • It cannot make the tested party selection. The “less complex entity” comparison is a professional judgement on the FAR — the AI can lay the two profiles side by side, the selection is the human’s (see tested party selection).
  • It cannot choose the method. The best method rule is applied to the comparability facts by the professional; the AI can summarise the candidates, the choice and its documentation are human (see how to choose a method).
  • It cannot set the range convention or the outlier treatment as policy. Those are firm methodology decisions (the IQR discussion), recorded as methodology, not computed per study.
  • It cannot sign off. The review states — manager review, partner sign-off — are human acts. A “final” study with no human sign-off in the trail is not a study, it is a computation.
  • It cannot replace the benefit analysis, the DEMPE analysis, or the intercompany agreement review. Those are substance judgements on contracts and facts; the AI drafts the first pass, the professional owns the conclusion (see routine vs entrepreneurial and intercompany agreements).

The governance layer (firm level)

The controls above are per-study; the governance layer is per-firm:

  1. Methodology first. The firm’s benchmarking methodology (filters, thresholds, range convention, adjustment policy) is written down and the AI layer operates within it — the AI applies the methodology, it does not choose it.
  2. Scope statement. A short, current statement of where AI is used in the TP work (screening assistance, enrichment, drafting support, risk scoring), what it is not used for (the boundary list), and where the human decision sits. This document is what the firm shows when the authority asks “how was this comparables set produced.”
  3. Data policy. The firm’s data (tested party financials, pricing positions, client data) is not used to train external models; the vendor contract says so in writing. This is the AI data policy question — the most under-negotiated clause in TP software contracts.
  4. Review cadence. Periodic sampling of AI-assisted screens against the human outcomes (the ground-truth records), with the sampling results feeding the methodology review.
  5. Incident path. A defined path for when an AI-assisted output is found wrong after sign-off: what is corrected, what is recorded, and whether the filed documentation needs a supplement. The path existing before the incident is what distinguishes governance from improvisation.

FAQ

If the AI does 95% of the screening work, is the remaining 5% “rubber stamping”? The test is not the percentage, it is the decision. If every company has a recorded human accept/reject with the AI’s evidence in front of the reviewer, the 95% is legitimate acceleration — the reviewer’s act is informed, company-by-company, and recorded. If the reviewer accepts a batch because “the AI said so,” without the per-company record, that is the gap: the decision layer collapsed into the recommendation layer.

Do the OECD or Indian authorities have a position on AI in TP? Neither has published an “AI in TP” standard yet; the authorities’ established positions are on evidence and contemporaneity — the documentation must support the conclusions, and the conclusions must be contemporaneous with the transactions. An AI-assisted workflow is defensible to the extent it meets those standards with more evidence (the source passages, the run records, the decision ledger) than the manual workflow did — which is exactly what the control architecture above is for.

How do we prove to a TPO that the AI layer was controlled? Three artefacts: the run ledger (what ran, on what inputs, with what configuration), the per-company evidence (the source passages behind each assessment), and the decision record (the human accept/reject/override with reasons, by named reviewer). The audit defense preparation that includes those three artefacts is the one the authority reads as a controlled process rather than a black box.

Is there a minimum standard for the override rationale? It must say what was changed relative to the AI’s assessment and why — against the comparability dimensions, not in general terms (“company doesn’t fit” is not a rationale; “revenue mix 60% services vs the tested party’s product line — different risk profile” is). The accept-reject defense guide has the standard and the TPO’s likely questions.

Run the screens as a study, not a spreadsheet

Quartyl applies the method, PLI and screening steps above as a pipeline — and keeps a documented reason for every exclusion.

Related docs

Book a Demo

Tell us what you'd like benchmarked

We'll confirm a 30-minute screen-share slot within one business day.

We reply within one business day. Your details are used only to arrange the demo — never shared or sold.