Skip to main content
Quartyl
TP Software & AIprofessional

Benchmarking Automation: What's Automated vs What Stays Human

A stage-by-stage automation map for the benchmarking study: the deterministic steps that run without a person, the AI-assisted steps that need evidence, and the judgement steps that must stay human.

Quartyl Team

Every stage of the benchmarking study falls into one of three automation classes, and the defensible workflow is built by classifying each stage honestly:

  • Deterministic — a computation or a rule with inputs and outputs; should run automatically, every time, identically.
  • AI-assisted — a reading or drafting step over company data; may run automatically with evidence, and produces a recommendation, not a decision.
  • Judgement — a professional decision on facts; a human acts, and the act is recorded.

The failure mode of over-automation is a workflow where the judgement class has been quietly reclassified as AI-assisted — the file looks automated and defensible, and then the TPO asks for the decision that was never made. This guide is the stage-by-stage map, the per-stage standards, and the boundaries.

The stage map

# Stage Class What runs automatically What must stay human
1 Scope & tested party setup Judgement (setup: deterministic) The profile structure, the parameter defaults, the validation checks The tested party selection and the FAR profile; the method/PLI choice
2 Search design AI-assisted (design: judgement) Code suggestions (NIC/NACE), the search string draft, the candidate count The scope decision — which codes, how inclusive, and why
3 Quantitative screening Deterministic The filter sequence, the thresholds, the pass/fail per company, the dump of the screen Setting the thresholds (methodology); reviewing the borderline set
4 Data enrichment AI-assisted (lookup: deterministic) GLEIF/LEI lookups, corporate linkage, website business-description capture, ratio computation Nothing routine — but the reliability of a thin source is a human call
5 Qualitative screening AI-assisted (decision: human) The per-company assessment draft, with the source passages cited The accept/reject/flag per company, with the reason recorded
6 Adjustments (working capital, etc.) Deterministic The computation from the comparable ratios, the tested party base, the method applied The adjustment policy (which adjustments, which method) as firm methodology
7 Range & statistics Deterministic The IQR/full-range bounds, the mid-point, the percentiles, the outlier handling per the recorded convention The range convention as firm methodology; the significance judgement on a thin set
8 Documentation & report AI-assisted (draft) The report generation from the study data — every number drawn from the study, not re-typed The professional’s review of the narrative; the sign-off
9 Review & sign-off Judgement The state transitions, the notifications, the chronology record The review itself — manager and partner acts, recorded

Read the map top to bottom and the shape of a defensible automation is visible: the deterministic core (stages 3, 6, 7, and the report generation) runs without a person; the AI-assisted stages produce evidenced recommendations; and the judgement stages (1, 2-scope, 5-decision, 9) carry a recorded human act that the file cannot exist without.

The deterministic core: what “automated” should mean

The deterministic stages are where automation is unambiguous and where most of the time saving is. The standard for each:

Quantitative screening (stage 3)

  • The filter sequence is fixed by the firm’s methodology — size, sector, geography, financial-ratio screens — and runs in the same order every time (see quantitative screening).
  • Every threshold is a parameter, visible on the study, and the pass/fail per company is recorded: which filter removed the company, and on what value. The self-audit checklist treats an unrecorded filter as a documentation gap by design.
  • The screen’s input data is snapshotted: the run is reproducible from the recorded inputs, not from “the database as of whenever.”

Adjustments (stage 6)

  • The working capital adjustment is computed from the comparable data — the ratio method or the days method, the tested party base, the target — with the method and the inputs recorded per study (see the working capital adjustment guide).
  • No adjustment value is a free-form cell: it is either computed or, where the firm’s methodology allows a judgmental adjustment (a documented country premium, an extraordinary event exclusion), it is recorded with the reason (see comparability adjustments).

Range & statistics (stage 7)

  • The range is computed from the adjusted PLIs under the recorded convention (the IQR rule, the outlier treatment, the significance threshold) — see the IQR vs full range and loss-making comparables guides for the conventions and their edge cases.
  • The output shows the derivation: the bounds, the mid-point, the tested party position, the set size, and the significance read.

The property the whole core shares: the output is a function of recorded inputs. Anyone with the study’s parameters and data can re-run it and get the same numbers. That reproducibility is what “automated” is worth — not the speed, the identity of the result.

The AI-assisted stages: the evidence standard

Stages 2, 4 and 5 (and the drafting at stage 8) are where AI adds real value and where the control standard matters. The standard, per stage:

Search design (stage 2)

  • The AI’s contribution is the suggestion layer: the NIC/NACE candidates for the tested party’s activity, the draft search string, the expected candidate count. The search design judgement — the scope, the inclusiveness, the trade-off between coverage and comparability — is the professional’s, recorded on the study.
  • The recorded search (the final string, the codes, the scope note) is what the documentation quotes; the suggestion trail is supporting detail.

Enrichment (stage 4)

  • The enrichment is sourced data: the GLEIF record (the LEI, the legal name, the linkage), the website extract (the business description, with the URL and the capture date), the computed ratios. Each item carries its source — this is the difference between enrichment and invention.
  • The thin-source flag: where the description is a one-line registry entry or the filing is stale, the record says so, so the qualitative screen (stage 5) prices the risk instead of being surprised by it.

Qualitative screening (stage 5)

  • The AI drafts the per-company assessment against the recorded dimensions (products, revenue mix, assets, functions) with the source passages — the filing line or the website paragraph behind each statement.
  • The human decision is per company: accept or reject, with the reason, recorded against the draft. A row the reviewer leaves undecided stays out of the statistics. The qualitative screening guide has the per-company standard; the accept-reject defense guide has what the TPO checks.
  • A batch-accept without the per-company record is the specific failure this stage exists to prevent — the automation is legitimate only while the decision layer is intact (see AI in transfer pricing for the full control architecture).

Documentation drafting (stage 8)

  • The report is generated from the study data — the benchmarking annex, the method rationale, the FAR summary — so the document and the working file are the same numbers by construction.
  • The AI’s role is the narrative first pass: the prose around the numbers. Every factual claim in the narrative maps to a study field or a cited source; the professional reviews and rewrites, and the sign-off is the human act that closes the stage.

The judgement stages: what must stay human

The map’s three judgement stages, and why automation stops at their border:

  • Tested party selection (stage 1). The “less complex entity” comparison is a professional judgement on the two FAR profiles — the assets owned, the risks controlled, the functions performed (see tested party selection). The tool can lay the profiles side by side and check the documentation; the selection and its reasoning are the human’s, recorded.
  • Search scope (stage 2, the design half). Over-inclusive searches dilute the set; under-inclusive ones starve it. The trade-off is a judgement on the comparability facts, documented as a design decision — not a parameter the optimiser tunes.
  • Review & sign-off (stage 9). The manager review and the partner sign-off are acts — the reviewer read the file, checked the overrides, and owns the result. A workflow whose sign-off is a button with no review record has automated the appearance of review; the audit trail is what separates the two.

The through-line: automation removes the mechanical steps between judgements; it does not remove the judgements. A fully “automated” study with no recorded human decision at stages 1, 5 or 9 is not faster — it is undefendable, because the file’s value was the decisions, and the decisions are not in it.

The time model (what automation actually saves)

Honest shape of the saving, for a standard single-year TNMM study:

Stage Manual time (typical) Automated time (typical) Where the saving is
Search & candidate pull Hours (database work) Minutes The pull and the initial screen run without a person
Quantitative screening Tens of minutes per screen, re-runs on edits Seconds Deterministic re-runs after any parameter change
Enrichment Hours (per-company lookups) Minutes GLEIF + website capture in the pipeline
Qualitative screening The largest block — per-company reading 30–60% of the manual time The AI draft shortens the per-company review; the decision is still per-company
Adjustments & range Minutes–hours (the side-sheet work) Seconds Computed, not pasted
Documentation Hours (assembling the annex) Minutes (generation) + the professional’s review The numbers are drawn from the study, not re-typed

The qualitative screen is where the saving is real but bounded — the per-company human decision does not disappear, it gets a better first pass. Any vendor claim that the qualitative screen “runs itself” is claiming the bounded step is unbounded; treat it as the anti-signal it is (see manual vs AI-assisted for the side-by-side).

FAQ

Can the whole study run unattended from upload to report? The pipeline can — upload through the deterministic screens, enrichment, adjustments and range, and report generation, without a person at the keyboard. The study cannot: the tested party selection, the search scope, the per-company qualitative decisions and the review sign-off are recorded human acts. The unattended run stops at each judgement stage and waits for the decision — that is the design, not a limitation.

What is the minimum automation that is worth adopting? The deterministic core first: the recorded quantitative screens, the computed working capital adjustment, the computed range, and the report generated from the study data. That core alone removes the side-sheet work and makes the file reproducible — and it is the foundation the AI-assisted layers (enrichment, screening drafts) are built on. Automating the AI layer before the deterministic core is in place puts the evidence on sand.

How do we keep the automation from drifting (different results across years)? Pin the methodology as data: the filter sequence, thresholds, adjustment policy and range convention are versioned firm settings, applied per study and recorded on the study. A year-over-year difference in the range should be traceable to a change in the data or a recorded change in the methodology — never to an untracked setting. The multi-year averaging discipline (consistency with the tested party, the same method across the periods) is the same principle applied to the periods.

Does automation change what the documentation must contain? No — it raises what the documentation can show. The same Rule 10D blocks, the same method rationale, the same accept/reject record; but the documentation can now show the filter sequence with the per-company results, the source passages behind each qualitative assessment, and the reproducible derivation of the range. The content standard is unchanged; the evidence standard is higher, which is the point (see TP documentation).

Run the screens as a study, not a spreadsheet

Quartyl applies the method, PLI and screening steps above as a pipeline — and keeps a documented reason for every exclusion.

Related docs

Book a Demo

Tell us what you'd like benchmarked

We'll confirm a 30-minute screen-share slot within one business day.

We reply within one business day. Your details are used only to arrange the demo — never shared or sold.