Manual vs AI-Assisted Benchmarking: Time, Cost and Defensibility
A side-by-side of the manual benchmarking workflow and the AI-assisted one: where the time actually goes, where quality improves, and the risks of uncontrolled automation.
The comparison firms actually need is not “manual vs AI” — it is manual vs AI-assisted-with-controls, because the uncontrolled AI workflow is not a real option: it is a workflow that fails at the first audit. This guide puts the two real workflows side by side — the traditional manual process and the controlled AI-assisted process — on the three axes that decide the buying and the design: time, cost, and defensibility. The honest result is that the saving is real but bounded, and that the defensibility difference goes in a direction people expect backwards.
The two workflows, side by side
| Stage | Manual workflow | AI-assisted (controlled) workflow |
|---|---|---|
| Search & candidate pull | Analyst runs the database search, exports the candidates, builds the working file | The search runs from the study parameters; the candidates and the initial screen land in the study automatically |
| Quantitative screening | Filters applied in the database or Excel; the pass/fail list maintained by hand; re-runs on any edit | The recorded filter sequence runs deterministically; every pass/fail and every re-run is in the study record |
| Enrichment | Per-company lookups: LEI/GLEIF, the website visit, the filing read, the ratios keyed | The ratio computation and the screens run in the pipeline; the website capture runs where it is enabled, and each capture carries its source |
| Qualitative screening | The analyst reads each company (filing, website, database text) and writes the assessment; the accept/reject with the reason | The AI drafts the per-company assessment with the source passages; the reviewer decides accept/reject/flag per company, with the reason, against the draft |
| Adjustments & range | Working capital adjustment in a side sheet (the ratios pulled, the days computed, the target applied); the range computed from the adjusted column | Computed from the recorded data and the recorded method; the range derivation shown with its inputs |
| Documentation | The annex assembled by copying the tables and the numbers into the document | The annex generated from the study; the narrative is the professional’s review |
| Review & sign-off | The manager reads the file (the working sheets and the document); the sign-off is a mark | The same review, on the study with its record — the screens, the per-company evidence, the overrides, the derivation — plus the sign-off act |
The structural difference: in the manual workflow the evidence is reconstructed — the working sheets exist, but the per-company source passages, the filter sequence with its per-company results, and the derivation of each number are in the analyst’s sheets and memory, and the documentation is a copy of them. In the controlled AI-assisted workflow the evidence is captured as the work happens — the pipeline records the screen, the enrichment captures the source, the screening decision is recorded against the draft, the range is derived from the recorded inputs. That single difference drives most of the defensibility column below.
Time: where it actually goes
For a standard single-year TNMM study (the common case: one tested party, one PLI, a set of a few dozen to a couple of hundred candidates), the honest shape of the time:
| Stage | Manual (typical) | AI-assisted (typical) | The saving |
|---|---|---|---|
| Search & pull + working file | 0.5–1 day | ~1 hour (setup + the run) | The mechanical setup; the design of the search is unchanged |
| Quantitative screening | 0.5–1 day (including re-runs) | Hours | The re-runs are seconds; the borderline review remains |
| Enrichment (per-company) | 1–2 days | 0.5 day (review of the captures) | The lookups and the website visits; the reading is shortened, not eliminated |
| Qualitative screening | 2–4 days (the largest block) | 1–2 days | The per-company assessment has an evidenced first pass; the decision is still per company |
| Adjustments & range | 0.5 day | ~1 hour | The side-sheet work becomes a computation |
| Documentation | 1–2 days | 0.5–1 day (the review) | The annex is generated; the narrative review remains |
| Total (analyst time) | ~6–11 days | ~3–5 days | Roughly half — and the saved time is the mechanical time, not the judgement time |
Three corrections to the usual expectations:
- The saving is ~50%, not 90%. The qualitative decision and the review do not compress at the same rate as the mechanical steps; any estimate that does is pricing the uncontrolled workflow (the one where the decisions are skipped), not the controlled one.
- The saved time is not senior time. What is saved is analyst mechanical time — the lookups, the re-runs, the copying. The senior review time is roughly unchanged; the senior reads a better file, which is a different value (below).
- The second year is cheaper than the first. The roll-forward (the profile, the methodology, the search design) is captured; the annual study re-runs against the current data. The manual workflow pays the full cost every year; the assisted one pays the judgement cost and a smaller mechanical cost (see documentation automation for the roll-forward mechanics).
Cost: the per-study view
The cost comparison that matters is cost per defensible study, and it has three terms the seat price hides:
| Term | Manual | AI-assisted | Note |
|---|---|---|---|
| Software / data access | The database licence (the comparable company feed) | The platform fee + the database access | The platform fee is the new line; it is priced per the workflow, not per lookup |
| Internal hours per study | The full analyst time above | ~Half the analyst time, at the same rate | The dominant term in most firms — the internal hours outweigh the licence |
| Rework / audit cost | The file reconstructed under audit pressure (the evidence gap found, the supplement prepared) | The lower probability of the evidence gap; the supplement, where needed, drawn from the record | The term everyone omits from the comparison and everyone feels in the TPO year |
The rework term is where the comparison is decided in practice: a manual file whose evidence was reconstructed at audit time costs the firm the senior hours to rebuild it, plus the credibility cost of the reconstruction being visible. The assisted file’s record — the screens, the sources, the decisions — is the audit file, prepared as the work was done. The audit defense guide has the proceeding side of this; the defending the accept-reject matrix guide has the specific record the TPO examines.
Defensibility: the backwards intuition
The intuition to correct: “AI makes the file less defensible because the authority won’t accept a machine’s work.” The actual position is the reverse, and it follows from what the defense is made of:
- The defense is the evidence, not the author. The TPO’s examination of the comparables set is an examination of the record: the search, the filters, the per-company reasons, the adjustments, the derivation. A record that was captured as the work happened — every source passage, every decision, every re-run — defends better than a record that was assembled after the work, which is what the manual file effectively is. The AI-assisted workflow’s structural property is exactly that capture.
- The human decision is still in the file — and it is recorded. The accept/reject per company, the override with its reason, the review and the sign-off: the controlled workflow does not remove the professional’s acts, it records them with more context than the manual workflow records. The authority’s question is “who decided, on what basis” — and the assisted file answers it row by row, with the source passages attached.
- The reproducibility is the difference. The manual file’s numbers are reproducible only by re-doing the work, in the analyst’s sheets, the way the analyst did it. The assisted file’s numbers are a function of recorded inputs — the same parameters and data produce the same screens, the same range. Where the authority re-runs a computation, the reproducible file is the one that survives the re-run.
Where the intuition is right: an uncontrolled AI workflow — the batch accept without the per-company record, the screening without the sources, the sign-off without the review — is the worst of both worlds: the speed of automation with none of the record. The AI controls guide defines the boundary precisely; the automation map shows which stages may run without a person and which carry a recorded human act.
The risks of uncontrolled automation (the list to govern against)
| Risk | The failure | The control that prevents it |
|---|---|---|
| Batch acceptance | The reviewer accepts the AI’s set without the per-company decision | The per-company accept/reject/flag with the reason is required for each company; the batch-accept path does not exist |
| Sourceless assessments | The screening “reason” is prose the AI produced, not a reading of a source | Every assessment carries the source passages; the assessment without a source is not displayed as a decision input |
| Thin-source confidence | A one-line registry entry screens with the same confidence as a deep filing | The source-thickness flag; the reviewer prices the risk it shows |
| Silent re-runs | A parameter change re-runs the screens and moves the range, unrecorded | Every re-run is in the study record; the range’s derivation shows its inputs and its version |
| The sign-off button | The partner signs a file no one walked | The review state requires the review record; the sign-off is an act against the file, logged |
| Drift between the file and the document | The document says one thing, the working file another | The document is generated from the study; the data blocks are not hand-edited |
| Training on firm data | The firm’s tested party data trains a model that serves others | The data policy in the contract: no training on firm data, tenant isolation, the residency terms (see security) |
The governance layer (the methodology as data, the scope statement, the sampling against outcomes, the incident path) is in the AI governance guide — the table above is what it exists to prevent, one row per failure mode.
FAQ
So the AI-assisted file is more defensible than the manual one? Always? Where the controls hold, yes — on the record’s quality, not on the professional’s judgement, which is the same in both. The precise claim: the assisted file’s evidence record is captured-as-work, complete and reproducible; the manual file’s is assembled-after-work, complete only to the extent the analyst’s sheets survived and were organized. The judgement layer (the tested party selection, the per-company decisions, the review) is the professional’s in both — the difference is how well the file shows it.
What is the realistic first-year adoption? The deterministic core first: the recorded quantitative screens, the computed adjustment and range, the generated annex. That is where the time saving is unambiguous and the defensibility gain is structural, and it does not depend on the AI layer at all. The AI-assisted screening and enrichment follow, with the control standard (the sources, the per-company decision) in place from the first assisted study — not retrofitted after the audit question.
How do we compare a vendor’s “AI does the screening” claim to this? Ask for the record, not the speed: show the per-company assessment with its source passages, show the human decision recorded against it, show the re-run history, show the export of the evidence for a proceeding. A vendor that demonstrates the speed and not the record is demonstrating the uncontrolled workflow — the one in the risk table above.
Does this change the firm’s benchmarking methodology? No — the methodology (the filters, the thresholds, the range convention, the adjustment policy) is the firm’s, unchanged; the workflow applies it with a record the manual process did not have. The one methodological addition the assisted workflow invites is writing the methodology down as data (the versioned firm settings), because the workflow needs it in a form the machine can apply — which turns out to be the form the documentation always wanted it in anyway.
Run the screens as a study, not a spreadsheet
Quartyl applies the method, PLI and screening steps above as a pipeline — and keeps a documented reason for every exclusion.
Related docs
Benchmarking Automation: What's Automated vs What Stays Human
A stage-by-stage automation map for the benchmarking study: the deterministic steps that run without a person, the AI-assisted steps that need evidence, and the judgement steps that must stay human.
Read docAI in Transfer Pricing: Use Cases, Controls and Governance
Where AI genuinely helps in TP work — screening, enrichment, drafting, risk — and the controls that keep it defensible: evidence capture, human decisions, override governance and the audit trail.
Read doc