Website Enrichment: Where Comparable Descriptions Actually Come From
By default Quartyl does not crawl comparable websites; descriptions come from the dump you upload. The optional Web Research step reads the open web.
A comparable’s financials say what it earned; its business description says what it sells. Be clear about what Quartyl does with that second half: by default it never visits the comparable’s website. This page covers where the description actually comes from, what the website column is, and what the Web Analysis sheet reports.
Default: no crawling, no scraping
The core analysis pipeline makes no outbound requests to comparable websites. There is no fetcher and no headless browser in the default path - the crawler this product once shipped with was removed. Everything the qualitative and AI screens say about a company traces back to text already present in your upload; no external entity record is joined onto a comparable at any point in the run.
The one exception is the optional Web Research step. It runs only for
tenants whose plan holds the advanced-AI-screening feature, the same one that
gates the deep pass. When it does run, it reads each candidate’s own website
and public records of its group structure, and feeds that evidence to the
deep-screening pass and the Web Analysis sheet — always labelled
Web research: so a reviewer can tell it apart from registry text.
Unless that step ran, everything below holds without qualification.
Two consequences matter when you read the output:
- The description is as current as the vendor file. If the dump’s description text is years stale, the run reasons over stale text. No report line implies a “checked at source” date, because nothing was checked at source (unless the optional Web Research step supplied a dated web source).
- Data egress is limited to the AI pass (and, when enabled, Web Research). The third-party calls in a default run are the AI-screening inference requests, and those carry company names and trade descriptions (see AI Screening). A default run contacts no registry and no company website.
What the website column is
The Website column in the grid and in the report is a dump column: the URL
field the data vendor supplied for that company, mapped by the column detector
when your file carries one. Quartyl stores it and displays it. In a default run
it does not open it, resolve it, check whether it is live, or record an HTTP
status for it - and no description text is attributed to it. (When the optional
Web Research step runs, it resolves a company’s site and reads it; those
findings arrive labelled Web research: and are never folded silently into the
vendor text.)
Treat it as a pointer for your own verification, not as evidence the platform verified.
Where the description actually comes from
The description the screens score is assembled from the description fields in the upload - the full overview, the business description / oneliner, the trade description, and the main products and services columns, whichever your file populates. That text is what the embedding comparison runs on and what the AI pass quotes from, against the nature-of-business statement you gave the study. A company with none of these is not scored on nothing: it is flagged for human review with the recorded reason “Missing/Insufficient Business Description” (see Qualitative Data).
What the Web Analysis sheet contains
The Excel output has a sheet named “Web Analysis”. The name refers to the qualitative half of the review. Its columns are:
| Column | Source |
|---|---|
S. No |
Row index |
Company Name |
The comparable as named in the dump |
BvD ID number |
The vendor identifier carried in the dump |
Country |
Country as recorded on the comparable |
External Sources |
The sourcing for any web finding: the company’s own site, each URL the findings rest on, the UTC date the evidence was captured, and the extraction confidence — with an explicit line where a site resolved but could not be read. Empty when Web Research did not run |
Products & Services |
The description text read back from the dump row, plus a labelled Web research: block when that step ran |
Corporate Structure |
The vendor independence indicator where the dump carries one, plus a labelled Web research: block (group structure, named parents, subsidiaries and shareholders, legal form, industry, employees, locations) when that step ran |
Functional Profile |
The functional summary the AI pass produced, where it ran, plus a labelled Web research: block when that step ran |
Result |
Final recommendation - the AI verdict where available, otherwise the qualitative recommendation |
Remarks |
The rationale recorded against that company: the AI reason, the rejection reason, or the qualitative rationale |
Note the fifth header. It says External Sources, not a scraping claim: the
only thing it holds is the URL behind a finding the run actually read, so a
reviewer can trace each web claim back to the page it came from. Where Web
Research did not run there is nothing to source and the cell stays empty - a
blank reads as “not found”, because no external registry data is joined onto the
comparable to fill it. Web findings never overwrite vendor text either; they are
appended under their own Web research: label.
Why a reviewer still verifies it
Because the description is vendor-supplied text rather than something the platform checked at source, the verification is yours:
- The text is on the row. Products & Services and the underlying dump columns travel into the report, so you can read exactly what the screen judged rather than a paraphrase of it.
- Quote-level traceability on the AI pass. Every AI verdict carries a verbatim evidence quote from the description it used, so any claim in a rationale can be matched back to the sentence behind it.
- Corrections are dispositions, not field edits. At review you change a comparable’s outcome through an accept / reject / flag override with a reason; there is no editor for the description fields themselves, since rewriting vendor text inside a finished run would break the audit link between the number and its source. See Reviewer Overrides and Rationale.
FAQ
What about the settings that look like a fetcher? The configuration still
carries a comment recording that the old crawler was removed, and the qualitative
constants file still defines weights and timeouts for a description-fetch
score that nothing performs (website text 0.60 versus dump text 0.40, a
10-second timeout, five parallel workers). No module imports them; they are
dead scoring constants, not a hidden feature. That is separate from the modern
Web Research step: when enabled, it genuinely reads company websites and
public group-structure sources, but it writes those findings into a distinct
web_research field the report renders under a Web research: label and feeds
to the deep-screening pass — it never backfills the qualitative description
score, and it does not touch the vendor text.
A comparable’s description looks wrong - what can I do? Decide the disposition: accept or reject it, with a reason that lands in the evidence record. The vendor text itself stays as uploaded, because the run’s numbers and its rationale were computed from it.
Does this weaken the comparability argument? It keeps it honest. In a default run the description is vendor-supplied text, so a defense file that says “the vendor description, cross-checked by us” is exactly right; nothing was checked at source, and no report line claims otherwise. When the optional Web Research step runs, an automated look at the company’s site does happen — and each web finding carries the source URL it came from and the UTC date it was captured, so you can answer “when, and from where” for every claim rather than hand-waving it. Either way the platform tells you which evidence is registry text and which is web text, instead of blending them.
See it working in your workspace
Sign in to run the steps above on a real study — or book a demo and we will walk the workflow end to end.
Related docs
Qualitative Data: The Per-Comparable Record and How You Act On It
What Quartyl stores in a comparable’s qualitative record — score, confidence and rationale — what is deliberately not in it, and how reviewer judgment enters as a disposition.
Read docAI Screening: How It Scores Companies and Cites Evidence
AI screening in Quartyl: LLM FAR analysis of each comparable against the tested party, verbatim evidence quotes, the outcome categories, human override and the gated deep pass.
Read doc