Skip to main content
Quartyl
Benchmarking in Quartylprofessional

Raw Dump Data: Inspection and Reuse After Screening

Raw dump data in Quartyl: the parquet persisted after quantitative screening, inspecting the numbers behind every filter decision, reuse across re-runs, and the per-tenant retention window.

Quartyl Team

The dump is the study’s raw data, preserved whole. At the start of the quantitative step of the pipeline, before any screening changes anything, the full uploaded file - normalized to standardized column names - is written as a parquet file to object storage. It is the numbers behind every filter decision, the rebuild source for the report, and a bounded, retained artifact rather than an in-memory byproduct.

What the dump is

  • The content - the complete comparable population as uploaded: every company row, all detected columns (the standardized financials, the year-suffixed multi-year columns, the description fields, the industry and website fields), nothing filtered out. Rejections are decisions made on top of this data; the data itself keeps everything.
  • The format and location - a normalized parquet at dumps/{tenant}/{study_id}/dump.parquet in the storage provider, with the storage key, the file size and a checksum recorded on the study. The raw rows are never stored as database JSON - the database keeps the key and the metadata, storage keeps the data.
  • The metadata - the row count, the column mapping (which original column mapped to which standard field, the output of the detector) and the detected fiscal years. That metadata is what makes the dump inspectable: you can see what the pipeline understood the file to be.

Plan feature: This capability requires the dump_retention_config plan feature (see Plan Features).

Inspecting the dump

The dump exists so a reviewer can answer “what number was that decision made on?” For a rejected comparable - “Below minimum revenue,” “High RPT,” “Incomplete Multi-Year Financials: missing data for year(s) 2023” - the dump holds the exact revenue, the exact RPT, the exact year columns, so the reason can be checked against the source instead of trusted on faith. The same is true in the other direction: an accepted company’s PLI, its per-year figures and its description fields are all in the dump, which is why the screen’s reasons are auditable rather than opaque.

The report rebuild reads the same artifact: when the study is approved and the Excel workbook is generated, the builder downloads the dump and rebuilds the full search-result and web-analysis sheets from it - on any worker host, because the key is a storage reference, not a local file path.

Reuse across re-runs

  • A re-screen is a fresh run on the same source file. Re-submitting a study (re-screen) re-reads the uploaded file and produces a new dump for that run, so each run’s numbers trace to that run’s artifact.
  • The dump outlives the run’s in-memory work. Pipeline steps mutate an in-memory results structure; the dump is the persistent source that survives step failures, worker crashes and retries, and that report generation later depends on.
  • Deleting a study deletes its dump with it. Orphan dumps - storage objects whose study row no longer exists or is soft-deleted - are reclaimed by the retention sweep, so storage does not accumulate data for studies that are gone.

Retention

The dump is retained on a per-tenant window, and the window is configurable:

Parameter Value
Default window 7 days
Configurable range 1 to 31 days, per tenant
Who sets it The platform superadmin, on the Add Firms console; assigning a plan carries the plan’s retention value to the tenant
When the clock starts When the study enters In Review - a study still awaiting review never loses its dump to the calendar
Enforcement A daily scheduled sweep lists the dump prefix, resolves each key’s tenant, and deletes objects older than that tenant’s window

Two properties matter in practice:

  • A deleted dump only blocks new rebuilds. The sweep removes the parquet; it never touches the study’s persisted analysis data or any report already generated. What changes is that a new Excel report can no longer be rebuilt from the full dump.
  • The per-tenant window is why a bucket lifecycle rule is not enough. Storage lifecycle rules are bucket-level and cannot read a per-tenant setting, which is why the sweep - not the bucket - is the enforcement mechanism.

FAQ

Can I download the dump? The dump is a system artifact referenced by storage key, not a user-facing download. Its contents surface where they matter: the per-company reasons in the grid, the details panels, and the report sheets rebuilt from it.

Why persist the dump before screening instead of after? So that whatever the screen decides - and whatever fails later in the pipeline - the complete pre-decision data is already safe. A dump saved after screening would only prove the decisions the screen made.

What if the retention window expires mid-review? It does not, by construction: the clock starts at In Review, and the default 7 days (or the tenant’s configured window) runs from there. A study in review keeps its dump until the window genuinely elapses, and the sweep is what enforces the boundary.

See it working in your workspace

Sign in to run the steps above on a real study — or book a demo and we will walk the workflow end to end.

Related docs

Book a Demo

Tell us what you'd like benchmarked

We'll confirm a 30-minute screen-share slot within one business day.

We reply within one business day. Your details are used only to arrange the demo — never shared or sold.