# Discover and download reports

This guide shows how to synchronise data products generated by the Clearing system as a partner. The download
process has four steps: **Authentication**, **Discovery** (identify the next dataset), **Preparation**
(request it as a formatted report), and **Download** (fetch the signed URL).

:::note
See [Authentication](/docs/authentication#partner-services) for how your organisation gets onboarded
and how to create a client to call the API.

Please refer to the [Clearing Reports API documentation](/apis/clearing-reports) for the full
endpoint reference — this guide focuses on example use, not exhaustive request/response schemas.
:::

## Concepts

- **Data product** — a recurring type of data feed, identified by a data product code (for example
  `SD-GL-1`). Each data product has one or more `dataProductVersionId`s over time — this is the
  identifier you actually use when calling the API (see Discovery below). Synchronising a data
  product means sequentially downloading all dataset IDs for its current `dataProductVersionId`.
- **Dataset** — one generated, immutable instance of a data product (a specific batch, with its
  own `id` and a human-readable filename, see `datasetName` and `reportName` below). Once
  produced, a dataset's content never changes — there is no need to re-check or re-download one
  you have already processed.
- **Report** — a formatted, downloadable rendering of a dataset (CSV, XLSX, PARQUET, ...),
  produced asynchronously via a download job.

For background on how a Settlement's hierarchical data ends up as the rows that populate a
dataset, see [Settlement Structure](/docs/partner-services/clearing/generic-settlements#settlement-structure) in
the Generic Settlements concept page.

## Clearing system's web interface

Before making your first API call, find your organisation's available `dataProductVersionId` values
by logging into the Clearing system's web interface. It can also be used for:

- Self-service search across the complete archive (5 years)
- Ad-hoc download of individual datasets, out of sequence

:::note
In the web interface, `dataProductVersionId` is the ID of a *rapportmal* (Norwegian for "report
template") — use this ID when calling the API.
:::

## Discovery: find the next dataset

Datasets are numbered with an always-increasing `id` (the sequence may have gaps). To synchronise
a data product, poll for its next dataset — using the data product's `dataProductVersionId` — after
the last `id` you have already processed. The dataset must have your organisation on its copy
list.

::endpoint[clearing-reports/getNextDataset]

| Response | Meaning |
|---|---|
| `200` | The next dataset id is returned |
| `204` | No dataset with `id > idAfter` yet, but others may come later since the version is still active |
| `410` | The data product version has been permanently disabled |

When synchronising a new data product with a backlog of datasets, the initial `idAfter` can be
identified via the Clearing system's web interface above. Alternatively, use `idAfter=0` together
with `fromDate` to limit how far back into the backlog you download.

:::warning
**Advance `idAfter` only after successfully processing a dataset — not just after calling
`next`.** `next` itself is a harmless, idempotent lookup; the real risk is downstream, since the
report-creation endpoint below creates a new download job every time it's called, with no
deduplication. A stalled `idAfter` will make you request — and generate — the same report over
and over. **Persist `idAfter` durably between synchronisation runs** — for example to a file or
database, not just in memory — or a job that restarts from scratch will re-walk, and re-download,
the entire history every time it runs. Track `idAfter` separately per `dataProductVersionId`,
especially during a version upgrade (see Versioning below).
:::

You can inspect a dataset's metadata before requesting a report for it. This is not necessary if
your only goal is to download the dataset, but it's useful for initial client configuration, and
the metadata can also be used to access the dataset via BigQuery as an alternative to downloading
it as a formatted report:

::endpoint[clearing-reports/getPartnerDatasetMetadata]

```json title="200 Ok"
{
  "id": 7232958,
  "dataProductVersion": 1108,
  "dataProductCode": "SD-GL-1",
  "orderDate": "2024-12-03",
  "orderBy": "CLEOS",
  "datasetType": "PARQUET",
  "datasetName": "SD-GL-1_ATB_AS_-_2024100772511_v1.1.1.parquet",
  "status": 1,
  "rows": 1340
}
```

| Response | Meaning |
|---|---|
| `200` | Metadata available |
| `403` | Your organisation is not on the dataset's copy list |

Legacy reports may return less information in this call, and are identified by a preformatted
`datasetType` like `XLSX` or `CSV`. This does **not** include a download link — request a
formatted report for that.

## Preparation: request a formatted report

Creates an asynchronous download job that will format the dataset and make it available as a
report via a signed Google bucket URL. `targetFormat` is optional — omit it to get the dataset's
native format.

::endpoint[clearing-reports/createDatasetDownloadJob]

```json title="201 Created"
{
  "id": "1170fca5-b0d6-4661-a002-112a21bde824",
  "status": 0,
  "datasetIDs": [7232958]
}
```

| Response | Meaning |
|---|---|
| `201` | Download job created |
| `403` | Your organisation is not on the dataset's copy list |

## Download: fetch the finished report

The report is created asynchronously, so the signed URL may not yet be available. Poll (or
long-poll with `waitFor`, max 40 seconds) until the job completes:

::endpoint[clearing-reports/getReportContents]

```json title="200 OK — report is ready"
{
  "id": "1170fca5-b0d6-4661-a002-112a21bde824",
  "status": 1,
  "reportName": "SD-GL-1_ATB_AS_-_2024100772511_v1.1.1.csv",
  "contentType": "text/csv",
  "signedBucketUrl": "https://storage.googleapis.com/...",
  "crc32c": "+1nWcg=="
}
```

| Response | Meaning |
|---|---|
| `200` | Available for download |
| `202` | Job in progress — try again |
| `403` | Your organisation is not on the dataset's copy list |
| `410` | Report formatting jobs are transient and are removed after a period of time, independent of status |

Download the file from `signedBucketUrl` directly, and verify it against the `crc32c` checksum.
These URLs are normally valid for 24 hours and may be shared with other users or systems during
that period — there is no authentication beyond knowledge of the URL itself, so treat it as
sensitive.

## Versioning: when a data product changes

Data products change over time. At some point a new `dataProductVersionId` is introduced for the
same data product code — for example to add a column. Clients cannot be expected to transition to
a new version immediately, so during a transition period **both versions are active** and
effectively produce duplicate data. The new version has its own, separate dataset ID sequence.

The table below illustrates a client synchronising `SD-GL-1` through a version upgrade from
`1152` to `1280`:

| Step | Call | `idAfter` sent | Dataset returned | Note |
|---|---|---|---|---|
| 1 | `next` (v1152) | `0` | `7349244` | Initial dataset on v1152 |
| 2 | `next` (v1152) | `7349244` | `7349245` | Sequential v1152 download continues |
| 3 | `next` (v1152) | `7349248` | `7349250` | v1280 has meanwhile become available (starting at `7349248`) — not yet used, since the client is still polling v1152. Note that a new version's ID sequence may start lower than the old one's |
| 4 | *(client switches to dataProductVersionId v1280)* | | | |
| 5 | `next` (v1280) | `7349248` | `7349287` | Client starts the v1280 sequence in ascending order |
| 6 | `next` (v1280) | `7349287` | `7349289` | |
| 7 | `next` (v1152) | `7349250` | `410 Gone` | v1152 is now end-of-life — further synchronisation of this version fails |
| 8 | `next` (v1280) | `7349298` | `7349299` | v1280 sequential download continues |

If a change to a data product is incompatible with the previous version, Entur will normally
implement it as an entirely different data product instead (for example `SD-GL-2`) rather than a
new version of `SD-GL-1`.

## Important to remember

- The `idAfter` cursor is client-managed, per `dataProductVersionId` — see the warnings above. This
  is the single most common integration mistake.
- A `signedBucketUrl` expires after a fixed period (typically 24 hours) — download promptly once
  a job completes.
- A download job is considered transient and is identified by the job, not the dataset — you can
  request the same dataset formatted again later if needed.

:::tip
A client-side reference implementation will be available on request as an example of automated
integration.
:::
