Entur Developer
Clearing

Discover and download reports

This guide shows how to synchronise data products generated by the Clearing system as a partner. The download process has four steps: Authentication, Discovery (identify the next dataset), Preparation (request it as a formatted report), and Download (fetch the signed URL).

See Authentication for how your organisation gets onboarded and how to create a client to call the API.

Please refer to the Clearing Reports API documentation for the full endpoint reference — this guide focuses on example use, not exhaustive request/response schemas.

Concepts

  • Data product — a recurring type of data feed, identified by a data product code (for example SD-GL-1). Each data product has one or more dataProductVersionIds over time — this is the identifier you actually use when calling the API (see Discovery below). Synchronising a data product means sequentially downloading all dataset IDs for its current dataProductVersionId.
  • Dataset — one generated, immutable instance of a data product (a specific batch, with its own id and a human-readable filename, see datasetName and reportName below). Once produced, a dataset's content never changes — there is no need to re-check or re-download one you have already processed.
  • Report — a formatted, downloadable rendering of a dataset (CSV, XLSX, PARQUET, ...), produced asynchronously via a download job.

For background on how a Settlement's hierarchical data ends up as the rows that populate a dataset, see Settlement Structure in the Generic Settlements concept page.

Clearing system's web interface

Before making your first API call, find your organisation's available dataProductVersionId values by logging into the Clearing system's web interface. It can also be used for:

  • Self-service search across the complete archive (5 years)
  • Ad-hoc download of individual datasets, out of sequence

In the web interface, dataProductVersionId is the ID of a rapportmal (Norwegian for "report template") — use this ID when calling the API.

Discovery: find the next dataset

Datasets are numbered with an always-increasing id (the sequence may have gaps). To synchronise a data product, poll for its next dataset — using the data product's dataProductVersionId — after the last id you have already processed. The dataset must have your organisation on its copy list.

ResponseMeaning
200The next dataset id is returned
204No dataset with id > idAfter yet, but others may come later since the version is still active
410The data product version has been permanently disabled

When synchronising a new data product with a backlog of datasets, the initial idAfter can be identified via the Clearing system's web interface above. Alternatively, use idAfter=0 together with fromDate to limit how far back into the backlog you download.

Advance idAfter only after successfully processing a dataset — not just after calling next. next itself is a harmless, idempotent lookup; the real risk is downstream, since the report-creation endpoint below creates a new download job every time it's called, with no deduplication. A stalled idAfter will make you request — and generate — the same report over and over. Persist idAfter durably between synchronisation runs — for example to a file or database, not just in memory — or a job that restarts from scratch will re-walk, and re-download, the entire history every time it runs. Track idAfter separately per dataProductVersionId, especially during a version upgrade (see Versioning below).

You can inspect a dataset's metadata before requesting a report for it. This is not necessary if your only goal is to download the dataset, but it's useful for initial client configuration, and the metadata can also be used to access the dataset via BigQuery as an alternative to downloading it as a formatted report:

JSONCode
{ "id": 7232958, "dataProductVersion": 1108, "dataProductCode": "SD-GL-1", "orderDate": "2024-12-03", "orderBy": "CLEOS", "datasetType": "PARQUET", "datasetName": "SD-GL-1_ATB_AS_-_2024100772511_v1.1.1.parquet", "status": 1, "rows": 1340 }
ResponseMeaning
200Metadata available
403Your organisation is not on the dataset's copy list

Legacy reports may return less information in this call, and are identified by a preformatted datasetType like XLSX or CSV. This does not include a download link — request a formatted report for that.

Preparation: request a formatted report

Creates an asynchronous download job that will format the dataset and make it available as a report via a signed Google bucket URL. targetFormat is optional — omit it to get the dataset's native format.

JSONCode
{ "id": "1170fca5-b0d6-4661-a002-112a21bde824", "status": 0, "datasetIDs": [7232958] }
ResponseMeaning
201Download job created
403Your organisation is not on the dataset's copy list

Download: fetch the finished report

The report is created asynchronously, so the signed URL may not yet be available. Poll (or long-poll with waitFor, max 40 seconds) until the job completes:

JSONCode
{ "id": "1170fca5-b0d6-4661-a002-112a21bde824", "status": 1, "reportName": "SD-GL-1_ATB_AS_-_2024100772511_v1.1.1.csv", "contentType": "text/csv", "signedBucketUrl": "https://storage.googleapis.com/...", "crc32c": "+1nWcg==" }
ResponseMeaning
200Available for download
202Job in progress — try again
403Your organisation is not on the dataset's copy list
410Report formatting jobs are transient and are removed after a period of time, independent of status

Download the file from signedBucketUrl directly, and verify it against the crc32c checksum. These URLs are normally valid for 24 hours and may be shared with other users or systems during that period — there is no authentication beyond knowledge of the URL itself, so treat it as sensitive.

Versioning: when a data product changes

Data products change over time. At some point a new dataProductVersionId is introduced for the same data product code — for example to add a column. Clients cannot be expected to transition to a new version immediately, so during a transition period both versions are active and effectively produce duplicate data. The new version has its own, separate dataset ID sequence.

The table below illustrates a client synchronising SD-GL-1 through a version upgrade from 1152 to 1280:

StepCallidAfter sentDataset returnedNote
1next (v1152)07349244Initial dataset on v1152
2next (v1152)73492447349245Sequential v1152 download continues
3next (v1152)73492487349250v1280 has meanwhile become available (starting at 7349248) — not yet used, since the client is still polling v1152. Note that a new version's ID sequence may start lower than the old one's
4(client switches to dataProductVersionId v1280)
5next (v1280)73492487349287Client starts the v1280 sequence in ascending order
6next (v1280)73492877349289
7next (v1152)7349250410 Gonev1152 is now end-of-life — further synchronisation of this version fails
8next (v1280)73492987349299v1280 sequential download continues

If a change to a data product is incompatible with the previous version, Entur will normally implement it as an entirely different data product instead (for example SD-GL-2) rather than a new version of SD-GL-1.

Important to remember

  • The idAfter cursor is client-managed, per dataProductVersionId — see the warnings above. This is the single most common integration mistake.
  • A signedBucketUrl expires after a fixed period (typically 24 hours) — download promptly once a job completes.
  • A download job is considered transient and is identified by the job, not the dataset — you can request the same dataset formatted again later if needed.

A client-side reference implementation will be available on request as an example of automated integration.

Last modified on
Did you find what you were looking for?