Discover and download reports
This guide shows how to synchronise data products generated by the Clearing system as a partner. The download process has four steps: Authentication, Discovery (identify the next dataset), Preparation (request it as a formatted report), and Download (fetch the signed URL).
See Authentication for how your organisation gets onboarded and how to create a client to call the API.
Please refer to the Clearing Reports API documentation for the full endpoint reference — this guide focuses on example use, not exhaustive request/response schemas.
Concepts
- Data product — a recurring type of data feed, identified by a data product code (for example
SD-GL-1). Each data product has one or moredataProductVersionIds over time — this is the identifier you actually use when calling the API (see Discovery below). Synchronising a data product means sequentially downloading all dataset IDs for its currentdataProductVersionId. - Dataset — one generated, immutable instance of a data product (a specific batch, with its
own
idand a human-readable filename, seedatasetNameandreportNamebelow). Once produced, a dataset's content never changes — there is no need to re-check or re-download one you have already processed. - Report — a formatted, downloadable rendering of a dataset (CSV, XLSX, PARQUET, ...), produced asynchronously via a download job.
For background on how a Settlement's hierarchical data ends up as the rows that populate a dataset, see Settlement Structure in the Generic Settlements concept page.
Clearing system's web interface
Before making your first API call, find your organisation's available dataProductVersionId values
by logging into the Clearing system's web interface. It can also be used for:
- Self-service search across the complete archive (5 years)
- Ad-hoc download of individual datasets, out of sequence
In the web interface, dataProductVersionId is the ID of a rapportmal (Norwegian for "report
template") — use this ID when calling the API.
Discovery: find the next dataset
Datasets are numbered with an always-increasing id (the sequence may have gaps). To synchronise
a data product, poll for its next dataset — using the data product's dataProductVersionId — after
the last id you have already processed. The dataset must have your organisation on its copy
list.
| Response | Meaning |
|---|---|
200 | The next dataset id is returned |
204 | No dataset with id > idAfter yet, but others may come later since the version is still active |
410 | The data product version has been permanently disabled |
When synchronising a new data product with a backlog of datasets, the initial idAfter can be
identified via the Clearing system's web interface above. Alternatively, use idAfter=0 together
with fromDate to limit how far back into the backlog you download.
Advance idAfter only after successfully processing a dataset — not just after calling
next. next itself is a harmless, idempotent lookup; the real risk is downstream, since the
report-creation endpoint below creates a new download job every time it's called, with no
deduplication. A stalled idAfter will make you request — and generate — the same report over
and over. Persist idAfter durably between synchronisation runs — for example to a file or
database, not just in memory — or a job that restarts from scratch will re-walk, and re-download,
the entire history every time it runs. Track idAfter separately per dataProductVersionId,
especially during a version upgrade (see Versioning below).
You can inspect a dataset's metadata before requesting a report for it. This is not necessary if your only goal is to download the dataset, but it's useful for initial client configuration, and the metadata can also be used to access the dataset via BigQuery as an alternative to downloading it as a formatted report:
Code
| Response | Meaning |
|---|---|
200 | Metadata available |
403 | Your organisation is not on the dataset's copy list |
Legacy reports may return less information in this call, and are identified by a preformatted
datasetType like XLSX or CSV. This does not include a download link — request a
formatted report for that.
Preparation: request a formatted report
Creates an asynchronous download job that will format the dataset and make it available as a
report via a signed Google bucket URL. targetFormat is optional — omit it to get the dataset's
native format.
Code
| Response | Meaning |
|---|---|
201 | Download job created |
403 | Your organisation is not on the dataset's copy list |
Download: fetch the finished report
The report is created asynchronously, so the signed URL may not yet be available. Poll (or
long-poll with waitFor, max 40 seconds) until the job completes:
Code
| Response | Meaning |
|---|---|
200 | Available for download |
202 | Job in progress — try again |
403 | Your organisation is not on the dataset's copy list |
410 | Report formatting jobs are transient and are removed after a period of time, independent of status |
Download the file from signedBucketUrl directly, and verify it against the crc32c checksum.
These URLs are normally valid for 24 hours and may be shared with other users or systems during
that period — there is no authentication beyond knowledge of the URL itself, so treat it as
sensitive.
Versioning: when a data product changes
Data products change over time. At some point a new dataProductVersionId is introduced for the
same data product code — for example to add a column. Clients cannot be expected to transition to
a new version immediately, so during a transition period both versions are active and
effectively produce duplicate data. The new version has its own, separate dataset ID sequence.
The table below illustrates a client synchronising SD-GL-1 through a version upgrade from
1152 to 1280:
| Step | Call | idAfter sent | Dataset returned | Note |
|---|---|---|---|---|
| 1 | next (v1152) | 0 | 7349244 | Initial dataset on v1152 |
| 2 | next (v1152) | 7349244 | 7349245 | Sequential v1152 download continues |
| 3 | next (v1152) | 7349248 | 7349250 | v1280 has meanwhile become available (starting at 7349248) — not yet used, since the client is still polling v1152. Note that a new version's ID sequence may start lower than the old one's |
| 4 | (client switches to dataProductVersionId v1280) | |||
| 5 | next (v1280) | 7349248 | 7349287 | Client starts the v1280 sequence in ascending order |
| 6 | next (v1280) | 7349287 | 7349289 | |
| 7 | next (v1152) | 7349250 | 410 Gone | v1152 is now end-of-life — further synchronisation of this version fails |
| 8 | next (v1280) | 7349298 | 7349299 | v1280 sequential download continues |
If a change to a data product is incompatible with the previous version, Entur will normally
implement it as an entirely different data product instead (for example SD-GL-2) rather than a
new version of SD-GL-1.
Important to remember
- The
idAftercursor is client-managed, perdataProductVersionId— see the warnings above. This is the single most common integration mistake. - A
signedBucketUrlexpires after a fixed period (typically 24 hours) — download promptly once a job completes. - A download job is considered transient and is identified by the job, not the dataset — you can request the same dataset formatted again later if needed.
A client-side reference implementation will be available on request as an example of automated integration.