This project read OGRA’s price-publications listing, found twenty-two documents on it, and built a record twenty-two days long. It then said, in several places and on this site, that OGRA overwrites its own page and that this archive is therefore the only durable copy of those documents. Both halves of that were wrong. This is the correction, and the account of the sweep that produced it.
Every claim below names the document id it was read from, so anybody can fetch the same bytes and check it. The sweep was carried out on 21 August 2026; what a regulator’s server answers is a fact about one day, and OGRA has undertaken nothing about any day after it.
The listing is a window; the documents do not go away
https://www.ogra.org.pk/download/?lang=en answers for an id far below anything the listing has shown for weeks. It answers for ids in the low hundreds — 140 serves a slide deck from a consultative session, 160 serves a file called test_job.pdf — and it goes on answering all the way up to the newest publication.
What the listing does is turn over. It shows about twenty documents, newest first, with no pagination and nothing behind it, so a publication drops out of view within a few weeks of being published. Nothing on ogra.org.pk then links to it. It is not withdrawn; it becomes unfindable.
That is a real loss and it is a different loss from the one this project was claiming. The only durable copy says the bytes are gone. They are not. What is gone is any way of finding them, and any assurance that they will keep being served: the id space is not a published interface, OGRA has never undertaken to keep it, and it can stop answering without notice or announcement.
And the floor cannot be found by bisection
The id space is not contiguous. 4300, 4330, 4331 and 4333 all answer 404 while 4332 and 4334 serve documents. A binary search over a space with holes in it finds a hole and calls it the end — which is how 4334 came to be believed the lowest resolving id when the space in fact runs several thousand ids further down.
This is the kind of error that leaves no symptom. A bisection that stops at a hole returns a number, the number looks like an answer, and nothing downstream can tell it from the real floor. The sweep that replaced it is slower by design — see the last section.
A filename is not an identifier
OGRA generates the Content-Disposition filename from the document’s effective date. Several documents therefore share one filename, and the run of ids from 14675 to 14695 answers twenty-one consecutive header requests with
and serves twenty-one different PDFs behind them, of quite different sizes: 14666 is 50,433 bytes, 14675 is 735,194, 14690 is 265,804.
So a filename dates a document and places it in a family. It never says which document it is. Nothing in this project may be keyed on one, and the catalogue records it as evidence rather than as an identifier. What identifies a document here is the SHA-256 of its bytes, which is also what the archive lists it under.
The families, and what is inside them
One download endpoint serves everything OGRA publishes. Sorted by what the filename says, and then — for the one family where the filename turned out not to settle it — by what is inside.
| Family | What it is | Readable? |
|---|---|---|
price_publication… | The daily price build-up this project reads — and, before it, one-page price-change notices under the same name | Text layer. The build-ups parse; the notices are refused |
prices_effective_dated_… | One oil marketing company’s retail outlets and the pump price at each. About two thirds of everything OGRA serves | Text layer, and nothing in it about what a price is made of |
detail_computation_ex_depot_sale_price_… | The build-up’s predecessor: import-parity working for every product OGRA quotes, and a full ex-depot build-up for kerosene, fortnightly | Photographs of paper. No text layer at all |
notification_petroleum_products_prices_… | The gazette notification putting the same prices in force, itemised for kerosene | Scan |
ifem_notification_… | The inland freight equalisation margin | Scan |
| Everything else | Wellhead price determinations, RLNG averages, SNGPL and SSGC revenue requirements, licensing notices, tenders, LPG statistics, staff notices | Not this project’s subject |
The price publication family is not one kind of document
Every one of the twenty-two on OGRA’s listing is the two-page build-up with the lettered rows — A to T for petrol, and A to W for diesel since the restructure of 20 August 2026. It has a text layer and the parser reads it.
Earlier members of the same family are not that at all. Document 14045, price_publication_effective_may_16_2026.pdf, is one page, and this is the whole of its text layer:
That is a price-change notice. It carries the two pump prices and says nothing about what is inside them. The parser refuses it, and refuses it correctly: this does not look like an OGRA price publication.
So the family is read rather than assumed. Every member of it is fetched, its text handed to the same parser the days are built with, and recorded as price_buildup or price_notice accordingly, with the digest of the bytes the reading was taken from. Only the build-ups are fed to the pipeline. The distinction is worth making rather than letting the notices fail: a document with no build-up in it and a build-up this project got wrong are different facts, and reporting both as “could not be read” would say the same thing about OGRA and about ourselves. Before the notices, the family ran weekly — one document each Saturday, dated the following day’s effective date.
The predecessor, which is the one that matters most
Document 10000, detail_computation_ex_depot_sale_price_effective_dated_january_16_2024.pdf, is two pages. Both are scanned images. The PDF has no text layer at all — pypdf extracts zero characters from either page, and each page holds exactly one JPEG. Read as pictures, they are:
- Annex-I — Ex-Refinery Sale Price Calculation
- The Arab Gulf FOB mean prices for the ten working days of the period, product by product: naphtha, HSFO 180 CST, kerosene, gas oil at four sulphur grades, gasoline 95 RON, gasoline 92 RON, and the selling exchange rate for each day. Below them the import parity working: FOB price, premium, C&F, the 7.5% tariff, C&F plus tariff, the conversion factor, and the price differential claim.
- Annex-II — Ex-Depot Sale Price Calculation
- A build-up for kerosene and kerosene alone: ex-refinery 178.20, IFEM 7.03, distributor margin 1.58, dealer margin nil, petroleum levy 0.05, price before sales tax 186.86, sales tax nil, maximum ex-depot sale price 186.86, direct and rail/depot columns side by side.
Document 10445, for 16 April 2024, is the same two annexes with April’s figures — ex-refinery 183.49, IFEM 7.96, maximum ex-depot sale price 193.08 — and is also entirely scanned.
So OGRA did publish the inside of a litre before 21 July 2026. It published the import-parity working for every product it quotes, and a full ex-depot build-up for kerosene, fortnightly. What it did not publish was a per-litre build-up for petrol or high-speed diesel — and it published none of it in a form any parser can read.
The gazette notification is the same story from the other side. Document 8020, dated Islamabad 15 August 2022 and effective the 16th, is one page and a scan. It is OGRA’s S.R.O. under the Petroleum Products (Petroleum Levy) Ordinance, and its table itemises, for kerosene oil: prescribed price 192.03, petroleum levy 10.00, inland freight margin 7.37, dealers commission nil, distributors margin 1.58, general sales tax nil, maximum ex-depot sale price 199.40. Rupees a litre, signed. Document 13937, for 1 May 2026, is the same shape and also a scan.
Those two families are what the scanned record reads, by machine, off the photographs — and it is kept structurally apart from everything else on this site for exactly the reason this section gives.
What this changes about the record
The record can go back further than 21 July 2026, and by less than it looks. The price publication family runs weekly before that date, but its earlier members are price-change notices with no build-up in them. Every document in the family is now fetched, read, and classified by what is in it rather than by what it is called, so the record extends by exactly the documents that carry a build-up and by no others.
Nothing before the build-up family can be parsed at all. The predecessor documents — the ex-depot computation, the gazette notifications, the IFEM determinations — are scans. There is no text layer to read, no way to check a figure lifted off an image against OGRA’s own arithmetic, and therefore no honest way for this project to publish a figure from one of them at the standard the rest of this site is held to. They are catalogued, their existence is stated here, and the separate, weaker record that does read them is marked as such on every row.
The archive’s value is real and is not the value that was claimed. OGRA serves the documents; it does not index them, does not link them, does not promise to keep serving them, and does not keep a copy that cannot change under you. This archive holds each publication by the date it takes effect, with the SHA-256 it had when a figure was read out of it, and refuses to overwrite one whose bytes have moved. That is a copy that stays findable and stays fixed. It is not the only copy in the world, and this site no longer says that it is.
How the map is kept
The sweep walks the id space with header requests and records one row per id: the filename OGRA served, this project’s reading of it, the date the filename carries, and when it was asked. It then opens every member of the price publication family whose bytes have not been read, and records what is actually inside.
It is deliberately slow — one request at a time, with a pause of nine tenths of a second, single-threaded — and entirely resumable. Those are the same decision. Fifteen thousand requests to a regulator’s website is a thing to do once, gently, and never again; a sweep that lost its place would do it twice. An id already answered is skipped without a request, so a second run against a complete catalogue makes no request at all.
A row is written only for a definite answer: a document, or a 404. A refusal, a server error, a timeout or a reset leaves the id unrecorded and the next sweep asks again. Recording an outage as an absence would answer that id permanently out of a row that was never about a document — and because the whole point of the catalogue is that it does not re-ask, nothing downstream would ever notice. Twenty consecutive answers of that kind stop the sweep outright, on the reasoning that the other end has changed its mind about us and the right response is to stop asking.