eklitrePrice alerts

Reading OGRA’s kerosene price scans

OGRA publishes kerosene’s build-up as photographs of paper. This is how a machine was made to read them, what checks a figure has to pass before it is published, what is deliberately left unread, and what this kind of reading cannot tell you.

Published 22 August 2026.

In this note
  1. The rule this record is built onNot a label on a figure. Separate tables, a separate file, a separate endpoint, and a mark on every day.
  2. OGRA is inconsistent about how it publishes theseA minority are typeset. Those are read exactly and carry no caveat at all, which is what makes the caveat mean something.
  3. What corroborates a figureThree identities on the page itself, and five quantities that must agree with a second document. Both, or the day is withheld.
  4. What is deliberately not readThe most valuable page in the archive is refused, and the reason is that it is the most valuable.
  5. How the reader was chosenTwo engines, one page, twelve figures, and why the version that read a figure is recorded beside it.
  6. What this cannot tell youThree limits of the measurement, stated because a rate quoted without them is a rate quoted wrongly.
  7. Where the documents are keptArchived, checked against the recorded digest on every fetch, and stated per document where a copy is not held.

OGRA publishes no machine-readable build-up for kerosene. What it publishes is an ex-depot sale price calculation and the gazette notification that puts the same prices in force, fortnightly, and it publishes most of them as photographs of paper: one JPEG a page, no text layer, nothing a parser can extract. This is the only part of this project whose figures are not read exactly, and everything below is about not letting that fact get lost.

The counts here were measured on 21 August 2026, over the record as it then stood. They are not live and they will not update: the current ones are on the scanned record, which reads them from the record at the moment you open it.

The rule this record is built on

Every other figure this project publishes was read from a text layer. A machine opened a document, asked for the characters it contains, and got them back exactly. Two runs return the same bytes; the parse either finds a figure or refuses.

An OCR figure is not that. It is a model’s opinion about what some ink says, and no amount of care makes the two the same kind of fact. So they are kept as a separate class, in four ways at once:

Their own tables
Five of them, holding the days, the readings, the documents, the messages and the second opinions — and not one row of any of them joins to the day record. There is no foreign key from these into it and none the other way, and that is asserted against the database’s own catalogue rather than left to code that could change, so it stays true whatever is done to the code.
Their own file and their own endpoint
data/scanned-buildup.json and /v1/scans. Nothing from here is in history.json, latest.json, daily.csv or any day record. /v1/meta counts them in a separate nested object, so that a caller adding up the numbers beside it cannot include these by accident.
A mark on every day, not a header
Each day carries reading, which is ocr or text_layer. A header is read once; a day is read every time.
A word that is unreachable from here
verified belongs to the text-layer record. A scanned day is corroborated, uncorroborated or withheld, and no code path can give it any other status however confident the reader was.

A label alone would not have been enough. The failure worth designing against is not somebody mislabelling a figure — it is a figure ending up inside a series somebody then quotes, at which point the label is three joins away and nobody is looking at it. A label can be dropped in transit. A separate table cannot.

OGRA is inconsistent about how it publishes these

Not all of these documents are photographs. OGRA typesets a minority of them, apparently at whim, in the same family and under the same filename pattern: 16 October 2024 is a real PDF with selectable text, and 1 September 2024 and 19 August 2026 either side of it are photocopies.

Where there is a text layer, this reads it and guesses at nothing. Those days are marked reading: "text_layer" and their figures are exactly what the document states — the same standard as everything else this site publishes. They carry no caveat at all. Recording them as OCR output would have been a lie in the cautious direction, which is still a lie, and a warning on a figure that does not need one teaches a reader to ignore the warning on figures that do.

The free measurement this seemed to promise does not exist. The hope was that on some dates one document would be typeset and the other photographed, which would give a known right answer and a model’s attempt at it side by side for nothing. It never happens: OGRA typesets both documents for an effective date or photographs both. The nearest cases are three dates where a typeset calculation sits beside a photographed notification, and there the model read the check rather than the figures — which measures something much less interesting.

That is a finding about how a regulator publishes, and it is not an excuse. It is why the right answer had to be produced by somebody reading the pictures by hand.

What corroborates a figure

Two checks. A day passes both or it is withheld — there is no partial publication and no per-figure escape, because a table that failed its own arithmetic is a table whose every cell is under suspicion. The check does not say which digit moved.

The page’s own arithmetic must close

The ex-depot table is a chain:

RowWhat it must equal
Ex-refinery
IFEM
(unlabelled)ex-refinery + IFEM
Distributor margin
Dealer margin
Petroleum Levy
Price before Sales Taxthe subtotal + both margins + the levy
Sales Tax
Max. ex-depot sale priceprice before sales tax + sales tax

Three identities, and every figure appears in two of them. Change the 9 in 269.95 to an 8 and three break at once. A misread digit almost never still sums. This is the same machinery and the same argument that catches a misparsed build-up in the text-layer record, and it is the reason Annex-I is deliberately not read.

A second document must agree

OGRA publishes two documents for most dates, photographed separately and read separately here. They state five quantities in common, and all five must agree:

QuantityCalculationGazette notification
IFEMrowcolumn (4)
Petroleum Levyrowcolumn (3)
Distributor marginrowcolumn (6)
Max. ex-depot sale pricerowcolumn (8)
Prescribed priceex-refinery + levy + both marginscolumn (2)

The last is the most valuable of the five: it ties four cells of one photograph to one cell of another, so a misreading in any of the four shows up even on a page whose own arithmetic happened to close.

The four direct comparisons must be exact. The rest of this project compares figures to within a paisa, because published figures round. That reasoning does not carry here: both documents print the same quantity to the same two decimals, so anything but equality means one was misread — and a last digit read as its neighbour, which a paisa of slack would hide exactly, is the likeliest OCR error there is. The compound comparison keeps the tolerance, because four figures each rounded to a paisa can legitimately sum a paisa away from one rounded once.

Where nothing can check it

The day is kept and marked unanswered — uncorroborated, with corroborated_by null. That is not “cleared” and must never render as it, which is why it takes a neutral tag on the scanned record rather than the sage a confirmation wears or the terracotta a refusal wears. The same distinction is load-bearing over the notified price series, where a contradiction with no corroboration means nobody answered it rather than somebody agreeing.

What is deliberately not read

Annex-I, the ex-refinery sale price calculation, on the first page of the same document. It carries the Arab Gulf FOB quotations for ten working days across eight products, the daily exchange rate, the premium, the 7.5% tariff, the conversion factor and the import parity price. It is the top of the petrol and diesel build-up and it is the most interesting thing in the entire scanned archive.

It is not read because its arithmetic does not close on the page. The ten daily quotations average into a figure the page states, and that average is the only check available; a misread digit in one of eighty quotation cells moves the mean by a hundredth and would pass unnoticed. Reading it would mean publishing figures whose only corroboration is the confidence of the model that read them — which is the standard this whole exercise exists to refuse.

It stays unread until a second source exists. It was declined deliberately — it was not overlooked. That sentence is here because the opposite assumption is the natural one: Annex-I sits on the front of documents this project already downloads, parses and archives, so anyone finding it unread will assume nobody got to it. The reason it is refused is precisely that it is valuable. It is the one place in this archive where OCR is most dangerous, not least: everywhere else a misread digit breaks a sum and the day is withheld, and there it would be published, confidently, with nothing to catch it.

How the reader was chosen

Tesseract was tried first. On the ex-depot table it returns rubbish: a row RapidOCR reads as Distributor margin 1.58 comes back from Tesseract as

Measured on the same page, RapidOCR read every one of its twelve figures correctly and Tesseract read none of them. Tesseract also needs a system second piece of software alongside it, which is one more thing that can differ from one machine to the next and quietly change what a page reads as. RapidOCR carries its own models, so it does not.

Which version read it is part of the reading. A machine that reads ink returns an interpretation rather than the bytes it was given, and a different version of it is a different interpretation of the same ink. So the engine and its version are recorded against every figure read with them, and one version is used throughout. A figure that would change under a newer reader is then something this record can show you, rather than something that quietly happened to it.

What this cannot tell you

The sample is not random. It is the dates OGRA happened to typeset one document and photograph the other. If OGRA’s scanner or its typesetting habits changed during the period, this measures whichever side of that the sample fell on. The sample size is reported with the figure and the figure should not be quoted without it.

It is blind to a correlated error. If the model misreads the same digit the same way on both documents for one date, that date agrees with itself and is published. The two are photographs of different pieces of paper set in different type, so it needs a coincidence — but “needs a coincidence” is not “cannot happen”.

A disagreement does not say which side was wrong. Where the typeset document is the notification and the photograph is the calculation, the model read the figures, so a disagreement is the model’s error and is counted as one. The reverse case is counted separately: there the model read the check, and an error costs a corroboration rather than corrupting a published figure.

And one thing this note is not about. What OGRA itemises in these documents is the maximum ex-depot sale price, which is not the pump price the rest of this site draws. They are a weaker kind of knowledge about a different quantity, and both halves of that matter.

Where the documents are kept

They are archived. The fifth rule of this project’s method is that the source is kept rather than linked: provenance by digest alone leans on OGRA continuing to serve a document, which is true today and guaranteed for no day after that, and that is exactly the dependency an archive exists to remove. So a copy of each document is kept here, and a figure can be checked against the page it was read from rather than against a link that may one day stop answering.

Not every copy is here yet, and the page tells you which. Where a document is held you can open it. Where it is not, the scanned record and the archive say so in as many words. Neither shows a link that would not resolve, and neither leaves you guessing which kind of row you are looking at.

Every fetch is checked against the record

A document is only written to the archive if its SHA-256 matches the digest recorded when the figures were read from it. A document that comes back different is not written: nothing is overwritten and the mismatch is recorded as a finding. OGRA having replaced a document this record quotes is a finding, not a download.

That check matters more here than anywhere else in this project. Every other figure could be re-derived by re-reading a text layer. These could not — if the picture changes underneath them, the readings describe a document that no longer exists at that address. So the check has been run over the whole archive: on 21 August 2026 every one of the 145 documents the archive then held was fetched again and every one matched the digest recorded for it. That is more documents than the readings cite — the archive counts the ones a published figure came from, and this counts everything kept, including the documents behind days that were withheld and dates that produced no reading at all.

Read next

  • The state’s share of a litreAbout Rs 106 of a Rs 349 litre of petrol — near a third — is the state: petroleum levy, climate support levy, customs duty and sales tax, on petrol and diesel alike. Each line itemised from OGRA’s build-up, and how the 2026 IMF measures built it.
  • The plan to stop setting the petrol pricePakistan set a tentative June 2027 target to deregulate petrol prices and let oil companies set the pump rate, with diesel to follow — what it would change for OGRA’s build-up.
  • The price here, and the price at the pumpOGRA’s notified price is the legal maximum, the dealer’s margin already inside the build-up, so a pump charging more than the petrol or diesel figure here is overcharging.

Know the new price before you reach the pump

OGRA notifies late at night. We post each new price as soon as we read it.