This page is not marketing — it is the product. For a closed score to be trusted in this field you need either a BlackRock-sized brand or an auditable methodology. The second is within our reach.
An official gazette is not what a state says, but what it does. It is published before it reaches the news, its date is exact, and it is binding. A speech shows intent; a regulation shows a committed resource. This system counts the second kind.
The first schema forced K/R as a binary: a document was either a capacity signal or a role signal. That was the model's most contestable assumption, and it was wrong. A “basing agreement” is both — it grows the host state's actual power and imposes an obligation on it. Under a binary schema one of those two opposed readings disappears.
Δ flow is a flow, not a stock. The Δ = K − R in the thesis layer is a stock: an actor's accumulated position, assigned by hand. The Δ measured on the dashboard is the net pressure produced within a window. They are not the same unit and sit in separate boxes.
Only their signs are compared. If the flow's sign matches the stock assumption it reads “supports”; if it is opposite, “contradicts”. This is exactly where the thesis is tested — a contradiction says either the assumption is wrong or the trend is turning.
Δ is hidden when label coverage is low. Δ is computed only from labelled documents; an unlabelled document contributes zero. At low coverage this does not merely make Δ incomplete — it makes it incomparable across time. If August is 100% labelled and March 20%, the series shows a fake August spike, precisely where fracture detection looks.
So below 85% coverage the figure is not shown as “small” — it is treated as invalid. The volume layer is unaffected: it needs no labels and is always complete, which is why fracture detection works regardless of labelling progress.
But coverage alone is not enough. Measured: in the United Kingdom's 90-day window, 16 of 16 documents were labelled — 100% coverage — yet only 9 of them were relevant. 100% coverage does not turn nine documents into a hundred. Δ now requires a second, independent condition: at least 12 relevant documents in the window.
The rest of the system measures states; this layer measures the instrument. It exists because of one finding: of 99 relevant labels, 49 carry r=0 and only 20 carry k=0. The labeller assigned zero on the role axis to half the documents, and on the capacity axis to far fewer. Since Δ = k − r, this is a systematic bias that makes every state look like a capacity builder — and indeed all four countries came out with positive Δ.
This is not a model error but a behaviour: “capacity” is something visible (an authority, a budget, a permit), whereas “role” is a theory-laden judgement. Writing zero on the axis you are unsure about is a reasonable reflex — but done 99 times in a row it produces an artefact, not a measurement. A human cannot do this; they get tired and ask “am I always writing zero in this field?” A machine does not ask.
On 20 August 2026 the same 40 documents were labelled blind twice more: once by the same model (test–retest), once by a different model (inter-rater). The sample is seeded, so all three rounds cover exactly the same documents.
| κ (quadratic weighted) | test–retest | inter-rater |
|---|---|---|
| relevance | +0.754 | +0.441 |
| field | +0.847 | +0.908 |
| k (capacity) | +0.773 | +0.779 |
| r (role) | +0.902 | +0.798 |
| k drift (per document) | −0.45 | −0.44 |
The expectation was that the role axis would come out weak, being a theory-laden judgement. Measurement said the opposite — r is more reliable than k in both rounds.
But this does not solve the r=0 problem — it changes its meaning. Across all four measurements, 55–60% of documents received r=0, and two different models write zero on the same documents. So r=0 is not noise but a stable default.
A caveat: two instruments from the same model family reading the same instruction text is not independent evidence. Whether the zeros come from the documents or from the definition of r can only be settled by a human round. Reliability is not validity: an instrument that makes the same error consistently scores high.
On the capacity axis both rounds scored 0.45 points lower per document than the base round. Two independent instruments drifting the same way suggests the problem lies not in the noise of a round but in the base round itself. That is not scatter but directional drift, and across a 45-document window its total exceeds Δ itself. So the two failure modes are tested separately:
Combining the two into one band was the first version's flaw, and measurement exposed it: Turkey's 90-day band came out as [−18.33 … −5.45] — the band excluded its own point estimate (+2.41). The figure was not wrong, its meaning was: the band was asserting “Δ's true value is −12”, and there is no basis for that claim; we do not know which round is right. Separating them is more conservative: a narrow band plus a large drift means the instrument is consistently unstable, and a reading that looks only at the band would call it sound. That is the most misleading case of all.
The band is not produced by a second implementation of Δ, but by calling the production aggregation function again with perturbed labels. A separate formula would silently drift from the very figure it claims to measure — the same decision as in the time machine.
The weakest link is not the scores but the relevance decision. With a different labeller, 8 of the 33 documents that entered Δ in the base round dropped out entirely (24%). That is a far larger lever than ±1 point shifts: a shifted score moves Δ by one unit, a dropped document erases its entire contribution. In the first version this sat on the “not modelled” list; a measured but unmodelled figure is more dangerous than an unmeasured one, because it is assumed to have been accounted for. It is now inside the band.
The band is fed by the inter-rater round when one exists; the test–retest round is a known lower bound and serves only as a fallback, labelled as such in the dashboard.
What is not modelled (the irrelevant→relevant direction, disagreement over field assignment, documents the filter never surfaced, and systematic bias shared by both labellers) widens the band further, so the published band is a lower bound on reality.
Fetch a source once, never fetch it again. The raw document, its download timestamp and its hash are kept permanently. The reason: when the taxonomy changes the archive is reprocessed, the source is not refetched. If a source site deletes its history or moves behind a paywall — as happened with Japan's Kanpō archive — this is all that remains. The archive is this project's real asset, not the model output.
Country colours on the globe are role archetypes: hegemon, revisionist, rising, bloc member, proxy, double-bound, isolated, buffer, periphery. If the thesis claims that states are positioned by their assigned role rather than by their power, then that is what the map should carry.
These assignments are assumptions, not measurements. The measured layer is separate and looks different on the globe: amber pillars (decision volume) and targeting links.
A single assignment is a simplification. India is both rising and double-bound, Iran both revisionist and isolated, Ukraine both buffer and proxy. Only the dominant one is shown — the same limit the binary K/R choice carried, and it must be stated as plainly.
Two images are produced: one to be looked at (archetype colours) and one to be read (each country filled with a unique code colour, no antialiasing). Clicking resolves the country by reading a pixel at the hit point's UV — running a point-in-polygon test in the browser across 177 countries would be both slow and wrong at the antimeridian.
| country | source | class | licence |
|---|---|---|---|
| United States | Federal Register (JSON API) | A | Public domain — 17 U.S.C. §105 |
| Türkiye | Resmî Gazete (HTML/PDF) | C | Unclear — legal opinion needed before republication |
| United Kingdom | legislation.gov.uk (Atom) | A | Open Government Licence v3.0 |
| Poland | Dziennik Ustaw (ELI API) | A | Official legal text — not copyrightable |
Class A: official API or bulk download · B: regular structure, no API · C: scraping + PDF/OCR · D: restricted access. Turkish content is archived and enters the measurement, but its full text is not republished here; only the title, source link and derived score are shown.
The Gazette was tried first for the UK, and measured: of the 516 notices published on 19 August 2026, 484 were corporate insolvencies, personal insolvencies and probate notices. The wrong source for measuring state behaviour. The real counterpart of the Federal Register is the statutory instrument stream — legislation.gov.uk.
A second trap found there: the feed returns 20 records per page and hides the rest
behind a rel="next" link. Without pagination, every day with more than 20 items
silently lost the remainder — more insidious than returning nothing, because partial data
looks entirely normal. It was caught by noticing that the daily maximum was exactly 20 on nine
separate days: a number repeating at the top of a distribution is a ceiling, not data.
This layer's job is not to decide correctly. Its job is to keep the obviously irrelevant half of the hundreds of daily documents away from the expensive layer. Its threshold is therefore deliberately loose: a wrong elimination is expensive, a wrong pass costs a few cents.
1.0;
if no positive term appears in the title the threshold rises to 2.2.
Every matched term is recorded — which decision was made and why stays auditable.Negative terms do not eliminate a document, they lower its score. A strong positive term can beat a negative one.
The raised body threshold was measured, not guessed: in the first version only 39% of passing documents were genuinely relevant, and the dominant failure was documents whose body mentioned a topic term while the title was about something else entirely — a car-rental regulation passed because the word “sanction” appeared in its text. Gazette titles are deliberately descriptive.
There is no ground truth; the hand-labelled gold set is the only truth we have. The filter is measured against it at every version. The two error types are not equal: a miss is expensive (that document never reaches the labelling layer), a false pass is cheap.
| version | precision | recall | what changed |
|---|---|---|---|
| r1 | 39% | 100% | first version — passes everything |
| r2 | 90% | 56% | negative list + body threshold; cut too much |
| r3 | 89% | 95% | gaps in the positive list closed |
r2's collapse was not caused by the threshold but by omissions in the positive list:
antidumping existed only in its hyphenated form (anti-dumping) and the
Federal Register does not hyphenate it. A single hyphen missed ten documents. Without the gold
set this would have been invisible — which is the whole point of having one.
Live calibration figures are in the dashboard's system tab, alongside the values actually in production.
In Türkiye, presidential decisions, international treaties and board rulings are published as scanned PDFs. Their content is an image; they appear to have no text layer. Our first threshold was simple: “a PDF yielding fewer than 180 characters is scanned.” It was wrong.
We measured the distribution across 83 PDFs and it was bimodal: Poland's genuine text PDFs yield 3,400–4,200 characters per page, Türkiye's scans 4–206. Nothing in between.
206 is not a coincidence. A scanned page does carry a text layer — but it is not content, it is the masthead: “13 August 2026 THURSDAY · Resmî Gazete · No: 33339”.
A total-character threshold mistook that boilerplate for content. A 22-page scanned court ruling counted as “has text” on the strength of 459 characters and never entered the OCR queue. Silently. Of 83 PDFs, 49 were effectively scanned while only 37 were flagged.
400
Measured per page, not in total. A wrong OCR is cheap (a few seconds of processing);
a missed OCR silently becomes a wrong label.Human review was initially triggered by the document average. Looking at the output showed that this was wrong: the Türkiye–Saudi Arabia visa-exemption agreement runs to 17 pages with an average confidence of 58.8. But the decision text is on the first page and is clean — the low score comes from the maps and coordinate tables on the following 16 pages.
70
Not the document average. In a gazette the decision text is always on the first page;
failing to read the annexes is not failing to read the decision.OCR output does not alter the raw archive; it is derived data, stored separately with its confidence. If the engine or language changes it is regenerated without returning to the source. The Turkish language pack is mandatory: without ğ, ş, ı, İ, ö, ü, ç the output silently breaks keyword matching.
Windows are divided by the number of days available, not by their nominal length, and fracture is not computed until the archive is at least 30 days deep. The earlier 21-day rule produced an artefact visible in the time series: Turkish fractures began exactly on the archive's 21st day, because a 30-day baseline computed from 21 days of data and still divided by 30 is roughly 30% too low, inflating the ratio by about 43%.
Gazette sites change structure without notice. When the ingest layer breaks it must not silently return empty: “zero documents today” is an alarm, not a normal day.
But some sources genuinely do not publish. The Federal Register does not publish at weekends or on federal holidays; legislation.gov.uk produces 0–9 documents a day. If we cannot tell the two apart we either miss real failures or get false alarms twice a week and stop reading alarms — the second being more dangerous. The calendar is therefore source-specific.
For Türkiye no fixed holiday calendar is hard-coded. On religious holidays the index page returns HTTP 200 with a notice instead of documents: “pursuant to Presidential Decree No. 10, the Official Gazette is not published today.” That notice is read directly — if the calendar changes the code need not, and “not published” never gets confused with “the parser broke”.
The claim is not “I measure the world correctly.” The claim is: I convert state behaviour into a traceable unit, and I show exactly how I convert it.
The taxonomy, weights and decay coefficient on this page are the values actually used in production — not a separate marketing text. Back to dashboard · thesis model · Türkçe