Verification

Predictions you can check. Recorded before the outcome.

Every SigningLab public forecast is committed with a cryptographic hash while the season is still open, and audited afterwards through a random sample that no one, including us, can choose. This page publishes the live commitments, explains the protocol, and shows a reproducible retrospective protocol rehearsal on the 2024 and 2025 seasons. Nothing here reveals the model; everything here can be checked without trusting us.

Live commitments
1,985
Signing verdicts for the 2026 season, across 15 leagues, each committed as a salted SHA-256 hash — published now, while outcomes are still unknown. Generation SL 2026.1 · committed 17 August 2026 · payloads contain only the binary verdict and public identifiers, never model internals.
Download: commitments_sl2026_1.json  ·  committed population (keys, no verdicts)  ·  key → commitment map  ·  protocol addendum v1.1  ·  verification README
Each entry is sha256( league|season|player|club|verdict|generation + "|" + salt )
File SHA-256: 311dbed3747f30f7014dd15a9589f6365a2c8ae8cec0e146f55fc94d389411f8

The salt (16 random bytes per record) prevents anyone from enumerating verdicts out of the hashes before the reveal. Altering, adding or removing a single verdict after publication would break the corresponding hash.

The protocol

Commit now. Sample by public randomness. Reveal what the coin chose.

Step 1 · Commit

Hashes before outcomes

While the season is open, every published verdict is committed as a salted SHA-256 hash. The commitment file is public and its own hash is fixed, so the set cannot be edited afterwards without detection.

Step 2 · Public randomness

A coin we cannot flip

On the announced audit date, the hash of the first Bitcoin block mined after that date becomes the sampling seed. No one can predict or influence it — not us, not anyone. The sample selection is reproducible by anybody with one line of code.

Step 3 · Sampled reveal

Statistical proof

We reveal payload and salt for the ~100 sampled entries only. Because no one knows in advance which entries will be drawn, every commitment must be honest: falsifying even 5% of the set would be caught with roughly 99% probability. A 100-case sample is fraud protection, not a precise accuracy estimate (about ±7pp on the aggregate); larger samples and full audits are available to qualified reviewers under NDA.

Why a sample and not the full set? Because publishing every labeled verdict would hand out exactly the training data needed to imitate the model. The sample gives auditors statistical proof of integrity while the verdict set itself remains the commercial product, available to clients in the app.

Announced audit anchor for generation SL 2026.1: the first Bitcoin block with timestamp on or after 1 July 2027, 00:00 UTC that remains in the canonical chain at a depth of 144 blocks (reorg protection), once the 2026 and 2026/27 season outcomes are complete. Selection rule, fixed now: score = sha256("SL-SAMPLE-v1|" + block_hash + "|" + commitment_hash), select exactly K = min(100, N) smallest scores over the deduplicated union of all chained batches published before the cutoff — computable by anyone from the public commitment files alone; a cumulative manifest of every batch is published before the beacon. Mandatory order: before the beacon is read, the outcomes of the full population are computed, frozen, and committed as a Merkle tree whose root is published on this page; each sampled entry is later revealed with its Merkle inclusion proof, so any auditor can verify that that specific outcome was already frozen — so outcomes cannot be interpreted differently after the sample is known. Only then are payload, salt and outcome revealed for the sampled entries.
Protocol rehearsal

A retrospective rehearsal, run on ourselves.

Before asking anyone to trust the protocol, we ran it against our own published seasons. The seeds are the first Bitcoin blocks of January 2025 and January 2026 — dates anchored to the seasons, fixed by the blockchain, and outside our control. From each season's published verdicts with known outcomes, the seed deterministically selects 100 cases: the 100 smallest values of sha256(block_hash|season|league|player|club). Anyone can re-run the draw and get the same 100 names.

SeasonSeed (Bitcoin block)Sample accuracyFull published setSample mixRecommendations rightRejections right
2024 #877259 · first block after 1 Jan 2025 UTC 86% (95% CI 78–91) 87.2% ✓ 46 YES / 54 NO 83% 89%
2025 #930341 · first block after 1 Jan 2026 UTC 86% (95% CI 78–91) 85.6% ✓ 46 YES / 54 NO 76% 94%

Two things the samples show. First, both landed inside the statistical interval of the full published set — the sampling estimator works. Second, almost half of each random sample is a positive recommendation: this is not a model that hides in "no" to farm easy accuracy. It recommends, and its recommendations hold up.

Download the sampled cases and the draw metadata:
sample_2024.csv  ·  sample_2025.csv  ·  draw_metadata.json  ·  full_population_2024.csv  ·  full_population_2025.csv
The full-population files (keys only, no verdicts or outcomes) let anyone confirm the 100 drawn cases really are the smallest scores of the whole set.

Honest limitation: the 2024 and 2025 verdicts were not hash-committed at the time they were made, so this retrospective rehearsal demonstrates the mechanism and the numbers, not proof of anteriority. Anteriority proof begins with the SL 2026.1 commitments above, which exist today while their outcomes do not.

The number that matters

Judge us on the players we say yes to.

Overall accuracy can flatter any cautious forecaster, because in our scoring a "do not sign" call is not counted as wrong when the deal ends neutral — advising against a signing that turns out merely average is not a mistake. So the honest test of a recommender is the side no cautious strategy can fake: the recommendations themselves. Across the published 2024 and 2025 seasons, the model issued 3,429 positive recommendations.

Strict success

35.4% vs 19.1%

Signings we recommended clearly outperformed their investment baseline at 1.85 times the rate of the market as a whole.

Failure rate

21.6% vs 49.2%

Recommended signings failed at less than half the base rate of all evaluated signings.

Missed gems

8.3%

Of the 5,130 signings the model advised against, only 8.3% went on to be strict successes — and 67.6% failed outright.

Verdict × outcome (2024+2025, n=8,559)Strict successNeutralFailureTotal
Recommended1,213 (35.4%)1,474 (43.0%)742 (21.6%)3,429
Advised against425 (8.3%)1,236 (24.1%)3,469 (67.6%)5,130
Whole population1,638 (19.1%)2,710 (31.7%)4,211 (49.2%)8,559

For completeness, the aggregate numbers on this same population and ruler: overall accuracy 86.4%; a strategy that rejected every signing would score 80.9% on this ruler (because neutral counts in favor of a negative call), an absolute gain of +5.5 points. We publish that baseline ourselves because aggregate accuracy flatters cautious forecasters — it is exactly why the recommendation-side numbers above, which no reject-everything strategy can imitate, are the figures we ask to be judged on.

Population: the signings SigningLab's models evaluate in its 21 tracked leagues, published seasons 2024 and 2025, outcomes measured at season close. The market hit rate quoted elsewhere on this site (~36%) uses the market convention — full credit for success, half credit for neutral — over all signings; it answers a different question than the model-accuracy figures and the two should not be subtracted from one another.

Boundaries

What stays private, and why you can still check us.

The model's architecture, features, weights, scores and per-league thresholds are proprietary and are neither committed nor revealed — they are the product of years of research and the basis of the service clubs pay for. None of that is needed to audit us: the commitments prove the verdicts existed before the outcomes, the public beacon proves we cannot choose what gets checked, and the sampled reveal lets anyone compare what we said against what happened.

Journalists and researchers who want to go deeper can request a supervised audit session — the same protocol, a larger sample, under NDA. Write to [email protected].