Public Evidence Chains for Rating Lists
A public evidence chain lets readers move from a rating claim back to the games, rules, decisions and publication stages that produced it.
A computer chess rating list should not be treated as an isolated table of engine names and Elo values. Every published rating is the end product of a longer process: engines are selected and configured; games are scheduled; openings are assigned; clocks and hardware are controlled; incidents are handled; results are written into PGN records; completed files are validated; event identities are closed; and a rating program transforms the accepted results into estimates of relative strength.
When this process is publicly documented, readers can move backwards from a rating claim to its supporting evidence. They can identify the relevant tournament or testing pool, inspect the downloadable games, understand the conditions under which they were played, distinguish provisional material from closed records and determine whether a published Elo value is being interpreted within its proper methodological limits.
That path is the public evidence chain.
For IJCCRL, this concept is more important than the publication of any single table. The project’s strongest methodological position is not that it produces one more ranking among many established computer-chess resources. Its distinctive position is that the different public surfaces of the IJCCRL Arena Series can operate as one connected publication system: live tournament observation, portable PGN evidence, public downloads, closed-event archives, winners registries and separate rating surfaces.
The result should be an architecture in which a reader does not have to accept an Elo figure on trust alone. The reader should be able to ask:
- Which games support this rating?
- Under what rules were those games played?
- Which engine versions were used?
- Was the event still provisional or already closed?
- Were crashes, timeouts or adjudications involved?
- Where is the downloadable PGN?
- Does the winner’s identity correspond to a conclusively closed final?
- Does the rating table include the event, exclude it or use only part of it?
- Has the historical record changed since the first publication?
A serious evidence chain does not guarantee that every methodological choice will be universally accepted. It does something more practical: it makes those choices visible enough to examine.
1. A Rating Table Is the Last Surface, Not the First
The visual simplicity of a rating list can conceal the complexity of its origin.
A published table may show:
- rank;
- engine name;
- version;
- Elo rating;
- number of games;
- score;
- draw percentage;
- average opponent;
- error estimate;
- likelihood of superiority;
- or another statistical indicator.
These fields are outputs. They are not the original evidence.
Before a number can appear in a rating table, games must exist. Before games can exist, a testing environment must be defined. Before the results can be accepted, the publication system must decide which games count, which do not count and how exceptional cases are handled.
This distinction is visible in established computer-chess resources. CCRL describes its work as running games between chess programs, collecting those games into a database and then computing and publishing a rating list. Its public structure includes testing conditions, downloadable PGN collections, historical data and rating tables rather than presenting Elo as a context-free number. 5 surface, for example, identifies its time-control equivalence, opening-book policy, pondering status, tablebase conditions and rating method. It also exposes the game volume supporting the table and distinguishes engines with larger and smaller samples. Those disclosures matter because they describe the environment in which the rating estimates have meaning. ates a related architecture from the tournament side. Its public system connects live viewing with schedules, crosstables, event information, PGN material, archived stages, engine data and documented rules. Its archive is not merely a list of champions; it is part of a wider event record. shed resources have different purposes, histories and methodologies. IJCCRL should not present its own workflow as a replacement for them. The methodological lesson is narrower and more useful: a rating claim becomes more intelligible when it remains connected to the evidence-producing process.
2. What a Public Evidence Chain Contains
A public evidence chain is a sequence of connected publication objects.
At minimum, it contains six layers:
- The live or source event
- The PGN game record
- The public download layer
- The closed historical archive
- The winners registry
- The rating publication
These layers answer different questions.
The live event answers:
What is happening now?
The PGN answers:
What moves, headers and results were recorded?
The download layer answers:
Where can the public obtain the evidence?
The archive answers:
What historical event did this material belong to?
The winners registry answers:
Who conclusively won the closed competition?
The rating list answers:
What estimate of relative performance emerges from the accepted game base?
Confusion begins when one surface is asked to perform all six functions.
A live crosstable is not yet a permanent archive. A ZIP file without context is not yet a complete historical record. A winner’s page is not a rating list. A rating table is not a substitute for its underlying PGN. An archive entry should not reproduce an entire downloads inventory. Each surface should remain specialised while linking to the others.
This division of labour makes the publication system easier to audit and easier to maintain.
3. The Live Event as the Source Surface
The live event is where the evidence chain begins.
A modern chess-engine broadcast may expose:
- the board position;
- move history;
- clocks;
- engine identities;
- evaluations;
- principal variations;
- depth;
- selective depth;
- nodes;
- speed;
- tablebase hits;
- standings;
- pairings;
- schedule;
- opening code;
- opening source;
- time control;
- hardware or resource configuration;
- and incident messages.
Not every field needs to be permanently stored in the final PGN, but the live surface establishes the immediate context in which the game was produced.
The IJCCRL Broadcast is designed as the routing surface for two distinct competitive environments: the Original UCI Track and the Derived Stockfish Track. The public architecture currently separates those lanes and connects them to Events, Archive, Winners, Downloads and the rating surfaces. on is part of the evidence chain. It is not merely a branding preference.
When two materially different engine populations are mixed without clear labelling, the historical meaning of their results becomes harder to interpret. A reader may not know whether a rating belongs to an original-engine field, a derived-engine field or a combined experimental pool. By preserving track identity at the live source, IJCCRL can carry the same identity through the PGN, download pack, archive entry, winners registry and final rating surface.
What the live surface can prove
A live surface can provide evidence that:
- a game was publicly observable;
- the announced engines were connected;
- moves appeared in sequence;
- a clock was active;
- a tournament stage was in progress;
- standings and schedules were associated with the event;
- and particular public metadata was displayed.
What the live surface cannot prove by itself
A live surface alone cannot prove that:
- every displayed game was later accepted into the rating database;
- no PGN correction was required;
- an apparent engine identity was the exact binary eventually documented;
- an incident had no effect on the final status;
- a provisional crosstable became the final official result;
- or the eventual rating calculation used every visible game.
The live event is therefore a source surface, not the final authority for every later claim.
4. PGN as Portable Evidence
PGN is the most important portable layer in the evidence chain.
A live interface may disappear, change design or overwrite its current state when the next game begins. A PGN record can preserve the game independently of the original broadcast interface.
At its simplest, the PGN stores:
- players or engine names;
- event;
- site;
- date;
- round;
- result;
- move sequence;
- and optional metadata or comments.
For computer-chess publication, additional headers may be useful when reliably available:
- exact engine version;
- architecture or binary profile;
- time control;
- opening identifier;
- tournament stage;
- track;
- adjudication status;
- termination reason;
- hardware identifier;
- thread count;
- hash allocation;
- tablebase configuration;
- or source-pack identifier.
The PGN is portable because it can be:
- downloaded;
- opened in multiple chess applications;
- parsed by scripts;
- checked for malformed games;
- divided by event or engine;
- used by rating software;
- compared against later versions;
- archived independently;
- and reprocessed when a publication method changes.
CCRL’s public game surface illustrates the long-term value of this portability. It makes its game collection available in compressed PGN form and provides selections by month, engine, ECO code and opening. The rating table is therefore connected to a game base that can be inspected and reorganised independently of the table itself. idence, but not infallible evidence
A PGN file should not be treated as automatically correct merely because it exists.
Possible problems include:
- malformed headers;
- inconsistent engine names;
- duplicate games;
- truncated movetext;
- impossible results;
- missing termination reasons;
- wrong event names;
- mixed tournament stages;
- duplicated mirrored legs;
- incorrect dates;
- encoding problems;
- non-standard result strings;
- or games copied into the wrong pack.
A public evidence chain therefore needs PGN validation, not just PGN availability.
Validation may include:
- syntax checks;
- duplicate detection;
- legal-move parsing;
- result consistency;
- engine-name normalisation;
- event and round validation;
- expected game-count comparison;
- mirrored-opening pair checks;
- crash and timeout classification;
- checksum generation;
- and comparison with the public crosstable.
The stronger the validation process, the easier it becomes to distinguish a raw export from a publication-ready evidence pack.
5. Downloads as the Public Access Layer
Evidence that exists only on the organiser’s private machine is not public evidence.
The download layer converts internal records into publicly accessible material. It is the bridge between “the organiser has the games” and “the reader can obtain the games”.
The IJCCRL Downloads surface currently exposes event-related ZIP archives and PGN materials across Original UCI and Derived Stockfish competitions. The inventory includes league material, knockout stages, finals, provisional rating packs and other tournament records. load entry should disclose more than a filename. Ideally, it should identify:
- event name;
- track;
- time control;
- stage;
- date of publication;
- game count;
- file type;
- compressed size;
- publication status;
- and any audit note required for interpretation.
For stronger integrity, a download entry may also include:
- SHA-256 checksum;
- internal pack ID;
- PGN schema or header version;
- rating-inclusion status;
- superseded-by reference;
- correction date;
- and associated archive entry.
Why checksums matter
Suppose an event pack is published and later a malformed header is corrected. If the file is silently replaced under the same name, readers cannot know whether they possess the original or corrected version.
A checksum solves part of that problem.
For example:
Pack ID: IJCCRL-2026-ORIG-BLITZ-LS-001
Version: 1.1
SHA-256: [published checksum]
Supersedes: version 1.0
Reason: corrected engine-name header in games 41–42
Rating impact: none
The point is not bureaucratic complexity. The point is to make change visible.
Raw, audited and rating-valid downloads
Not every downloadable PGN pack has the same status.
A practical classification is:
Raw export
Directly produced by the tournament environment. Useful for immediate preservation but not yet fully checked.
Validated pack
Parsed, deduplicated and checked against expected event structure.
Audited pack
Validated and accompanied by incident, adjudication or exclusion notes where required.
Rating-valid pack
Accepted as input to a defined rating calculation.
Historical pack
Preserved for documentary reasons but excluded from the active rating base.
This vocabulary helps readers understand that public availability and rating eligibility are related but not identical.
6. The Archive as a Historical Object
An archive is not merely a folder where old material is placed.
A proper archive transforms a completed competition into a stable historical object.
The IJCCRL Archive defines its role as the canonical index for closed events, preserving event identity, track separation, time control, winner status, download routing and rating context without duplicating the full downloads inventory or Elo tables. ion is methodologically sound.
A download page answers:
Which files can I obtain?
An archive entry answers:
What was this event?
A useful archive entry may include:
- canonical event title;
- season or cycle;
- track;
- time control;
- format;
- stage;
- start and closure dates;
- engine field;
- scheduled and accepted game counts;
- final status;
- winner if applicable;
- links to evidence packs;
- link to the winners registry;
- link to the relevant rating surface;
- incident summary;
- and publication revision history.
Historical identity must remain stable
Event naming often changes during an active tournament. Internal filenames may contain abbreviations, spelling differences or provisional labels. Those variations are operationally understandable, but the archive needs one canonical identity.
For example, several files might refer to the same event using:
ORIG CLASSICAL LEAGUE
IJCCRL ELITE RR
CLASSICAL ORIGINAL TRACK 2026
II CLASSICAL ARENA SERIES LEAGUE
The archive should map those operational labels to one stable public title.
This prevents the same tournament from appearing to be four unrelated events.
The archive should preserve corrections
Historical integrity does not mean pretending mistakes never happened.
If an archive record is corrected, the responsible approach is to preserve:
- what changed;
- when it changed;
- why it changed;
- who authorised the correction;
- and whether the change affected results or ratings.
A correction log strengthens the historical object because it shows that the publication system can identify and repair errors without silently rewriting its past.
7. Winners as Closed-Result Identity
A winners registry answers a narrow but important question:
Who conclusively won the event?
That question should be answered only after the competition is closed.
The IJCCRL Winners page currently defines itself as the official winners registry of the IJCCRL Arena Series and requires strict separation between the Original UCI Track and Derived Stockfish Track. It also states that champions are published only after an event or final is closed, audited and ready for publication. tects the evidence chain from premature closure.
During a live event, one engine may lead the standings. During a knockout final, one side may be mathematically close to victory. A provisional post may discuss those facts, but the winners registry should not identify a champion until the result is conclusive under the event rules.
A winner is not the same as the highest-rated engine
Tournament victory and rating leadership answer different questions.
A winner’s page records:
- the outcome of a particular competition;
- under a particular format;
- against a defined field or opponent;
- during a particular period.
A rating list estimates:
- relative performance across an accepted game network;
- under defined calculation and inclusion rules.
An engine can win a final without becoming number one in the long-term rating list. An engine can lead a rating list and lose a short knockout match. Neither result invalidates the other.
The evidence chain should preserve this distinction.
What a winner’s entry should contain
A strong winner record may include:
- canonical tournament title;
- closure date;
- winning engine;
- exact version;
- author or project attribution;
- track;
- time control;
- format or stage;
- final score;
- opponent or field;
- archive link;
- download link;
- and audit status.
The entry does not need to reproduce every game or every rating. Its purpose is identity and closure.
8. Ratings as Estimates, Not Isolated Claims
A rating is an estimate produced from a defined result set.
It is not an intrinsic property permanently attached to an engine.
The same engine may receive different ratings when tested under:
- different time controls;
- different hardware;
- different thread counts;
- different opening suites;
- different tablebase conditions;
- different opponent pools;
- different version-normalisation rules;
- different game-exclusion policies;
- or different rating models.
This is why ratings from independent lists should not be casually compared as though they belonged to one universal scale.
The current IJCCRL rating lists hub explicitly separates Original UCI and Derived Stockfish tracks, divides publication by Classical, Blitz and Bullet, and distinguishes provisional or rebuilding surfaces from claims of scientific closure. essary part of a public evidence chain. Readers need to know not only the number, but the scope of the number.
Every published rating should answer five questions
1. Which games were used?
The table should identify its input base directly or indirectly.
2. Which games were excluded?
Crashes, timeouts, duplicates, malformed games, abandoned tests or incompatible events may require separate treatment.
3. Which rating method was used?
Examples include Ordo, BayesElo or another documented method.
4. How was the scale anchored?
A rating list may fix one engine, normalise the pool or use another reference convention.
5. What is the publication status?
The result may be:
- provisional;
- event-stage;
- experimental;
- under rebuild;
- closed;
- or historical.
Without those answers, an Elo figure can look more exact than the evidence permits.
9. Incidents, Adjudications and Exclusions Are Part of the Evidence
Public evidence chains must document abnormal cases, not only normal games.
Computer-chess events can encounter:
- engine crashes;
- GUI crashes;
- server failures;
- communication loss;
- illegal moves;
- stalled processes;
- time forfeits;
- disconnected remote hardware;
- corrupted PGN output;
- tablebase adjudications;
- manual adjudications;
- duplicate scheduling;
- wrong engine binaries;
- incorrect settings;
- or opening assignment errors.
These cases influence what enters the evidence base.
TCEC’s documented rules illustrate why incident policy belongs to the public architecture. Its rules distinguish time losses, server failures, remote-machine disconnections, adjudication conditions and engine bugs, and its rating policy states that certain categories of games are discarded. principle is not that every organisation must adopt the same rules. The principle is that the rules must be known before readers interpret the outputs.
Incident documentation should separate observation from decision
A useful incident record may contain:
Observed fact
Engine process stopped responding at move 63.
System response
Tournament manager recorded a timeout loss.
Audit classification
Engine-side failure; no host-level interruption detected.
Publication decision
Game retained in tournament standings but excluded from the long-term rating pool.
This structure prevents a factual observation from being confused with a methodological decision.
10. Public Does Not Always Mean Reproducible
Public evidence and reproducible evidence are related but not identical.
A system may publish:
- the PGN;
- event conditions;
- engine identities;
- rating method;
- and final table.
That gives the reader substantial public evidence.
Full reproduction may still be impossible without:
- exact binaries;
- exact neural-network files;
- source revisions;
- operating-system details;
- CPU and memory configuration;
- opening-suite version;
- random seeds;
- tournament-manager configuration;
- tablebase files;
- command-line parameters;
- and original logs.
A responsible publication should therefore describe its actual level of reproducibility.
A useful scale is:
Level 1 — Observable
The games were publicly visible.
Level 2 — Portable
The PGN is downloadable.
Level 3 — Auditable
Rules, incidents, versions and inclusion decisions are documented.
Level 4 — Recalculable
The accepted PGN and rating configuration allow readers to reproduce the published table.
Level 5 — Re-runnable
The complete environment is documented well enough to repeat the underlying tournament with materially equivalent conditions.
Most public computer-chess systems can realistically aim for Levels 2–4. Level 5 is much harder because software, hardware and operating environments change.
The publication should state the level it actually supports rather than imply perfect reproducibility.
11. The IJCCRL Evidence-Chain Diagram
The IJCCRL publication system can be represented as the following evidence chain:
ENGINE SUBMISSION / SELECTION
│
├── engine identity
├── version and binary
├── track classification
├── UCI settings
└── eligibility status
│
▼
EVENT DEFINITION
│
├── time control
├── hardware/resources
├── opening suite
├── mirrored-pair policy
├── tablebases
├── adjudication rules
└── tournament format
│
▼
LIVE EVENT / BROADCAST
│
├── moves
├── clocks
├── pairings
├── standings
├── engine telemetry
└── visible incidents
│
▼
RAW GAME RECORDS
│
├── PGN export
├── tournament-manager logs
├── incident logs
└── configuration snapshot
│
▼
VALIDATION AND AUDIT
│
├── syntax validation
├── legal-move parsing
├── duplicate detection
├── expected-game verification
├── mirrored-opening verification
├── identity normalisation
├── crash/timeout classification
└── exclusion decisions
│
▼
PUBLIC DOWNLOAD PACK
│
├── validated PGN
├── pack version
├── checksum
├── game count
├── status note
└── correction history
│
├───────────────────────────────┐
▼ ▼
CLOSED EVENT ARCHIVE WINNERS REGISTRY
│ │
├── canonical identity ├── champion
├── track ├── final score
├── time control ├── author attribution
├── closure status └── publication status
└── evidence routing
│ │
└───────────────┬───────────────┘
▼
RATING INPUT SET
│
├── accepted games
├── excluded games
├── engine-name mapping
├── version policy
├── rating method
└── scale configuration
│
▼
COMPUTER CHESS RATING LIST
│
├── Elo estimates
├── game counts
├── uncertainty
├── publication date
└── provisional/final status
│
▼
HISTORICAL REVISION LOG
The diagram is linear for clarity, but the real system contains feedback loops.
For example:
- PGN validation may reveal that an event result needs correction.
- An archive review may identify a duplicated pack.
- A rating run may reveal inconsistent engine-name mappings.
- A winner entry may need author attribution updated.
- A revised pack may require rating recalculation.
The evidence chain must therefore support controlled revision without losing the history of earlier publications.
12. Stable Identifiers Make the Chain Stronger
Human-readable titles are essential, but machine-readable identifiers make evidence chains more reliable.
Consider assigning identifiers to:
- events;
- event stages;
- engines;
- engine versions;
- opening suites;
- PGN packs;
- rating runs;
- and publication revisions.
For example:
Event ID:
IJCCRL-2026-ORIG-CLASSICAL-S2
Stage ID:
IJCCRL-2026-ORIG-CLASSICAL-S2-FINAL
Engine ID:
OBSIDIAN
Engine Version ID:
OBSIDIAN-DEV-16.15
Pack ID:
IJCCRL-2026-ORIG-CLASSICAL-S2-FINAL-PGN-V1
Rating Run ID:
IJCCRL-ORIG-CLASSICAL-RATING-2026-07-R3
A stable identifier solves several problems.
First, it prevents minor title variations from breaking the chain.
Second, it allows a rating run to point to an exact PGN pack rather than a generic event name.
Third, it allows the archive to preserve multiple revisions without pretending they are identical files.
Fourth, it makes automated publication safer. A script can verify IDs more reliably than prose titles containing punctuation, abbreviations or inconsistent date formats.
13. Evidence Chains Need Versioning
A public publication system is rarely static.
Engine versions change. PGN headers are corrected. New games are added. A rating list may be recalculated under updated inclusion rules. A tournament originally marked provisional may later be closed.
Every important object should therefore have a version or publication date.
PGN pack version
v1.0 — original audited publication
v1.1 — corrected event header
v2.0 — game exclusions changed after audit
Rating-run version
R1 — provisional League Stage
R2 — completed regular phase
R3 — corrected engine-name mapping
Archive revision
2026-07-10 — event closed
2026-07-12 — winner attribution expanded
2026-07-14 — rating link added
Versioning should distinguish cosmetic corrections from evidence-changing corrections.
A spelling correction may have no statistical effect. Removing a duplicated game may change the Elo calculation. The revision note should say which kind of change occurred.
14. Reader Paths Through the Evidence Chain
A well-designed system should support different types of reader.
The casual follower
The casual follower may begin with a winner.
Their path is:
Winner
→ event archive
→ final score
→ selected games or download
The rating reader
The rating reader begins with an Elo table.
Their path is:
Rating list
→ publication note
→ accepted game base
→ PGN pack
→ rules and exclusions
The engine author
An engine author may begin with an unexpected result.
Their path is:
Rating row
→ game selection
→ individual PGNs
→ event configuration
→ incident record
The historian
A historian begins with an old event.
Their path is:
Archive
→ canonical event identity
→ winner
→ downloadable pack
→ contemporary rating context
The technical auditor
An auditor begins with the evidence pack.
Their path is:
PGN checksum
→ pack version
→ validation report
→ rating-run configuration
→ published table
The same public system should support all these routes without forcing every page to contain every detail.
15. Common Breaks in an Evidence Chain
Evidence chains usually fail through disconnection rather than total absence.
Break 1: A rating table without game links
The table may be readable, but its evidence cannot be located.
Break 2: A PGN file without event identity
The moves exist, but readers cannot reliably determine the tournament, track or stage.
Break 3: A download pack with no version
Readers cannot tell whether the file has changed.
Break 4: A winner published before closure
The registry becomes vulnerable to reversals or tiebreak changes.
Break 5: An archive that duplicates the downloads page
The historical identity of events is buried under filenames.
Break 6: Mixed tracks
Original UCI and derived-engine results become difficult to interpret independently.
Break 7: Mixed time controls
Bullet, Blitz and Classical evidence is collapsed into a misleading general claim.
Break 8: Silent exclusions
The rating table uses fewer games than the public tournament produced, but the difference is unexplained.
Break 9: Silent corrections
A PGN or rating table changes without a revision note.
Break 10: Engine-name fragmentation
One engine version appears under several spellings, creating duplicate identities in rating software.
Break 11: Version collapse without explanation
Multiple materially different engine versions are aggregated under one name, but the list does not state the policy.
Break 12: Dead links
The evidence technically existed, but the path from claim to source has decayed.
These failures are avoidable through disciplined publication architecture.
16. A Practical Publication Contract
Every IJCCRL rating publication can be accompanied by a compact evidence contract.
For example:
Rating surface:
Original UCI Track — Blitz
Publication status:
Provisional regular-phase estimate
Rating date:
2026-07-22
Accepted game base:
396 games
Source events:
IJCCRL Arena Series 2026 — Original UCI Track — Blitz League Stage
PGN pack:
IJCCRL-2026-ORIG-BLITZ-LS-PGN-V1.2
Rating method:
Ordo [version/configuration stated]
Engine mapping:
One row per tested version
Excluded records:
2 duplicate exports; 1 aborted host-failure game
Opening policy:
Mirrored openings; each line played with colours reversed
Hardware:
Declared event configuration
Archive status:
Event closed / rating publication provisional
Revision history:
R1 initial publication
This contract can appear above or below the table without overwhelming the reader.
Its purpose is to answer the highest-value questions immediately.
17. Evidence Chains and Publication Quality
A public evidence chain also improves the quality of the editorial content surrounding the data.
An article about a rating list becomes more useful when it includes:
- original analysis;
- clear sourcing;
- a complete explanation of scope;
- evidence of practical expertise;
- and links to the records behind its conclusions.
Google’s public guidance on helpful content emphasises original information or analysis, comprehensive treatment, clear sourcing, first-hand expertise and content created primarily to help an intended audience. It also recommends making authorship and production context understandable where readers would expect them. r-chess publication, the evidence chain is a direct way to demonstrate those qualities.
The value does not come from repeating a target phrase. It comes from enabling the reader to complete a real task:
- understand the rating;
- locate the games;
- inspect the conditions;
- identify the event;
- distinguish provisional from closed material;
- and evaluate whether the claim is appropriately limited.
18. IJCCRL’s Position Among Established Resources
IJCCRL should describe its evidence-chain model as its own publication method.
It should not claim that:
- CCRL lacks evidence;
- TCEC lacks archives;
- established lists are obsolete;
- IJCCRL replaces existing international resources;
- or one workflow is universally superior.
CCRL has a long-established rating-and-database model with public testing conditions, downloadable games, historical tables and large-scale statistical output. ture tournament-and-broadcast model connecting live competition, event stages, rules, crosstables, PGN records and archived seasons. odological article should therefore make a restrained claim:
IJCCRL is developing a connected publication workflow for its own events, in which live observation, PGN evidence, public downloads, archive entries, winner identities and rating estimates remain navigable as parts of one system.
That is strong enough.
It is an original architectural position because it describes how the IJCCRL surfaces relate to one another. It does not depend on diminishing the work of other organisations.
19. Minimum Evidence Standard for an IJCCRL Rating Claim
Before publishing or materially updating a rating surface, the following questions should be answered.
Event definition
- Is the event identified canonically?
- Is the track stated?
- Is the time control stated?
- Is the opening policy known?
- Is the expected game count known?
- Are the participating engine versions known?
Game integrity
- Does the PGN parse legally?
- Are duplicate games removed?
- Are results internally consistent?
- Do mirrored pairs match the event policy?
- Are crashes and timeouts classified?
- Are excluded games listed?
Public access
- Is the accepted PGN downloadable?
- Does the pack have a version?
- Is its game count stated?
- Is a checksum available?
- Is the publication status visible?
Historical closure
- Is the event active, provisional or closed?
- Is there a canonical archive entry?
- Is the winner conclusively identified when applicable?
- Are corrections documented?
Rating calculation
- Is the rating method stated?
- Is the input pack identified?
- Is the engine-name mapping documented?
- Is the rating scale or anchor explained?
- Are uncertainty and sample maturity visible?
- Is the list described as provisional when necessary?
A rating claim that cannot answer these questions may still be useful as an early indicator. It should not be presented with stronger authority than its evidence supports.
20. Frequently Asked Questions
What is a public evidence chain in computer chess?
A public evidence chain is the documented path connecting a chess-engine claim to the records that support it. It may connect a rating table to an accepted PGN pack, event rules, incident decisions, archive entry, winner record and original live competition.
Is publishing PGN files enough?
No. PGN publication is a central step, but readers also need to know which event produced the games, whether the file was validated, whether any games were excluded and whether the pack was actually used for the published rating calculation.
Should every tournament affect the rating list?
No. A tournament may be historically valid without being methodologically compatible with a particular rating pool. Different time controls, tracks, hardware conditions, experimental formats or incomplete records may justify separate publication or exclusion.
Is a tournament winner automatically the strongest engine?
No. A winner’s registry records the outcome of a specific competition. A rating list estimates relative strength from a defined network of accepted results. The two surfaces answer different questions.
Why separate Original UCI and Derived Stockfish tracks?
The separation preserves the competitive and historical identity of materially different engine populations. It prevents results from being conflated and makes each rating surface easier to interpret within its intended scope.
Can a public rating list still be provisional?
Yes. Public means accessible, not necessarily final. A provisional list can be useful when it clearly states its game base, scope, uncertainty and incomplete status.
Why preserve old rating versions?
Historical versions show how the evidence base and estimates changed over time. They help readers distinguish genuine performance changes from methodological revisions, added games, renamed engines or corrected records.
What happens when a PGN error is found?
The file should be corrected through a visible version change. The publication should explain the error, identify the affected games and state whether the correction changes standings, winner identity or ratings.
Conclusion
A public evidence chain lets readers move from a rating claim back to the games and rules that produced it.
That simple principle has wide consequences.
It means the live event should expose enough context to identify what is being played. It means PGN files should preserve the games in portable form. It means downloads should provide stable public access. It means archives should transform completed competitions into coherent historical objects. It means winners should be published only after conclusive closure. It means ratings should be presented as estimates derived from declared evidence, not as isolated declarations of absolute strength.
For IJCCRL, the evidence chain can become the organising principle that connects all major publication surfaces:
Live event
→ PGN record
→ validated download
→ archive entry
→ winner identity
→ rating input
→ published estimate
→ historical revision
The strength of this system does not depend on claiming perfection.
It depends on visibility.
Readers should be able to see where the games came from, how they were handled, what was accepted, what was excluded and how the final rating claim was produced. When an error is corrected, the correction should be visible. When a list is provisional, that status should be explicit. When tracks or time controls are methodologically distinct, they should remain separate.
A rating table can always be copied.
An evidence chain is harder to reproduce because it reflects the architecture, discipline and historical continuity of the publication behind the table.
That is why public evidence chains matter for computer chess rating lists. They turn a numerical claim into a navigable body of evidence.
Sources and References
- TCEC Archive and public tournament interface
- CCRL 40/15 Rating List and testing summary
- CCRL testing conditions and methodology
- CCRL downloadable PGN game database
- Google Search Central: Creating helpful, reliable, people-first content

Jorge Ruiz Centelles
Filólogo y amante de la antropología social africana
