Skip to content
Home » News » How to Read Computer Chess Rating List Method Notes

How to Read Computer Chess Rating List Method Notes

computer chess rating list method notes

Table of Contents

Computer Chess Rating List Method Notes

Method notes are the instruction manual for interpreting a rating list. They define what population was measured, how the ratings were calculated, which games entered the dataset, what hardware and time control were used, how openings and tablebases were handled, and which engines or versions were included or excluded. Without those notes, a rating table may still display ranks and Elo values, but the reader cannot determine exactly what those numbers describe.

A computer chess rating list is not a universal thermometer of engine strength. It is a statistical model built from a particular collection of games played under particular conditions. The table is the visible output; the method notes describe the experiment that produced it.

This distinction matters whenever a rating is quoted, compared or used as evidence. A statement such as “Engine A is rated 25 Elo above Engine B” is incomplete unless both engines were measured on the same connected scale, under the same list conditions, with enough games to support that resolution. A statement such as “Engine C is the strongest engine” is even more demanding: it requires the reader to establish the relevant time control, hardware class, opponent population, version policy and statistical uncertainty.

Good method notes do not make every rating list equally strong. Nor do they eliminate uncertainty. Their purpose is to expose the assumptions and boundaries that allow the reader to interpret the result responsibly.

This guide explains how to read those notes in a disciplined order.

Method Notes Define the Object Being Measured

A rating table usually attracts attention through its most visible fields:

  • rank;
  • engine name;
  • rating;
  • error estimate;
  • number of games;
  • score;
  • draw percentage;
  • likelihood-of-superiority information.

Those fields answer questions about the output. They do not fully describe the input.

The method notes answer a different set of questions:

  • Which games were accepted?
  • Which engine versions were grouped or separated?
  • Was the list based on one processor class or several?
  • Was thinking on the opponent’s time disabled?
  • Did every engine receive the same opening treatment?
  • Were endgame tablebases available?
  • Were games adjudicated?
  • Were crashes and time losses retained?
  • Was the rating scale anchored to a named engine or an arbitrary numerical base?
  • Does the table show all tested versions or only one representative version per family?
  • Are ratings provisional, historical or current?

These details define the statistical object represented by the table.

Consider a public list described as “all engines, best versions only.” That phrase immediately changes the meaning of the ranking. It is not a complete historical list of every tested binary. It is a filtered family-level view in which older or weaker releases may be hidden from the main table.

Now consider a second list described as “all engines and versions.” That table serves a different purpose. It may be more useful for historical comparison, regression tracking or identifying changes between releases, but it will contain many related versions whose results are not independent in the ordinary sense.

Neither view is automatically better. They answer different questions. The method notes tell the reader which question the publisher has chosen to answer.

The same principle applies to hardware classes. A list restricted to single-CPU engines is different from a list that combines single- and multi-CPU configurations. A table restricted to public releases is different from one that includes private development versions. A list of original UCI engines is different from one that includes derivative families under a separate publication track.

Before reading the first Elo value, identify the object being ranked.

1. Rating Method and Numerical Base

The first technical question is how the ratings were calculated.

Several tools and models are used in computer chess. BayesElo is a well-known program by Rémi Coulom that reads game records in PGN format and produces estimated ratings. Ordo is another established rating program. Its documentation describes a model related to Elo that calculates the competitors together and can anchor a named engine to a selected numerical value.

Knowing the program name is useful, but it is not sufficient. Readers should also look for the settings that define the scale.

The numerical base is a coordinate, not a physical constant

Relative rating differences carry the principal competitive meaning. The absolute number assigned to the population depends on how the scale is positioned.

Suppose a connected set of engines has estimated relative ratings of:

  • Engine A: +40
  • Engine B: +10
  • Engine C: −50

A publisher could add 3000 to every value and publish:

  • Engine A: 3040
  • Engine B: 3010
  • Engine C: 2950

Another publisher could use a different anchor and report:

  • Engine A: 3540
  • Engine B: 3510
  • Engine C: 3450

The internal differences are unchanged. Engine A remains 30 points above Engine B and 90 points above Engine C. The absolute numbers differ because the numerical origin differs.

Method notes should therefore disclose one of the following:

  • the engine used as an anchor;
  • the rating assigned to that engine;
  • the average or pool value used as a base;
  • a historical scale-maintenance rule;
  • or an explanation that the scale is provisional and may be rebased.

Without this information, readers may mistake absolute values for directly transferable measurements.

Ratings from separate lists are not automatically on the same scale

Two lists can both use BayesElo and still produce numbers that should not be compared directly. They may differ in:

  • rating base;
  • opponent population;
  • time control;
  • hardware;
  • draw model;
  • included versions;
  • book policy;
  • tablebase access;
  • date range;
  • connectivity between engines.

Using the same rating software does not make the datasets equivalent.

A safe comparison begins with differences inside one connected list. Cross-list comparison requires much more caution. The reader must establish that the two scales have a meaningful bridge, such as a sufficiently connected set of common engines tested under comparable conditions.

Read the uncertainty columns with the rating

A rating estimate should not be detached from its uncertainty.

Public computer-chess tables often display a central rating with positive and negative error values. CCRL, for example, presents columns for Elo and accompanying “+” and “−” values, together with score, average opponent, draw rate, games and LOS.

The responsible unit of interpretation is therefore not merely:

Engine A: 3600

It is closer to:

Engine A: estimated at 3600 on this list’s scale, under these conditions, with this uncertainty and this game base.

When two engines are separated by fewer points than the uncertainty around their estimates, the table may not support a confident ordinal claim. The printed rank can still be mechanically correct according to the central estimates, but “ranked first” is not always equivalent to “demonstrably stronger.”

Look for the reference date

Ratings change when games are added, corrected, excluded or recalculated. Method notes should identify:

  • the calculation date;
  • the game-database cutoff;
  • the number of games included;
  • and, ideally, the rating-tool version or command profile.

A date is part of the result. Quoting an Elo value without its calculation date can silently mix different states of the database.

2. Games, Sample Size and Rating Maturity

The next question is how much evidence supports each estimate.

Game count is one of the most visible indicators of maturity, but it must be interpreted carefully. More games usually improve statistical precision, yet the usefulness of those games depends on their structure.

Total games and games per engine are different measures

A list may contain millions of games overall while a newly added engine has completed only a small test set.

CCRL’s public header, for example, reports the total database size and separately distinguishes engines with at least 200 games through bold formatting. That threshold is a publication convention for the table. It should not be misread as a universal scientific law stating that 199 games are unusable and 200 games are final.

Readers should inspect:

  • total games in the database;
  • games played by the engine being quoted;
  • number of distinct opponents;
  • distribution of games across opponents;
  • colour balance;
  • opening-pair balance;
  • recency of the games;
  • and whether the result belongs to a connected rating network.

An engine with 1,000 games concentrated against two opponents may be less informative for broad placement than an engine with a well-distributed schedule across a representative pool.

Rating maturity is not binary

A practical publication vocabulary might include:

  • initial — too little evidence for stable placement;
  • provisional — useful as an early estimate, but expected to move;
  • developing — connected to the pool with increasing opponent coverage;
  • mature — supported by substantial, diversified evidence;
  • historical — retained for archive purposes but no longer current;
  • superseded — replaced by a newer release in the main family view;
  • excluded — removed from the active calculation under a published rule.

These categories are editorial states, not universal mathematical definitions. Their value comes from being explained and applied consistently.

Read the opponent distribution

An Elo estimate is relational. It is derived from results against other rated participants.

A list can become weakly connected if groups of engines rarely play across the boundary between them. For example:

  • a top group plays mostly within the top group;
  • a middle group plays mostly within the middle group;
  • older engines play only historical opponents;
  • private engines appear briefly and then disappear;
  • one hardware class has little contact with another.

The rating software may still produce a complete table, but the uncertainty between weakly connected regions can be greater than the headline numbers suggest.

Method notes should therefore explain how new engines are integrated. Useful information includes:

  • calibration opponents;
  • gauntlet composition;
  • round-robin structure;
  • minimum opponent diversity;
  • requirements for promotion into the main table;
  • and procedures for reconnecting historical or isolated subgroups.

Draw rate affects information density

Modern engine games often have high draw rates, especially among strong engines under controlled opening conditions. A large number of games does not necessarily produce the same information as the same number of games in a more decisive pool.

This does not mean draws are useless. Draws contain information about relative performance. It means the reader should not use game count as a complete proxy for precision.

A list displaying draw percentage helps readers understand the competitive environment. Very high draw rates can make small estimated differences difficult to resolve, particularly when engines are closely matched.

Version turnover can destabilise apparent maturity

An engine family may have thousands of historical games, while the current binary has only a few hundred. Method notes should clarify whether ratings are assigned to:

  • each exact version;
  • the best version of a family;
  • a merged family identity;
  • development builds;
  • or rolling replacements.

Combining results from distinct versions can produce a larger sample but blur the identity of the measured program. Keeping every version separate preserves identity but can fragment the data.

Again, there is no universally correct answer. The publication must disclose the policy.

3. Time Control and Hardware

A rating belongs to its time-control and hardware environment.

Engine rankings can change when more time, more threads, larger hash tables, different processor instructions or different accelerators are available. A rating list should therefore state enough information for the reader to understand the resource model.

Decode the time-control notation

A notation such as “40/15” may mean 40 moves in 15 minutes, followed by another block of time after move 40. CCRL’s 40/15 documentation explicitly describes a repeating control: engines receive 15 minutes for the first 40 moves, then another 15 minutes for the next 40 moves.

That is not equivalent to every other format that might casually be called “15-minute chess.”

Compare:

  • 40 moves in 15 minutes, repeating;
  • 15 minutes for the whole game;
  • 10 minutes plus 5 seconds per move;
  • 5 minutes plus 2 seconds per move;
  • a fixed number of nodes;
  • a fixed search depth;
  • a fixed movetime;
  • a benchmark-normalised equivalent control.

Each format creates a different search-budget distribution.

Increment-based controls reduce the risk of complete time exhaustion in long games. Repeating controls reward engines that manage a sequence of time-control boundaries. Sudden-death controls place greater pressure on time management. Fixed-node tests remove much of the hardware-speed variable but do not reproduce normal wall-clock tournament behaviour.

The method notes should define the notation rather than assuming the reader knows what it means.

Equivalent hardware is not identical hardware

Distributed rating projects often run games on multiple machines. To maintain a broadly comparable time budget, they may calibrate each machine against a reference processor or benchmark engine.

CCRL states that its 40/15 control is normalised to an Intel i7-4770K and that Stockfish 10 is used as the benchmark for determining an equivalent time control on a particular machine.

This is important information, but the phrase “equivalent time control” should not be interpreted as “all hardware effects have disappeared.”

Processors can differ in:

  • cache hierarchy;
  • memory latency;
  • branch prediction;
  • instruction-set support;
  • core topology;
  • simultaneous multithreading;
  • thermal behaviour;
  • operating-system scheduling;
  • compiler optimisation;
  • and engine-specific scaling.

A benchmark can normalise gross speed, but it may not produce perfect equivalence for every engine architecture.

Readers should therefore distinguish:

  • identical-hardware testing, where every game uses the same machine profile;
  • calibrated distributed testing, where machine speed is normalised against a benchmark;
  • mixed-resource testing, where engines may receive different resources by design;
  • and open hardware lists, where the hardware itself is part of the comparison.

Threads and hash must be explicit

An engine tested with four threads should not be silently compared with another tested with one thread unless that is the declared purpose of the list.

Relevant notes include:

  • number of threads or CPUs;
  • whether logical or physical cores are counted;
  • hash-table size;
  • NUMA policy;
  • large-page use;
  • GPU model for neural engines;
  • neural-network file;
  • and any engine-specific resource exceptions.

CCRL’s 40/15 conditions state that hash should normally be equal within a match or tournament, with documented exceptions for multi-CPU engines or engines that cannot use a selected size.

That kind of exception is not automatically a flaw. It becomes a problem only when it is hidden or applied inconsistently.

Pondering changes the resource model

“Ponder on” means an engine can think during the opponent’s clock time. “Ponder off” means it cannot.

A list with pondering enabled measures more than nominal clock time. It also measures:

  • prediction of the opponent’s move;
  • reuse of analysis;
  • thread scheduling during the opponent’s turn;
  • and the effective hardware time consumed by each game.

CCRL’s main 40/15 header states “Ponder off,” making the resource model easier to interpret.

When method notes omit ponder status, the reader should not assume it was disabled.

4. Opening Policy and Tablebases

Opening conditions can influence both fairness and the kind of strength a list measures.

A rating list based on starting-position games is not identical to one based on a curated opening suite. A list using deep, unbalanced openings measures different properties from one using short generic book lines.

Identify the opening source

Method notes should answer:

  • Was a book or suite used?
  • How were positions selected?
  • How long were the forced lines?
  • Were positions filtered?
  • Were colours reversed?
  • Were openings repeated across engine pairs?
  • Were highly unbalanced or decisive positions excluded?
  • Were engines allowed to use their own books?
  • Was book learning enabled?

CCRL’s 40/15 conditions allow a generic opening book, limit the line to 12 moves per side, require the same book within a match or tournament, disable engine-specific books when possible, and turn off book and position learning.

Each clause has interpretive consequences.

A common book reduces the influence of proprietary opening preparation. A line-length limit restricts how much of the game is predetermined. Disabling learning prevents one engine from adapting its book behaviour over repeated games.

Switched-side testing is important, but not magical

When the same opening is played twice with colours reversed, the design controls some of the opening’s colour bias. If one side of an opening is favourable, both engines receive an opportunity to play that side.

However, a mirrored pair does not eliminate every source of variance.

The two games may still differ because of:

  • stochastic search;
  • thread scheduling;
  • transposition-table state;
  • randomisation;
  • timing noise;
  • engine-specific opening transitions;
  • or nondeterministic neural inference.

Mirroring is a strong design feature, but it should not be described as perfect cancellation.

Readers should check whether the list reports:

  • pair ratio;
  • completion of both games in each pair;
  • colour balance;
  • handling of interrupted pairs;
  • and whether the opening advanced only after the pair closed.

Opening difficulty changes what the rating represents

A neutral general book tends to estimate performance from relatively conventional positions. A suite of deliberately unbalanced openings may increase decisiveness and reduce the draw rate, but it can also place greater weight on:

  • defensive resilience;
  • tactical recovery;
  • conversion technique;
  • or performance in unusual structures.

Neither approach is inherently invalid. The method notes should explain the objective.

The word “balanced” also requires care. It can refer to:

  • engine-evaluated balance;
  • historical human score;
  • absence of a forced result;
  • symmetrical material;
  • or a target evaluation interval.

Readers should look for the actual selection rule.

Tablebase notes require more than a yes-or-no label

“Tablebases enabled” is incomplete.

A useful note should identify:

  • tablebase family, such as Syzygy or another format;
  • maximum number of pieces;
  • whether pawnful and pawnless files are available;
  • probe depth;
  • probe limit;
  • 50-move-rule handling;
  • storage medium;
  • and whether adjudication used tablebase results independently of engine probing.

CCRL’s list header mentions 3-, 4- and 5-piece endgame tablebases, while its fuller 40/15 conditions state that 4-, 5- or 6-piece tablebases may be used. This illustrates a practical reading rule: consult the dated list header and the detailed conditions page together, and resolve any difference in wording before making a precise claim about a particular dataset.

A brief tablebase label may describe the usual environment rather than every historical game. Publication dates and revision dates therefore matter.

Separate engine access from adjudication

There are two distinct roles for tablebases:

  1. Engine-side access
    The engine probes a tablebase during search and uses the result to choose moves.
  2. Tournament-side adjudication
    The tournament controller or auditor declares a result because the position is theoretically won, lost or drawn.

These roles can produce different effects.

If engines have access, tablebases are part of playing strength under the test configuration. If only the tournament controller uses them, they affect game termination and possibly the recorded result without directly informing the engine’s search.

Method notes should not merge those two policies into one vague statement.

5. Engine Status, Versions and Exclusions

A rating table is also an editorial selection.

The publisher decides which engines are eligible, how they are named, whether private versions are visible and when a result should be removed from the active list.

Exact binary identity matters

A serious entry should distinguish, where applicable:

  • engine family;
  • exact version;
  • release date or development identifier;
  • architecture;
  • thread count;
  • network file;
  • default or modified settings;
  • public, private, commercial or open-source status.

“Stockfish,” “Stockfish 18,” “Stockfish development build” and a modified Stockfish-derived engine are not interchangeable identities.

Likewise, an engine compiled for AVX2 is not necessarily identical in execution profile to an SSE4.1 binary, even when both implement the same source version. They may be functionally equivalent at the algorithmic level, but the published test identity should remain explicit.

Best-version views hide historical depth

CCRL’s primary 40/15 list is labelled “best versions only,” while the same site offers a complete list containing all engines and versions and filtered views for public, free, open-source, single-CPU and multi-CPU categories.

This is a useful publication design because it separates two reading tasks:

  • current family-level comparison;
  • historical version-level research.

A reader quoting a rank must identify which view was used.

An engine may appear fifth in the best-version view but much lower in a complete table crowded with related releases. The numbers have not necessarily changed; the selection rule has.

Public and private status affects reproducibility

A private engine can produce genuine games and a statistically valid estimate inside a list. Yet independent reproduction may be impossible if the binary and configuration are unavailable.

Method notes should therefore distinguish statistical inclusion from reproducibility.

Useful status labels include:

  • public release;
  • open source;
  • freeware;
  • commercial;
  • private;
  • development build;
  • author-supplied test build;
  • withdrawn;
  • obsolete;
  • superseded.

These labels do not determine playing strength. They tell the reader what can be independently inspected or reproduced.

Inclusion is not necessarily a legal or originality judgment

CCRL explicitly states that its role is not to determine the originality of engines and that inclusion or non-inclusion should not be interpreted as a statement about an engine’s status.

This is an important example of editorial restraint.

A list may include or exclude an engine for many reasons:

  • technical instability;
  • inability to disable a private book;
  • repeated crashes;
  • protocol failures;
  • insufficient games;
  • unavailable binary;
  • duplication of a family;
  • author request;
  • test-capacity priorities;
  • incompatible hardware;
  • or publication policy.

Readers should not invent a stronger conclusion than the notes support.

Understand “killed,” removed and inactive entries

Some rating projects maintain separate records of engines removed from active testing or calculation. CCRL publishes a list of “killed” engines with their accumulated game counts.

The term must be interpreted according to the project’s own documentation. It may refer to database maintenance rather than a declaration that the engine is fraudulent or technically worthless.

Before quoting a removed engine’s old rating, ask:

  • Why was it removed?
  • Are its games still present in historical calculations?
  • Was it replaced by a newer version?
  • Was the entry disconnected from the active pool?
  • Did removal trigger a recalculation?
  • Is the old rating still visible only for archival purposes?

6. How to Read the Main Statistical Columns

Method notes define the environment, while the table columns summarise performance inside that environment.

Rating

The central estimate on the list’s own numerical scale.

It should be read relationally, not as an absolute universal strength value.

Positive and negative error values

These indicate uncertainty around the estimate according to the rating model and reporting convention.

They are essential when the Elo gap between engines is small.

Score

The percentage of available points obtained in the included games:

  • win = 1 point;
  • draw = 0.5 points;
  • loss = 0 points.

Score is meaningful only in relation to opponent strength.

Average opponent

A summary of the average rating of the engine’s opposition.

A high score against a weak average opponent does not imply the same performance as the same score against an elite pool.

Draw rate

The proportion of games drawn.

This helps describe the testing environment and the information density of the result.

Games

The number of included games for that entry.

Readers should still investigate opponent diversity and recency.

LOS

LOS commonly refers to likelihood of superiority. It is a probability-like statistic derived from the rating model and game evidence.

It should not be translated casually into a universal claim. A high LOS between two entries means the available model and dataset support an ordering with a particular degree of confidence. It does not prove that one engine will win a future short match.

Rank

Rank is the ordering of central estimates after applying the table’s selection rules.

It is not an independent statistic. When estimates overlap heavily, several neighbouring ranks may be practically unresolved.

7. Do Not Compare Tables Before Comparing Their Notes

The most common misuse of rating lists is to copy two Elo numbers from different sources and subtract them.

Suppose:

  • Engine A is rated 3650 on List X.
  • Engine B is rated 3600 on List Y.

It does not follow that Engine A is 50 Elo stronger.

The lists may differ in:

  • base;
  • time control;
  • hardware;
  • engine pool;
  • rating software;
  • draw handling;
  • version selection;
  • book suite;
  • tablebases;
  • adjudication;
  • date;
  • and publication objective.

Even two tables from the same organisation may be intentionally separate. CCRL, for example, publishes distinct environments for its medium control, blitz, Fischer Random Chess and other categories. The method notes are what prevent the reader from collapsing them into one unsupported scale.

The safest hierarchy of comparison is:

  1. two engines on the same table and same calculation date;
  2. two versions on a connected complete-version table;
  3. two tables from the same project with an explicit bridge;
  4. two external lists only after detailed methodological reconciliation.

The further down that hierarchy the comparison moves, the more conditional the conclusion must become.

8. Five Questions to Ask Before Quoting a Rating

Before publishing, repeating or debating an Elo value, answer these five questions.

Question 1: What exactly is being rated?

Identify:

  • engine and version;
  • binary architecture;
  • thread count;
  • network or evaluation file;
  • configuration;
  • public or private status;
  • and whether the table shows an exact version or a best-of-family representative.

Do not quote a family name when the result belongs to a specific binary.

Question 2: Under what playing conditions was it measured?

Identify:

  • time control;
  • repeating or sudden-death clock;
  • increment;
  • hardware or calibration method;
  • threads;
  • hash;
  • pondering;
  • opening policy;
  • tablebase access;
  • and adjudication.

A rating without conditions is an incomplete measurement.

Question 3: How mature is the estimate?

Check:

  • number of games;
  • uncertainty;
  • number and diversity of opponents;
  • connectivity to the main pool;
  • colour balance;
  • mirrored-pair completion;
  • recency;
  • and provisional status.

Do not replace this evaluation with a single arbitrary game-count threshold.

Question 4: How was the scale calculated and anchored?

Check:

  • rating program;
  • relevant settings;
  • rating base;
  • anchor engine;
  • treatment of draws;
  • database cutoff;
  • and whether a recalculation followed exclusions or corrections.

Do not assume that equal numerical values from different lists represent equal strength.

Question 5: What qualifications or exclusions apply?

Check:

  • best versions only or all versions;
  • public-only or mixed status;
  • original and derived tracks;
  • crash and timeout policy;
  • removed engines;
  • incomplete games;
  • manual corrections;
  • and historical versus active status.

A rating can be accurately copied yet still be misleading if these qualifications are omitted.

9. A Worked Reading Example

Imagine a hypothetical table with this header:

Ponder off. Generic opening book, maximum 10 moves. Five-piece tablebases. Equivalent to 5 minutes plus 2 seconds per move on the reference CPU. Ratings calculated with Ordo from 80,000 games. Base engine fixed at 3400. Best public versions only. Engines with fewer than 300 games are provisional.

The table shows:

  • Engine A — 3512 ± 9 — 1,600 games
  • Engine B — 3507 ± 11 — 1,100 games
  • Engine C — 3490 ± 28 — 180 games

A superficial reading says:

  1. Engine A is first.
  2. Engine B is second.
  3. Engine C is third.
  4. Engine A is five Elo stronger than Engine B.

A method-aware reading is more precise.

Step 1: Define the population

The table contains best public versions only.

It does not describe:

  • private builds;
  • older versions;
  • every tested binary;
  • or engines excluded by the public-release rule.

The ranking applies to this filtered population.

Step 2: Define the resource environment

Pondering is disabled. Engines use a time control normalised to a reference CPU. The result describes performance under that clock and calibration method.

It does not establish the same ordering at:

  • bullet;
  • long classical;
  • fixed nodes;
  • different thread counts;
  • or another processor architecture.

Step 3: Define the opening and endgame environment

A common generic book controls the first part of the game. Five-piece tablebases affect some endings.

The table is not a pure starting-position test, nor is it a no-tablebase test.

Step 4: Define the scale

The list was calculated with Ordo and anchored by assigning 3400 to a selected base engine.

The absolute values belong to that scale. They should not be compared numerically with an unrelated list unless a valid bridge exists.

Step 5: Read the uncertainty

Engine A’s central estimate is five points above Engine B’s, but their uncertainty intervals overlap substantially.

The safe conclusion is:

Engine A has the higher central estimate on this table, but the available rating evidence does not support describing the five-point gap as a clearly resolved strength difference.

Step 6: Read maturity

Engine C has only 180 games and is explicitly provisional. Its 3490 estimate has a much larger uncertainty.

The safe conclusion is not that Engine C is definitively 17 points weaker than Engine B. It is:

Engine C is provisionally placed below A and B, but its estimate is less mature and may move materially as the schedule expands.

Step 7: Preserve the date

The calculation date must accompany any quotation. If Engine C later completes 1,000 games, its rating and uncertainty may change even without a new binary.

This worked example shows why method notes are not secondary prose. They determine almost every responsible sentence that can be written about the table.

10. Applying the Same Reading Discipline to IJCCRL

IJCCRL separates its publication architecture into rating surfaces, event records, downloads, archive material, winners and audit documentation. The current hub explicitly distinguishes the Original UCI Track from the Derived Stockfish Track and keeps Classical, Blitz and Bullet ratings separate.

Readers entering through the main chess engines ratings lists platform should therefore treat each table as a track- and time-control-specific publication surface.

The current computer chess rating lists hub provides the status and scope of the available lists. A provisional event-stage table should not be quoted as though it were a closed universal scale, and a knockout winner should not automatically replace a league-based Elo calculation.

For the supporting methodology, the rating publication rules and audit evidence explain the competition grammar, clocks, opening-pair requirements, tablebase policy, termination rules and publication conditions.

The same five-question test applies:

  1. Which track and time control?
  2. Which exact engines and versions?
  3. How many games and opponents?
  4. What calculation method and rating base?
  5. What publication status and audit qualifications?

The purpose of this structure is not to make every table immune to revision. It is to make the scope of each result visible.

11. What Good Method Notes Should Contain

A reader-focused method note does not need to become a full research paper, but it should expose the essential contract of the list.

A strong minimum specification includes:

Rating calculation

  • rating software and version;
  • principal command settings;
  • scale base or anchor;
  • calculation date;
  • database cutoff;
  • uncertainty convention;
  • LOS convention, if published.

Game population

  • total games;
  • games per engine;
  • minimum publication threshold;
  • provisional-status rule;
  • opponent-selection policy;
  • disconnected-pool handling;
  • duplicated-game detection;
  • treatment of interrupted or corrupted records.

Time and hardware

  • clock notation;
  • increment or repeating control;
  • reference hardware;
  • calibration procedure;
  • threads;
  • hash;
  • pondering;
  • operating-system or architecture constraints;
  • GPU resources where applicable.

Openings

  • source book or suite;
  • line depth;
  • selection method;
  • switched-side policy;
  • pair-completion rule;
  • colour balance;
  • engine-book policy;
  • learning policy.

Tablebases and adjudication

  • tablebase format;
  • piece count;
  • probe policy;
  • 50-move-rule treatment;
  • engine-side access;
  • controller-side adjudication;
  • draw and win adjudication thresholds;
  • manual-intervention policy.

Engine identity

  • exact version;
  • binary architecture;
  • network or evaluation file;
  • UCI settings;
  • public/private/commercial/open-source label;
  • family-grouping rule;
  • inclusion and exclusion policy;
  • crash and timeout policy.

Publication state

  • provisional, official, historical or superseded;
  • revision history;
  • corrections;
  • removed games;
  • archived versions;
  • links to supporting PGN or audit material.

A list that omits some of these elements may still be useful. The reader should simply reduce the strength of the conclusions accordingly.

Conclusion

Method notes are the instruction manual for interpreting a computer chess rating list.

They tell the reader what the rating scale means, how it was calculated, how much evidence supports each engine, what resources were available, how openings and endings were controlled, and why particular engines or games appear—or do not appear—in the table.

The central lesson is simple: never quote the number before reading the conditions.

A rating is not an engine property that exists independently of the test. It is an estimate produced by a connected body of results under a declared methodology. Change the time control, hardware, opponent pool, opening policy, version policy or numerical base, and the published value may change.

Responsible reading therefore proceeds in this order:

  1. identify the exact list and date;
  2. establish the rating method and base;
  3. inspect game volume, uncertainty and opponent coverage;
  4. decode time control and hardware;
  5. examine opening and tablebase policy;
  6. verify engine identity and publication status;
  7. read exclusions and qualifications;
  8. only then quote the rating.

Method notes do not remove uncertainty. They make the uncertainty interpretable.

That is what turns a ranking table into a usable technical publication.

Sources and Technical References

  • CCRL 40/15 rating list and statistical table. Public example of list headers, game totals, rating columns, version filters and maturity indicators.
  • CCRL 40/15 testing conditions. Public documentation of time-control calibration, repeating clocks, tablebases, pondering, hash and opening-book policy.
  • Bayesian Elo Rating, Rémi Coulom. Official BayesElo description and documentation.
  • Ordo, Miguel A. Ballicora. Official repository and description of the rating program.

Jorge Ruiz

Jorge Ruiz Centelles

Filólogo y amante de la antropología social africana