tournament performance vs chess engine ratings
A computer-chess tournament and a chess engine rating list may use many of the same games, but they do not answer the same question.
A tournament asks which engine obtained the best result under a fixed event structure. That structure may be a round robin, league, gauntlet, knockout cup, semifinal or long head-to-head final. It has a defined field, schedule, time control, hardware profile, opening policy, tablebase configuration, adjudication system and tiebreak procedure.
A long-term rating list asks a broader statistical question: given the accepted games in a connected population, what relative-strength estimates best explain the observed results under the chosen rating model?
The tournament winner is therefore an event fact. The rating is a model-based estimate.
An engine can win a competition without becoming the highest-rated engine in the corresponding long-term list. Conversely, an engine can remain first in a rating table after losing a final. Neither outcome is logically inconsistent. The apparent contradiction arises only when an event title, a performance rating and a longitudinal Elo estimate are treated as interchangeable measurements.
For readers of chess engines ratings lists, the essential discipline is to keep these evidential layers separate. A ranking table can be statistically legitimate and still be quoted incorrectly if its time control, engine identities, uncertainty, opponent pool or publication status are omitted. A tournament result can be completely valid while supporting a much narrower conclusion than the phrase “the strongest engine” suggests.
This article explains the distinction from statistical, experimental and editorial perspectives.
1. Tournament titles are event outcomes
A tournament title belongs to a specific competition.
If Engine A wins a 100-game final against Engine B, the historically correct statement is that Engine A won that final. The statement remains true regardless of whether Engine A later falls below Engine B in a rating list, loses a rematch or is superseded by a new version.
The title records realised performance within the event’s rules.
Those rules determine the competition’s evidential scope:
- which engine versions participated;
- how many games were played;
- which openings were assigned;
- whether positions were reversed with colours switched;
- which time control and hardware were used;
- how crashes and time losses were treated;
- which adjudication rules applied;
- and how ties were resolved.
TCEC, for example, prohibits engine updates once an event has started. This freezes the competitors for that particular competition and protects the identity of the event result. Its current rules also distinguish the competition result from a separate rating system that aggregates a much larger body of games.
The winner has demonstrated superiority in the operational sense required by the tournament: it scored enough points, or survived enough knockout matches, to finish first.
That achievement should not be weakened. A title is not “less real” because finite-event variance exists. Sporting and competitive history is made of realised outcomes, not only expected values.
The error is instead to enlarge the claim beyond the competition.
“Engine A won the event” is directly supported.
“Engine A was the best-performing engine in this event” is normally supported, subject to the event’s tiebreak rules.
“Engine A is universally stronger than every other engine” is not established by the title alone.
The distinction is similar to the difference between winning a championship and leading a statistical rating system. The first records who prevailed in a bounded contest. The second estimates relative strength across a broader population and time horizon.
2. Long-term ratings estimate latent relative strength
A rating is not a trophy, and it is not a direct measurement of an internal quantity located inside the program.
Ratings are inferred from pairwise outcomes.
In the simplest conceptual model, each engine has an underlying strength parameter. The probability of one engine outperforming another depends on the difference between their parameters. Wins, draws and losses provide evidence from which those relative values are estimated.
This places engine rating systems within the wider statistical field of paired-comparison models. Bradley–Terry-type models and their extensions are widely used to infer latent rankings from repeated pairwise results. Statistical research also emphasises that dependencies, ties, covariates and schedule structure can complicate the basic independent-comparison model.
The word latent is important. Strength is not observed directly. What is observed is a sequence of results generated under particular conditions.
The rating calculation then asks: which set of relative strengths makes those results reasonably likely?
BayesElo reads game records, commonly in PGN form, and can produce ratings, opponent statistics, confidence intervals and likelihood-of-superiority information. Its documentation also allows the numerical scale to be shifted by assigning a chosen value to a reference engine.
Ordo likewise calculates engine ratings from a PGN result set. It considers the results together rather than updating engines only in chronological sequence, and it allows the average of the pool or a named reference engine to define the displayed numerical base.
These properties demonstrate why an engine’s printed Elo is not a universal constant.
The meaningful evidence is primarily relational:
- Engine A is estimated above Engine B within the same connected calculation.
- The estimated gap is a certain number of points.
- The estimate has a certain uncertainty.
- The result applies to the game population and conditions used.
The absolute numerical value depends partly on the scale’s anchoring convention. Adding 300 points to every engine would change none of the internal differences or expected scores.
Long-term ratings therefore describe a structured population, not an isolated engine in all possible environments.
3. Tournament performance is a finite sample
Every tournament contains a limited number of games.
Even a long computer-chess event represents only a sample of all games that could have been played between its participants under the same or alternative conditions.
Suppose that Engine A has a genuine expected score of 52.9 percent against Engine B under a particular testing environment. In the conventional Elo expectation formula, that corresponds approximately to a 20-point rating advantage:
[
E_A=\frac{1}{1+10^{-\Delta R/400}}
]
With (\Delta R=20), the expected score is about 0.529.
Over 100 games, the expected score is therefore approximately 52.9–47.1.
That expected advantage is fewer than three points across the entire match.
The realised result does not have to equal its expectation. Engine A might score 56, 52, 49 or another value. The rating difference determines the centre of a probability distribution, not a guaranteed final score.
The exact distribution is more complicated than a simple independent-binomial model because engine matches contain draws, colour assignments, mirrored openings, correlated game pairs, technical incidents and possible non-independence. Nevertheless, the central principle remains: a modest expected advantage can be overwhelmed by finite-sample variation.
Statistical work on Elo-type systems also shows that the treatment of draws is not a trivial detail. The assumptions embedded in the rating model influence how results are translated into strength estimates, especially in draw-heavy environments.
Consequently, the lower-rated engine can win a match without invalidating the rating list.
The result may be surprising, but it is not impossible under the model.
A rating says that one outcome is more probable or that one expected score is higher. It does not say that every finite contest will reproduce the ordering.
4. Performance rating is local to the event
A tournament performance rating is an attempt to describe the strength implied by an engine’s score against the opposition it faced during one event.
It is a local summary.
If Engine A scores 65 percent against opponents whose established average rating is 3400, its performance rating will be above 3400 by an amount determined by the chosen conversion model.
That figure depends on at least four elements:
- the engine’s event score;
- the opponents’ assigned ratings;
- the numerical scale on which those ratings exist;
- the rating formula or software used.
A performance rating is therefore not independent evidence floating above the tournament. It inherits the assumptions and calibration of the opposition ratings.
If the opponents are miscalibrated, weakly connected or taken from another incompatible scale, the performance figure may look precise while lacking a defensible interpretation.
Performance rating also does not become permanent merely because it is numerically expressed as Elo.
An engine can produce a very high performance in one tournament because it played unusually well, received favourable opening outcomes, exploited particular opponents, avoided its weakest stylistic matchups or benefited from ordinary variance.
When more games are added, the estimate often moves toward the engine’s broader long-term level. This is not evidence that the tournament calculation was dishonest. It is a normal consequence of using a small sample to estimate a long-run parameter.
The event performance remains historically meaningful.
The appropriate statement is:
Engine A produced a performance of approximately X on this event’s scale against this opposition.
The inappropriate statement is:
Engine A permanently became an X-rated engine in every environment.
5. Long-term ratings aggregate more evidence—and more heterogeneity
A longitudinal rating list can reduce sampling uncertainty by incorporating many more games than a single event.
CCRL’s 40/15 project, for example, publishes a large distributed game database together with ratings, game counts, errors, scores, average opposition and draw percentages. Its methodology also specifies a calibrated time control, reference hardware, pondering policy, hash treatment, opening conditions and tablebase use.
The scientific advantage of aggregation is evident: more observations can improve the precision of relative-strength estimates and connect engines across a broader population.
However, aggregation introduces its own interpretive problems.
A long-term rating database may combine:
- several tournaments;
- repeated matches;
- development and release versions;
- different opponent distributions;
- multiple seasons;
- hardware generations;
- time-control categories;
- revised opening suites;
- technical wins and losses;
- and engines whose playing strength changed over time.
The publisher must decide what the rating unit represents.
Is the entry an exact binary?
Is it the most recent version?
Is it the best version tested?
Is it an engine family combining all releases?
Is it a historical identity whose rating accumulates across development generations?
TCEC’s current rating policy is an instructive example. The official rules state that the rating list is updated with official games, ignores version numbers and aggregates games under the engine identity, including results from different periods.
That policy creates a family-level longitudinal signal.
It does not answer the same question as a list that gives Stockfish 17, Stockfish 18 and a dated development build separate entries.
Family aggregation improves continuity and sample size but blurs exact-version identity. Version-specific publication preserves reproducibility but fragments the evidence and may leave new releases provisional for longer.
Neither approach is inherently illegitimate. The method note must tell the reader which object is being rated.
6. Tournament format changes the meaning of victory
Not all tournament winners are produced by the same competitive mechanism.
Round robin and league events
In a round robin, engines accumulate points across several opponents. The winner normally demonstrates the strongest realised performance across the field.
This format samples breadth. An engine must perform against different search styles, evaluation architectures and tactical or positional profiles.
However, the final ordering still depends on the schedule, number of cycles, openings, opponent field and tiebreak rules. If the event is short, small score differences may be highly sensitive to one game pair.
Gauntlets
In a gauntlet, one engine or a small set of engines faces a chosen opponent group.
This can be an effective calibration method, especially for testing a new version. It is less suitable for declaring a complete hierarchy among all participants because many opponents may never play each other.
Knockout cups
A knockout measures survival through a sequence of matches.
The champion may eliminate several strong opponents, but the bracket matters. Two top engines may meet early while another finalist follows a different path.
Knockouts also concentrate variance. A small number of critical game pairs or tiebreak games can determine advancement.
The title is fully valid, but it should not be treated as a maximum-likelihood ranking of the entire field.
Long head-to-head finals
A long final estimates one specific matchup more intensively.
It can provide strong evidence about Engine A versus Engine B under the final’s conditions, especially when opening pairs are carefully controlled.
It says less about either engine’s performance against the rest of the population.
Swiss systems and incomplete leagues
These formats increase field size while limiting total games, but the participants do not all face identical opposition.
Standings remain competition results. Converting them into ratings requires accounting for opponent strength and schedule structure.
Tournament design is therefore not merely an organisational detail. It determines which comparative question the result can answer.
7. Why the favourite can lose
There are several reasons why a higher-rated engine can lose a tournament or final without exposing a contradiction.
Ordinary sampling variation
The observed score fluctuates around the expected score.
When the true gap is small, the lower-rated engine may win a substantial minority of finite matches.
Opening allocation
Computer-chess events commonly use prepared openings to create varied and potentially decisive positions.
A mirrored design gives each engine both sides of the same opening, but the two games are not guaranteed to produce symmetric outcomes. Search nondeterminism, time use and position-specific strengths can affect the pair.
A small number of opening pairs can have a disproportionate effect in a short match.
Stylistic interaction
A scalar rating assumes that much of comparative strength can be represented by one dimension.
Real competitive systems may contain interaction effects. Engine A may perform especially well against B, B against C and C against A. Research on paired-comparison and Elo systems shows that intransitivity or model misspecification can affect the interpretation of a single scalar ranking.
Such effects do not make ratings useless. They mean that the rating is an approximation of average population performance, not a complete description of every matchup.
Time management and technical reliability
Crashes, stalls, disconnections and time losses are realised event outcomes.
Whether they should count in a playing-strength rating depends on the publication objective. A practical competition may reasonably treat reliability as part of competitive performance. A research list interested primarily in search quality may classify technical losses separately.
The policy must be declared before interpretation.
Version-specific form
The tournament binary may differ from the version represented by a family-level long-term rating.
An updated engine may be stronger or weaker than its accumulated historical identity. A newly introduced bug can also affect one event without immediately transforming the long-term rating.
Hardware and resource interaction
An engine that is favoured on one thread may scale less effectively on many threads. A GPU-dependent engine may change relative position according to the accelerator, network and time control.
CCRL’s detailed testing conditions demonstrate why hardware reference, calibration, hash, pondering and tablebases belong to the interpretation of its list.
A tournament defeat under one resource model does not automatically invalidate a rating obtained under another.
8. Draws affect statistical resolution
Modern top-engine competitions can be highly draw-heavy.
A draw is not statistically empty. It contributes evidence about the relationship between two engines. Repeated draws against stronger or weaker opposition influence the estimated rating.
Nevertheless, high draw rates make very small strength differences harder to resolve.
If two engines draw most games, the decisive score margin across a moderate match may be produced by only a few results. One additional win can change the observed event lead substantially.
This is one reason why long matches, large databases and uncertainty estimates are important.
BayesElo includes confidence-interval and likelihood-of-superiority functionality, while CCRL displays error information alongside ratings and games.
The academic SSDF publication record provides a concrete example of responsible interpretation. In its 2023 list, newly tested engines were reported with game counts and the explicit observation that further games were needed to reduce error bars. The list thereby distinguished a promising central estimate from a fully mature placement.
The same restraint should apply to tournament-versus-rating claims.
A one-point event victory may decide a championship. It does not necessarily resolve a long-term Elo ordering.
The event title and the statistical uncertainty can coexist:
- the champion is known;
- the long-run strength difference may remain uncertain.
9. Rank, expected score and probability of winning a match are different
Three related concepts are often collapsed into one.
Rating rank
This is the order of the engines’ central rating estimates.
Expected score
This is the modelled average score of one engine against another over repeated games under the rating assumptions.
Probability of winning a finite match
This is the probability that one engine finishes ahead after a defined number of games, including any draw and tiebreak structure.
A 55 percent expected game score does not mean a 55 percent chance of winning every match length.
In a very long match, the stronger engine’s probability of finishing ahead normally increases because random fluctuations average out.
In a very short match, the probability may remain much closer to 50 percent than readers expect.
Draws, paired openings and tiebreak rules further complicate the conversion.
A rating list normally publishes the first two concepts—rank and relative expected performance. It does not automatically publish an exact probability for every possible tournament format.
Therefore, the statement “Engine A was favourite” should not be translated into “Engine A could not legitimately lose.”
Being favourite means that the model assigned a higher expectation, not certainty.
10. Event results can provide new rating evidence
Winning a tournament does not increase an engine’s rating by ceremonial award.
Ratings change only if the event games enter the accepted database and the calculation is rerun.
This distinction matters because different publications use different inclusion rules.
A project may include:
- every completed official game;
- only games from rating-valid events;
- only mirrored pairs;
- only games at specified time controls;
- or only games passing a post-event audit.
A title can therefore exist independently from rating eligibility.
For example, an exhibition cup may produce a legitimate winner while remaining excluded from the main rating calculation because it used experimental hardware, an unusual opening format or too few games.
Conversely, league games may enter a rating database even when the engine did not win the event.
The rating system does not reward prestige. It processes accepted results.
Chess.com’s 2020 computer-rating publication illustrates another valid architecture: ratings were derived from performances in its competition environment and were presented as useful indicators for that platform, while the publisher cautioned against treating them as the definitive estimate of each engine’s universal strength.
That editorial qualification is methodologically important.
A rating can be useful, carefully calculated and still remain environment-specific.
11. Performance claims require exact engine identity
The sentence “Stockfish won” may be sufficient for a casual event report, but it is not enough for a technical comparison.
A reproducible record should identify:
- exact version or commit;
- executable architecture;
- number of threads;
- hash size;
- external neural-network file;
- non-default UCI options;
- tablebase configuration;
- and hardware profile.
The need becomes especially important when comparing a tournament result with a long-term family rating.
Suppose a family-level rating combines several years of Stockfish versions, but the final was played by one dated development binary. The event winner is that binary. The long-term rating belongs to the aggregation policy.
The two identities overlap, but they are not identical statistical objects.
The same issue applies when multiple evaluation backends or neural files are used.
A binary-network combination can differ materially from another combination carrying the same general engine name.
Publication should therefore distinguish:
- event identity, meaning the exact competitor entered;
- rating identity, meaning the entity to which the database assigns games;
- family identity, meaning the broader project or lineage.
Without that separation, a result can be technically correct but semantically misleading.
12. A worked hypothetical example
Consider a fictional 16-engine league followed by a 50-game final.
Engine Atlas begins the event with a long-term rating of 3505.
Engine Borealis begins at 3490.
The 15-point difference suggests only a modest expected advantage for Atlas. It does not imply domination.
During the league:
- Atlas scores 62 points;
- Borealis scores 61.5;
- both qualify for the final;
- Borealis has performed particularly well against the strongest half of the field;
- Atlas has achieved a slightly better overall score.
The final uses 25 mirrored opening pairs.
Borealis wins 26–24.
What can be said?
Valid event statements
Borealis won the final.
Borealis is the tournament champion.
Borealis scored 52 percent in the final.
Valid rating statements before recalculation
Atlas entered the final with the higher long-term rating.
The pre-event list estimated Atlas slightly above Borealis within its database.
The rating did not guarantee the final result.
Statements requiring recalculation
Borealis is now higher-rated than Atlas.
The final increased Borealis by a particular number of Elo points.
Atlas is no longer the statistical favourite.
These claims cannot be made until the games are accepted and the rating model is rerun.
Suppose the recalculated list becomes:
- Atlas: 3503 ± 8;
- Borealis: 3497 ± 9.
Borealis remains the champion, while Atlas retains the higher central long-term estimate.
There is no contradiction.
The final asks:
Who scored more in these 50 games?
The list asks:
Which relative-strength values best explain the larger accepted database?
Suppose instead that Borealis dominates 34–16. After inclusion, its rating may move above Atlas. Even then, the event result and the rating calculation remain separate records. The title is determined by the final score; the rating change is determined by the model and all accepted games.
13. Winners pages and rating pages serve different evidential roles
A technically mature publication system should not force one page to perform every function.
Winners
A winners page should answer:
- who won;
- which event;
- which track;
- which time control;
- which opponent or field;
- what score;
- and when the result became official.
It is a historical title register.
Archive
An archive should preserve the event as a closed object:
- participant list;
- stage structure;
- dates;
- final standings;
- winner;
- audit status;
- and links to evidence.
Downloads
The download surface should expose the final evidence package:
- PGN;
- result files;
- checksums;
- rating outputs;
- event configuration;
- and audit report where available.
Rating Lists
The rating page should publish:
- rating method;
- numerical base;
- engine identity policy;
- rating;
- uncertainty;
- games;
- opponent context;
- calculation date;
- and provisional or final status.
Blog
The blog can connect the layers in a readable event narrative.
It may explain that the champion defeated the pre-event rating favourite, but it should not transform that observation into an unsupported universal-strength claim.
This division is not bureaucratic. It protects meaning.
14. The IJCCRL publication architecture
Within IJCCRL, tournament performance and rating evidence should move through a defined public chain.
The current chess engine rating lists provide the numerical surfaces.
The Rules and Audit page defines the competition and publication grammar.
The Downloads area preserves accessible evidence packages.
The Archive records closed events and historical snapshots.
The Winners surface records championship identity.
This architecture allows a reader to move from a claim to its supporting layer.
For example:
Engine X won the IJCCRL Blitz Final.
The link should lead to Winners or the archived final.
Engine X is rated first in the Original UCI Blitz list.
The link should lead to the relevant rating table.
The final passed the mirrored-opening and audit requirements.
The link should lead to Rules & Audit and the event report.
The calculation used these 400 accepted games.
The link should lead to the downloadable PGN and rating output.
The architecture also prevents one event result from silently replacing the long-term statistical record.
15. Observation, calculation and interpretation
Every public rating claim contains three layers.
Observation
The observation is what happened:
- a game result;
- a final score;
- an engine crash;
- an opening assignment;
- a completed mirrored pair;
- or an event standing.
Calculation
The calculation transforms accepted observations:
- total points;
- performance rating;
- Elo estimate;
- confidence interval;
- likelihood of superiority;
- or recalculated standings.
Interpretation
The interpretation is the editorial sentence:
- “Engine A won the tournament.”
- “Engine B remains higher rated.”
- “The rating gap is not statistically resolved.”
- “The event suggests improvement but does not establish a universal ordering.”
Keeping the layers separate prevents the interpretation from becoming stronger than the evidence.
A corrected PGN may change the calculation.
A recalculation may change the rank.
A more precise explanation may change only the interpretation.
A trustworthy publication preserves all three.
16. How to quote tournament and rating evidence correctly
Before publishing a comparison, ask five questions.
1. Which exact engine competed?
Use the version, build or commit, not only the family name.
2. What exactly did it win?
Name the event, stage, field and time control.
3. What does the rating represent?
Identify the project, rating tool, numerical base, engine aggregation policy and calculation date.
4. How much evidence supports the rating?
Check game count, uncertainty, opponent coverage and publication status.
5. Are the tournament and rating conditions compatible?
Compare hardware, threads, hash, pondering, openings, tablebases and adjudication.
A responsible sentence might read:
Borealis 4.2 won the 2026 Original UCI Blitz Final by 26–24. Atlas 7.1 nevertheless retained the higher central estimate in the event-close long-term rating calculation, whose uncertainty intervals substantially overlapped.
This wording preserves both records.
An irresponsible sentence would read:
Borealis proved that every rating list was wrong.
The second claim is much broader than the evidence.
17. Reader checklist
Before quoting a tournament winner or a chess engine rating:
- Separate the championship title from the long-term rating.
- Identify the exact engine version or family policy.
- Check the tournament length and format.
- Examine time control, hardware and opening conditions.
- Read the rating method and numerical base.
- Check games, opponents and uncertainty.
- Determine whether event games entered the rating database.
- Preserve the calculation or snapshot date.
- Distinguish technical losses from ordinary game results.
- Use the correct canonical source: Winners, Archive, Downloads or Rating Lists.
Frequently asked questions
Can a lower-rated engine legitimately win a final?
Yes. Ratings express expected relative performance, not certainty. Finite matches contain sampling variation, and small rating differences imply only small expected-score advantages.
Does winning a tournament automatically increase an engine’s Elo?
No. The games must first satisfy the list’s inclusion policy and be added to the rating database. The rating must then be recalculated.
Is a tournament performance rating permanent?
No. It is a local estimate derived from one event’s score and opposition. It may differ substantially from a long-term rating.
Can the tournament champion remain below the runner-up in the rating list?
Yes. The title is determined by the event score. The rating is estimated from the broader accepted dataset.
Does losing to a lower-rated engine prove the rating system is inaccurate?
No. An unexpected result may be entirely compatible with the probabilities implied by the rating difference. Repeated systematic prediction failures would be more relevant evidence against the model.
Is a long tournament always better than a rating list?
They serve different purposes. A long head-to-head match can provide strong evidence about one pairing. A connected rating list provides broader comparisons across a population.
Should crashes count in ratings?
That depends on the declared objective. A practical competition may treat operational reliability as part of performance. A research list may publish technical outcomes separately. The policy must be explicit and consistent.
Can ratings from TCEC, CCRL, SSDF, Chess.com and IJCCRL be compared directly?
Not by simple numerical subtraction. Their pools, bases, hardware, time controls, inclusion rules and identity policies differ.
What should be quoted when reporting the event?
Quote the winner and final score for event history. Quote the dated rating table, uncertainty and games for long-term statistical placement.
Conclusion
A tournament result and a long-term chess engine rating are complementary forms of evidence.
The tournament records what happened in a bounded competitive event. It identifies the champion under a declared field, schedule, resource profile, opening policy and rule set.
The performance rating summarises what the engine’s event score implies relative to the assigned ratings of its opponents.
The long-term rating estimates relative strength from a broader connected population of accepted games.
These layers can agree. The highest-rated engine may enter as favourite, win the event and strengthen its position in the recalculated table.
They can also diverge. A lower-rated engine may win a final while the favourite retains the higher long-term estimate. That is not an inconsistency. It is the expected relationship between a realised finite outcome and a probabilistic longitudinal model.
Responsible computer-chess publication therefore avoids two opposite errors.
The first is to dismiss tournament winners because rating uncertainty exists. Titles are real historical outcomes.
The second is to treat every title as proof of a permanent universal hierarchy. Ratings, event formats and testing conditions impose narrower boundaries.
The correct public record preserves both:
- who won;
- and what the larger body of evidence estimates.
That distinction allows chess engine tournaments to remain meaningful competitive events while rating lists remain disciplined statistical publications.
Sources and technical references
- Cattelan, Manuela. “Models for Paired Comparison Data: A Review with Emphasis on Dependent Data.” Statistical Science, 27(3), 412–433.
https://doi.org/10.1214/12-STS396 - Szczecinski, Leszek, and Aymen Djebbi. “Understanding and Pushing the Limits of the Elo Rating Algorithm.”
https://arxiv.org/abs/1910.06081 - Hamilton, Alexander H. “The Impact of Intransitivity on the Elo Rating System.”
https://pmc.ncbi.nlm.nih.gov/articles/PMC12742789/ - Coulom, Rémi. “Bayesian Elo Rating.” Official BayesElo documentation and software.
https://www.remi-coulom.fr/Bayesian-Elo/ - Ballicora, Miguel A. “Ordo: Ratings for Chess and Other Games.” Official source repository.
https://github.com/michiguel/Ordo - Top Chess Engine Championship. Official TCEC Rules.
https://wiki.chessdom.org/Rules - TCEC Games Archive. Official public game repository.
https://github.com/TCEC-Chess/tcecgames - Computer Chess Rating Lists. “CCRL 40/15 Testing Conditions.”
https://computerchess.org.uk/4040/about.html - Computer Chess Rating Lists. “CCRL 40/15 Downloads and Statistics.”
https://computerchess.org.uk/ccrl/4040/ - Sandin, Lars. “The SSDF Chess Engine Rating List, 2023-05.” ICGA Journal, 45(1), 28–30.
https://doi.org/10.3233/ICG-230231 - Chess.com ComputerChessChamps. “Chess.com Computer Ratings: November 2020.” Contextual industry publication.
https://www.chess.com/article/view/chess-com-computer-ratings-nov-2020

Jorge Ruiz Centelles
Filólogo y amante de la antropología social africana
