“AI beat professionals,” “a game was solved,” and “this program computes a strategy” are three different claims. The research becomes easier to navigate when you keep the exact game and the kind of evidence beside each result.
Read the question before the algorithm.
For each paper, write the number of players, betting structure, objective and evaluation method. A two-player zero-sum model, a six-player chip game and a tournament with payouts are not interchangeable. A result about one does not automatically become a result about the others.
Begin with the abstract and the definition of the task. When the prerequisites are unfamiliar, use the introductory CFR teaching material or the books and courses guide. You do not need to implement an algorithm before understanding what its claim is limited to.
Regret and sampling: a foundation, not a certificate.
The 2007 CFR paper introduces counterfactual regret as a way to work with large imperfect-information games. The 2009 Monte Carlo CFR paper develops sampling variants and explains expected-update and regret properties. These are useful starting points for understanding why repeatedly updating strategies can be more than an arbitrary heuristic.
The word “sampling” does not mean every sampled run has reached a strong strategy. To assess a concrete run, you still need its game definition, implementation, stopping criterion and evidence. A method's convergence statement is not the same as a finite-run guarantee for an unexamined program.
Keep the milestone's game in its name.
| Research | Scope to retain | Do not conclude |
|---|---|---|
| Cepheus | The Alberta publication archive describes heads-up limit hold'em as essentially weakly solved. | That all no-limit or tournament poker is solved. |
| Libratus | The authors' IJCAI record describes architecture and reported performance in heads-up no-limit hold'em. | That its result covers six-player tournament payouts. |
| DeepStack | The original paper concerns deep reasoning in heads-up no-limit poker. | That any neural poker model has the same demonstrated performance. |
| Pluribus | CMU's research account describes six-player no-limit experiments and limited-lookahead search. | That empirical strength proves every multiplayer strategy is an equilibrium or covers every payout structure. |
The sources below preserve these distinctions. The institutional Pluribus summary is an accessible account, not a substitute for the full paper and supplement when checking a technical claim. The Alberta archive contains many papers; its inclusion does not mean every one was re-read for this guide.
Ask how performance was measured.
A reported match result depends on the game, opponents, sample and evaluation procedure. AIVAT is linked because it studies variance reduction for evaluating agents in imperfect-information games. Reducing measurement noise is a different contribution from improving the playing strategy itself.
When you see a win rate or error bar, ask what the observational unit is, how the experiment was designed and what assumptions enter the estimator. Do not read “statistically significant against these opponents” as “best possible against every opponent.” Those are different questions even when the reported experiment is sound.
Separate an environment, an evaluator and a solver.
A hand evaluator tells you which legal hand wins. A game environment implements actions and state changes. A solver seeks a strategy under an objective. A training interface presents exercises against a reference. One component working does not validate the entire chain.
OpenSpiel is a documented framework for research in games. The directory also links PokerKit, the Poker Hand History specification and a hand-history dataset. Read the specific game support, version, license and data provenance before using them. No repository or dataset is described here as independently executed or audited.
Leave a reproducible reading note.
Record the title, version, original publication link, exact game, main claim, evidence and one unresolved question. Separate the authors' claims from your inferences. A useful outcome is knowing which result does not apply to your problem.
This is a public research-reading map. It does not publish private solver code, training data or unvalidated charts. Study and experimentation belong away from active play and must follow the relevant operator's rules.
A six-player AI won a reported match. Is every six-player tournament now solved?
No. You must retain the experiment's game and payoff structure. Demonstrated performance in a particular chip game does not establish a solution for every tournament payout configuration, stack arrangement or opponent.
Original publication records and research documentation
Public abstracts, institutional descriptions and documentation pages provide the entry points. This guide is not a proof audit or experimental reproduction.
NeurIPS / Zinkevich and colleagues
Counterfactual regret minimization: original paper
The 2007 publication introducing counterfactual regret and a method for solving large imperfect-information games.
Before you use it: Publication page and abstract read, not a proof audit or reproduction. Equilibrium conclusions require the paper's game assumptions.
Public page read · · Full directory listing
NeurIPS / Lanctot and colleagues
Monte Carlo CFR: sampling and regret
The 2009 paper introducing a family of sampling-based CFR algorithms, including outcome and external sampling.
Before you use it: Publication page and abstract read. Expected updates and convergence guarantees do not certify a particular finite training run.
Public page read · · Full directory listing
University of Alberta CPRG
University of Alberta computer poker publications
An original research-group archive linking papers and abstracts on Cepheus, DeepStack, evaluation and game solving.
Before you use it: Historical publication archive, not a current product list. The Cepheus result concerns heads-up limit hold'em, not every form of poker; linked papers were not all audited.
Public page read · · Full directory listing
IJCAI / Noam Brown and Tuomas Sandholm
Libratus: architecture and heads-up no-limit poker
The authors' 2017 proceedings page outlines blueprint computation, nested subgame solving and self-improvement.
Before you use it: Abstract and publication record read. The reported game is heads-up no-limit hold'em; this is not a multiplayer tournament solution or an implementation test.
Public page read · · Full directory listing
Carnegie Mellon University
Pluribus: six-player poker research explained
The research institution's 2019 account of Pluribus, its six-player experiments and limited-lookahead search.
Before you use it: Institutional research summary, not the full proof or a reproduced experiment. Reported performance is distinct from solving every multiplayer game or payout format.
Public page read · · Full directory listing
Moravcik and coauthors / arXiv
DeepStack: imperfect-information reasoning in poker
Research on recursive reasoning, decomposition, and learned value estimates in heads-up no-limit hold'em.
Before you use it: Abstract and metadata inspected. Heads-up research is not a six-player tournament solution or a verified product recommendation.
Public page read · · Full directory listing
Burch, Schmid, Moravcik, and Bowling / arXiv
AIVAT: reducing variance in agent evaluation
The authors' paper on evaluating agents using value estimates and known strategies to reduce outcome variance.
Before you use it: Abstract and metadata inspected, not a full-paper audit or replication. The listed method has information and modeling requirements.
Public page read · · Full directory listing
OpenSpiel
OpenSpiel documentation
A research framework for games and learning algorithms, including imperfect-information settings.
Before you use it: A general research toolkit, not a ready-made tournament strategy recommendation.
Public page read · · Full directory listing
Prepared with AI assistance. Published . Worked examples are fictional, not Nolan's hands or results. Sources and review limits accompany each listing. Use study material before or after play, not as live assistance.