Research code + experiments exploring search and reinforcement learning for Blokus Duo. This project compares classical planning (Minimax, MCTS) and learned components (CNN evaluation/policy) across large-scale simulation.
Understand what makes agents strong at spatial, combinatorial games by combining:
- principled search algorithms,
- learned evaluation functions,
- and rigorous simulation-based evaluation.
This repository is meant to be:
- a research/engineering artifact (reproducible experiments, analysis notebooks),
- a portfolio piece showcasing applied RL + search,
- and a foundation for future work (stronger training loops, faster simulation, better evaluation).
- Ran hundreds of thousands of simulations across multiple agent matchups
- Compared:
- Minimax variants
- MCTS (with/without neural guidance)
- PPO (baseline RL approach)
The README and notebooks summarize findings; the paper provides a deeper narrative.
- Main paper: https://docs.google.com/document/d/1t95HiT_Vk48AHe5BSArWpV6ObkvVYfmjIF5mwMdtEJw/edit?usp=sharing
(Consider also exporting a PDF intodocs/so the repo is self-contained.)
Recommended to add:
docs/images/board.png— sample board statedocs/images/pieces.png— pieces overview (you already havepieces*.png)docs/images/results.png— headline plot from comparisons
Example:
docs/images/
├── board.png
├── pieces.png
└── results.png