PokerBench
Small language model training and evaluation for poker decision-making under uncertainty.
2025.08 - 2025.12
Overview
Poker is a challenging testbed for language models because it requires strategic reasoning under uncertainty, hidden-information reasoning, and precise numeric action prediction. This project explores whether small language models can learn poker decision-making behavior through task-specific fine-tuning.
Method
We fine-tuned Qwen3, LLaMA-3.2, and Gemma variants on PokerBench, an instruction-action dataset of solver-labeled poker decisions. The project included zero-shot and few-shot baseline evaluation, QLoRA-based fine-tuning, and metric-based comparison using Action Accuracy and Exact Match.
Results
QLoRA fine-tuning substantially improved small language model performance compared with zero-shot baselines. Fine-tuned SLMs achieved strong Action Accuracy and Exact Match, showing that task-specific fine-tuning can help small models learn structured poker decision-making behavior under uncertainty.
Tech Stack
Python · PyTorch · Hugging Face Transformers · QLoRA · LoRA · bitsandbytes