PokerBench

Small language model training and evaluation for poker decision-making under uncertainty.

2025.08 - 2025.12

Overview

Poker is a challenging testbed for language models because it requires strategic reasoning under uncertainty, hidden-information reasoning, and precise numeric action prediction. This project explores whether small language models can learn poker decision-making behavior through task-specific fine-tuning.

Method

We fine-tuned Qwen3, LLaMA-3.2, and Gemma variants on PokerBench, an instruction-action dataset of solver-labeled poker decisions. The project included zero-shot and few-shot baseline evaluation, QLoRA-based fine-tuning, and metric-based comparison using Action Accuracy and Exact Match.

Figure 1. Fine-tuning and evaluation pipeline. PokerBench scenarios are formatted for QLoRA fine-tuning and evaluated with Action Accuracy and Exact Match.

Results

QLoRA fine-tuning substantially improved small language model performance compared with zero-shot baselines. Fine-tuned SLMs achieved strong Action Accuracy and Exact Match, showing that task-specific fine-tuning can help small models learn structured poker decision-making behavior under uncertainty.

Figure 2. Fine-tuned SLMs vs larger baselines. The comparison highlights the performance-resource trade-off across model scales.
Figure 3. Zero-shot vs fine-tuned performance. Fine-tuning substantially improves small language model accuracy on structured poker decisions.

Tech Stack

Python · PyTorch · Hugging Face Transformers · QLoRA · LoRA · bitsandbytes