♜REGICIDESOLO / RESEARCH LAB LOCAL PLAYGROUND
ONE HAND. TWELVE ENEMIES.

A little strategy.
A little regicide.

THE CASTLE
♜

Health
Attack
Shield

Your hand

Select a card to play.
♥HealDiscard → deck bottom
♦DrawRefill up to eight
♠ShieldPersists this enemy
♣StrikeDouble damage
Remembered cards & known positions
IDEAS → EXPERIMENTS → EVIDENCE

The research bench.

Compare policies on shared deals. Keep conclusions proportional to the evidence.

View the strategy performance report ↗ · MC speed & resource report ↗ · 95% target study ↗

HYPOTHESIS 01

Keep the cards flowing.

Does stronger preference for drawing and healing improve survival and wins? Change one factor at a time, then compare on the same development deals.

Weights affect heuristic play, rollout continuations, and search candidate ranking. Changes apply to this browser session.

Learning from Monte Carlo

“MC-taught linear” and “MC-taught neural” use learned action preferences to decide where to spend simulations. Hints show the linear model’s main contributions. Both still finish simulated games under the full rules. Compare the models and inspect their training →

Mixed rollout exploration

Applies to “Monte Carlo · mixed policies”. Choose a random policy for a whole simulated game, or occasionally choose a random legal move inside a guided simulation. These affect imagined futures only.

Search allocation

For guided-allocation and MC-taught agents. Higher values spend more trials on alternatives; zero chooses greedily after initial samples.

Search stops at the simulation cap or this time ceiling, whichever comes first. Zero disables the time limit.

Search considers the top candidates by heuristic score. A larger shortlist broadens exploration but gives each action fewer samples.

CURRENT POSITION

Where do agents disagree?

Compare recommendations on the exact same observation. Hidden deck order is never sent to an agent.

Interpretation matters.

Heuristic scores and policy preferences are not win probabilities. Rollouts estimate a particular continuation policy. Search is approximate. No agent here is claimed optimal.

REPRODUCIBLE RUNS

Experiment ledger

Only fresh headless games enter evaluations. Rewound, hinted, and manually played sessions are kept separate.

WORKING NOTES

Research log

Keep hypotheses separate from findings. Notes stay in this project.

Run an experiment

Open the strategy performance notebook report →

From the project folder, run these commands. Results include configurations, version fingerprints, timings, intervals, and replayable failures.

python3 -m regicide evaluate --agents random heuristic --games 100 --split dev
python3 -m regicide evaluate --agents heuristic rollout tree hybrid --games 1000 --offset 10000 --budget 128 --time-ms 100 --workers 2
python3 -m regicide train --episodes 300 --output models/policy.json
python3 -m regicide evaluate --agents heuristic learned --model models/policy.json --split heldout --games 100

Both budgets are ceilings: search stops at the simulation cap or time ceiling, whichever comes first (one rollout can overshoot). Deterministic simulation budgets are preferable for exact reproducibility. See README for ablations and learning details.

BASE GAME · OFFICIAL SOLO VARIANT

Know the table.

The turn

  1. Play one card, a legal matching-number combination, or an Ace with one other card.
  2. Resolve mandatory suit powers, excluding the enemy’s own suit. Hearts resolve before Diamonds.
  3. Deal damage. If the enemy dies, skip retaliation. An exact kill puts the royal on top of your draw pile.
  4. If the enemy survives, discard enough value to cover its remaining attack. Spade shielding persists until that enemy dies.

Combos use 2–4 cards of one number with total value at most 10. Aces are worth 1 and can only be alone or paired. Jacks, Queens and Kings in hand are worth 10, 15 and 20.

Two chances to refresh

Use a Jester before playing or before damage payments. Discard your hand and refill up to eight. It bypasses Diamond immunity for the refill, but never cancels any immunity. It does not heal the discard pile.

A win with zero, one or two Jesters used earns Gold, Silver or Bronze respectively. The initial research objective rewards every win equally.

Information conventions

We show the top of the discard, remembered cards, and possible identities after healing. We never reveal healed identities, future enemy order, or unseen draw order. Solo yielding is disabled. Consecutive Jesters are allowed before committing to the phase.

Discard inspection and consecutive Jester timing are documented conventions pending further publisher clarification. All comparisons record this rules version.

Sources & scope

Rules verified against the publisher’s official rulebook and publisher clarification on damage discards. See docs/RULES.md for decisions and unresolved edges.

This is an independent local research project for the original Regicide, by Paul Abrahams, Luke Badger and Andy Richdale. It does not implement Regicide Legacy or Inferno. No official artwork is used.