Learn · Rulebooks
Test a rulebook
Replay a category corpus through the pack and read the variance it produces.
Checked against the product on · written for people publishing rule packs
A pack is measured on steadiness: given the same event twice, does it apply the same labels. The harness replays a corpus of sessions through the pack several times and counts the sessions whose runs disagree.
The measure is the set of labels a pass proposes. A session counts as unsteady when any repeat produces a different set.
| Grade | Sessions that disagree | Reads as |
|---|---|---|
| S | Under 0.1% | The pack answers the same way every time. |
| A | Under 1% | Steady across a large corpus. |
| B | Under 5% | Steady on the common cases. |
| C | Under 15% | Steady on the clear cases. |
| D | 15% or more | The pack answers differently on the same input. |
Run the harness
Pick the category your pack belongs to.
A category holds the templates and the replay corpus the harness uses.
Pick a template and a repeat count.
Start the run.
The replay happens in a sandbox. Nothing is written to any work area.
Read the variance and the sessions that disagreed.
Each one names the labels that moved between runs.
The report card
A grade is set at publication. A report card is how the pack is doing in production afterwards, across every work area that installed it. The report card is visible to buyers and leaves the grade alone.