Skip to content
Docs / Playtest your model

Playtest your model

Before your model plays for real, let your own game play with it. canonopy playtest runs on your machine: the only thing of ours that runs there is the finished model you downloaded, and only the numbers your program prints come back.

bash
canonopy playtest my-bot@latest --cmd "python my_game.py" --episodes 20

It needs canonopy-runtime (see Running it yourself).

How your game talks to the model

  1. The model is served at a small local endpoint, on 127.0.0.1 only. Your program sends the same JSON you send Jev and gets the usual answer:
python
import json, os, urllib.request

def decide(state):
    req = urllib.request.Request(os.environ["CANONOPY_DECIDE_URL"], data=json.dumps({"state": state}).encode(),
                                 headers={"Content-Type": "application/json"})
    return json.loads(urllib.request.urlopen(req).read())["answers"]
  1. Your command runs once per episode (or once for all of them with --once), with CANONOPY_DECIDE_URL, CANONOPY_EPISODE (1, 2, …), CANONOPY_EPISODES and CANONOPY_SEED set.
  2. At the end of each episode it prints one JSON line of numbers: anything you measure.
python
print("CANONOPY_RESULT " + json.dumps({"score": 1240, "deaths": 1, "stuck": 0}))

The CANONOPY_RESULT prefix is optional: without it, the episode's last line that is a JSON object counts. Everything else your program prints is shown, untouched.

What comes back

The numbers go to your workspace (POST /v1/domains/{domain}/play-results), summarised per number: the average, the lowest and the highest. The advice shows your last playtest, and you can compare after a retrain.

POST/v1/domains/{domain}/play-results
GET/v1/domains/{domain}/play-results

Improve from play (optional): --keep-situations 200 also sends a sample of up to that many states your game sent to the model (300 at most). Your next retrain answers them with your model's own instructions and learns them, so it gets better where your game really goes.

flagwhat it does
--episodes Nepisodes to play (default 10)
--oncerun your command once; it plays every episode and prints a line for each
--timeout Sseconds per episode (default 600)
--port Pthe local port (default: any free one)
--keep-situations Nalso send a sample of the states your game hit
--no-sendprint the results, send nothing
--name MODELwith a local .zip: the model to send the results to

Without the CLI: canonopy-runtime playtest MODEL.zip --cmd "..." --episodes 20 prints the same results, and your coding agent can send them with report_play_results.