Production deployment
The bot and API run on a machine you own, behind a Cloudflare Tunnel. There is no
public IP, no forwarded port and no cloud compute bill. cloudflared dials out to
Cloudflare’s edge and traffic returns down that connection.
Browser ──https──> Cloudflare edge [TLS, WAF, rate limiting, challenge] ↕ outbound tunnel — no inbound port on the host cloudflared ──http──> api:8787 ──> db:5432 bot ──outbound WS──> DiscordThis keeps two properties that a serverless split would lose:
- The Discord gateway stays a long-lived connection, with no HTTP-interactions rewrite.
- The spend cap, concurrency semaphore and rate limiter stay in-memory values in one process rather than becoming distributed state. A budget period adds one small table beside the cap, not a service.
See docs/ARCHITECTURE.md for the pipeline itself.
Hostnames, paths and the image name below are placeholders (judge.example.com,
/path/to/mtg-judgebot, ghcr.io/<owner>/<repo>). Substitute your own.
1. Prerequisites
Section titled “1. Prerequisites”- A host that stays on, with Docker and the compose plugin. The stack runs in about
200 MB RSS (Postgres ~157 MB, api and bot a few MB each), so 2 GB of RAM is ample.
It never has to build: CI publishes the image and the host pulls it (§8). That
matters because
cargo build --releaseacross ten crates plus a Vite build wants ~4 GB and a CPU that a NAS does not have. - Compose syntax here is held to what older bundled versions accept. Synology’s
Container Manager ships v2.20, which predates the
env_filelong form..env.deploymust exist on any machine running thetunnelprofile, and only there. - No Rust toolchain on the host. The image carries its own migrations.
botandapiapply pending ones at startup (JUDGE_AUTO_MIGRATE, on by default).docker compose run --rm refresh migrateis the explicit form, for an empty database or an operator who opted out. Restoring a dump (§2) brings the schema and the_sqlx_migrationsledger with it. - A domain whose DNS is hosted on Cloudflare. Tunnel hostnames resolve only for
records in the same Cloudflare account, so third-party DNS cannot CNAME to
<uuid>.cfargotunnel.com. The free plan requires moving the whole zone.
2. Data migration
Section titled “2. Data migration”A new instance has nothing to migrate. Skip to §3, and load the data in §4 with
docker compose run --rm refresh init (the documentation site’s “Requirements and first
run” page is the walkthrough). This section is
for moving an instance that already has data to another host.
When moving, do this before anything else. Restore a dump rather than re-ingesting. A cold rebuild re-parses the CR and the Scryfall bulk file and re-embeds every rule through the embedding provider, which costs money.
# old hostdocker exec judgebot-db pg_dump -U judgebot -Fc judgebot > judgebot.dump
# new hostdocker compose up -d dbdocker exec -i judgebot-db pg_restore -U judgebot -d judgebot --clean --if-exists \ < judgebot.dumpCheck the restore before moving on:
docker compose exec -T db psql -U judgebot -d judgebot \ -c "select (select count(*) from cards) cards, (select count(*) from rulings) rulings, (select count(*) from rules where embedding is not null) embedded, (select provider||'/'||model||'/'||dimensions from embedding_space) space;"embedded being 0 means the vector leg is off. Re-run scripts/refresh-data.sh embed
rather than shipping a degraded retriever.
The vector leg is also off when space is empty, or names a model other than the one
the bot is configured with. The bot’s vector legs stay dark until the row and the
configuration agree. It logs an error-level line at startup and on the change, and it
never mixes a column. The fix depends on the cause:
- The row is mislabelled:
UPDATE embedding_space SET model = .... - The model is changing:
scripts/refresh-data.sh reembed --yes. This is paid, because it embeds every row again. - The row is already right: the same command only fills whatever is still empty.
3. Create the tunnel
Section titled “3. Create the tunnel”In Cloudflare Zero Trust → Networks → Tunnels, create a tunnel (remotely managed) and add a public hostname:
| Field | Value |
|---|---|
| Subdomain | judge |
| Domain | example.com |
| Service | http://api:8787 |
api is the compose service name. cloudflared resolves it on the compose network, so
the API never needs a published port. Cloudflare creates the proxied
judge CNAME <uuid>.cfargotunnel.com record for you. It must stay proxied
(orange cloud), unlike every other record in the zone.
Copy the connector token into .env.deploy as TUNNEL_TOKEN.
4. Configure and start
Section titled “4. Configure and start”Configuration lives in two separate files:
cp .env.example .env # app config: API keys, DISCORD_TOKEN, GUILD_ID, # JUDGE_OPERATOR_DISCORD and JUDGE_OPERATOR_EMAIL (both required here)cp .env.deploy.example .env.deploy # deploy credentials: TUNNEL_TOKEN, R2_*.env is the env_file for bot, api and refresh. .env.deploy is read only by
cloudflared and scripts/backup-db.sh. A token that can rewrite the tunnel or
delete every backup therefore never enters the environment of the internet-facing API.
Both files are gitignored.
Model credentials belong in .env: ANTHROPIC_API_KEY, every api_key_env a
judge.toml names, and the cloud doors’ AWS_*/GOOGLE_APPLICATION_CREDENTIALS.
bot, api and refresh read them, and nothing else does.
In .env, set:
COMPOSE_PROFILES=tunnel # `docker compose up -d` now includes cloudflaredAPI_CLIENT_IP=cloudflare # rate-limit on CF-Connecting-IPJUDGE_MAX_USD=... # the backstop for anonymous trafficJUDGE_BUDGET_PERIOD=month # one budget for bot and api, kept across restartsJUDGE_ALERT_WEBHOOK=... # told when the cap trips, or a refresh or backup failsWithout JUDGE_BUDGET_PERIOD the cap is per process and per lifetime: bot and api
can each spend JUDGE_MAX_USD, and every restart (a redeploy, a crash loop) starts them
again from zero. On a host that runs unattended, set a period. judge-cli stats shows
what each day cost.
If COMPOSE_PROFILES in .env doesn’t take effect on an older Compose, pass
--profile tunnel instead.
Then bring it up. A restored database (§2) already carries the schema and the migration ledger, so there is nothing to migrate:
docker compose up -dcurl -s localhost:8787/api/healthStarting from an empty database, docker compose up -d is also enough: bot
and api create the schema at startup. The explicit form is for a look at what is
about to happen, or for JUDGE_AUTO_MIGRATE=false. It runs from the same image and
needs nothing but Docker (run starts db if it is not up):
docker compose pulldocker compose run --rm --pull missing refresh migrate # never falls back to building on the hostsqlx migrate run --source crates/bot/migrations from a workstation, over an SSH
tunnel to the loopback-bound Postgres, writes the same ledger with the same checksums.
It takes neither the refresh-job lock nor the ahead check, so prefer the container form
on a live host.
Use a full docker compose up -d whenever the compose file’s db service changes.
up -d --build bot api leaves db alone, so a changed port binding or healthcheck
would otherwise persist indefinitely.
Other model providers
Section titled “Other model providers”This step is optional. Without a judge.toml the containers run Anthropic direct with
ANTHROPIC_API_KEY and Voyage with VOYAGE_API_KEY, as .env has them. A
judge.toml runs the judge on other providers: a cheap model for extraction, Claude
through your own cloud account, a local Ollama, an OpenAI-compatible gateway.
judge.example.toml documents every knob, and the README’s “Choosing a model” is the
short version. Write the file and name it in .env:
JUDGE_CONFIG=./judge.toml # a host path: what `cargo run` reads, and what compose mountsdocker-compose.yml bind-mounts that file read-only into bot, api and refresh at
/etc/judgebot/judge.toml and points the containers’ JUDGE_CONFIG there, so the one
variable serves the host and the containers. With it blank, compose mounts the tracked
judge.example.toml instead, so that the mount has a source. The loader reads
nothing it was not pointed at, and the setup stays the .env one. Three consequences:
-
The
api_key_envof every provider a stage names must be set in.env, including forrefresh. Each binary resolves all three stages ([models.extract],[models.synth],[models.embed]) at load. Sorefreshfails its nightlyembedstep on a chat key it never uses rather than run half a configuration. A provider table no stage names is parsed but its key is never read..envis theenv_filefor all three containers, so one line there covers them. -
The cloud doors take credentials from the platform chain, not the file. For
claude-platform-on-awsandbedrock, either putAWS_ACCESS_KEY_ID,AWS_SECRET_ACCESS_KEY(andAWS_SESSION_TOKEN) in.env, or mount a credentials file and name it. The containers run asnobodywith no home directory, so the default~/.awslocation does not exist:# docker-compose.override.yml (gitignored like judge.toml; compose merges it in by itself)services:bot: &awsvolumes: ["/home/you/.aws:/etc/aws:ro"]environment:AWS_SHARED_CREDENTIALS_FILE: /etc/aws/credentialsAWS_CONFIG_FILE: /etc/aws/configAWS_PROFILE: judgebotapi: *awsFor
vertex, do the same with a service-account JSON andGOOGLE_APPLICATION_CREDENTIALS=/etc/gcp/sa.json. An IAM role scoped to invoking the model is enough, because nothing here manages infrastructure.refreshneeds none of this: embeddings are Voyage or OpenAI-compatible, never a cloud door. None of it goes in.env.deploy, whichbot/apido not read. -
The startup log tells you what resolved. Every binary logs one
config=... extract=... synth=... embed=... cap=$...line.bot,api,evalandjudge-climake chat calls, so they then logcloud credentials resolvedper cloud provider. The chain is probed once at startup, so a host with no credentials exits there naming the provider and the door.refreshmakes no chat call and never probes. Those four binaries then logembedding space matches the databaseorembedding space mismatch(below).docker compose run --rm --entrypoint judge-cli api configprints the resolution as JSON, secrets redacted.
A changed judge.toml is read at the next start. That means docker compose restart bot api, not up -d. Compose recreates a container only when its configuration or image
changed, and the content of a bind-mounted file is neither. up -d prints Running
and leaves the old configuration in place. (Changing JUDGE_CONFIG itself in .env
does change the configuration, and up -d recreates.) The next refresh run picks the
file up on its own. Changing [models.embed] is the one edit that needs the database
moved too (§7).
An older Compose may reject the ${JUDGE_CONFIG:+…} interpolation in
docker-compose.yml at parse time. If so, set the containers’ side by hand in the same
override file: environment: {JUDGE_CONFIG: /etc/judgebot/judge.toml} on bot, api
and refresh, with the mount left as it is.
MCP endpoint
Section titled “MCP endpoint”This step is optional. judge-api can serve the judge’s tool surface (crates/agent)
to an MCP client over the same tunnel, at /mcp. Turning it on takes two things: the
--mcp interface and an MCP_TOKEN. --mcp without a token is refused at startup.
A token without --mcp logs a startup warning, so a token in .env never reads as an
endpoint that is not there. There is no anonymous mode. Every request must carry
Authorization: Bearer <MCP_TOKEN> or gets a 401 before the protocol sees it. Behind
the token are:
judge: the full pipeline with model spend, under the sameJUDGE_MAX_USDandJUDGE_CONCURRENCYas the web page.- The agent-driven sessions: no model calls, database work only.
- The read-only lookups.
API_INTERFACES=--api --web --mcp # the api container's front doors; without --mcp the token only warnsMCP_TOKEN=<openssl rand -base64 32> # at least 24 characters, or the API refuses to startMCP_ALLOWED_HOSTS=judge.example.com,localhost # Host values accepted: the tunnel's hostname, plus # localhost for curl on the host; the list replaces the defaultMCP_ALLOWED_HOSTS matters. The MCP transport validates Host against a loopback-only
default (a DNS-rebinding guard), and cloudflared forwards the public hostname. An
empty list therefore means every /mcp request is refused with 403 while /api/judge
keeps working. Then run docker compose up -d api and, from a workstation:
claude mcp add --transport http judge https://judge.example.com/mcp --header "Authorization: Bearer <MCP_TOKEN>"The per-IP rate limit of /api/judge does not apply to /mcp, where the token is the
identity. Instead judge runs through /mcp are capped per window
(MCP_JUDGE_LIMIT per MCP_JUDGE_WINDOW_SECS, default 20 an hour), on top of the
shared JUDGE_CONCURRENCY slots and JUDGE_MAX_USD cap. That cap is the blast radius
of a leaked token: about MCP_JUDGE_LIMIT × $0.12 an hour. It is never the whole spend
cap or every judge slot at once, so the public page keeps working.
Sessions and lookups make no chat-model call. They do embed the question or the search text with the configured embedder when there is one. That costs fractions of a cent and is uncapped, the only paid upstream without a cap.
Rotate a leaked token by changing .env and restarting api. Extending the edge
rate-limiting rule of §5 to /mcp costs nothing. A Cloudflare Access policy in front
of /mcp (service token) keeps unauthenticated traffic off the origin, and the bearer
check stays as the second layer.
A verdict an agent persists through a session is kept as history for that agent’s own thread. It is never shown to Discord or web askers as a prior-call example. Nobody can rate it, because there is no Discord message to vote on, and the answer text is the outside agent’s. Only the citations were validated.
A shell on the host can use the same tools without the network:
docker compose run --rm --entrypoint judge-cli api card "Blood Moon". The api
service’s entrypoint is judge-api, so run needs --entrypoint.
Card-symbol emoji
Section titled “Card-symbol emoji”The bot draws {W} as a picture using application emoji, which belong to the
Discord application rather than to any server. Discord stores them, so they survive
redeploys and restores. Upload them once per application, not once per deploy:
cargo run --release -p judge-ingest -- emoji # needs DISCORD_TOKEN; no databaseThe command is idempotent. It uploads only the symbols that are missing, so re-run it
after Scryfall adds one. Skipping it is safe: the bot logs a warning at startup
and falls back to writing {W} as text. The web page needs none of this, because it
loads the symbols from Scryfall’s CDN.
Client address for rate limiting
Section titled “Client address for rate limiting”The per-IP limiter needs an address the caller cannot choose, because /api/judge is
anonymous and every request costs Anthropic tokens.
X-Forwarded-For is not that address. Cloudflare appends the connecting address
to a caller-supplied X-Forwarded-For rather than replacing it, so its first hop is
whatever the caller wrote. A client sending X-Forwarded-For: 1.2.3.4 and incrementing
it per request would mint a fresh rate-limit allowance every time. It would do so
through the tunnel, which is the trusted path. Loopback binding does not help, because
the forged header rides in over the tunnel like any other.
Cloudflare sets CF-Connecting-IP on every request and the client cannot forge it, so
that is what API_CLIENT_IP=cloudflare buckets on. crates/api/src/http.rs
never consults X-Forwarded-For.
Leave API_CLIENT_IP=peer for any deployment where Cloudflare is not the sole ingress.
CF-Connecting-IP is trustworthy only when nothing can reach the origin directly.
Behind another reverse proxy
Section titled “Behind another reverse proxy”The tunnel is a choice, not a requirement. api publishes on 127.0.0.1:8787, and any
reverse proxy on the host can terminate TLS in front of it. Two things need care.
Answers take up to a minute, so the proxy’s read timeout must be longer than that. Caddy’s default is unlimited. nginx’s is 60 seconds.
Every request now arrives from the proxy, so API_CLIENT_IP=peer puts all visitors in
one rate-limit bucket. The per-address limit works again if the proxy tells judge-api
the client’s address in the one header it reads, overwriting whatever the client sent,
with API_CLIENT_IP=cloudflare. That is sound for the same reason as behind Cloudflare:
the header is set by something you control, and the origin listens on loopback only, so
nothing off the host can reach it. It holds only while that proxy is the public edge:
put Cloudflare in front of it as well and {remote_host} is Cloudflare’s address, so use
the tunnel setup instead.
judge.example.com { reverse_proxy 127.0.0.1:8787 { header_up CF-Connecting-IP {remote_host} }}server { server_name judge.example.com; # plus your TLS configuration location / { proxy_pass http://127.0.0.1:8787; proxy_set_header Host $host; proxy_set_header CF-Connecting-IP $remote_addr; proxy_read_timeout 120s; }}If the proxy does not set that header, leave API_CLIENT_IP=peer and rate limit at the
proxy instead (nginx’s limit_req), because the in-process limit then counts everyone
together. With --mcp, add the public hostname to MCP_ALLOWED_HOSTS as in the tunnel
setup. §5’s edge rules are Cloudflare’s, so the in-process limit, the concurrency slots
and the spend budget are what stand between an anonymous page and the model bill.
5. Edge spend protection
Section titled “5. Edge spend protection”With hosting at $0 the LLM bill is the entire bill. The in-process limiter is the backstop, and the edge is the front line.
- Rate limiting rule on
/api/judge: matchAPI_RATE_LIMIT/API_RATE_WINDOW_SECS(default 4 per 300s), or set it slightly tighter. The free plan includes one rule. - Managed Challenge as a WAF custom rule on the HTML document request, not on
/api/judge. A challenge served to anXHRcannot be solved byfetch, so challenging the API path breaks the page. Challenging the document gates a visitor once and subsequent/api/judgecalls carry thecf_clearancecookie.
Full Turnstile with server-side siteverify is stronger, but needs a token in the
POST body and a verification call inside judge_route before any spending. It is
worth it only if the edge rules prove insufficient.
With a budget period set, a capped instance comes back by itself when the period turns.
To resume sooner, raise JUDGE_MAX_USD and docker compose up -d. The period’s spend is
in the database, so the restart does not reset it.
6. Weekly backups to R2
Section titled “6. Weekly backups to R2”- Create an R2 bucket.
- Create an API token scoped to Object Read & Write on that bucket only.
- Fill in the
R2_*values in.env.deploy. - Install the cron entry:
crontab -e15 4 * * 0 /path/to/mtg-judgebot/scripts/backup-db.sh >> ~/judgebot-backup.log 2>&1Log to somewhere the running user can write. A >> into root-owned /var/log fails
before the script starts, leaving a backup that looks configured and never runs.
On Synology DSM, do not use crontab -e. DSM manages /etc/crontab in its own
format and can overwrite hand-edited user crontabs. Use Control Panel → Task
Scheduler → Create → Scheduled Task → User-defined script and set User: root
(Container Manager’s Docker socket is root-only). Give it absolute paths, since
the scheduler runs with a minimal environment:
/volume1/homes/<you>/mtg-judgebot/scripts/backup-db.sh \ >> /volume1/homes/<you>/judgebot-backup.log 2>&1If docker isn’t found, prefix the task with PATH=/usr/local/bin:$PATH. Tick the
task’s email-on-error option so a failing backup is noisy rather than silent.
scripts/backup-db.sh works in this order:
- It dumps and gzips the database.
- It refuses to upload anything under
BACKUP_MIN_BYTES, so a stub never becomes the newest restore point. - It uploads.
- Only then does it prune past
BACKUP_KEEP_DAYS.
Weekly runs at the default 60 days keep about eight restore points, far inside R2’s 10 GB free tier.
Run it once by hand to confirm credentials, then do a restore drill. An untested backup is not a backup. The drill restores the object that landed in R2, not a local copy:
scripts/backup-db.sh # take onescripts/backup-db.sh list # newest lastscripts/backup-db.sh fetch judgebot-<stamp>.dump.gz
docker compose exec -T db createdb -U judgebot restoretestgunzip -c judgebot-<stamp>.dump.gz \ | docker compose exec -T db pg_restore -U judgebot -d restoretestdocker compose exec -T db psql -U judgebot -d restoretest -c "select count(*) from cards;"docker compose exec -T db dropdb -U judgebot restoretestrm judgebot-<stamp>.dump.gzThe card count should match section 2. *.dump.gz is gitignored.
7. Scheduled data refresh
Section titled “7. Scheduled data refresh”Scryfall publishes new bulk data daily and Wizards ships a Comprehensive Rules
release with most sets. judge-ingest refresh brings the database up to date in one
unattended run. It runs inside the same image as bot and api: a third entrypoint,
compose service refresh, off by default behind the refresh profile. Its steps, in
order:
cards: Scryfall oracle cards, printed names and rulings. These are upserts, so cards the bot already knows are refreshed in place.rules latest: reads Wizards’ rules page, compares the linkedMagicCompRules <date>.txtagainstmax(rules.cr_version)and loads it only when the version differs.retire: re-checks every stored call’s citations against the data just loaded (see below).embed: only rows whose text changed. The CR loader nulls the embedding of those rows alone, so a new CR costs the embedder a few hundred rules, not all of them.emoji: uploads any card symbol Scryfall added. Skipped whenDISCORD_TOKENis unset.
Each step runs even if an earlier one failed, and the exit status is non-zero if any did.
With a judge.toml, refresh reads the same file bot/api do (compose mounts it
from JUDGE_CONFIG, §4). Its embed step writes the vector space the bot queries, and
refuses when the two disagree.
scripts/refresh-data.sh is the cron entry point. It takes a lock so two runs never
overlap, then runs docker compose run --rm --pull missing refresh. That reuses the
image docker compose pull already fetched and never builds on the host. Any argument
is passed through as the judge-ingest subcommand, so scripts/refresh-data.sh rules latest is a CR-only check.
Install it beside the backup, on the same scheduler. §6 has the Synology notes: run as root, absolute paths, tick email-on-error.
crontab -e30 5 * * * /path/to/mtg-judgebot/scripts/refresh-data.sh >> ~/judgebot-refresh.log 2>&1Daily is right for cards, because Scryfall corrects Oracle text and adds rulings
between sets. It also bounds how long a new CR goes unnoticed to a day. Most days it
costs one Scryfall download and nothing else. The run takes a few minutes on a NAS,
most of it parsing the default_cards file. It does not disturb the running bot: every
load is one transaction, so retrieval sees the old data or the new, never a mix.
Prior calls follow the data they cite. The retire step asks of every stored call the
question that admitted it: does each cited rule, ruling or Oracle text still exist and
still contain the quote? A call that fails is retired (calls.retired_at, with the
offending citation in retired_reason) and leaves retrieval. One whose citations hold
again later is restored. The effects:
- A CR release retires only the calls whose cited rules changed.
- An Oracle erratum retires the calls about that card, because each call remembers a fingerprint of its context cards’ text.
- A reworded ruling retires the calls that quoted it.
A rule that only moved is followed, as when Wizards inserts a keyword and the rest
of the section shifts by one. The rules step matches old and new rules
by text with rule numbers masked out. It rewrites the citations and answers of the
calls that cite the rule before the retirement check runs, so those calls stay live.
Every rule renumbered and call relocated is logged. To see what a run did:
select retired_reason, count(*) from calls where retired_at is not null group by 1;The retriever’s vector leg is blind to re-embedded rules for the minute between the
rules and embed steps. If embed fails (the embedder down, rate-limited) those
rules stay unembedded and the next night’s run picks them up, since embed always
fills every NULL.
Run it once by hand after installing, and expect the log to end with
refresh step ok five times. A one-off manual load also works from a
workstation (cargo run --release -p judge-ingest -- rules <url>). That is also how
to force a re-parse of an already-loaded version: delete the cached txt first.
Embedding model change
Section titled “Embedding model change”Vectors from two models cannot share a column. The database therefore records which
model’s vectors it holds (embedding_space, one row: provider kind, model, width), and
every reader and writer checks it first. A bot whose [models.embed] names a different
model or width does not mix. It logs embedding space mismatch; vector legs off and
answers from the curated map and full-text search alone until the two agree.
judge-ingest reembed moves the database to a new model. It pays the provider for
every rule, glossary entry and stored call again. That is why it is a dry run by
default, and why to take a backup first.
scripts/backup-db.sh # a restore point holding the old vectors (§6)$EDITOR judge.toml # [models.embed]: the new provider/model/dimensionsscripts/refresh-data.sh reembed # dry run: what is stored, what would be cleared, # rows, a rough cost; probes the new model once; # exits non-zero having changed nothingscripts/refresh-data.sh reembed --yes # one transaction: retype vector(N), rebuild the # HNSW indexes, clear every vector, rewrite the # row — then the ordinary embed loopdocker compose restart bot api # bot/api read judge.toml once, at startup; # `up -d` would see nothing to do (§4)The running bot/api hold the [models.embed] they started with, so a restart is
required at some point. They re-read the space row on every request. Their vector legs
go dark the moment the row disagrees with their configuration and come back the moment
it agrees, in either order:
- Restarting before
--yesdarkens them from the restart until the switch. - Restarting after
--yesdarkens them from the switch until the restart.
Either way there is one dark window and no mixing. The order above keeps it short.
The refill is resumable. If the embed loop dies (rate limit, a provider outage),
scripts/refresh-data.sh embed or the next nightly run fills whatever is still NULL.
Retrieval degrades to the other legs for the rows not yet embedded. reembed --yes
resumes too: with the row already switched it has nothing to switch, so it fills the
empty rows and pays for nothing twice.
Clearing and re-buying every vector in the same space takes --clear. The row does not
change, so the running bot logs no mismatch and its vector leg answers from nothing
until the refill finishes.
Resume outside the nightly refresh window (the cron above). Two refills at once each
buy the same batch, and one of them then fails on rows the other already filled.
If only the model name differs from the row (same provider, same width), the dry run
says so. When the configured model did produce the stored vectors,
UPDATE embedding_space SET model = ... relabels them for nothing, and --yes would
buy them all again.
The probe is what makes --yes safe to type. A wrong key, URL or model name, or a
model whose width is not the configured dimensions, fails before anything is cleared.
That matters because after the switch the only ways back are paying for the old space again or the
restore drill.
reembed needs the new judge.toml (JUDGE_CONFIG in .env) and the new provider’s
api_key_env in .env. refresh-data.sh passes its arguments through to
judge-ingest inside the refresh container, which already has both. The dry run’s
cost line is an order of magnitude at a generic list price, not a quote: ~$0.15 per
million tokens, chars-to-tokens at 4:1.
8. Redeploying
Section titled “8. Redeploying”The host never builds. .github/workflows/publish-image.yml builds on every push to
main that touches the image. That includes data/, since crates/core/build.rs
generates the Category enum from data/categories.yaml. The workflow pushes to
ghcr.io/<owner>/<repo> (the repository it runs in) as latest plus an
immutable sha-<short> tag. The image is a manifest list for linux/amd64 and
linux/arm64, each built on a runner of its own architecture, so an ARM host (a
Raspberry Pi, an ARM NAS, Apple silicon under Docker Desktop) pulls the same tag.
The compose file pulls JUDGE_IMAGE, which defaults to the upstream package. A fork
sets it to its own package in .env once its first workflow run has published. A fork
also sets JUDGE_SOURCE_URL to its repository. The source offer every interface
makes (the web footer, /help and /license, the MCP instructions) then points users
at the code that is running, which is what the AGPL asks of anyone serving a modified
version.
The workflow stamps the image with the commit it built (JUDGE_COMMIT, a
Docker build argument the Dockerfile hands to crates/bot/build.rs), so the offer
names the revision. A local docker compose build has no .git in its context.
It must pass the commit itself, together with whether the tree matches it, or the offer
reads “commit unknown”:
JUDGE_COMMIT=$(git rev-parse HEAD) JUDGE_DIRTY=$(git diff-index --quiet HEAD || echo 1) \ docker compose up -d --build bot apiA JUDGE_COMMIT that is not a commit id (a branch name, a tag) fails the build rather
than becoming “commit unknown” on every interface.
A GitHub release whose tag is vX.Y.Z adds version tags to the image already built for
that commit, without rebuilding it: X.Y.Z, X.Y and, from 1.0 on, X. The
release is the image that has been running as latest, down to the platform
digests. A host that prefers to move on releases rather than on every push pins one:
JUDGE_IMAGE_TAG=1.0 # in .env: follows 1.0.x patch releases; 1.0.0 pins one exactlyCONTRIBUTING.md says how a release is cut. The ordinary deploy:
git pull # runbook + compose changesdocker compose pulldocker compose up -ddocker compose up -d --build still works on a machine with the CPU and RAM for it.
build: . is retained for local development.
Releases with a migration
Section titled “Releases with a migration”bot and api apply pending migrations at startup, before anything else touches
the database. The ordinary deploy above is therefore complete for a release whose
commit adds a file under crates/bot/migrations/. The first of the two to start
migrates and the other finds nothing pending. The calls-rewrite advisory lock
serialises them, and sqlx’s own migrator lock is a second layer for a concurrent
sqlx migrate run. The log says schema migrated with the versions.
A migration that fails exits the process. Under restart: unless-stopped that is a
crash-loop with the reason in docker compose logs bot. It is loud on purpose: a bot
running against the wrong schema would answer questions and quietly fail to persist
them.
The migration holds the same advisory lock the refresh job’s CR load, retirement
pass and embedding writes take, so those wait for it and it waits for them. (The
Scryfall card/rulings upsert takes no lock and needs none: one transaction, no
calls rows.) A deploy that lands during the nightly refresh therefore sits at
startup until the CR load finishes, which is minutes on a NAS. It logs one warning
line, another job holds the calls rewrite lock ... waiting, in
docker compose logs bot. That is a wait, not a hang.
The lock does not stop the other service, because compose recreates bot and api
independently. So stop both first for a migration that is not additive. One that
rewrites calls rows (20260902000001 did, moving ruling citations to content keys)
must not race a call being persisted by the old binary, and ALTER TABLE waits
on any in-flight query. The release notes in the commit say when this applies:
git pulldocker compose pulldocker compose stop bot api # only when the release notes say the migration rewrites rowsdocker compose up -dTo move the schema by hand instead, set JUDGE_AUTO_MIGRATE=false in .env. It
reaches bot/api at their next recreate, which up -d does because the env file
changed. Then run the explicit form from the new image:
docker compose pulldocker compose stop bot api # only when the migration rewrites rowsdocker compose run --rm --pull missing refresh migrate # prints what it applies; refuses a changed filedocker compose up -dPrivate package pulls
Section titled “Private package pulls”The upstream package is public and needs no login. A fork’s package inherits the
fork’s visibility, so a private fork’s host must log in once with a classic PAT
carrying only read:packages. On Synology, log in as root, since that is the user
Container Manager and the Task Scheduler run as:
echo "$GHCR_TOKEN" | docker login ghcr.io -u <github-user> --password-stdinMaking the package public instead (GHCR package settings, independent of repo visibility) removes the login step.
Rolling back
Section titled “Rolling back”Every build leaves an immutable tag, so a bad deploy is a one-line revert. Take the
sha-<short> from the workflow run summary, or the version of the last good release:
JUDGE_IMAGE_TAG=sha-abc1234 # in .env; or a release, e.g. 1.0.0docker compose pull && docker compose up -dClear JUDGE_IMAGE_TAG to return to latest.
An older tag starting against a newer schema logs database is ahead of this binary
and does not migrate. Whether it then works depends on the migration. An additive one
(a new table, a nullable column) is harmless to the old binary. One that changed a
column the old queries use is not. Rolling back across it means restoring the
pre-release dump too (scripts/backup-db.sh fetch, §6), so take the weekly
backup by hand right before such a deploy. judge-ingest migrate
refuses an ahead database rather than guessing.
cloudflared and db are untouched by a code deploy. The tunnel reconnects on its
own if the connector restarts.
9. Troubleshooting
Section titled “9. Troubleshooting”| Symptom | Cause |
|---|---|
| 502 from the public hostname | api is down, or the tunnel’s service is not http://api:8787 |
| Tunnel healthy, hostname NXDOMAIN | the judge record is grey-clouded; it must be proxied |
/mcp answers 401 |
wrong or missing Authorization: Bearer <MCP_TOKEN> |
/mcp answers 403 while /api/health is fine |
the public hostname is not in MCP_ALLOWED_HOSTS |
/mcp answers 405 (a browser GET shows the web page) |
--mcp is not in API_INTERFACES, so /mcp is just another page path |
api exits naming --mcp and MCP_TOKEN |
--mcp with no token to gate it; set MCP_TOKEN or drop the flag |
/mcp 404s and the log warns about MCP_TOKEN |
the token is set but --mcp is not in API_INTERFACES |
api exits naming --web and index.html |
--web with no built page at WEB_DIST; drop --web or rebuild the image |
The page 404s but /api/health is fine |
--web is not in API_INTERFACES; the startup log line lists what is on and what is off |
| Everyone shares one rate-limit bucket | API_CLIENT_IP=peer behind the tunnel — every request looks like the cloudflared container |
| Rate limiting never triggers | API_CLIENT_IP=cloudflare while something other than Cloudflare can reach the origin, so CF-Connecting-IP is caller-supplied |
bot or api restart-loops naming JUDGE_OPERATOR_DISCORD / JUDGE_OPERATOR_EMAIL |
the contact that surface must show is unset or malformed in .env; set it and docker compose up -d |
| Bot online, web page dead | expected if only api failed — the gateway is a separate outbound connection |
cloudflared restart-loops on startup |
COMPOSE_PROFILES=tunnel with TUNNEL_TOKEN empty or stale in .env.deploy |
| Members are told the bot “hit its spending cap” | JUDGE_MAX_USD is spent for the process or the period (judge-cli stats shows the days); raise it and docker compose up -d, or wait for the period to turn |
| A refresh or backup failed and nobody noticed | set JUDGE_ALERT_WEBHOOK in .env; both scripts post there on a non-zero exit |
| Backup cron silently never runs | log path not writable by your user, or .env.deploy missing |
Refresh exits another refresh is running with nothing running |
a previous run was killed before removing .refresh.lock in the repo root; rmdir it |
| Refresh loads the CR every night | rules.cr_version disagrees with the file name on Wizards’ page — check the current comprehensive rules release log line for published vs stored |
| Refresh runs but the bot still cites the old CR | it does not: retrieval reads the database live; check the run actually finished (refresh step ok for rules and embed) |
JUDGE_CONFIG=/etc/judgebot/judge.toml: file not found at startup |
JUDGE_CONFIG in .env names a host file that does not exist; Docker mounted an empty directory in its place (and created a root-owned one on the host — sudo rmdir it) |
providers.X: NAME (api_key_env) is not set at startup |
the key was exported in the shell that ran cargo run but never written to .env, which is all the containers read; or, from refresh alone, only the embed provider’s key was set because “refresh only embeds” — refresh resolves the chat stages too, so the extract/synth providers’ keys must be in .env as well |
Edited judge.toml, docker compose up -d, nothing changed |
up -d recreates only on a configuration or image change and a bind-mounted file’s content is neither; docker compose restart bot api (§4) |
providers.X (...): no credentials at startup |
a cloud door with an empty chain: no AWS_* in .env, no mounted credentials file, or a mounted file the nobody user cannot read |
embedding space mismatch; vector legs off |
[models.embed] names a model or width other than the one the database holds; reembed (§7) to move the data, or change the file back |
no price for X/Y at startup |
a model on an openai provider without [models.<stage>.pricing]; add one (USD per million tokens) or pricing = "free" on the provider |
MTG Judgebot is unofficial Fan Content permitted under the Fan Content Policy. Not approved or endorsed by Wizards of the Coast. Portions of the materials used are property of Wizards of the Coast. ©Wizards of the Coast LLC.
The Comprehensive Rules come from Wizards of the Coast. Card data, rulings and card symbols come from Scryfall, which is not affiliated with this project. Rule links go to the Yawgatog mirror. License and attribution.