Operations
This page tells you how to deploy Platform, roll it back, move it to Postgres, and turn off routes. To run Platform on one RunPod pod, see Getting started.
Compose services
docker-compose.yml defines these services. Each one restarts unless stopped and keeps 3 log files of 10 MB.
| Service | Command | Profile |
|---|---|---|
api | alembic upgrade head, then gunicorn with 1 worker and 8 threads on port 8000 | default |
bot | python3 bot_main.py. Run exactly one, or each scheduled post goes out more than once | default |
dashboard | The dashboard build (Dockerfile.dashboard) on port 5001 | default |
web | The old officer app (Dockerfile.web) on port 5000, for admin.thesoda.io | default |
postgres | Postgres 16 with pgvector | postgres |
worker | python3 worker_main.py | postgres |
mcp | python3 mcp_main.py on port 8001 | mcp |
site | The landing and docs site (Dockerfile.site, Next.js) on 127.0.0.1:3010, for platform.thesoda.io | site |
The API mounts ./data (SQLite database and JWT keys), ./.env and ./google-secret.json.
Caution: make sure that google-secret.json is on the host, also if it is empty. If it is not there, the API container does not start.
The API uses one gunicorn worker because the one-time sign-in codes are in process memory. Increase --workers only after those codes move to the database.
Container limits
Each container has a memory and a CPU cap, so one process cannot take the whole server. The values come from .env, with these defaults:
| Service | Memory | CPUs | Settings |
|---|---|---|---|
api | 1536m | 1.5 | API_MEM_LIMIT, API_CPUS |
bot | 768m | 1 | BOT_MEM_LIMIT, BOT_CPUS |
worker | 1536m | 1 | WORKER_MEM_LIMIT, WORKER_CPUS |
mcp | 768m | 1 | MCP_MEM_LIMIT, MCP_CPUS |
web | 256m | 0.5 | WEB_MEM_LIMIT, WEB_CPUS |
dashboard | 256m | 0.5 | DASHBOARD_MEM_LIMIT, DASHBOARD_CPUS |
site | 384m | 0.5 | SITE_MEM_LIMIT, SITE_CPUS |
postgres | 1g | 1 | POSTGRES_MEM_LIMIT, POSTGRES_CPUS |
A container that goes over its memory cap is stopped and starts again (restart: unless-stopped). A new value applies when the container is made again, for example at the next deploy. The Superadmin page shows the memory use and cap of the API container. These caps are for each service, not each org: Limits caps what each org uses.
Images and CI
| Workflow | Does |
|---|---|
check.yml | On each push and PR: make ci and bandit; migrations and tests on Postgres 16; the dashboard tests and build |
images.yml | Builds the API, web and dashboard images on each PR. On main it pushes them to GHCR as ghcr.io/<owner>/<repo>-api, -web and -dashboard |
cd.yml | Deploys the example SoDA server after check.yml passes on main. It runs only in asusoda/platform |
The Dockerfiles, the compose file and CI pull base images (Python, Node, pgvector) from mirror.gcr.io, Google's copy of Docker Hub. It has no pull limit for anonymous users, so CI does not stop with 429 Too Many Requests. Platform's own images go to GHCR.
Dockerfile.api uses uv sync --frozen. If uv.lock does not agree with pyproject.toml, the build fails. Commit the two files together. The dashboard image gets VITE_API_URL and VITE_SITE_URL at build time (repository variables in CI), so a change to them needs a new build.
Deploy
cd.yml connects to the server over SSH and runs these commands in the repo folder:
DEPLOY_FROM="$(git rev-parse HEAD)" # the commit that runs now
git fetch origin main && git checkout main && git reset --hard origin/main
make backup # copy data/user.db to data/backups/, keep the last 14
make deploy DEPLOY_FROM="$DEPLOY_FROM"
make healthcd.yml updates the checkout before make backup. Thus the server uses the Makefile of the new commit, also when the old Makefile has no backup target. If make deploy or make health fails, cd.yml runs make rollback.
This diagram shows the deploy.
flowchart TD push["Push to main"] --> check["check.yml passes"] check --> cd["cd.yml connects to the server over SSH"] cd --> reset["git reset --hard origin/main"] reset --> backup["make backup"] backup --> deploy["make deploy: build changed images, alembic upgrade head, up -d"] deploy -->|fails| rollback["make rollback"] deploy -->|passes| health["make health"] health -->|fails| rollback health -->|passes| done["Deploy done"]
make deploy does these steps:
- Get
origin/mainand find the files changed sinceDEPLOY_FROM. IfDEPLOY_FROMis empty, it uses the commit that was checked out before the fetch. - Select the images to build. A change in
web/orDockerfile.webbuildsweb. A change indashboard/orDockerfile.dashboardbuildsdashboard. A change insite/,docs/orDockerfile.sitebuildssite. A change in a compose file or theMakefilebuilds all four. A change in.github/or another.mdfile builds nothing. All other changes buildapi. - Tag the current images as
:previous. - Build the changed images. The old containers keep running during the build.
- Run
uv run alembic upgrade headon the host. If it fails, the deploy stops and the old containers keep running. - If it built
api,webordashboard, recreateapi,bot,webanddashboardtogether withup -d --remove-orphans, andmcptoo when its container exists. Then wait up to 60 seconds for each to be healthy and check that each runs the new image. They are recreated together becauseweb,dashboard,botandmcprequire theapicontainer: podman-compose 1.0.6 recreates a service's dependencies with it, and it cannot replaceapiwhile other containers require it.--remove-orphansremoves containers of services that are no longer indocker-compose.yml, so an old container does not keep a port. - If it built
siteand thesoda-sitecontainer exists, recreatesitealone withup -d --no-deps. It does not requireapi, so the other containers keep running.
Caution: set VITE_API_URL in .env before you build the dashboard. If it is empty, Dockerfile.dashboard stops the build and the deploy fails.
Upgrade an existing deployment
Do the first deploy after a large upgrade by hand. For an upgrade from asusoda/platform at 579a6a84 or older, do these steps on the server:
- Pause CD, or do the deploy before the next push to
main. - Run
make backup. If the oldMakefilehas nobackuptarget, copydata/user.dbtodata/user.db.pre-upgrade. Keepdata/jwt_*.pem. - Run
uv run alembic current. The result must bea1b2c3d4e5f6, the last SoDA migration. - Add
VITE_API_URLto.env, for examplehttps://api.thesoda.io. KeepREACT_APP_API_URLforweb/. - If the database has no
alembic_versiontable (create_allmade it), runuv run alembic stamp a1b2c3d4e5f6. If you do not,alembic upgrade headstops with "table already exists". - Keep the current commit for a rollback:
OLD=$(git rev-parse HEAD). - Run
git fetch origin main, thengit reset --hard origin/main. - Run
make deploy DEPLOY_FROM="$OLD". IfDEPLOY_FROMis not set,make deployfinds no changed files and builds nothing. - Run
flask --app main config checkin the API container. It must showmigrations at headand noFAIL.
LeetCode: the API now posts the daily question as a job and keeps the "posted today" record in the leetcode_daily table. If you deploy after LEETCODE_DAILY_TIME on a day that already had a post, the channel gets a second post. Deploy before that time to prevent it.
Roll back
make rollback tags soda-internal-api:previous as latest and starts the containers again.
Caution: the rollback does not change the dashboard image or the database. If the failed deploy ran a migration, run uv run alembic downgrade -1, or copy back the file that make backup wrote to data/backups/.
Caution: make rollback alone does not work across the upgrade from a1b2c3d4e5f6. Migration b7d9f1a3c5e8 renames users.asu_id to student_id, and the old image has no alembic/ folder. To go back:
- Run
uv run alembic downgrade a1b2c3d4e5f6, or copy the backup back todata/user.db. - Check out the old commit:
git checkout "$OLD". - Run
make build, thenmake up.
Move to Postgres
The API reads DATABASE_URL. CI runs the tests on SQLite and Postgres 16. Do these steps on a staging server first.
- Add
POSTGRES_PASSWORDto.env. Start the database:docker compose --profile postgres up -d postgres. - Make the schema:
DATABASE_URL=postgresql://platform:<password>@localhost:5432/platform uv run alembic upgrade head. - Stop the writers:
docker compose stop api bot. - Copy the data:
uv run python deploy/copy_sqlite_to_postgres.py sqlite:///./data/user.db <postgres url>. The script refuses tables that have rows. It stops with an error if a row count or an org's points total is different. - Set
DATABASE_URL=postgresql://platform:<password>@postgres:5432/platformin.env. Rundocker compose --profile postgres up -d. Theworkerservice starts and runs the jobs. - Keep
data/user.dbfor two weeks or more. To go back, removeDATABASE_URLand restart.
Host the docs site
The site service serves site/ with next start, because the site uses redirects and sends a page as Markdown when the client asks for it. It listens only on 127.0.0.1:3010, so a reverse proxy must forward the domain to it.
- Build and start it once:
docker compose --profile site up -d --no-deps site. podman-compose 1.0.6 ignores profiles: usepodman-compose -f docker-compose.yml up -d --no-deps site. - Point the domain at the server, and forward it in the reverse proxy to
http://127.0.0.1:3010.
After this, make deploy rebuilds the site when site/ or docs/ changes, and make health checks it.
Command-line tools
Run these in the API container (make shell) or on your machine with the same .env:
flask --app main config check # settings, database and migrations; exits non-zero on a fault
flask --app main org list
flask --app main org create --name "Robotics Club" --prefix robotics --guild-id <id> --officer-role-id <id> --on points,storefront
flask --app main org modules robotics --on calendar
flask --app main jobs list
flask --app main jobs run calendar.sync_all
flask --app main jobs run points.import_event_csv -a org_prefix=robotics -a event_name=X -a event_points=5 -a file_content=...jobs run runs the job in the shell process, not through the queue. The audit log records it.
Turn off routes
DISABLED_ROUTES is a comma-separated list of path prefixes. If a request path starts with one of them, the API returns the same 404 as an unknown route. The route stays in the code and in tests/contract/routes.txt. If DISABLED_ROUTES is empty or not set, all routes are on.
Caution: end a folder prefix with /. If you do not, /api/bot also turns off /api/botstatus.
- Set
DISABLED_ROUTESin the server's.env. - Restart the API:
docker compose restart api. The API reads.envwhen it starts. - Make sure that a turned-off path returns 404:
curl -i https://<api host>/api/public/getnextevent.
The example AIS server sets these prefixes, because the routes are broken:
| Prefix | Fault |
|---|---|
/api/public/getnextevent | The view returns no response, so each call returns 500 |
/api/bot/ | The game routes read current_app.auth_bot, which gunicorn never sets. Some also call db_connect methods that do not exist |
DISABLED_ROUTES=/api/public/getnextevent,/api/bot/Health, logs and errors
GET /healthreturnsstatus,commitandstarted_at.commitcomes from theGIT_COMMIT_HASHbuild argument, so it shows the image, not the files on disk.make logsshows the last 50 lines.make logs-followfollows them.- With
LOG_FORMAT=json(set in compose), each line is a JSON object withts,level,logger,msgand the request fields (route,status,org,reason).LOG_FORMAT=textgives colored lines. - If
SENTRY_DSNis set, the API, bot, job worker and MCP server send errors to Sentry, with aservicetag (api,bot,worker,mcp) and the commit as the release. They also send log lines atSENTRY_LOGS_LEVEL(defaultWARNING) and above, and traces forSENTRY_TRACES_SAMPLE_RATEof requests (default0.1).SENTRY_PROFILES_SAMPLE_RATE(default0) turns on profiles.SENTRY_ENVIRONMENT(defaultproduction) names the environment. - Platform keeps its own error log, with no outside service. Each process (
api,bot,worker,mcp) records every log line at ERROR or above in theerror_groupstable, with the stack trace, the org and the route. Each group also keeps the context of its last event: the request method and path (no query string), the job name, the logger and line, the host, process, thread, release (GIT_COMMIT_HASH) and Python version. The dashboard sends browser errors and API calls that got no answer or a status of 500 or more. Repeats of one error add to one group. - Officers see their org's errors on Activity, Errors, and resolve or delete them there, one at a time or as a selection. A resolved error opens again when it happens again. A deleted error that happens again starts a new group. The superadmin page shows the errors of every org and the errors with no org, such as a failed job.
- Discord alerts: an officer turns on the Event webhooks module and adds a webhook with the Errors event. See Webhooks.
ERROR_WEBHOOK_URLin.envgets every new error of every org and of the server. Each new or returning error posts one message, at most 30 for each process in an hour. - Sentry is optional. Set
SENTRY_DSNonly if you want Sentry in addition to the error log.
Caution: do not delete data/jwt_private.pem or data/jwt_public.pem. If you delete them, every officer must sign in again.