Operations
This page tells you how to deploy Platform, roll it back, move it to Postgres, turn off routes and run Hermes Agent as an org app. To run Platform on one RunPod pod, see Getting started.
Compose services
docker-compose.yml defines these services. Each one restarts unless stopped and keeps 3 log files of 10 MB.
| Service | Command | Profile |
|---|---|---|
api | alembic upgrade head, then gunicorn with 1 worker and 8 threads on port 8000 | default |
bot | python3 bot_main.py. Run exactly one, or each scheduled post goes out more than once | default |
dashboard | The dashboard build (Dockerfile.dashboard) on port 5000 | default |
postgres | Postgres 16 with pgvector | postgres |
worker | python3 worker_main.py | postgres |
mcp | python3 mcp_main.py on port 8001 | mcp |
The API mounts ./data (SQLite database and JWT keys), ./.env and ./google-secret.json.
Caution: make sure that google-secret.json is on the host, also if it is empty. If it is not there, the API container does not start.
The API uses one gunicorn worker because the one-time sign-in codes are in process memory. Increase --workers only after those codes move to the database.
Images and CI
| Workflow | Does |
|---|---|
check.yml | On each push and PR: make ci and bandit; migrations and tests on Postgres 16; the dashboard tests and build |
images.yml | Builds the API and dashboard images on each PR. On main it pushes them to GHCR as ghcr.io/<owner>/<repo>-api and -dashboard |
hermes-image.yml | Builds deploy/hermes when it changes. On main it pushes ghcr.io/<owner>/<repo>-hermes |
cd.yml | Deploys the example SoDA server after check.yml passes on main. It runs only in asusoda/platform |
Dockerfile.api uses uv sync --frozen. If uv.lock does not agree with pyproject.toml, the build fails. Commit the two files together. The dashboard image gets VITE_API_URL and VITE_SITE_URL at build time (repository variables in CI), so a change to them needs a new build.
Deploy
cd.yml connects to the server over SSH and runs these targets in the repo folder:
make discard-local-changes # git reset --hard
make backup # copy data/user.db to data/backups/, keep the last 14
make deploy
make healthIf make deploy or make health fails, it runs make rollback.
make deploy does these steps:
- Get
origin/mainand find the changed files. - Select the images to build. A change in
dashboard/orDockerfile.dashboardbuildsdashboard. A change in a compose file or theMakefilebuilds both. A change in.github/or a.mdfile builds nothing. All other changes buildapi. - Run
uv run alembic upgrade headon the host. If it fails, the deploy stops and the old containers keep running. - Tag the current images as
:previous. - Build and start the changed services, then wait up to 60 seconds for each to be healthy. When it builds
api, it also startsbotagain, because the bot uses the same image.
Roll back
make rollback tags soda-internal-api:previous as latest and starts the containers again.
Caution: the rollback does not change the dashboard image or the database. If the failed deploy ran a migration, run uv run alembic downgrade -1, or copy back the file that make backup wrote to data/backups/.
Move to Postgres
The API reads DATABASE_URL. CI runs the tests on SQLite and Postgres 16. Do these steps on a staging server first.
- Add
POSTGRES_PASSWORDto.env. Start the database:docker compose --profile postgres up -d postgres. - Make the schema:
DATABASE_URL=postgresql://platform:<password>@localhost:5432/platform uv run alembic upgrade head. - Stop the writers:
docker compose stop api bot. - Copy the data:
uv run python deploy/copy_sqlite_to_postgres.py sqlite:///./data/user.db <postgres url>. The script refuses tables that have rows. It stops with an error if a row count or an org's points total is different. - Set
DATABASE_URL=postgresql://platform:<password>@postgres:5432/platformin.env. Rundocker compose --profile postgres up -d. Theworkerservice starts and runs the jobs. - Keep
data/user.dbfor two weeks or more. To go back, removeDATABASE_URLand restart.
Command-line tools
Run these in the API container (make shell) or on your machine with the same .env:
flask --app main config check # settings, database and migrations; exits non-zero on a fault
flask --app main org list
flask --app main org create --name "Robotics Club" --prefix robotics --guild-id <id> --officer-role-id <id> --off points,storefront
flask --app main org modules robotics --on calendar
flask --app main jobs list
flask --app main jobs run calendar.sync_all
flask --app main jobs run points.import_event_csv -a org_prefix=robotics -a event_name=X -a event_points=5 -a file_content=...jobs run runs the job in the shell process, not through the queue. The audit log records it.
Turn off routes
DISABLED_ROUTES is a comma-separated list of path prefixes. If a request path starts with one of them, the API returns the same 404 as an unknown route. The route stays in the code and in tests/contract/routes.txt. If DISABLED_ROUTES is empty or not set, all routes are on.
Caution: end a folder prefix with /. If you do not, /api/bot also turns off /api/botstatus.
- Set
DISABLED_ROUTESin the server's.env. - Restart the API:
docker compose restart api. The API reads.envwhen it starts. - Make sure that a turned-off path returns 404:
curl -i https://<api host>/api/public/getnextevent.
The example AIS server sets these prefixes, because the routes are broken:
| Prefix | Fault |
|---|---|
/api/public/getnextevent | The view returns no response, so each call returns 500 |
/api/bot/ | The game routes read current_app.auth_bot, which gunicorn never sets. Some also call db_connect methods that do not exist |
DISABLED_ROUTES=/api/public/getnextevent,/api/bot/Health, logs and errors
GET /healthreturnsstatus,commitandstarted_at.commitcomes from theGIT_COMMIT_HASHbuild argument, so it shows the image, not the files on disk.make logsshows the last 50 lines.make logs-followfollows them.- With
LOG_FORMAT=json(set in compose), each line is a JSON object withts,level,logger,msgand the request fields (route,status,org,reason).LOG_FORMAT=textgives colored lines. - If
SENTRY_DSNis set, the API, bot, job worker and MCP server send errors to Sentry, with aservicetag (api,bot,worker,mcp) and the commit as the release. They also send log lines atSENTRY_LOGS_LEVEL(defaultWARNING) and above, and traces forSENTRY_TRACES_SAMPLE_RATEof requests (default0.1).SENTRY_PROFILES_SAMPLE_RATE(default0) turns on profiles.SENTRY_ENVIRONMENT(defaultproduction) names the environment. - Platform keeps its own error log, with no outside service. Each process (
api,bot,worker,mcp) records every log line at ERROR or above in theerror_groupstable, with the stack trace, the org and the route. The dashboard sends browser errors and API calls that got no answer or a status of 500 or more. Repeats of one error add to one group. - Officers see their org's errors on Activity, Errors, and resolve them there. A resolved error opens again when it happens again. The superadmin page shows the errors of every org and the errors with no org, such as a failed job.
- Discord alerts: an officer sets a Discord webhook on Activity, Errors (org secret
error_webhook_url).ERROR_WEBHOOK_URLin.envgets every new error of every org and of the server. Each new or returning error posts one message, at most 30 for each process in an hour. - Sentry is optional. Set
SENTRY_DSNonly if you want Sentry in addition to the error log.
Caution: do not delete data/jwt_private.pem or data/jwt_public.pem. If you delete them, every officer must sign in again.
Hermes Agent
deploy/hermes/ runs Hermes Agent as an org app on RunPod. Hermes talks to members in Discord and uses the org's tools through the MCP server. Any agent that uses MCP connects the same way. See RunPod apps for app manifests and deploys.
| File | Holds |
|---|---|
Dockerfile | The official Hermes image at a fixed version, started as hermes gateway run |
platform-config.sh | Runs at each start. Writes the model and the Platform MCP server into /opt/data/config.yaml and keeps the rest of the file |
app.example.json | The app manifest, with the values to fill in |
- Run the MCP server where Hermes can reach it. On a pod from Getting started, add
8001/httpto the pod ports. - Make a machine token of kind
agentwith the scopes Hermes can use, for exampleorg:readandknowledge:read. Addagents:readandagents:writeonly if Hermes keeps member memories. - Make a Discord app and bot for Hermes. It is not the org's Platform bot.
- Make a RunPod network volume of 10 GB in one data center. Hermes keeps its config, memories, sessions and skills in
/opt/dataon it. - Save the org secrets in the table below.
- Copy
app.example.jsonand fill in the volume id, its data center, the MCP URL and the Discord role that can talk to Hermes. Register it withPUT /api/apps/hermesand the body{"manifest": {...}}. - Deploy a tag that the workflow pushed:
POST /api/apps/hermes/deploywith{"tag": "<commit sha>"}. The health check reads/healthon port 8642.
| Org secret | Value |
|---|---|
app_hermes_discord_token | The Hermes Discord bot token |
app_hermes_openrouter_key | The model provider token. Another provider needs its own env name, such as ANTHROPIC_API_KEY |
app_hermes_platform_token | The machine token from step 2 |
app_hermes_api_key | A long random string that protects the Hermes API on port 8642 |
The manifest env sets HERMES_PROVIDER and HERMES_MODEL, PLATFORM_MCP_URL and PLATFORM_TOKEN (without both, Hermes has no Platform tools), and DISCORD_ALLOWED_ROLES, DISCORD_ALLOWED_USERS or DISCORD_ALLOWED_CHANNELS (set one or more).
Caution: anyone with API_SERVER_KEY has full use of the agent, including its terminal. Anyone with access to the org's RunPod account can read the pod env. Revoke the machine token on the dashboard Tokens page to stop Hermes from using Platform.
To keep the memories and skills of a Hermes that runs on a laptop, copy its ~/.hermes folder to the network volume before the first deploy. Do not copy .env; put its secrets in org secrets. When the pod is healthy, stop the old gateway (systemctl --user stop hermes-gateway), or the same Discord bot runs two times.
Agents on the platform pod
To save the cost of more pods, the platform pod can also run Sparky (engine and Discord bot) and Hermes. deploy/runpod/start.sh starts each one when its token is set. Each agent restarts 10 seconds after it stops. Both use a hosted model API, so the pod needs no GPU.
| File | Does |
|---|---|
sparky.sh | Copies engine, discord and sparky.toml from the public image SPARKY_IMAGE:SPARKY_TAG (default ghcr.io/ashworks1706/sparkyai-rust:main) to /workspace/sparky/bin with image_files.py, then runs both with this platform as the store |
hermes.sh | Installs Hermes HERMES_VERSION from its source tag to /workspace/hermes, writes the model and the MCP server into its config, and runs hermes gateway run as the user hermes |
image_files.py | Copies files out of a public image without a container runtime. It downloads again only when the image digest changes |
Set these in the pod env, then restart the pod.
| Agent | Variable | Value |
|---|---|---|
| Sparky | SPARKY_DISCORD__TOKEN, SPARKY_DISCORD__GUILD_ID | The Sparky bot token and the server id |
| Sparky | SPARKY_PLATFORM__TOKEN | A machine token of kind agent with agents:read, agents:write, knowledge:read, accounts:link, accounts:token. Sparky also uses the MCP server at http://127.0.0.1:8001/mcp: every scope you add gives it those tools. Every member can use read tools, so add only reads that all members may see. Write tools need the Manage Server permission in Discord and a confirmation |
| Sparky | SPARKY_MODEL__BASE_URL, SPARKY_MODEL__API_KEY, SPARKY_MODEL__NAME | Any OpenAI-compatible chat API. The summary model is the same unless SPARKY_SUMMARY__* is set |
| Sparky | SPARKY_TAG | Optional. An image tag (commit sha) in place of main |
| Hermes | HERMES_ENV_DISCORD_BOT_TOKEN | The Hermes bot token. Hermes gets each HERMES_ENV_* variable without the prefix |
| Hermes | HERMES_ENV_DISCORD_ALLOWED_ROLES | The Discord roles that can talk to Hermes |
| Hermes | HERMES_ENV_OPENROUTER_API_KEY, HERMES_PROVIDER, HERMES_MODEL | The model provider key, the provider and the model. Another provider needs its own key name, such as HERMES_ENV_ANTHROPIC_API_KEY |
| Hermes | HERMES_PLATFORM_TOKEN | A machine token of kind agent. Hermes uses the MCP server at http://127.0.0.1:8001/mcp. For read access give org:read and knowledge:read. To let Hermes run the org like an officer, add activity:read, settings:write, integrations:manage, knowledge:write, apps:read, apps:manage, apps:deploy, alerts:manage and compute:manage. Then limit who can talk to Hermes with HERMES_ENV_DISCORD_ALLOWED_ROLES |
Sparky sends no query vector unless SPARKY_EMBEDDING__BASE_URL is set. The platform then embeds each query with the Embeddings integration, so set the Embeddings card to the model that embedded the org's knowledge. The engine listens on 127.0.0.1:8080 only, and run_sandbox is off because the pod cannot run containers.
Hermes keeps its config, memories and skills in /workspace/hermes/home. To keep those of a Hermes that runs on a laptop, copy its ~/.hermes folder there before the first start (not .env), then chown -R hermes:hermes /workspace/hermes/home. Stop the laptop gateway when the pod one runs.
Caution: Hermes can run shell commands on the pod. It runs as the user hermes with only its own env, and /workspace/data and the platform checkout are closed to it. To remove the shell tools from Discord, set platform_toolsets.discord in its config.yaml, for example to [web, vision, skills, todo].
Use your own GPU later
The model settings are URLs, so a GPU changes only the pod env. Run the Sparky RunPod image (ghcr.io/ashworks1706/sparkyai-runpod) on a GPU pod with SPARKY_MODELS_API_KEY set and port 8000/http and 8001/http open. llama-server then serves chat on 8000 and embeddings on 8001 at the pod proxy URLs.
- On the platform pod, set
SPARKY_MODEL__BASE_URL=https://<gpu pod>-8000.proxy.runpod.net/v1,SPARKY_MODEL__API_KEYto theSPARKY_MODELS_API_KEYvalue, andSPARKY_MODEL__NAMEto the GGUF name, for exampleQwen/Qwen3-4B-GGUF:Q4_K_M. - Point the Embeddings integration at
https://<gpu pod>-8001.proxy.runpod.net/v1with the same key. If its model name changes, embed the knowledge again. - To give Hermes the same model, set
HERMES_PROVIDER=custom,HERMES_BASE_URLto the chat URL,HERMES_MODEL, and the key inHERMES_ENV_OPENAI_API_KEY.
Set SPARKY_POD_ROLE=models on the GPU pod so it runs only the model servers. Without it the image also runs its own engine and bot, and the same bot token must not run on two pods.