Compare commits

...

13 Commits

Author SHA1 Message Date
developer 24017bcb7f Merge pull request 'Promote development → master (deploy v2: versioning + visible jobs + rollback)' (#106) from development into master 2026-06-22 06:01:03 +00:00
developer 40d8f06588 Merge pull request 'Deploy v2 — release versioning + visible deploy jobs + manual rollback' (#105) from feature/release-versioning into development
CI / changes (push) Successful in 2s
CI / unit (push) Successful in 11s
CI / integration (push) Successful in 15s
CI / ui (push) Successful in 57s
CI / changes (pull_request) Successful in 2s
CI / gate (push) Successful in 0s
CI / unit (pull_request) Successful in 10s
CI / integration (pull_request) Successful in 16s
CI / ui (pull_request) Successful in 56s
CI / deploy (push) Successful in 1m19s
CI / gate (pull_request) Successful in 0s
CI / deploy (pull_request) Has been skipped
2026-06-22 05:40:55 +00:00
Ilia Denisov c59e522732 feat(deploy): visible prod-deploy jobs + manual prod-rollback
CI / changes (pull_request) Successful in 2s
CI / unit (pull_request) Successful in 11s
CI / integration (pull_request) Successful in 16s
CI / ui (pull_request) Successful in 57s
CI / gate (pull_request) Successful in 0s
CI / deploy (pull_request) Successful in 1m23s
- prod-deploy.yaml is now four visible sequential jobs (build -> deploy-main ->
  deploy-bot -> verify) so the rollout stages show in the Actions UI; the
  per-service rolling stays in the deploy-main log.
- prod-rollback.yaml: a separate manual workflow_dispatch. Leave target_version
  blank to roll back to the previous deployed version (the host now tracks
  DEPLOYED_TAG + PREVIOUS_TAG), or pick a release tag. Re-deploys an already
  published image rolling + health-gated, image-only (no rebuild, no DB migration).
- prod-deploy.sh tracks the previous tag (commit_tag) for the blank-input rollback.
- Docs: ARCHITECTURE §13 + deploy/README runbook cover versioning + rollback.
2026-06-22 07:37:08 +02:00
Ilia Denisov 8d45ae6e3b feat: stamp the build version into every service
pkg/version.Version (default "dev") is set at link time via -ldflags from each
service Dockerfile's VERSION build-arg, which the deploy passes as the git tag
(git describe --tags). It surfaces as the OpenTelemetry service.version resource
attribute (so Grafana/Tempo are version-aware), alongside the SPA's existing
About version. Adds the VERSION build-arg to the backend/gateway/validator/bot
compose builds and a serviceResource test covering service.name + service.version.
2026-06-22 07:28:27 +02:00
developer 2c4f4b10dc Merge pull request 'Promote development → master (initial production release: pre-release line + Stage 18)' (#104) from development into master 2026-06-22 05:05:48 +00:00
developer 520a9092fe Merge pull request 'Stage 18 — prod contour deploy (two-host registry rollout, rolling + auto-rollback)' (#103) from feature/prod-contour-deploy into development
CI / changes (push) Successful in 2s
CI / unit (push) Successful in 11s
CI / integration (push) Successful in 17s
CI / ui (push) Successful in 57s
CI / changes (pull_request) Successful in 2s
CI / gate (push) Successful in 0s
CI / deploy (push) Successful in 1m4s
CI / unit (pull_request) Successful in 10s
CI / integration (pull_request) Successful in 15s
CI / ui (pull_request) Successful in 56s
CI / gate (pull_request) Successful in 0s
CI / deploy (pull_request) Has been skipped
2026-06-22 04:59:46 +00:00
Ilia Denisov 9f970495ee fix(deploy): guard cd and split DOCKER_GID assignment (shellcheck)
CI / changes (pull_request) Successful in 2s
CI / unit (pull_request) Successful in 10s
CI / integration (pull_request) Successful in 16s
CI / ui (pull_request) Successful in 56s
CI / gate (pull_request) Successful in 0s
CI / deploy (pull_request) Successful in 1m7s
cd $COMPOSE_DIR now aborts on failure instead of deploying from the wrong dir;
DOCKER_GID is declared then exported so the subshell exit isn't masked.
2026-06-22 00:35:20 +02:00
Ilia Denisov 3d9ba3ac3d docs(deploy): bake Stage 18 prod-deploy decisions into the live docs
CI / changes (pull_request) Successful in 2s
CI / unit (pull_request) Successful in 10s
CI / integration (pull_request) Successful in 16s
CI / ui (pull_request) Successful in 57s
CI / gate (pull_request) Successful in 0s
CI / deploy (pull_request) Successful in 1m8s
- ARCHITECTURE §13 prod bullet -> the realized mechanism: registry transport,
  two-host, rolling + auto-rollback, migration maintenance window, node_exporter,
  the undersized launch; the contour paragraph notes node_exporter + the
  telegram-local profile.
- deploy/README gains a prod rollout runbook (how to run, migrations/restore, cert
  rotation, sizing/monitoring, the full PROD_ set) + node_exporter row, the
  telegram-local profile note, and the soft AWG_CONF note.
- PLAN Stage 18 records the resolved open details and the remaining live cutover
  (pending erudit-game.ru DNS); the tracker reads 'machinery built; cutover pending DNS'.
- PRERELEASE TX/AG note the prod wiring is built.
2026-06-22 00:30:30 +02:00
Ilia Denisov 171b71b7e0 feat(deploy): manual prod-deploy pipeline with rolling rollback (Stage 18)
A workflow_dispatch-only rollout from master (confirm=deploy):

- .gitea/workflows/prod-deploy.yaml builds + pushes the images to the registry,
  ships the compose/config/certs/env over SSH, deploys the main host via
  prod-deploy.sh, then the bot host, then verifies the public site.
- deploy/prod-deploy.sh rolls the main stack one service at a time in dependency
  order (postgres->backend->gateway->landing->validator->caddy), health-checking
  after each; any failure rolls the whole stack back to the previous tag. A schema
  migration adds a maintenance window: the backend (sole writer) is stopped for a
  consistent pg_dump before migrating; image rollback stays DB-safe (expand-contract),
  the dump is kept for a manual restore.
- prod overlay: pull the four main images from the registry by tag.
- Runtime secrets reach the host via a sourced env.sh (single-quoted values keep the
  bcrypt hash's literal $ intact, unlike a --env-file).
2026-06-22 00:25:09 +02:00
Ilia Denisov 2b399d0838 feat(deploy): prod compose split + host-memory monitoring (Stage 18)
Split the contour across the two prod hosts and retune for the small main host:

- Gate vpn+bot to the telegram-local profile. The CI test deploy now passes
  --profile telegram-local so the test contour still brings them; the prod main
  host omits both, and the prod bot runs standalone from docker-compose.bot.yml.
- docker-compose.prod.yml (main-host overlay): publish caddy 80/443 (no host
  caddy in prod; caddy owns ACME) and gateway 9443 (the remote bot dials in over
  mTLS); GOMAXPROCS=2, smaller memory caps and 7d Prometheus retention for the
  2 vCPU / 1.9 GiB host. It launches deliberately undersized; resize reactively.
- docker-compose.bot.yml: standalone bot for the tg host (no VPN, OTLP off since
  otelcol is unreachable from there, dials the main host's bot-link).
- Add node_exporter + a Prometheus scrape so host memory pressure (the OOM
  signal on the tight main host), not just per-container docker_stats, is visible.
- Soften AWG_CONF to a default: only the profiled vpn sidecar consumes it, and
  compose interpolates profiled-out services too, so prod must not require it.
2026-06-22 00:12:43 +02:00
Ilia Denisov f5f45e7afb feat(deploy): Ansible provisioning for prod hosts (Stage 18)
Idempotent playbooks under deploy/ansible/ prepare both production hosts:
docker-ce + compose plugin, a non-sudo deploy service account holding the CI
deploy key, key-only sshd, default-deny ufw, fail2ban, unattended upgrades and
chrony. The main host also opens 80/443/9443 and creates the external edge
network; the tg host verifies direct Bot API egress (the no-VPN assumption).

The application is deployed separately by the prod-deploy workflow (later
phase), running as the deploy account this playbook provisions.
2026-06-21 23:54:57 +02:00
developer b54cb8878d Merge pull request 'fix(ui): retry Mini App launch on backend failure; hide account linking' (#102) from feature/tg-boot-retry-hide-linking into development
CI / changes (push) Successful in 2s
CI / unit (push) Has been skipped
CI / integration (push) Has been skipped
CI / ui (push) Successful in 57s
CI / gate (push) Successful in 0s
CI / deploy (push) Successful in 1m6s
2026-06-21 19:38:37 +00:00
Ilia Denisov e336638ca8 fix(ui): retry Mini App launch on backend failure; hide account linking
CI / changes (pull_request) Successful in 2s
CI / unit (pull_request) Has been skipped
CI / integration (pull_request) Has been skipped
CI / ui (pull_request) Successful in 57s
CI / gate (pull_request) Successful in 0s
CI / deploy (pull_request) Successful in 1m8s
Inside Telegram, a failed initData authentication (e.g. the backend down
during a deploy) dropped the user onto the web login screen — the /app/
experience, which has no place inside the Mini App. bootstrap now retries the
launch a few times in silence and then renders a dedicated boot-error screen
with a Retry button (new BootError.svelte, app.bootError), never falling back
to the web sign-in. A blocked account is still terminal and goes straight to
the blocked screen.

The profile "Link an account" section (email + Telegram link) is hidden while
sign-in is provider-only; the anonymous /app/ guest whose upgrade path this is
comes later. The flow is kept wired (`hidden` on .emailbox) and its two e2e
specs are skipped, both to be re-enabled together.

Adds i18n boot.* copy (en/ru), a mock authTelegram failure hook plus an e2e
covering the retry screen, and bakes both behaviours into FUNCTIONAL(.md/_ru).
2026-06-21 21:23:27 +02:00
41 changed files with 1572 additions and 57 deletions
+5 -2
View File
@@ -301,8 +301,11 @@ jobs:
# App version for the About screen: the git tag if present, else the short SHA
# (the test checkout is shallow/untagged, so this is the SHA here — fine).
export APP_VERSION="$(git -C "$GITHUB_WORKSPACE" describe --tags --always 2>/dev/null || echo dev)"
docker compose --ansi never build --progress plain
docker compose --ansi never up -d --remove-orphans
# The telegram-local profile brings the bot + its VPN sidecar; prod runs the
# bot on its own host instead (deploy/docker-compose.bot.yml), and the prod
# main host omits both. Without the profile they would not start here.
docker compose --ansi never --profile telegram-local build --progress plain
docker compose --ansi never --profile telegram-local up -d --remove-orphans
# The config-only services bind-mount the reseeded config dir. A plain `up -d`
# leaves them on the previous bind mount (the dir was rm'd + recreated), so a
# changed Caddyfile or Grafana dashboard is ignored — force-recreate them to
+266
View File
@@ -0,0 +1,266 @@
# Manual production rollout. Runs ONLY from master, ONLY on workflow_dispatch with
# confirm=deploy (development->master is merged + green first; this is the separate,
# deliberate prod step). Visible sequential jobs from most to least significant:
# build -> deploy-main -> deploy-bot -> verify
# The per-service rolling (postgres->backend->gateway->landing->validator->caddy),
# health-gating and auto-rollback live in deploy/prod-deploy.sh on the main host and
# show in the deploy-main log. Manual post-deploy rollback is prod-rollback.yaml.
# See deploy/README.md (prod runbook).
name: prod-deploy
run-name: "prod deploy ${{ github.sha }}"
on:
workflow_dispatch:
inputs:
confirm:
description: 'Type "deploy" to confirm a production rollout from master.'
required: true
default: ""
permissions:
contents: read
env:
NO_COLOR: "1"
DOCKER_CLI_HINTS: "false"
REGISTRY: docker.iliadenisov.ru/developer
jobs:
build:
if: ${{ github.ref == 'refs/heads/master' && inputs.confirm == 'deploy' }}
runs-on: ubuntu-latest
defaults:
run:
shell: bash
outputs:
tag: ${{ steps.ver.outputs.tag }}
env:
PROD_REGISTRY_USER: ${{ vars.PROD_REGISTRY_USER }}
PROD_REGISTRY_PASSWORD: ${{ secrets.PROD_REGISTRY_PASSWORD }}
VITE_TELEGRAM_BOT_ID: ${{ vars.PROD_VITE_TELEGRAM_BOT_ID }}
VITE_TELEGRAM_LINK: ${{ vars.PROD_VITE_TELEGRAM_LINK }}
VITE_TELEGRAM_GAME_CHANNEL_NAME: ${{ vars.PROD_VITE_TELEGRAM_GAME_CHANNEL_NAME }}
VITE_GATEWAY_URL: ${{ vars.PROD_VITE_GATEWAY_URL }}
POSTGRES_PASSWORD: ${{ secrets.PROD_POSTGRES_PASSWORD }}
GM_BASICAUTH_HASH: ${{ secrets.PROD_GM_BASICAUTH_HASH }}
TELEGRAM_MINIAPP_URL: ${{ vars.PROD_TELEGRAM_MINIAPP_URL }}
DICT_VERSION: ${{ vars.PROD_DICT_VERSION }}
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Compute version tag
id: ver
run: echo "tag=$(git describe --tags --always)" >> "$GITHUB_OUTPUT"
- name: Registry login
run: echo "$PROD_REGISTRY_PASSWORD" | docker login "${REGISTRY%%/*}" -u "$PROD_REGISTRY_USER" --password-stdin
- name: Build and push images
working-directory: deploy
run: |
export TAG="${{ steps.ver.outputs.tag }}" APP_VERSION="${{ steps.ver.outputs.tag }}" SCRABBLE_CONFIG_DIR=.
# The four main-stack images via compose (reuses the build args, incl. VERSION);
# the bot separately, since it is profiled out of the prod compose.
docker compose -f docker-compose.yml -f docker-compose.prod.yml build
docker compose -f docker-compose.yml -f docker-compose.prod.yml push backend gateway landing validator
docker build -f ../platform/telegram/Dockerfile --target bot --build-arg VERSION="$TAG" -t "$REGISTRY/scrabble-telegram-bot:$TAG" ..
docker push "$REGISTRY/scrabble-telegram-bot:$TAG"
deploy-main:
needs: build
runs-on: ubuntu-latest
defaults:
run:
shell: bash
env:
TAG: ${{ needs.build.outputs.tag }}
PROD_REGISTRY_USER: ${{ vars.PROD_REGISTRY_USER }}
PROD_REGISTRY_PASSWORD: ${{ secrets.PROD_REGISTRY_PASSWORD }}
PROD_SSH_KEY: ${{ secrets.PROD_SSH_KEY }}
PROD_SSH_KNOWN_HOSTS: ${{ secrets.PROD_SSH_KNOWN_HOSTS }}
MAIN_HOST: ${{ vars.PROD_MAIN_HOST }}
POSTGRES_PASSWORD: ${{ secrets.PROD_POSTGRES_PASSWORD }}
GM_BASICAUTH_HASH: ${{ secrets.PROD_GM_BASICAUTH_HASH }}
GRAFANA_ADMIN_PASSWORD: ${{ secrets.PROD_GRAFANA_ADMIN_PASSWORD }}
TELEGRAM_BOT_TOKEN: ${{ secrets.PROD_TELEGRAM_BOT_TOKEN }}
PROD_BOTLINK_CA: ${{ secrets.PROD_BOTLINK_CA }}
PROD_BOTLINK_GATEWAY_CERT: ${{ secrets.PROD_BOTLINK_GATEWAY_CERT }}
PROD_BOTLINK_GATEWAY_KEY: ${{ secrets.PROD_BOTLINK_GATEWAY_KEY }}
GM_BASICAUTH_USER: ${{ vars.PROD_GM_BASICAUTH_USER }}
GRAFANA_ROOT_URL: ${{ vars.PROD_GRAFANA_ROOT_URL }}
CADDY_SITE_ADDRESS: ${{ vars.PROD_CADDY_SITE_ADDRESS }}
LOG_LEVEL: ${{ vars.PROD_LOG_LEVEL }}
DICT_VERSION: ${{ vars.PROD_DICT_VERSION }}
POSTGRES_DB: ${{ vars.PROD_POSTGRES_DB }}
POSTGRES_USER: ${{ vars.PROD_POSTGRES_USER }}
TELEGRAM_MINIAPP_URL: ${{ vars.PROD_TELEGRAM_MINIAPP_URL }}
steps:
- uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Set up SSH
run: |
mkdir -p ~/.ssh && chmod 700 ~/.ssh
printf '%s\n' "$PROD_SSH_KEY" > ~/.ssh/id_deploy && chmod 600 ~/.ssh/id_deploy
printf '%s\n' "$PROD_SSH_KNOWN_HOSTS" > ~/.ssh/known_hosts
- name: Determine previous tag and migration
run: |
ssh_main() { ssh -i ~/.ssh/id_deploy -o BatchMode=yes "deploy@$MAIN_HOST" "$@"; }
PREV_TAG="$(ssh_main 'cat /opt/scrabble/DEPLOYED_TAG 2>/dev/null || echo none')"
MIGRATION=0
if [ "$PREV_TAG" != none ]; then
if ! git cat-file -e "$PREV_TAG^{commit}" 2>/dev/null; then
MIGRATION=1
elif git diff --name-only "$PREV_TAG..$TAG" -- backend/internal/postgres/migrations/ | grep -q .; then
MIGRATION=1
fi
fi
{ echo "PREV_TAG=$PREV_TAG"; echo "MIGRATION=$MIGRATION"; } >> "$GITHUB_ENV"
echo "prev=$PREV_TAG migration=$MIGRATION"
- name: Render main env + certs
run: |
umask 077
mkdir -p stage/certs-main
cat > stage/env.sh <<EOF
export REGISTRY='$REGISTRY'
export SCRABBLE_CONFIG_DIR='/opt/scrabble'
export POSTGRES_DB='${POSTGRES_DB:-scrabble}'
export POSTGRES_USER='${POSTGRES_USER:-scrabble}'
export POSTGRES_PASSWORD='$POSTGRES_PASSWORD'
export GM_BASICAUTH_USER='${GM_BASICAUTH_USER:-gm}'
export GM_BASICAUTH_HASH='$GM_BASICAUTH_HASH'
export GRAFANA_ADMIN_PASSWORD='$GRAFANA_ADMIN_PASSWORD'
export GRAFANA_ROOT_URL='$GRAFANA_ROOT_URL'
export CADDY_SITE_ADDRESS='$CADDY_SITE_ADDRESS'
export LOG_LEVEL='${LOG_LEVEL:-info}'
export DICT_VERSION='$DICT_VERSION'
export APP_VERSION='$TAG'
export TELEGRAM_BOT_TOKEN='$TELEGRAM_BOT_TOKEN'
export TELEGRAM_MINIAPP_URL='$TELEGRAM_MINIAPP_URL'
export GATEWAY_ABUSE_BAN_ENABLED='true'
EOF
printf '%s\n' "$PROD_BOTLINK_CA" > stage/certs-main/ca.crt
printf '%s\n' "$PROD_BOTLINK_GATEWAY_CERT" > stage/certs-main/gateway.crt
printf '%s\n' "$PROD_BOTLINK_GATEWAY_KEY" > stage/certs-main/gateway.key
chmod 644 stage/certs-main/*
- name: Deploy the main host
run: |
ssh_main() { ssh -i ~/.ssh/id_deploy -o BatchMode=yes "deploy@$MAIN_HOST" "$@"; }
ssh_main 'mkdir -p /opt/scrabble/compose'
tar -C deploy -czf - docker-compose.yml docker-compose.prod.yml prod-deploy.sh \
| ssh_main 'tar -C /opt/scrabble/compose -xzf -'
tar -C deploy -czf - caddy otelcol prometheus tempo grafana \
| ssh_main 'tar -C /opt/scrabble -xzf -'
tar -C stage -czf - certs-main \
| ssh_main 'rm -rf /opt/scrabble/certs && mkdir -p /opt/scrabble/certs && tar -C /opt/scrabble/certs --strip-components=1 -xzf -'
scp -i ~/.ssh/id_deploy -o BatchMode=yes stage/env.sh "deploy@$MAIN_HOST:/opt/scrabble/env.sh"
echo "$PROD_REGISTRY_PASSWORD" | ssh_main "docker login ${REGISTRY%%/*} -u $PROD_REGISTRY_USER --password-stdin"
ssh_main "TAG='$TAG' PREV_TAG='$PREV_TAG' MIGRATION='$MIGRATION' bash /opt/scrabble/compose/prod-deploy.sh"
deploy-bot:
needs: [build, deploy-main]
runs-on: ubuntu-latest
defaults:
run:
shell: bash
env:
TAG: ${{ needs.build.outputs.tag }}
PROD_REGISTRY_USER: ${{ vars.PROD_REGISTRY_USER }}
PROD_REGISTRY_PASSWORD: ${{ secrets.PROD_REGISTRY_PASSWORD }}
PROD_SSH_KEY: ${{ secrets.PROD_SSH_KEY }}
PROD_SSH_KNOWN_HOSTS: ${{ secrets.PROD_SSH_KNOWN_HOSTS }}
TG_HOST: ${{ vars.PROD_TG_HOST }}
MAIN_HOST: ${{ vars.PROD_MAIN_HOST }}
TELEGRAM_BOT_TOKEN: ${{ secrets.PROD_TELEGRAM_BOT_TOKEN }}
TELEGRAM_PROMO_BOT_TOKEN: ${{ secrets.PROD_TELEGRAM_PROMO_BOT_TOKEN }}
PROD_BOTLINK_CA: ${{ secrets.PROD_BOTLINK_CA }}
PROD_BOTLINK_BOT_CERT: ${{ secrets.PROD_BOTLINK_BOT_CERT }}
PROD_BOTLINK_BOT_KEY: ${{ secrets.PROD_BOTLINK_BOT_KEY }}
LOG_LEVEL: ${{ vars.PROD_LOG_LEVEL }}
TELEGRAM_MINIAPP_URL: ${{ vars.PROD_TELEGRAM_MINIAPP_URL }}
TELEGRAM_GAME_CHANNEL_ID: ${{ vars.PROD_TELEGRAM_GAME_CHANNEL_ID }}
TELEGRAM_CHAT_ID: ${{ vars.PROD_TELEGRAM_CHAT_ID }}
TELEGRAM_BOT_USERNAME: ${{ vars.PROD_TELEGRAM_BOT_USERNAME }}
TELEGRAM_BOT_LINK: ${{ vars.PROD_VITE_TELEGRAM_LINK }}
steps:
- uses: actions/checkout@v4
- name: Set up SSH
run: |
mkdir -p ~/.ssh && chmod 700 ~/.ssh
printf '%s\n' "$PROD_SSH_KEY" > ~/.ssh/id_deploy && chmod 600 ~/.ssh/id_deploy
printf '%s\n' "$PROD_SSH_KNOWN_HOSTS" > ~/.ssh/known_hosts
- name: Render bot env + certs
run: |
umask 077
mkdir -p stage/certs-bot
cat > stage/env.bot.sh <<EOF
export SCRABBLE_CONFIG_DIR='/opt/scrabble'
export BOT_IMAGE='$REGISTRY/scrabble-telegram-bot:$TAG'
export BOTLINK_GATEWAY_ADDR='$MAIN_HOST:9443'
export TELEGRAM_BOT_TOKEN='$TELEGRAM_BOT_TOKEN'
export TELEGRAM_MINIAPP_URL='$TELEGRAM_MINIAPP_URL'
export TELEGRAM_GAME_CHANNEL_ID='$TELEGRAM_GAME_CHANNEL_ID'
export TELEGRAM_CHAT_ID='$TELEGRAM_CHAT_ID'
export TELEGRAM_PROMO_BOT_TOKEN='$TELEGRAM_PROMO_BOT_TOKEN'
export TELEGRAM_BOT_USERNAME='$TELEGRAM_BOT_USERNAME'
export TELEGRAM_BOT_LINK='$TELEGRAM_BOT_LINK'
export LOG_LEVEL='${LOG_LEVEL:-info}'
EOF
printf '%s\n' "$PROD_BOTLINK_CA" > stage/certs-bot/ca.crt
printf '%s\n' "$PROD_BOTLINK_BOT_CERT" > stage/certs-bot/bot.crt
printf '%s\n' "$PROD_BOTLINK_BOT_KEY" > stage/certs-bot/bot.key
chmod 644 stage/certs-bot/*
- name: Deploy the bot host
run: |
ssh_tg() { ssh -i ~/.ssh/id_deploy -o BatchMode=yes "deploy@$TG_HOST" "$@"; }
ssh_tg 'mkdir -p /opt/scrabble/compose'
tar -C deploy -czf - docker-compose.bot.yml | ssh_tg 'tar -C /opt/scrabble/compose -xzf -'
tar -C stage -czf - certs-bot \
| ssh_tg 'rm -rf /opt/scrabble/certs && mkdir -p /opt/scrabble/certs && tar -C /opt/scrabble/certs --strip-components=1 -xzf -'
scp -i ~/.ssh/id_deploy -o BatchMode=yes stage/env.bot.sh "deploy@$TG_HOST:/opt/scrabble/env.bot.sh"
echo "$PROD_REGISTRY_PASSWORD" | ssh_tg "docker login ${REGISTRY%%/*} -u $PROD_REGISTRY_USER --password-stdin"
ssh_tg 'set -a; . /opt/scrabble/env.bot.sh; set +a; cd /opt/scrabble/compose;
docker compose -f docker-compose.bot.yml pull;
docker compose -f docker-compose.bot.yml up -d'
ssh_tg 'for i in $(seq 1 20); do
s=$(docker inspect -f "{{.State.Status}}" scrabble-telegram-bot 2>/dev/null || echo missing)
r=$(docker inspect -f "{{.State.Restarting}}" scrabble-telegram-bot 2>/dev/null || echo true)
if [ "$s" = running ] && [ "$r" = false ]; then
c1=$(docker inspect -f "{{.RestartCount}}" scrabble-telegram-bot); sleep 5
c2=$(docker inspect -f "{{.RestartCount}}" scrabble-telegram-bot)
[ "$c1" = "$c2" ] && { echo "bot healthy"; exit 0; }
fi
sleep 3
done
echo "bot not healthy:"; docker logs --tail 80 scrabble-telegram-bot; exit 1'
verify:
needs: [deploy-main, deploy-bot]
runs-on: ubuntu-latest
defaults:
run:
shell: bash
env:
PROD_SSH_KEY: ${{ secrets.PROD_SSH_KEY }}
PROD_SSH_KNOWN_HOSTS: ${{ secrets.PROD_SSH_KNOWN_HOSTS }}
MAIN_HOST: ${{ vars.PROD_MAIN_HOST }}
CADDY_SITE_ADDRESS: ${{ vars.PROD_CADDY_SITE_ADDRESS }}
steps:
- name: Set up SSH
run: |
mkdir -p ~/.ssh && chmod 700 ~/.ssh
printf '%s\n' "$PROD_SSH_KEY" > ~/.ssh/id_deploy && chmod 600 ~/.ssh/id_deploy
printf '%s\n' "$PROD_SSH_KNOWN_HOSTS" > ~/.ssh/known_hosts
- name: Verify the public site
run: |
domain="${CADDY_SITE_ADDRESS%% *}"
ssh -i ~/.ssh/id_deploy -o BatchMode=yes "deploy@$MAIN_HOST" "for i in \$(seq 1 20); do
if curl -fsS -k --resolve $domain:443:127.0.0.1 https://$domain/ -o /dev/null &&
curl -fsS -k --resolve $domain:443:127.0.0.1 https://$domain/app/ -o /dev/null &&
docker run --rm --network scrabble-internal alpine:3.20 wget -q -T 5 -O /dev/null http://backend:8080/readyz; then
echo 'public site + /app/ + backend healthy'; exit 0
fi
sleep 5
done
echo 'public verify failed; recent caddy + gateway + backend logs:'
docker logs --tail 40 scrabble-caddy; docker logs --tail 40 scrabble-gateway; docker logs --tail 40 scrabble-backend
exit 1"
+223
View File
@@ -0,0 +1,223 @@
# Manual production rollback. Runs ONLY from master, ONLY on workflow_dispatch with
# confirm=rollback. Re-deploys an already-published image tag (no build): leave
# target_version blank to roll back to the previously deployed version (read from the
# main host), or set it to a specific release tag from the Releases page. The
# re-deploy is the same rolling, health-gated path as prod-deploy (TAG=target,
# MIGRATION=0 — rollback is image-only and never migrates the DB; image rollback is
# DB-safe under the expand-contract rule). See deploy/README.md (prod runbook).
name: prod-rollback
run-name: "prod rollback ${{ inputs.target_version || 'previous' }}"
on:
workflow_dispatch:
inputs:
confirm:
description: 'Type "rollback" to confirm a production rollback.'
required: true
default: ""
target_version:
description: "Release tag to roll back to (blank = the previous deployed version)."
required: false
default: ""
permissions:
contents: read
env:
NO_COLOR: "1"
DOCKER_CLI_HINTS: "false"
REGISTRY: docker.iliadenisov.ru/developer
jobs:
rollback-main:
if: ${{ github.ref == 'refs/heads/master' && inputs.confirm == 'rollback' }}
runs-on: ubuntu-latest
defaults:
run:
shell: bash
outputs:
target: ${{ steps.resolve.outputs.target }}
env:
PROD_REGISTRY_USER: ${{ vars.PROD_REGISTRY_USER }}
PROD_REGISTRY_PASSWORD: ${{ secrets.PROD_REGISTRY_PASSWORD }}
PROD_SSH_KEY: ${{ secrets.PROD_SSH_KEY }}
PROD_SSH_KNOWN_HOSTS: ${{ secrets.PROD_SSH_KNOWN_HOSTS }}
MAIN_HOST: ${{ vars.PROD_MAIN_HOST }}
POSTGRES_PASSWORD: ${{ secrets.PROD_POSTGRES_PASSWORD }}
GM_BASICAUTH_HASH: ${{ secrets.PROD_GM_BASICAUTH_HASH }}
GRAFANA_ADMIN_PASSWORD: ${{ secrets.PROD_GRAFANA_ADMIN_PASSWORD }}
TELEGRAM_BOT_TOKEN: ${{ secrets.PROD_TELEGRAM_BOT_TOKEN }}
PROD_BOTLINK_CA: ${{ secrets.PROD_BOTLINK_CA }}
PROD_BOTLINK_GATEWAY_CERT: ${{ secrets.PROD_BOTLINK_GATEWAY_CERT }}
PROD_BOTLINK_GATEWAY_KEY: ${{ secrets.PROD_BOTLINK_GATEWAY_KEY }}
GM_BASICAUTH_USER: ${{ vars.PROD_GM_BASICAUTH_USER }}
GRAFANA_ROOT_URL: ${{ vars.PROD_GRAFANA_ROOT_URL }}
CADDY_SITE_ADDRESS: ${{ vars.PROD_CADDY_SITE_ADDRESS }}
LOG_LEVEL: ${{ vars.PROD_LOG_LEVEL }}
DICT_VERSION: ${{ vars.PROD_DICT_VERSION }}
POSTGRES_DB: ${{ vars.PROD_POSTGRES_DB }}
POSTGRES_USER: ${{ vars.PROD_POSTGRES_USER }}
TELEGRAM_MINIAPP_URL: ${{ vars.PROD_TELEGRAM_MINIAPP_URL }}
INPUT_TARGET: ${{ inputs.target_version }}
steps:
- uses: actions/checkout@v4
- name: Set up SSH
run: |
mkdir -p ~/.ssh && chmod 700 ~/.ssh
printf '%s\n' "$PROD_SSH_KEY" > ~/.ssh/id_deploy && chmod 600 ~/.ssh/id_deploy
printf '%s\n' "$PROD_SSH_KNOWN_HOSTS" > ~/.ssh/known_hosts
- name: Resolve rollback target
id: resolve
run: |
ssh_main() { ssh -i ~/.ssh/id_deploy -o BatchMode=yes "deploy@$MAIN_HOST" "$@"; }
CURRENT="$(ssh_main 'cat /opt/scrabble/DEPLOYED_TAG 2>/dev/null || echo none')"
if [ -n "$INPUT_TARGET" ]; then
TARGET="$INPUT_TARGET"
else
TARGET="$(ssh_main 'cat /opt/scrabble/PREVIOUS_TAG 2>/dev/null || echo none')"
fi
if [ -z "$TARGET" ] || [ "$TARGET" = none ]; then
echo "no rollback target (no PREVIOUS_TAG on the host and no target_version input)"; exit 1
fi
if [ "$TARGET" = "$CURRENT" ]; then
echo "target $TARGET is already the deployed version; nothing to do"; exit 1
fi
echo "rolling back: current=$CURRENT -> target=$TARGET"
echo "target=$TARGET" >> "$GITHUB_OUTPUT"
{ echo "TARGET=$TARGET"; echo "CURRENT=$CURRENT"; } >> "$GITHUB_ENV"
- name: Render main env + certs
run: |
umask 077
mkdir -p stage/certs-main
cat > stage/env.sh <<EOF
export REGISTRY='$REGISTRY'
export SCRABBLE_CONFIG_DIR='/opt/scrabble'
export POSTGRES_DB='${POSTGRES_DB:-scrabble}'
export POSTGRES_USER='${POSTGRES_USER:-scrabble}'
export POSTGRES_PASSWORD='$POSTGRES_PASSWORD'
export GM_BASICAUTH_USER='${GM_BASICAUTH_USER:-gm}'
export GM_BASICAUTH_HASH='$GM_BASICAUTH_HASH'
export GRAFANA_ADMIN_PASSWORD='$GRAFANA_ADMIN_PASSWORD'
export GRAFANA_ROOT_URL='$GRAFANA_ROOT_URL'
export CADDY_SITE_ADDRESS='$CADDY_SITE_ADDRESS'
export LOG_LEVEL='${LOG_LEVEL:-info}'
export DICT_VERSION='$DICT_VERSION'
export APP_VERSION='$TARGET'
export TELEGRAM_BOT_TOKEN='$TELEGRAM_BOT_TOKEN'
export TELEGRAM_MINIAPP_URL='$TELEGRAM_MINIAPP_URL'
export GATEWAY_ABUSE_BAN_ENABLED='true'
EOF
printf '%s\n' "$PROD_BOTLINK_CA" > stage/certs-main/ca.crt
printf '%s\n' "$PROD_BOTLINK_GATEWAY_CERT" > stage/certs-main/gateway.crt
printf '%s\n' "$PROD_BOTLINK_GATEWAY_KEY" > stage/certs-main/gateway.key
chmod 644 stage/certs-main/*
- name: Roll the main host back
run: |
ssh_main() { ssh -i ~/.ssh/id_deploy -o BatchMode=yes "deploy@$MAIN_HOST" "$@"; }
ssh_main 'mkdir -p /opt/scrabble/compose'
tar -C deploy -czf - docker-compose.yml docker-compose.prod.yml prod-deploy.sh \
| ssh_main 'tar -C /opt/scrabble/compose -xzf -'
tar -C deploy -czf - caddy otelcol prometheus tempo grafana \
| ssh_main 'tar -C /opt/scrabble -xzf -'
tar -C stage -czf - certs-main \
| ssh_main 'rm -rf /opt/scrabble/certs && mkdir -p /opt/scrabble/certs && tar -C /opt/scrabble/certs --strip-components=1 -xzf -'
scp -i ~/.ssh/id_deploy -o BatchMode=yes stage/env.sh "deploy@$MAIN_HOST:/opt/scrabble/env.sh"
echo "$PROD_REGISTRY_PASSWORD" | ssh_main "docker login ${REGISTRY%%/*} -u $PROD_REGISTRY_USER --password-stdin"
# Image-only rollback: no migration window (TAG=target, MIGRATION=0). A failed
# rollback's auto-revert returns to the current version (PREV_TAG=$CURRENT).
ssh_main "TAG='$TARGET' PREV_TAG='$CURRENT' MIGRATION=0 bash /opt/scrabble/compose/prod-deploy.sh"
rollback-bot:
needs: rollback-main
runs-on: ubuntu-latest
defaults:
run:
shell: bash
env:
TARGET: ${{ needs.rollback-main.outputs.target }}
PROD_REGISTRY_USER: ${{ vars.PROD_REGISTRY_USER }}
PROD_REGISTRY_PASSWORD: ${{ secrets.PROD_REGISTRY_PASSWORD }}
PROD_SSH_KEY: ${{ secrets.PROD_SSH_KEY }}
PROD_SSH_KNOWN_HOSTS: ${{ secrets.PROD_SSH_KNOWN_HOSTS }}
TG_HOST: ${{ vars.PROD_TG_HOST }}
MAIN_HOST: ${{ vars.PROD_MAIN_HOST }}
TELEGRAM_BOT_TOKEN: ${{ secrets.PROD_TELEGRAM_BOT_TOKEN }}
TELEGRAM_PROMO_BOT_TOKEN: ${{ secrets.PROD_TELEGRAM_PROMO_BOT_TOKEN }}
PROD_BOTLINK_CA: ${{ secrets.PROD_BOTLINK_CA }}
PROD_BOTLINK_BOT_CERT: ${{ secrets.PROD_BOTLINK_BOT_CERT }}
PROD_BOTLINK_BOT_KEY: ${{ secrets.PROD_BOTLINK_BOT_KEY }}
LOG_LEVEL: ${{ vars.PROD_LOG_LEVEL }}
TELEGRAM_MINIAPP_URL: ${{ vars.PROD_TELEGRAM_MINIAPP_URL }}
TELEGRAM_GAME_CHANNEL_ID: ${{ vars.PROD_TELEGRAM_GAME_CHANNEL_ID }}
TELEGRAM_CHAT_ID: ${{ vars.PROD_TELEGRAM_CHAT_ID }}
TELEGRAM_BOT_USERNAME: ${{ vars.PROD_TELEGRAM_BOT_USERNAME }}
TELEGRAM_BOT_LINK: ${{ vars.PROD_VITE_TELEGRAM_LINK }}
steps:
- uses: actions/checkout@v4
- name: Set up SSH
run: |
mkdir -p ~/.ssh && chmod 700 ~/.ssh
printf '%s\n' "$PROD_SSH_KEY" > ~/.ssh/id_deploy && chmod 600 ~/.ssh/id_deploy
printf '%s\n' "$PROD_SSH_KNOWN_HOSTS" > ~/.ssh/known_hosts
- name: Render bot env + certs
run: |
umask 077
mkdir -p stage/certs-bot
cat > stage/env.bot.sh <<EOF
export SCRABBLE_CONFIG_DIR='/opt/scrabble'
export BOT_IMAGE='$REGISTRY/scrabble-telegram-bot:$TARGET'
export BOTLINK_GATEWAY_ADDR='$MAIN_HOST:9443'
export TELEGRAM_BOT_TOKEN='$TELEGRAM_BOT_TOKEN'
export TELEGRAM_MINIAPP_URL='$TELEGRAM_MINIAPP_URL'
export TELEGRAM_GAME_CHANNEL_ID='$TELEGRAM_GAME_CHANNEL_ID'
export TELEGRAM_CHAT_ID='$TELEGRAM_CHAT_ID'
export TELEGRAM_PROMO_BOT_TOKEN='$TELEGRAM_PROMO_BOT_TOKEN'
export TELEGRAM_BOT_USERNAME='$TELEGRAM_BOT_USERNAME'
export TELEGRAM_BOT_LINK='$TELEGRAM_BOT_LINK'
export LOG_LEVEL='${LOG_LEVEL:-info}'
EOF
printf '%s\n' "$PROD_BOTLINK_CA" > stage/certs-bot/ca.crt
printf '%s\n' "$PROD_BOTLINK_BOT_CERT" > stage/certs-bot/bot.crt
printf '%s\n' "$PROD_BOTLINK_BOT_KEY" > stage/certs-bot/bot.key
chmod 644 stage/certs-bot/*
- name: Roll the bot host back
run: |
ssh_tg() { ssh -i ~/.ssh/id_deploy -o BatchMode=yes "deploy@$TG_HOST" "$@"; }
ssh_tg 'mkdir -p /opt/scrabble/compose'
tar -C deploy -czf - docker-compose.bot.yml | ssh_tg 'tar -C /opt/scrabble/compose -xzf -'
tar -C stage -czf - certs-bot \
| ssh_tg 'rm -rf /opt/scrabble/certs && mkdir -p /opt/scrabble/certs && tar -C /opt/scrabble/certs --strip-components=1 -xzf -'
scp -i ~/.ssh/id_deploy -o BatchMode=yes stage/env.bot.sh "deploy@$TG_HOST:/opt/scrabble/env.bot.sh"
echo "$PROD_REGISTRY_PASSWORD" | ssh_tg "docker login ${REGISTRY%%/*} -u $PROD_REGISTRY_USER --password-stdin"
ssh_tg 'set -a; . /opt/scrabble/env.bot.sh; set +a; cd /opt/scrabble/compose;
docker compose -f docker-compose.bot.yml pull;
docker compose -f docker-compose.bot.yml up -d'
verify:
needs: [rollback-main, rollback-bot]
runs-on: ubuntu-latest
defaults:
run:
shell: bash
env:
PROD_SSH_KEY: ${{ secrets.PROD_SSH_KEY }}
PROD_SSH_KNOWN_HOSTS: ${{ secrets.PROD_SSH_KNOWN_HOSTS }}
MAIN_HOST: ${{ vars.PROD_MAIN_HOST }}
CADDY_SITE_ADDRESS: ${{ vars.PROD_CADDY_SITE_ADDRESS }}
steps:
- name: Set up SSH
run: |
mkdir -p ~/.ssh && chmod 700 ~/.ssh
printf '%s\n' "$PROD_SSH_KEY" > ~/.ssh/id_deploy && chmod 600 ~/.ssh/id_deploy
printf '%s\n' "$PROD_SSH_KNOWN_HOSTS" > ~/.ssh/known_hosts
- name: Verify the public site
run: |
domain="${CADDY_SITE_ADDRESS%% *}"
ssh -i ~/.ssh/id_deploy -o BatchMode=yes "deploy@$MAIN_HOST" "for i in \$(seq 1 20); do
if curl -fsS -k --resolve $domain:443:127.0.0.1 https://$domain/ -o /dev/null &&
docker run --rm --network scrabble-internal alpine:3.20 wget -q -T 5 -O /dev/null http://backend:8080/readyz; then
echo 'rolled-back site healthy'; exit 0
fi
sleep 5
done
echo 'verify failed'; docker logs --tail 40 scrabble-caddy; docker logs --tail 40 scrabble-backend; exit 1"
+23 -13
View File
@@ -51,7 +51,7 @@ independent (see ARCHITECTURE §9.1).
| 15 | Dual Telegram bots & language-gated variants | **done** |
| 16 | Deploy infra & test contour (Dockerfiles, gateway static UI, compose, observability) | **done** |
| 17 | Test-contour verification & defect fixes | **done** |
| 18 | Prod contour deploy (SSH export/import, manual after merge) | todo |
| 18 | Prod contour deploy (registry, two-host, rolling + auto-rollback; manual after merge) | machinery built; first cutover pending DNS |
| 19 | User feedback (in-app submit + attachment, admin review/reply, account roles) | **done** |
Scaffolding is incremental: `go.work` lists only existing modules; each stage
@@ -413,18 +413,28 @@ raw list is kept here as the record of what the first contour run surfaced.
"что-то пошло не так". при этом "new -> эрудит" работает. Попробуй посмотреть в логах сейчас, может что-то есть. Или как-то иначе проанализируй, или давай вместе будем смотреть, если не получится.
### Stage 18 — Prod contour deploy
Scope: the **production contour** on a remote host over SSH. Deploy by **container export/import**
(`docker save``scp`/ssh → `docker load``docker compose up` on the remote), the SSH key + host IP
in Gitea secrets; **strictly manual** (`workflow_dispatch`) after `development` is merged to `master`
(the Stage 16 branch model: `feature/* → development → master`, merge gated green). Two-contour config
uses **`TEST_`/`PROD_` secret/variable prefixes** — Gitea 1.26 has no deployment environments (verified:
the `environments` API 404s), so a flat prefixed namespace is the convention.
Reuses the Stage 16 `deploy/docker-compose.yml` as-is, mapping the **`PROD_`** set onto the same
unprefixed compose vars. **No host caddy on prod**, so the contour's own caddy terminates TLS — set
`CADDY_SITE_ADDRESS` to the prod domain so caddy does its own ACME (the Caddyfile is already
parameterised for this; the test contour leaves it `:80` behind the host caddy).
Open details (re-interview): export/import vs a registry trade-off; prod domain/cert source (ACME vs a
provided cert) at the contour caddy; prod VPN; rollback.
Scope: the **production contour** on **two remote hosts** over SSH — main (full stack, `erudit-game.ru`)
and tg (the bot only). Resolved open details (re-interviewed):
- **Transport: a registry** (not export/import) — build + push to `docker.iliadenisov.ru`, the hosts pull by tag.
- **Cert: ACME** at the contour caddy (`CADDY_SITE_ADDRESS=erudit-game.ru www.erudit-game.ru`, no host caddy).
- **No prod VPN** — the bot host has native Bot API egress (verified `api.telegram.org` → 200).
- **Rollback** — rolling per-service deploy (least → most dependent), health-gated, auto-rollback to the
previous image tag; a maintenance window + consistent `pg_dump` only on a schema migration
(expand-contract keeps the auto-rollback image-only; the dump is a manual safety net).
**Strictly manual** (`workflow_dispatch` from `master`, `confirm=deploy`) after `development → master`
is merged green. `TEST_`/`PROD_` prefixed Gitea secrets/variables (Gitea 1.26 has no deployment
environments — the `environments` API 404s). Hosts are provisioned by **`deploy/ansible/`** (docker, a
non-sudo `deploy` user with the CI key, key-only sshd, ufw, fail2ban). The main host is **launch-sized**
(2 vCPU / 1.9 GiB): `docker-compose.prod.yml` trims the R7 limits (`GOMAXPROCS=2`, smaller caps, 7d
Prometheus retention) and adds `node_exporter` for host-memory monitoring (launch undersized, resize at
Selectel reactively). `vpn`+`bot` are gated to a `telegram-local` compose profile (test only); the prod
bot runs standalone from `docker-compose.bot.yml`. `GATEWAY_ABUSE_BAN_ENABLED=true`.
**Built:** `deploy/ansible/` (both hosts provisioned + verified), the compose split + `node_exporter`,
`.gitea/workflows/prod-deploy.yaml` + `deploy/prod-deploy.sh`, the full `PROD_` secret/variable set.
**Remaining (acceptance):** the **first live cutover** — waits on the `erudit-game.ru` DNS delegation
(`A`/`www` → the main host) that ACME requires; then run the workflow and verify the public site end-to-end.
### Stage 19 — User feedback *(done)*
A user→operator feedback channel, sequenced after the numbered stages but shipped **before** the Stage 18
+2 -2
View File
@@ -39,8 +39,8 @@ the edge before prod. Each phase maps back to the owner's raw pre-release TODO l
| FM | First-move tile draw (official rules): each seated player draws a tile, the one closest to "A" leads (a blank beats every letter), ties re-drawing until a single leader; **honest per-draw `crypto/rand` entropy**, not the bag seed, so the **record** (`game_setup_draws`, migration `00013`) — not a seed — is the only account of the outcome, kept for future **tournaments** (designed as a discrete per-tile "player N draws" step). Friend/AI draws at create; **auto-match draws at *open*** against a synthetic `uuid.Nil` opponent whose draw rows are back-filled on join, so the opener's seat is fixed up front and the existing open-game pre-move is preserved (no reseating, no play-gating). Admin `/_gm/games/:id` gains the recorded draw list + a simple **step-by-step board replay** (`ReplayTimeline`). | owner ad-hoc | **done** |
| SB | Single Telegram bot + per-user variant preferences: the two per-language bots collapse into **one** (drop `accounts.service_language`, `supported_languages`, the `*_EN`/`*_RU` env vars and game-language push routing — the single bot renders in the recipient's `preferred_language`); New Game variant gating moves to a profile **`variant_preferences`** set (default Erudit only, Erudit-first, server-enforced on the caller's auto-match/vs-AI/invitation-create paths, an invited friend may accept any variant); env vars collapse to unsuffixed `TELEGRAM_BOT_TOKEN`/`TELEGRAM_GAME_CHANNEL_ID`/`VITE_TELEGRAM_LINK`/`VITE_TELEGRAM_GAME_CHANNEL_NAME` and `GATEWAY_DEFAULT_SUPPORTED_LANGUAGES` is removed; wire drops `service_language`/`supported_languages` (Session, ValidateInitDataResponse) + the push `language` routing field and adds `variant_preferences` to Profile/UpdateProfile. | owner ad-hoc | **done** |
| DV | Dictionary version hygiene: CI + image/compose seed track the current release (`v1.2.1`); a **seed-drift guard** records the flat dir's seed in an authoritative `.seed_version` marker so a bumped build seed on a live volume is ignored (it can't relabel live bytes — which would mis-serve the dictionary + void games pinned to the prior label); `DICT_VERSION` is the fresh-volume seed only, a live contour migrates through the admin console | owner ad-hoc | **done** |
| TX | Telegram egress off the main host: split the connector into a home **validator** (Mini App / Login-Widget HMAC, no VPN, no Bot API — so game login no longer depends on Telegram being reachable) and a remote **bot** (Bot API long-poll + `sendMessage`) that holds **no inbound port** and dials the gateway over a reverse **mTLS bot-link** (`pkg/proto/botlink/v1`); the gateway funnels out-of-app push (fire-and-forget, at-most-once) and the backend admin broadcasts (a relay that awaits the bot's ack) down the link. The bot is Telegram-rate-limited; **one bot now**, with seams (a bot registry + `owns_updates` + command ids) for N later; **no webhook** (rejected: one URL per token, adds inbound + a static address). The **unified test contour** runs the split (the bot keeps its VPN sidecar and dials the gateway by its internal name; certs from `deploy/gen-certs.sh`). The **prod** wiring — the bot on a separate host (no VPN), the gateway bot-link port published, `PROD_` certs with scheduled rotation, an SSH deploy of both hosts together — is the **deferred final stage** (Stage 18). | owner ad-hoc | **done** (code + test contour; prod wiring Stage 18) |
| AG | Anti-abuse IP ban + honeypot/honeytoken (prod-only): a fail2ban-style in-memory `ratelimit.Banlist` keyed by client IP, fed by sustained rate-limiter rejections (the IP-keyed public/email/admin classes — the user class stays the soft-flag's concern), a **honeypot** decoy path (the contour caddy tags `/.env`, `/.git`, `/wp-*`, … with `X-Scrabble-Honeypot` and routes them to the gateway), and a **honeytoken** (`GATEWAY_HONEYTOKEN`, a planted bearer). The `abuseGuard` edge middleware refuses a banned IP with **429** before any work — closing the R3 gap that the static SPA/landing was outside the token bucket. Off by default — it keys by the real client IP the shared-NAT test contour does not expose (detection still logs there); enabled in prod via `GATEWAY_ABUSE_BAN_ENABLED`. Operators see + lift bans on the console **Throttled** page; the gateway syncs its active set to the backend (`/api/v1/internal/bans/sync`, `internal/banview`) every 30 s and applies operator unbans. | owner ad-hoc | **done** (code + test contour; ban enabled in prod Stage 18) |
| TX | Telegram egress off the main host: split the connector into a home **validator** (Mini App / Login-Widget HMAC, no VPN, no Bot API — so game login no longer depends on Telegram being reachable) and a remote **bot** (Bot API long-poll + `sendMessage`) that holds **no inbound port** and dials the gateway over a reverse **mTLS bot-link** (`pkg/proto/botlink/v1`); the gateway funnels out-of-app push (fire-and-forget, at-most-once) and the backend admin broadcasts (a relay that awaits the bot's ack) down the link. The bot is Telegram-rate-limited; **one bot now**, with seams (a bot registry + `owns_updates` + command ids) for N later; **no webhook** (rejected: one URL per token, adds inbound + a static address). The **unified test contour** runs the split (the bot keeps its VPN sidecar and dials the gateway by its internal name; certs from `deploy/gen-certs.sh`). The **prod** wiring — the bot on a separate host (no VPN), the gateway bot-link port published, `PROD_` certs, an SSH deploy of both hosts together — is **built in Stage 18** (the two-host registry rollout; first cutover pending the `erudit-game.ru` DNS). | owner ad-hoc | **done** (code + test contour; prod wiring built — Stage 18) |
| AG | Anti-abuse IP ban + honeypot/honeytoken (prod-only): a fail2ban-style in-memory `ratelimit.Banlist` keyed by client IP, fed by sustained rate-limiter rejections (the IP-keyed public/email/admin classes — the user class stays the soft-flag's concern), a **honeypot** decoy path (the contour caddy tags `/.env`, `/.git`, `/wp-*`, … with `X-Scrabble-Honeypot` and routes them to the gateway), and a **honeytoken** (`GATEWAY_HONEYTOKEN`, a planted bearer). The `abuseGuard` edge middleware refuses a banned IP with **429** before any work — closing the R3 gap that the static SPA/landing was outside the token bucket. Off by default — it keys by the real client IP the shared-NAT test contour does not expose (detection still logs there); enabled in prod via `GATEWAY_ABUSE_BAN_ENABLED`. Operators see + lift bans on the console **Throttled** page; the gateway syncs its active set to the backend (`/api/v1/internal/bans/sync`, `internal/banview`) every 30 s and applies operator unbans. | owner ad-hoc | **done** (code + test contour; ban on in prod via Stage 18 — machinery built, cutover pending DNS) |
| CM | Channel-chat moderation + promo bot: a second standalone bot in the bot container answers `/start` with a localized message + a **URL** button into the **main** bot's Mini App (`?startapp`; a `web_app` button would sign initData with the promo token, which the main validator rejects). The **main** bot gates write access in a channel's linked discussion chat. The chat **allows sending by default** and the bot only restricts (Telegram intersects the chat default with the per-user permission, so a per-user grant cannot exceed a deny-by-default group): it **mutes** a member who is not registered or is admin-suspended or holding a new **`chat_muted`** role, and **un-mutes** an eligible one it had muted, for a member currently in the chat (a `getChatMember` guard, since bots cannot list members). Eligibility = `registered AND NOT suspended AND NOT chat_muted` (the game suspension dominates), resolved once in the backend and reached two ways: the bot's `ResolveChatEligibility` on a `chat_member` event over the existing mTLS bot-link, and a backend `chat_access_changed` event → gateway → `ChatGate` command (emitted on block/unblock, a `chat_muted` change, a first registration, or a temporary-block expiry via a sweeper; idempotent). No schema change — `chat_muted` reuses `account_roles`. | owner ad-hoc | **done** |
| → | Stage 18 — prod contour deploy | — | see [`PLAN.md`](PLAN.md) |
+3 -1
View File
@@ -33,7 +33,9 @@ COPY backend ./backend
# Reduce the workspace to what the backend needs: backend + pkg. loadtest and the
# gateway replace it requires are not in this context, so drop both.
RUN go work edit -dropuse=./gateway -dropuse=./platform/telegram -dropuse=./loadtest -dropreplace=scrabble/gateway@v0.0.0
RUN CGO_ENABLED=0 GOOS=linux go build -trimpath -o /out/backend ./backend/cmd/backend
# VERSION (the deploy passes the git tag) is stamped into the binary via the linker.
ARG VERSION=dev
RUN CGO_ENABLED=0 GOOS=linux go build -trimpath -ldflags "-X scrabble/pkg/version.Version=${VERSION}" -o /out/backend ./backend/cmd/backend
# --- runtime -----------------------------------------------------------------
FROM gcr.io/distroless/static-debian12:nonroot
+69 -3
View File
@@ -17,11 +17,12 @@ operational reference for **every environment variable**.
| `backend` | built (`backend/Dockerfile`) | Domain service; bakes in the DAWG dictionaries; runs migrations at boot. |
| `postgres` | `postgres:17-alpine` | Database (named volume, `pg_isready` healthcheck). |
| `validator` | built (`platform/telegram/Dockerfile`, target `validator`) | Telegram HMAC validator (no VPN, no Bot API); internal gRPC at `validator:9091`. Game login depends only on this. |
| `vpn` + `bot` | sidecar + built (`platform/telegram/Dockerfile`, target `bot`) | Telegram bot; egresses through the AmneziaWG sidecar; holds no inbound port — dials the gateway bot-link (mTLS) at `gateway:9443`. |
| `vpn` + `bot` | sidecar + built (`platform/telegram/Dockerfile`, target `bot`) | Telegram bot, gated to the **`telegram-local`** profile; egresses through the AmneziaWG sidecar and dials the gateway bot-link (mTLS) at `gateway:9443`. The test contour activates the profile; the prod **main** host omits it and runs the bot standalone on its **own host** (`docker-compose.bot.yml`, no VPN — native Bot API egress). |
| `otelcol` | `otel/opentelemetry-collector-contrib` | OTLP/gRPC `:4317` → Prometheus scrape (`:9464`) + Tempo. |
| `prometheus` | `prom/prometheus` | Metrics, 15d retention. |
| `prometheus` | `prom/prometheus` | Metrics, 15d retention (7d in prod). |
| `tempo` | `grafana/tempo` | Traces, 72h retention. |
| `grafana` | `grafana/grafana` | Dashboards (provisioned), anonymous-admin behind caddy's `/_gm/grafana`. |
| `node_exporter` | `quay.io/prometheus/node-exporter` | Host CPU/memory/disk metrics (Prometheus job `node`); the OOM signal on the tight prod main host (2 vCPU / 1.9 GiB). |
Networking: inter-service traffic is on the private `internal` network
(project-scoped DNS); only `caddy` joins the shared external `edge` network so the
@@ -59,7 +60,6 @@ compose binds from this directory.
| Variable | Gitea kind | Purpose |
| --- | --- | --- |
| `POSTGRES_PASSWORD` | secret | Postgres password (also embedded in `BACKEND_POSTGRES_DSN`). |
| `AWG_CONF` | secret | AmneziaWG config for the VPN sidecar (the bot's only Telegram egress in the test contour). **Must not contain a `DNS=` line** — it hijacks the shared netns's resolv.conf and breaks the bot resolving `otelcol` / `gateway`. Without it, Docker's resolver handles `otelcol`, `gateway` and `api.telegram.org`. |
| `GM_BASICAUTH_HASH` | secret | bcrypt hash gating `/_gm` (admin console + Grafana). Generate with `docker run --rm caddy:2-alpine caddy hash-password --plaintext '<pw>'`. |
| `TELEGRAM_MINIAPP_URL` | variable | The Mini App URL the bot hands out in deep links / buttons. |
@@ -67,6 +67,13 @@ compose binds from this directory.
secret) and the bot (Bot API). It defaults to empty in compose, but both **fail at
boot** when it is empty.
**Conditionally — `AWG_CONF`** (secret): the AmneziaWG config for the VPN sidecar, needed
only when the `telegram-local` profile runs (the test contour and local runs with the
bot). It is **not** `:?`-guarded — compose interpolates profiled-out services too, so the
prod main host (no VPN) must not require it. It **must not contain a `DNS=` line** — that
hijacks the shared netns's resolv.conf and breaks the bot resolving `otelcol` / `gateway`;
without it Docker's resolver handles `otelcol`, `gateway` and `api.telegram.org`.
## Optional variables (with defaults)
| Variable | Gitea kind | Default | Purpose |
@@ -110,6 +117,65 @@ collector's / gateway's internal IP is fine (connected route), but its `AWG_CONF
which resolves `otelcol`, `gateway` and `api.telegram.org`. `GATEWAY_ADMIN_*` is
intentionally **unset** — caddy owns `/_gm` in the contour.
## Production rollout
Prod runs on **two hosts** (main = full stack + ACME on the domain; tg = the bot only,
native Bot API, no VPN), one-time provisioned by **[`ansible/`](ansible/)** (docker, a
non-sudo `deploy` user holding the CI key, key-only sshd, default-deny ufw, fail2ban).
Re-run `ansible/` after a host resize — it is idempotent.
**To roll out:** merge `development → master` (CI green), then run the **`prod-deploy`**
workflow manually (Gitea → Actions → prod-deploy → run from `master`, input
`confirm=deploy`). It builds + pushes the images to the registry, ships the
compose/config/certs/env over SSH, deploys the main host with `prod-deploy.sh` (rolling,
health-gated, **auto-rollback to the previous tag**), then the bot host, then probes the
public site. After `master` is green this workflow is the **only** thing that touches
prod — nothing auto-deploys there. It runs four visible jobs: **build → deploy-main →
deploy-bot → verify** (the per-service rolling shows in the deploy-main log).
**Versioning.** Each release is a git tag `vX.Y.Z` on `master`; the deploy stamps
`git describe --tags` into every image tag, every binary (`-ldflags``pkg/version`
the `service.version` telemetry attribute) and the SPA About screen. Tag the release
before running the deploy:
```sh
git tag -a v1.0.0 -m v1.0.0 && git push origin v1.0.0
```
**Manual rollback** (any time after a successful deploy). Run the **`prod-rollback`**
workflow (Gitea → Actions → prod-rollback, `confirm=rollback`). Leave `target_version`
blank to roll back to the previously deployed version (read from the host's
`PREVIOUS_TAG`), or set it to a release tag from the **Releases** page. It re-deploys
that already-published image rolling + health-gated — no rebuild, no DB migration
(image rollback is DB-safe under the expand-contract rule). The registry keeps every
release tag, so any prior release is reachable.
**Migrations** must be **expand-contract** (backward-compatible; goose is forward-only):
the automatic rollback is image-only and never restores the DB. A deploy that changes
`backend/internal/postgres/migrations/` opens a maintenance window — the backend (sole
writer) is stopped for a consistent `pg_dump` into `/opt/scrabble/dumps` before the new
backend migrates. **Manual DB restore** (only if a migration was destructive):
`docker exec -i scrabble-postgres psql -U scrabble -d scrabble -c 'DROP SCHEMA backend CASCADE'`,
then pipe the dump into the same `psql`, and redeploy the matching old tag.
**bot-link cert rotation:** regenerate (`deploy/gen-certs.sh /tmp/c --force`), reset the
five `PROD_BOTLINK_*` secrets from `/tmp/c`, and re-run the workflow — both hosts redeploy
together with the fresh CA.
**Sizing / monitoring:** the main host launches undersized (2 vCPU / 1.9 GiB); the prod
overlay trims limits + `GOMAXPROCS=2` + 7d Prometheus retention, and `node_exporter` feeds
host memory to Grafana (`/_gm/grafana/`). Watch host memory and resize at Selectel when
players arrive.
**`PROD_` Gitea set** (mirrors `TEST_`, mapped onto the unprefixed names above) — secrets:
`PROD_{POSTGRES_PASSWORD, GM_BASICAUTH_HASH, GRAFANA_ADMIN_PASSWORD, TELEGRAM_BOT_TOKEN,
TELEGRAM_PROMO_BOT_TOKEN, REGISTRY_PASSWORD, SSH_KEY, SSH_KNOWN_HOSTS, BOTLINK_CA,
BOTLINK_GATEWAY_CERT, BOTLINK_GATEWAY_KEY, BOTLINK_BOT_CERT, BOTLINK_BOT_KEY}`; variables:
`PROD_{REGISTRY_USER, MAIN_HOST, TG_HOST, CADDY_SITE_ADDRESS, GM_BASICAUTH_USER,
GRAFANA_ROOT_URL, LOG_LEVEL, DICT_VERSION, TELEGRAM_MINIAPP_URL, TELEGRAM_GAME_CHANNEL_ID,
TELEGRAM_CHAT_ID, TELEGRAM_BOT_USERNAME, VITE_TELEGRAM_BOT_ID, VITE_TELEGRAM_LINK,
VITE_TELEGRAM_GAME_CHANNEL_NAME}`.
## Host-side setup (outside this repo)
- **`edge` network** must exist on the host (`docker network create edge`).
+49
View File
@@ -0,0 +1,49 @@
# Prod host provisioning (Stage 18)
Idempotent Ansible that prepares the two production hosts. It installs Docker, a
non-sudo `deploy` service account, SSH hardening, a default-deny firewall,
fail2ban, unattended security upgrades and time sync. It does **not** deploy the
application — that is `.gitea/workflows/prod-deploy.yaml`'s job, running as the
`deploy` account this playbook creates.
Hosts are referenced by `~/.ssh/config` aliases (`scrabble-main-ops`,
`scrabble-tg-ops`), so no IPs or key paths live in the repo.
## Prerequisites (controller)
- `ansible` with the bundled collections (`community.general`, `community.docker`,
`ansible.posix`).
- The two hosts reachable as root via the ssh-config aliases, host keys already
accepted into `known_hosts` (`host_key_checking = True`).
## One-time: the CI deploy key
The CI prod-deploy workflow logs into the hosts as `deploy` using a dedicated
key. Generate it once on the controller, authorize its public half via the
playbook, and store its private half **only** in the Gitea `PROD_SSH_KEY` secret:
```sh
ssh-keygen -t ed25519 -N '' -C scrabble-ci-deploy \
-f ~/.ssh/scrabble_ci_deploy_ed25519
# private half -> Gitea secret PROD_SSH_KEY (set via API); never commit it
```
## Run
```sh
cd deploy/ansible
ansible-playbook site.yml
```
The playbook reads the public key from `~/.ssh/scrabble_ci_deploy_ed25519.pub` by
default; override with `-e deploy_ci_pubkey_path=/path/to/key.pub`. Re-running is
safe (idempotent) and survives a host resize.
## What each host gets
- **both** (`common`): docker-ce + compose plugin, `daemon.json` (live-restore,
10m×3 log rotation), `deploy` user (docker group, no sudo), key-only sshd,
`ufw` default-deny incoming + allow SSH, fail2ban sshd jail, unattended
upgrades, chrony, `/opt/scrabble/{config,certs,dumps,images}`.
- **main**: `ufw` opens 80/443/9443; the external `edge` docker network.
- **tg**: verifies direct `api.telegram.org` egress (the no-VPN assumption).
+11
View File
@@ -0,0 +1,11 @@
[defaults]
inventory = inventory.ini
roles_path = roles
interpreter_python = /usr/bin/python3
host_key_checking = True
stdout_callback = yaml
deprecation_warnings = False
retry_files_enabled = False
[ssh_connection]
pipelining = True
+21
View File
@@ -0,0 +1,21 @@
---
# Service account the CI prod-deploy workflow uses to drive docker on the hosts.
# Membership in the docker group is root-equivalent (docker socket access), which
# is all the deploy workflow needs; the account is deliberately not given sudo.
deploy_user: deploy
# Public half of the dedicated CI deploy SSH key, read from the controller at run
# time. The private half is generated on the controller during provisioning and
# stored ONLY in the Gitea PROD_SSH_KEY secret; it is never committed. Override the
# path with -e deploy_ci_pubkey_path=/path/to/key.pub if the key lives elsewhere.
deploy_ci_pubkey_path: "{{ lookup('env', 'HOME') }}/.ssh/scrabble_ci_deploy_ed25519.pub"
deploy_ci_pubkey: "{{ lookup('file', deploy_ci_pubkey_path) }}"
# Base directory the deploy workflow rsyncs compose files, config, certs and dumps
# into. Owned by deploy_user so the workflow needs no elevation.
scrabble_base_dir: /opt/scrabble
# Docker daemon json-file log rotation, mirroring the compose x-logging anchor so
# the host's own containers (and any ad-hoc runs) rotate identically.
docker_log_max_size: "10m"
docker_log_max_file: "3"
+19
View File
@@ -0,0 +1,19 @@
# Production inventory for Stage 18.
#
# Hosts resolve through the operator's ~/.ssh/config aliases, so HostName (public
# IP), User and IdentityFile live there — no IPs or key paths are committed here.
# scrabble-main-ops -> main stack host (public IP, domain erudit-game.ru)
# scrabble-tg-ops -> Telegram bot host (direct Bot API egress, no VPN)
[main]
scrabble-main-ops
[tg]
scrabble-tg-ops
[prod:children]
main
tg
[prod:vars]
ansible_user=root
@@ -0,0 +1,15 @@
---
- name: restart docker
ansible.builtin.service:
name: docker
state: restarted
- name: reload sshd
ansible.builtin.service:
name: ssh
state: reloaded
- name: restart fail2ban
ansible.builtin.service:
name: fail2ban
state: restarted
+167
View File
@@ -0,0 +1,167 @@
---
# Common baseline applied to both prod hosts: Docker engine, a non-sudo deploy
# service account, SSH hardening, a default-deny firewall, fail2ban, unattended
# security upgrades and time sync. Every task is idempotent.
- name: Install base packages
ansible.builtin.apt:
name:
- ca-certificates
- curl
- gnupg
- ufw
- fail2ban
- unattended-upgrades
- chrony
state: present
update_cache: true
cache_valid_time: 3600
# --- Docker engine (official repo; trixie is published upstream) ---------------
- name: Create apt keyring directory
ansible.builtin.file:
path: /etc/apt/keyrings
state: directory
mode: "0755"
- name: Install Docker apt GPG key
ansible.builtin.get_url:
url: https://download.docker.com/linux/debian/gpg
dest: /etc/apt/keyrings/docker.asc
mode: "0644"
- name: Add Docker apt repository
ansible.builtin.apt_repository:
repo: >-
deb [arch=amd64 signed-by=/etc/apt/keyrings/docker.asc]
https://download.docker.com/linux/debian {{ ansible_distribution_release }} stable
filename: docker
state: present
- name: Install Docker engine and the compose plugin
ansible.builtin.apt:
name:
- docker-ce
- docker-ce-cli
- containerd.io
- docker-buildx-plugin
- docker-compose-plugin
state: present
update_cache: true
- name: Configure the Docker daemon (live-restore + log rotation)
ansible.builtin.template:
src: daemon.json.j2
dest: /etc/docker/daemon.json
mode: "0644"
notify: restart docker
- name: Enable and start Docker
ansible.builtin.service:
name: docker
enabled: true
state: started
# --- Deploy service account ----------------------------------------------------
- name: Create the deploy service account
ansible.builtin.user:
name: "{{ deploy_user }}"
groups: docker
append: true
shell: /bin/bash
create_home: true
- name: Ensure the deploy .ssh directory
ansible.builtin.file:
path: "/home/{{ deploy_user }}/.ssh"
state: directory
owner: "{{ deploy_user }}"
group: "{{ deploy_user }}"
mode: "0700"
- name: Authorize the CI deploy SSH key (exclusive)
ansible.builtin.copy:
dest: "/home/{{ deploy_user }}/.ssh/authorized_keys"
content: "{{ deploy_ci_pubkey }}\n"
owner: "{{ deploy_user }}"
group: "{{ deploy_user }}"
mode: "0600"
# --- SSH hardening -------------------------------------------------------------
- name: Harden sshd (key-only auth)
ansible.builtin.template:
src: sshd-hardening.conf.j2
dest: /etc/ssh/sshd_config.d/10-scrabble-hardening.conf
mode: "0644"
validate: sshd -t -f %s
notify: reload sshd
# --- Firewall (default deny incoming) ------------------------------------------
# SSH is allowed before the policy flips so enabling ufw never locks us out.
- name: Allow SSH through the firewall
community.general.ufw:
rule: allow
name: OpenSSH
- name: Default-deny incoming, allow outgoing
community.general.ufw:
direction: "{{ item.direction }}"
policy: "{{ item.policy }}"
loop:
- { direction: incoming, policy: deny }
- { direction: outgoing, policy: allow }
- name: Enable the firewall
community.general.ufw:
state: enabled
# --- fail2ban ------------------------------------------------------------------
- name: Configure the fail2ban sshd jail
ansible.builtin.template:
src: jail.local.j2
dest: /etc/fail2ban/jail.local
mode: "0644"
notify: restart fail2ban
- name: Enable and start fail2ban
ansible.builtin.service:
name: fail2ban
enabled: true
state: started
# --- Unattended security upgrades + time sync ----------------------------------
- name: Enable unattended upgrades
ansible.builtin.copy:
dest: /etc/apt/apt.conf.d/20auto-upgrades
mode: "0644"
content: |
APT::Periodic::Update-Package-Lists "1";
APT::Periodic::Unattended-Upgrade "1";
- name: Enable and start chrony
ansible.builtin.service:
name: chrony
enabled: true
state: started
# --- Deploy directories --------------------------------------------------------
- name: Create the scrabble base directories
ansible.builtin.file:
path: "{{ scrabble_base_dir }}/{{ item }}"
state: directory
owner: "{{ deploy_user }}"
group: "{{ deploy_user }}"
mode: "0750"
loop:
- ""
- config
- certs
- dumps
- images
@@ -0,0 +1,8 @@
{
"live-restore": true,
"log-driver": "json-file",
"log-opts": {
"max-size": "{{ docker_log_max_size }}",
"max-file": "{{ docker_log_max_file }}"
}
}
@@ -0,0 +1,9 @@
# Managed by Ansible (deploy/ansible).
[DEFAULT]
bantime = 1h
findtime = 10m
maxretry = 5
backend = systemd
[sshd]
enabled = true
@@ -0,0 +1,6 @@
# Managed by Ansible (deploy/ansible). Key-only authentication.
# root stays reachable by key (prohibit-password) for provisioning re-runs.
PasswordAuthentication no
PermitRootLogin prohibit-password
PubkeyAuthentication yes
KbdInteractiveAuthentication no
+18
View File
@@ -0,0 +1,18 @@
---
# Main stack host: public web + bot-link ports and the external 'edge' network
# the compose stack attaches caddy to.
- name: Open public web and bot-link ports
community.general.ufw:
rule: allow
port: "{{ item }}"
proto: tcp
loop:
- "80" # HTTP (ACME challenge + redirect to HTTPS)
- "443" # HTTPS (caddy edge)
- "9443" # bot-link mTLS (remote bot dials in; mutual TLS gates access)
- name: Ensure the external 'edge' docker network exists
community.docker.docker_network:
name: edge
state: present
+19
View File
@@ -0,0 +1,19 @@
---
# Telegram bot host: holds no inbound port beyond SSH (the bot dials out to the
# Bot API and into the main host's bot-link). We only verify direct Bot API
# egress here, since the "no VPN" decision depends on it.
- name: Verify direct Telegram Bot API egress (no VPN on this host)
ansible.builtin.uri:
url: https://api.telegram.org/
method: GET
status_code: [200, 301, 302, 401, 404] # any HTTP reply proves reachability
timeout: 10
register: tg_egress
failed_when: false
- name: Report Telegram reachability
ansible.builtin.debug:
msg: >-
api.telegram.org reachable:
{{ (tg_egress.status | default(0) | int) > 0 }} (status {{ tg_egress.status | default('none') }})
+31
View File
@@ -0,0 +1,31 @@
---
# Stage 18 host provisioning. Idempotent: safe to re-run after a host resize.
# Prepares hosts only (docker, hardening, service account, firewall); the
# application is deployed separately by .gitea/workflows/prod-deploy.yaml.
- name: Common baseline (both hosts)
hosts: prod
become: true
pre_tasks:
- name: Require a well-formed CI deploy public key
ansible.builtin.assert:
that:
- deploy_ci_pubkey | length > 0
- deploy_ci_pubkey is search('^(ssh|ecdsa)-')
fail_msg: >-
deploy_ci_pubkey is empty or malformed. Generate the key first
(see deploy/ansible/README.md) or override deploy_ci_pubkey_path.
roles:
- common
- name: Main stack host
hosts: main
become: true
roles:
- main
- name: Telegram bot host
hosts: tg
become: true
roles:
- tg
+53
View File
@@ -0,0 +1,53 @@
# Production Telegram bot host descriptor (standalone — NOT an overlay). Run only on
# the bot host:
# docker compose -f docker-compose.bot.yml up -d
#
# The bot egresses to the Bot API directly (no VPN sidecar) and dials the main host's
# published bot-link :9443 over mTLS. It exports no telemetry — otelcol lives on the
# main host and is unreachable from here — so observe it via `docker logs` on this host.
# Values come from the prod-deploy workflow (PROD_ secrets/variables); BOT_IMAGE is the
# pushed registry tag and BOTLINK_GATEWAY_ADDR is the main host's <ip>:9443.
name: scrabble-bot
services:
bot:
container_name: scrabble-telegram-bot
image: ${BOT_IMAGE:?set BOT_IMAGE to the registry tag}
restart: unless-stopped
logging:
driver: json-file
options:
max-size: "10m"
max-file: "3"
environment:
TELEGRAM_BOT_TOKEN: ${TELEGRAM_BOT_TOKEN:?set TELEGRAM_BOT_TOKEN}
TELEGRAM_GAME_CHANNEL_ID: ${TELEGRAM_GAME_CHANNEL_ID:-}
TELEGRAM_CHAT_ID: ${TELEGRAM_CHAT_ID:-}
TELEGRAM_PROMO_BOT_TOKEN: ${TELEGRAM_PROMO_BOT_TOKEN:-}
TELEGRAM_BOT_USERNAME: ${TELEGRAM_BOT_USERNAME:-}
TELEGRAM_BOT_LINK: ${TELEGRAM_BOT_LINK:-}
TELEGRAM_MINIAPP_URL: ${TELEGRAM_MINIAPP_URL:?set TELEGRAM_MINIAPP_URL}
# Real Bot API in prod (the test contour pins TELEGRAM_TEST_ENV=true instead).
TELEGRAM_TEST_ENV: "false"
TELEGRAM_API_BASE_URL: ${TELEGRAM_API_BASE_URL:-}
TELEGRAM_OWNS_UPDATES: "true"
# Dials the main host's published bot-link. ServerName stays `gateway` (the cert
# SAN), so TLS validation is independent of the dial address.
TELEGRAM_GATEWAY_ADDR: ${BOTLINK_GATEWAY_ADDR:?set BOTLINK_GATEWAY_ADDR (main:9443)}
TELEGRAM_BOTLINK_SERVER_NAME: gateway
TELEGRAM_BOTLINK_TLS_CERT: /certs/bot.crt
TELEGRAM_BOTLINK_TLS_KEY: /certs/bot.key
TELEGRAM_BOTLINK_TLS_CA: /certs/ca.crt
TELEGRAM_LOG_LEVEL: ${LOG_LEVEL:-info}
TELEGRAM_SERVICE_NAME: scrabble-telegram-bot
# No telemetry export: otelcol is on the main host, unreachable from here.
TELEGRAM_OTEL_TRACES_EXPORTER: none
TELEGRAM_OTEL_METRICS_EXPORTER: none
GOMAXPROCS: "1"
volumes:
- ${SCRABBLE_CONFIG_DIR:-.}/certs:/certs:ro
deploy:
resources:
limits:
cpus: "1.0"
memory: 256M
+98
View File
@@ -0,0 +1,98 @@
# Production main-host overlay, applied on top of docker-compose.yml on the main host:
# docker compose -f docker-compose.yml -f docker-compose.prod.yml up -d
#
# It (1) publishes caddy 80/443 — there is no host caddy in prod, so the contour caddy
# owns the edge and does its own ACME on CADDY_SITE_ADDRESS — and the gateway bot-link
# :9443 the remote bot dials in over mTLS; and (2) retunes the R7 limits down for the
# 2 vCPU / 1.9 GiB host (GOMAXPROCS=2, smaller memory caps, shorter Prometheus
# retention). The contour launches deliberately undersized at zero players; the added
# node_exporter + Grafana watch host memory so it can be resized at Selectel when
# traffic arrives.
#
# The bot + its VPN sidecar are absent here (the telegram-local profile is not
# activated); the prod bot runs on its own host from docker-compose.bot.yml.
services:
caddy:
ports:
- "80:80"
- "443:443"
deploy:
resources:
limits:
memory: 96M
gateway:
# Prod pulls the pushed image by tag instead of building locally; the base
# build: section stays dormant because the deploy always pulls first.
image: ${REGISTRY:?set REGISTRY}/scrabble-gateway:${TAG:?set TAG}
ports:
- "9443:9443"
environment:
# 2 vCPU host: align the Go scheduler with the cgroup quota (R7's 3 needs 3 cores).
GOMAXPROCS: "2"
deploy:
resources:
limits:
cpus: "2.0"
memory: 384M
backend:
image: ${REGISTRY:?set REGISTRY}/scrabble-backend:${TAG:?set TAG}
deploy:
resources:
limits:
memory: 384M
postgres:
deploy:
resources:
limits:
memory: 384M
validator:
image: ${REGISTRY:?set REGISTRY}/scrabble-telegram-validator:${TAG:?set TAG}
deploy:
resources:
limits:
memory: 96M
landing:
image: ${REGISTRY:?set REGISTRY}/scrabble-landing:${TAG:?set TAG}
deploy:
resources:
limits:
memory: 64M
otelcol:
deploy:
resources:
limits:
memory: 256M
prometheus:
command:
- --config.file=/etc/prometheus/prometheus.yml
- --storage.tsdb.retention.time=7d
deploy:
resources:
limits:
memory: 256M
tempo:
deploy:
resources:
limits:
memory: 384M
grafana:
deploy:
resources:
limits:
memory: 256M
postgres_exporter:
deploy:
resources:
limits:
memory: 64M
+38 -1
View File
@@ -71,6 +71,8 @@ services:
# Seed dictionary for a FRESH volume; the per-contour value comes from the
# deploy env (Gitea TEST_/PROD_DICT_VERSION). See the volume note below.
DICT_VERSION: ${DICT_VERSION:-v1.2.1}
# Build version stamped into the binary (git tag; see pkg/version).
VERSION: ${APP_VERSION:-dev}
restart: unless-stopped
logging: *default-logging
depends_on:
@@ -132,6 +134,8 @@ services:
VITE_TELEGRAM_GAME_CHANNEL_NAME: ${VITE_TELEGRAM_GAME_CHANNEL_NAME:-}
VITE_GATEWAY_URL: ${VITE_GATEWAY_URL:-}
VITE_APP_VERSION: ${APP_VERSION:-dev}
# Go binary version (the SPA's VITE_APP_VERSION is the same git tag).
VERSION: ${APP_VERSION:-dev}
restart: unless-stopped
logging: *default-logging
depends_on: [backend]
@@ -218,6 +222,8 @@ services:
context: ..
dockerfile: platform/telegram/Dockerfile
target: validator
args:
VERSION: ${APP_VERSION:-dev}
restart: unless-stopped
logging: *default-logging
environment:
@@ -240,14 +246,22 @@ services:
networks: [internal]
# --- Telegram bot (egress via the VPN sidecar in test; dials the gateway) ---
# vpn + bot are gated to the `telegram-local` profile: the test contour runs them
# locally (CI passes --profile telegram-local), the prod main host omits them, and
# the prod bot runs on its own host from deploy/docker-compose.bot.yml.
vpn:
container_name: scrabble-telegram-vpn
image: docker.iliadenisov.ru/developer/amneziawg-sidecar:latest
profiles: ["telegram-local"]
restart: unless-stopped
logging: *default-logging
privileged: true
environment:
AWG_CONF: ${AWG_CONF:?set AWG_CONF}
# Required by the vpn sidecar, which is gated to the telegram-local profile.
# Compose can't scope a `:?` guard to a profile (interpolation runs for
# profiled-out services too) and the prod main host has no VPN, so this is a soft
# default; the test contour always supplies TEST_AWG_CONF and the sidecar validates it.
AWG_CONF: ${AWG_CONF:-}
networks:
internal:
aliases: [telegram]
@@ -255,10 +269,13 @@ services:
bot:
container_name: scrabble-telegram-bot
image: scrabble-telegram-bot:latest
profiles: ["telegram-local"]
build:
context: ..
dockerfile: platform/telegram/Dockerfile
target: bot
args:
VERSION: ${APP_VERSION:-dev}
restart: unless-stopped
logging: *default-logging
depends_on: [vpn]
@@ -444,6 +461,26 @@ services:
memory: 128M
networks: [internal]
# node_exporter exports host CPU/memory/disk metrics. The prod main host runs a tight
# 1.9 GiB budget, so host memory pressure — not just per-container docker_stats — is
# what warns before an OOM. Prometheus scrapes it at :9100 (see prometheus.yml).
node_exporter:
container_name: scrabble-node-exporter
image: quay.io/prometheus/node-exporter:v1.8.2
restart: unless-stopped
logging: *default-logging
command:
- --path.rootfs=/host
- --collector.filesystem.mount-points-exclude=^/(sys|proc|dev|host)($|/)
pid: host
volumes:
- /:/host:ro,rslave
deploy:
resources:
limits:
memory: 64M
networks: [internal]
networks:
internal:
name: scrabble-internal
+143
View File
@@ -0,0 +1,143 @@
#!/usr/bin/env bash
# Production main-host deploy driver. Runs ON the main host, invoked over SSH by
# .gitea/workflows/prod-deploy.yaml as the deploy user (which must already be
# `docker login`ed to the registry). It pulls the images at the new tag and rolls
# the stack ONE service at a time in dependency order (least -> most dependent),
# health-checking after each; any failure rolls the whole stack back to the
# previously deployed tag.
#
# A schema migration adds a maintenance window: the backend (the only writer) is
# stopped so a consistent pg_dump is taken before the new backend migrates forward.
# Image rollback alone is safe under the expand-contract migration rule, so the
# automatic rollback never touches the database; the dump is kept for a MANUAL
# restore if a migration turned out to be destructive (see deploy/prod/README.md).
#
# Required env (exported by the workflow over SSH):
# REGISTRY registry namespace, e.g. docker.iliadenisov.ru/developer
# TAG new image tag (the deployed git SHA)
# PREV_TAG previously deployed tag, or "none" on the first deploy
# MIGRATION "1" when the deploy carries a schema migration, else "0"
# Optional: COMPOSE_DIR ENV_FILE DUMP_DIR STATE_FILE POSTGRES_USER POSTGRES_DB
set -uo pipefail
# Runtime compose vars (POSTGRES_*, GM_*, GRAFANA_*, CADDY_*, TELEGRAM_*, REGISTRY,
# SCRABBLE_CONFIG_DIR, ...) come from a shell-sourceable env file the workflow writes
# with single-quoted values. Exporting them into the process environment lets compose
# interpolate ${...} without re-parsing the value — a plain --env-file would mangle the
# literal '$' in the bcrypt GM_BASICAUTH_HASH.
ENV_FILE="${ENV_FILE:-/opt/scrabble/env.sh}"
# shellcheck disable=SC1090
[ -f "$ENV_FILE" ] && . "$ENV_FILE"
REGISTRY="${REGISTRY:?REGISTRY required (env.sh)}"
TAG="${TAG:?TAG required}"
PREV_TAG="${PREV_TAG:-none}"
MIGRATION="${MIGRATION:-0}"
COMPOSE_DIR="${COMPOSE_DIR:-/opt/scrabble/compose}"
DUMP_DIR="${DUMP_DIR:-/opt/scrabble/dumps}"
STATE_FILE="${STATE_FILE:-/opt/scrabble/DEPLOYED_TAG}"
# The prior deployed tag, preserved on every successful deploy so prod-rollback can
# target "the previous version" with no operator input.
PREV_STATE_FILE="${PREV_STATE_FILE:-/opt/scrabble/PREVIOUS_TAG}"
PG_USER="${POSTGRES_USER:-scrabble}"
PG_DB="${POSTGRES_DB:-scrabble}"
cd "$COMPOSE_DIR" || { echo "compose dir $COMPOSE_DIR missing"; exit 1; }
export REGISTRY
# otelcol joins the host docker group to read the socket; the GID varies per host.
DOCKER_GID="$(getent group docker | cut -d: -f3)"
export DOCKER_GID
dc() { docker compose -f docker-compose.yml -f docker-compose.prod.yml "$@"; }
use_tag() { export TAG="$1"; }
# --- health probes (one-off containers on the contour networks, like CI) --------
_probe() { docker run --rm --network "$1" alpine:3.20 wget -q -T 5 -O /dev/null "$2"; }
health_backend() { for _ in $(seq 1 20); do _probe scrabble-internal http://backend:8080/readyz && return 0; sleep 3; done; return 1; }
health_landing() { for _ in $(seq 1 20); do _probe scrabble-internal http://landing:80/ && return 0; sleep 3; done; return 1; }
health_postgres() { for _ in $(seq 1 30); do [ "$(docker inspect -f '{{.State.Health.Status}}' scrabble-postgres 2>/dev/null)" = healthy ] && return 0; sleep 2; done; return 1; }
health_running() { # health_running <container>: running, not restarting, stable restart count
local n="$1" s r c1 c2
for _ in $(seq 1 20); do
s="$(docker inspect -f '{{.State.Status}}' "$n" 2>/dev/null || echo missing)"
r="$(docker inspect -f '{{.State.Restarting}}' "$n" 2>/dev/null || echo true)"
if [ "$s" = running ] && [ "$r" = false ]; then
c1="$(docker inspect -f '{{.RestartCount}}' "$n")"; sleep 5
c2="$(docker inspect -f '{{.RestartCount}}' "$n")"
[ "$c1" = "$c2" ] && return 0
fi
sleep 3
done
return 1
}
roll() { # roll <service> <health-cmd...>
local svc="$1"; shift
echo ">>> rolling $svc -> $TAG"
dc up -d --no-build --no-deps "$svc" || return 1
"$@" || { echo "!!! $svc failed health check"; return 1; }
echo "<<< $svc healthy"
}
rollback() {
echo "########## ROLLBACK -> $PREV_TAG ##########"
if [ "$PREV_TAG" = none ]; then
echo "no previous tag (first deploy): cannot roll back; leaving the stack up for inspection."
return
fi
use_tag "$PREV_TAG"
dc up -d --no-build --remove-orphans
echo "rolled back to $PREV_TAG."
[ "$MIGRATION" = 1 ] && echo "NOTE: the DB is forward-migrated; a pre-deploy dump is in $DUMP_DIR — restore manually ONLY if the migration was destructive (see deploy/README.md, prod runbook)."
}
commit_tag() {
# Record the just-deployed tag as current, preserving the prior one as previous.
[ -f "$STATE_FILE" ] && cp "$STATE_FILE" "$PREV_STATE_FILE"
echo "$TAG" > "$STATE_FILE"
}
mkdir -p "$DUMP_DIR"
echo "=== prod deploy: tag=$TAG prev=$PREV_TAG migration=$MIGRATION ==="
use_tag "$TAG"
dc pull
# First deploy: nothing to roll from; bring the whole stack up and gate on health.
if [ -z "$(docker ps -aq -f name=scrabble-backend)" ]; then
echo "first deploy: bringing the whole stack up"
dc up -d --no-build --remove-orphans || { echo "compose up failed"; exit 1; }
health_backend || { echo "backend not ready"; exit 1; }
health_landing || { echo "landing not ready"; exit 1; }
commit_tag
echo "first deploy healthy ($TAG)."
exit 0
fi
# Migration deploy: freeze writes and snapshot a consistent dump before migrating.
if [ "$MIGRATION" = 1 ]; then
echo "migration deploy: opening maintenance window (stopping the backend = the only writer)"
dc stop backend
dump="$DUMP_DIR/pre-$TAG-$(date +%Y%m%d-%H%M%S).sql"
if ! docker exec scrabble-postgres pg_dump -U "$PG_USER" -d "$PG_DB" -n backend > "$dump"; then
echo "pg_dump failed; restarting the old backend and aborting"
dc start backend
exit 1
fi
echo "consistent dump: $dump"
fi
# Roll one service at a time, least -> most dependent; any failure rolls everything back.
roll postgres health_postgres || { rollback; exit 1; }
roll backend health_backend || { rollback; exit 1; }
roll gateway health_running scrabble-gateway || { rollback; exit 1; }
roll landing health_landing || { rollback; exit 1; }
roll validator health_running scrabble-telegram-validator || { rollback; exit 1; }
roll caddy health_running scrabble-caddy || { rollback; exit 1; }
# Observability + node_exporter: bring up the remainder and pick up any config changes.
dc up -d --no-build --remove-orphans || { rollback; exit 1; }
# Final internal sanity before committing the new tag.
health_backend || { rollback; exit 1; }
commit_tag
echo "=== deploy healthy ($TAG) ==="
+5
View File
@@ -18,3 +18,8 @@ scrape_configs:
- job_name: postgres_exporter
static_configs:
- targets: ["postgres_exporter:9187"]
# Host-level metrics (memory/CPU/disk). Matters most on the prod main host's tight
# 1.9 GiB budget, where total host memory is the OOM-proximity signal.
- job_name: node
static_configs:
- targets: ["node_exporter:9100"]
+34 -12
View File
@@ -1056,8 +1056,10 @@ plaintext relay (`GATEWAY_BOTLINK_RELAY_ADDR`) the backend admin console calls.
The full contour (`deploy/docker-compose.yml`) runs one `gateway`, one `backend`,
one Postgres, the static `landing`, the Telegram `validator` and `bot` (+ the bot's VPN
sidecar) and the **observability stack**
OTel Collector (OTLP/gRPC ingest → Prometheus metrics + Tempo traces) and Grafana
sidecar — the `bot`+`vpn` pair is gated to a `telegram-local` compose profile so the prod
main host can omit them) and the **observability stack**
OTel Collector (OTLP/gRPC ingest → Prometheus metrics + Tempo traces), a `node_exporter`
for host CPU/memory (the prod main host's OOM signal), and Grafana
with provisioned datasources and dashboards. All services export OTLP to the
collector; the bot shares the VPN sidecar's netns, so its `AWG_CONF` must not
carry a `DNS=` directive (that would hijack resolv.conf and stop it resolving
@@ -1081,16 +1083,36 @@ Two contours, two secret/variable prefixes (`TEST_` / `PROD_`):
generated by `deploy/gen-certs.sh` before `compose up`; the bot keeps its VPN sidecar
for Telegram egress and dials the gateway by its internal name, so the bot-link stays
on the internal network.
- **Prod**: a manual SSH deploy after `development → master`. There is no
host caddy, so the contour ships its own caddy terminating TLS — set
`CADDY_SITE_ADDRESS` to the domain and the caddy does its own ACME. The **bot runs
on a separate host** with native Telegram access (no VPN), deployed by SSH alongside
the main app (rolled together so the bot-link protocol versions never skew); the
gateway **publishes** the bot-link port and the certificates come from `PROD_`
secrets — a long-lived CA with leaves rotated by a scheduled job. The bot dials the
gateway's public bot-link endpoint and holds no inbound port; login is unaffected if
that host or the link is down. *(This prod wiring is the deferred final stage; the
code and the unified test contour land first — see `PRERELEASE.md`.)*
- **Prod**: a **manual** rollout — `.gitea/workflows/prod-deploy.yaml`, `workflow_dispatch`
only (from `master`, `confirm=deploy`), run after `development → master` is merged green.
It builds and pushes the images to the registry (`docker.iliadenisov.ru`), then deploys
over SSH onto **two hosts** provisioned by `deploy/ansible/` (docker, a non-sudo `deploy`
service account holding a dedicated CI key, key-only sshd, default-deny ufw, fail2ban):
the **main host** runs the full stack (`docker-compose.yml` + `docker-compose.prod.yml`),
the **bot host** runs only the bot (`docker-compose.bot.yml`, no VPN — native Bot API
egress, telemetry off). There is no host caddy, so the contour caddy terminates TLS —
`CADDY_SITE_ADDRESS` is the domain and caddy does its own ACME. The gateway **publishes**
the bot-link `:9443`; the remote bot dials it over mTLS (certs from `PROD_BOTLINK_*`,
ServerName `gateway`, so TLS validation is independent of the public dial address), holds
no inbound port, and login is unaffected if that host or the link is down.
`deploy/prod-deploy.sh` rolls the main stack **one service at a time in dependency order**
(postgres → backend → gateway → landing → validator → caddy), health-checking after each;
any failure **rolls the whole stack back to the previous image tag**. A **schema migration**
adds a maintenance window: the backend (the sole writer) is stopped for a consistent
`pg_dump` before the new backend migrates forward — image rollback stays DB-safe under the
expand-contract migration rule, and the dump is kept for a manual restore. The workflow runs
four visible jobs (build → deploy-main → deploy-bot → verify). Releases are git tags
`vX.Y.Z`; the version is stamped into the image tag, every binary (`-ldflags``pkg/version`
→ the `service.version` telemetry attribute) and the SPA About screen. A separate manual
**`prod-rollback`** workflow re-deploys any prior release tag (blank input = the previous
deployed version, tracked on the host) over the same rolling, health-gated path — image-only,
no DB migration. The main host is
intentionally **launch-sized** (2 vCPU / 1.9 GiB): the prod overlay trims the R7 limits
(`GOMAXPROCS=2`, smaller caps, 7d Prometheus retention) and a **node_exporter** feeds
host-memory metrics to Grafana so it can be resized reactively as players arrive.
`GATEWAY_ABUSE_BAN_ENABLED=true` in prod (the per-IP ban is meaningful only with real
client IPs). The `vpn`+`bot` pair is gated to a `telegram-local` compose profile the test
contour activates; the prod main host omits it.
## 14. CI & branches
+8 -1
View File
@@ -31,7 +31,10 @@ ephemeral guest. The gateway validates the credential once and mints a thin
session token; the backend resolves it to an internal `user_id`. A **Telegram Mini
App** launch authenticates from the platform's signed `initData`, themes the UI to
the Telegram colours, and — on first contact — seeds the new account's interface
language from the Telegram client. Telegram runs a **single bot**: every player uses
language from the Telegram client. If a launch cannot reach the backend (for example during a
deployment), the Mini App retries quietly and then shows a small "couldn't load" screen with a
**Retry** button, rather than dropping to the web sign-in, which has no place inside Telegram.
Telegram runs a **single bot**: every player uses
the same bot, and all of its chat and out-of-app notifications are written in the
player's own **interface language** (en/ru). A separate optional **promo bot** can run alongside the
main one — its only job is to answer `/start` with a short message and a button that opens the
@@ -56,6 +59,10 @@ reconnect), and pending reads resume on their own — the interface stays usable
flashing a red banner each time.
### Accounts, linking & merge
_Sign-in is currently provider-only, so the in-profile linking UI is temporarily hidden; it
returns once the anonymous `/app/` guest (whose upgrade path this is) ships. The flow below
describes it for when it does._
First platform contact auto-provisions a durable account. From the profile a player
links an email (via a confirm code) or their Telegram (via the web sign-in); a guest
who links their first identity becomes a durable account. The "already taken" status
+8 -1
View File
@@ -32,7 +32,10 @@ top-1 подсказку, безлимитную проверку слова с
session-токен; backend сопоставляет его с внутренним `user_id`. Запуск **Telegram
Mini App** авторизует по подписанным `initData` платформы, перекрашивает интерфейс
в цвета Telegram и — при первом контакте — задаёт язык интерфейса нового аккаунта по
языку Telegram-клиента. Telegram держит **единого бота**: все игроки пользуются одним
языку Telegram-клиента. Если запуск не может достучаться до бэкенда (например, во время
деплоя), Mini App тихо повторяет попытки, а затем показывает небольшой экран «не удалось
загрузить» с кнопкой **Повторить**, вместо того чтобы сбрасывать на веб-вход, которому внутри
Telegram не место. Telegram держит **единого бота**: все игроки пользуются одним
и тем же ботом, а весь его чат и внеприложенческие уведомления пишутся на **языке
интерфейса** самого игрока (en/ru). Рядом с основным может работать отдельный опциональный
**промо-бот** — его единственная задача отвечать на `/start` коротким сообщением и кнопкой,
@@ -57,6 +60,10 @@ Mini App** авторизует по подписанным `initData` плат
рабочим вместо красного баннера каждый раз.
### Аккаунты, привязка и слияние
_Вход сейчас только через провайдера, поэтому UI привязки в профиле временно скрыт; он
вернётся, когда появится анонимный `/app/`-гость (для апгрейда которого он и нужен). Описание
ниже — на этот случай._
Первый контакт с платформы заводит постоянный аккаунт. Из профиля игрок
привязывает email (по confirm-коду) или свой Telegram (через веб-вход); гость,
привязавший первую личность, становится постоянным аккаунтом. Факт «личность уже
+3 -1
View File
@@ -70,7 +70,9 @@ RUN rm gateway/internal/webui/dist/landing.html
# Reduce the workspace to what the gateway needs: gateway + pkg (loadtest is not in
# this context; its scrabble/gateway replace targets ./gateway, which is present here).
RUN go work edit -dropuse=./backend -dropuse=./platform/telegram -dropuse=./loadtest
RUN CGO_ENABLED=0 GOOS=linux go build -trimpath -o /out/gateway ./gateway/cmd/gateway
# VERSION (the deploy passes the git tag) is stamped into the binary via the linker.
ARG VERSION=dev
RUN CGO_ENABLED=0 GOOS=linux go build -trimpath -ldflags "-X scrabble/pkg/version.Version=${VERSION}" -o /out/gateway ./gateway/cmd/gateway
# --- runtime -----------------------------------------------------------------
FROM gcr.io/distroless/static-debian12:nonroot AS gateway
+13 -3
View File
@@ -29,6 +29,8 @@ import (
"go.opentelemetry.io/otel/sdk/resource"
sdktrace "go.opentelemetry.io/otel/sdk/trace"
"go.opentelemetry.io/otel/trace"
"scrabble/pkg/version"
)
// Exporter selectors supported per signal.
@@ -95,9 +97,7 @@ func New(ctx context.Context, cfg Config) (*Runtime, error) {
return nil, err
}
res, err := resource.New(ctx, resource.WithAttributes(
attribute.String("service.name", cfg.ServiceName),
))
res, err := serviceResource(ctx, cfg)
if err != nil {
return nil, fmt.Errorf("telemetry: build resource: %w", err)
}
@@ -122,6 +122,16 @@ func New(ctx context.Context, cfg Config) (*Runtime, error) {
return &Runtime{tracerProvider: tracerProvider, meterProvider: meterProvider}, nil
}
// serviceResource builds the OpenTelemetry resource describing this service: its
// service.name and the service.version stamped into the binary at build time
// (pkg/version, set from the git tag by the deploy).
func serviceResource(ctx context.Context, cfg Config) (*resource.Resource, error) {
return resource.New(ctx, resource.WithAttributes(
attribute.String("service.name", cfg.ServiceName),
attribute.String("service.version", version.Version),
))
}
// TracerProvider returns the runtime tracer provider, or the global one when r is
// not initialised.
func (r *Runtime) TracerProvider() trace.TracerProvider {
+21
View File
@@ -4,6 +4,8 @@ import (
"context"
"testing"
"time"
"scrabble/pkg/version"
)
// TestConfigValidate covers the supported and rejected exporter selections.
@@ -82,3 +84,22 @@ func TestNilRuntime(t *testing.T) {
t.Errorf("nil runtime Shutdown: %v", err)
}
}
// TestServiceResource checks the resource carries service.name and the embedded
// service.version (pkg/version, stamped at build time).
func TestServiceResource(t *testing.T) {
res, err := serviceResource(context.Background(), DefaultConfig("svc"))
if err != nil {
t.Fatalf("serviceResource: %v", err)
}
attrs := map[string]string{}
for _, kv := range res.Attributes() {
attrs[string(kv.Key)] = kv.Value.AsString()
}
if attrs["service.name"] != "svc" {
t.Errorf("service.name = %q, want svc", attrs["service.name"])
}
if attrs["service.version"] != version.Version {
t.Errorf("service.version = %q, want %q", attrs["service.version"], version.Version)
}
}
+10
View File
@@ -0,0 +1,10 @@
// Package version exposes the build version stamped into every Scrabble service
// binary. The default is "dev"; release builds override it through the linker
// (`go build -ldflags "-X scrabble/pkg/version.Version=<value>"`), wired from the
// VERSION build-arg in each service Dockerfile, which the deploy sets to the git
// tag (`git describe --tags`). It surfaces as the OpenTelemetry service.version
// resource attribute (see pkg/telemetry) and the SPA About screen.
package version
// Version is the build version, "dev" unless overridden at link time.
var Version = "dev"
+4 -2
View File
@@ -19,8 +19,10 @@ COPY platform/telegram ./platform/telegram
# Reduce the workspace to what the platform needs: only pkg + platform/telegram.
RUN go work edit -dropuse=./backend -dropuse=./gateway -dropuse=./loadtest -dropreplace=scrabble/gateway@v0.0.0 -dropreplace=scrabble-solver
RUN CGO_ENABLED=0 GOOS=linux go build -trimpath -o /out/validator ./platform/telegram/cmd/validator
RUN CGO_ENABLED=0 GOOS=linux go build -trimpath -o /out/bot ./platform/telegram/cmd/bot
# VERSION (the deploy passes the git tag) is stamped into both binaries via the linker.
ARG VERSION=dev
RUN CGO_ENABLED=0 GOOS=linux go build -trimpath -ldflags "-X scrabble/pkg/version.Version=${VERSION}" -o /out/validator ./platform/telegram/cmd/validator
RUN CGO_ENABLED=0 GOOS=linux go build -trimpath -ldflags "-X scrabble/pkg/version.Version=${VERSION}" -o /out/bot ./platform/telegram/cmd/bot
# --- validator (home) --------------------------------------------------------
FROM gcr.io/distroless/static-debian12:nonroot AS validator
+5 -2
View File
@@ -238,7 +238,10 @@ test('profile edit disables Save and flags an invalid display name', async ({ pa
await expect(save).toBeEnabled();
});
test('link account: a taken email opens the irreversible merge confirmation', async ({ page }) => {
// Account linking is hidden in Profile.svelte while we target provider sign-in (the anonymous
// /app/ guest who upgrades by linking comes later). The flow is kept wired; re-enable these two
// specs together with the `.emailbox` section.
test.skip('link account: a taken email opens the irreversible merge confirmation', async ({ page }) => {
await loginLobby(page);
await openProfile(page);
@@ -258,7 +261,7 @@ test('link account: a taken email opens the irreversible merge confirmation', as
await expect(page.getByText('Merge accounts?')).toBeHidden();
});
test('link account: the Telegram web sign-in control is offered in a browser', async ({ page }) => {
test.skip('link account: the Telegram web sign-in control is offered in a browser', async ({ page }) => {
await loginLobby(page);
await openProfile(page);
await expect(page.getByRole('button', { name: 'Link Telegram' })).toBeVisible();
+24
View File
@@ -83,6 +83,30 @@ test('tg-fullscreen header keeps a constant native-nav gap as the font scales',
expect(large.overflows).toBe(false);
});
test('inside Telegram, a failed launch shows the retry screen, not the web login', async ({ page }) => {
// initData carrying the mock's "bootfail" sentinel makes authTelegram reject, simulating a
// backend outage during launch (e.g. a deploy rolling). The Mini App must surface its own
// boot-error/retry screen and never fall back to the web (guest/email) login.
await page.addInitScript(() => {
Object.assign(window, {
Telegram: {
WebApp: {
initData: 'query_id=bootfail&user=%7B%22id%22%3A1%7D&auth_date=1&hash=deadbeef',
initDataUnsafe: {},
ready() {},
expand() {},
},
},
});
});
await page.goto('/');
// After the silent retries, the boot-error screen with its Retry button shows…
await expect(page.getByRole('button', { name: 'Retry' })).toBeVisible();
// …and the web login (guest) is never shown inside Telegram.
await expect(page.getByRole('button', { name: /guest/i })).toHaveCount(0);
});
test('outside Telegram, the /telegram/ entry redirects to the site root', async ({ page }) => {
await page.goto('/telegram/');
+6 -1
View File
@@ -18,6 +18,7 @@
import CommsHub from './game/CommsHub.svelte';
import Feedback from './screens/Feedback.svelte';
import Blocked from './screens/Blocked.svelte';
import BootError from './screens/BootError.svelte';
onMount(() => {
void bootstrap();
@@ -83,6 +84,10 @@
{#if !routeIsLobby}
<div class="splash">{t('common.loading')}</div>
{/if}
{:else if app.bootError}
<!-- A Mini App launch that failed to authenticate (e.g. the backend was down mid-deploy):
show the retry screen instead of falling back to the web login. -->
<BootError />
{:else if app.blocked}
<Blocked />
{:else}
@@ -123,7 +128,7 @@
<StaleInviteModal />
<WelcomeRedeemModal />
{#if routeIsLobby && !app.splashDone && !app.blocked}
{#if routeIsLobby && !app.splashDone && !app.blocked && !app.bootError}
<Splash />
{/if}
+58 -8
View File
@@ -18,6 +18,7 @@ import {
telegramDisableVerticalSwipes,
telegramHaptic,
telegramLaunch,
type TelegramLaunch,
telegramOnEvent,
telegramRequestFullscreen,
telegramSetChrome,
@@ -41,6 +42,10 @@ export interface Toast {
export const app = $state<{
ready: boolean;
/** Inside a Mini App, set when the launch failed to authenticate after its retries (e.g. the
* backend was down during a deploy). App.svelte then renders the boot-error retry screen
* instead of the web login — a Mini App has no manual sign-in to fall back to. */
bootError: boolean;
/** Whether the lobby's first cold load has settled (success or error). The loading splash
* (components/Splash.svelte) watches it to know when to dismiss; set by screens/Lobby. */
lobbyReady: boolean;
@@ -90,6 +95,7 @@ export const app = $state<{
resync: number;
}>({
ready: false,
bootError: false,
lobbyReady: false,
splashDone: false,
streamAlive: false,
@@ -563,14 +569,7 @@ export async function bootstrap(): Promise<void> {
// listener above then re-syncs the safe-area insets. Desktop keeps the bot's full-size
// window. No-op on clients predating Bot API 8.0.
telegramRequestFullscreen();
try {
await adoptSession(await gateway.authTelegram(launch.initData));
// A blocked account skips deep-link routing — the blocked screen overlays every route.
if (!app.blocked) await routeStartParam(launch.startParam);
} catch (err) {
handleError(err);
navigate('/login');
}
await bootTelegram(launch);
app.ready = true;
return;
}
@@ -585,6 +584,57 @@ export async function bootstrap(): Promise<void> {
app.ready = true;
}
// Inside a Mini App the only identity is the Telegram session, so a failed launch must never fall
// back to the web login screen. A transient backend outage (a deploy rolling over) is retried a
// few times in silence; only then does the boot-error screen surface, from which Retry re-runs the
// same path (retryTelegramBoot).
const TELEGRAM_BOOT_RETRIES = 2;
const TELEGRAM_BOOT_RETRY_MS = 1200;
function delay(ms: number): Promise<void> {
return new Promise((resolve) => setTimeout(resolve, ms));
}
/**
* bootTelegram authenticates a Mini App launch from its initData and routes any deep-link start
* parameter, retrying a few times on a transient failure before raising the boot-error screen
* (app.bootError). A blocked account is terminal — it switches straight to the blocked screen
* without retrying.
*/
async function bootTelegram(launch: TelegramLaunch): Promise<void> {
for (let attempt = 0; ; attempt++) {
try {
await adoptSession(await gateway.authTelegram(launch.initData));
// A blocked account skips deep-link routing — the blocked screen overlays every route.
if (!app.blocked) await routeStartParam(launch.startParam);
app.bootError = false;
return;
} catch (err) {
if (err instanceof GatewayError && err.code === 'account_blocked') {
await enterBlocked();
return;
}
if (attempt >= TELEGRAM_BOOT_RETRIES) {
app.bootError = true;
return;
}
await delay(TELEGRAM_BOOT_RETRY_MS);
}
}
}
/**
* retryTelegramBoot re-attempts the Mini App launch from the boot-error screen's Retry button. It
* clears the error and shows the loading state again, then runs the same retrying boot; on success
* the app renders normally, otherwise the boot-error screen returns.
*/
export async function retryTelegramBoot(): Promise<void> {
app.bootError = false;
app.ready = false;
await bootTelegram(telegramLaunch());
app.ready = true;
}
/**
* routeStartParam navigates a Telegram deep-link start parameter to its target: a
* specific game, the friends screen with a friend-code redemption, or the lobby
+3
View File
@@ -11,6 +11,9 @@ export const en = {
'blocked.temporary': 'Your account is blocked until {until}.',
'blocked.reason': 'Reason:',
'boot.errorTitle': "Couldn't load the game",
'boot.errorBody': 'Please try again in a moment.',
'common.back': 'Back',
'common.cancel': 'Cancel',
'common.ok': 'OK',
+3
View File
@@ -12,6 +12,9 @@ export const ru: Record<MessageKey, string> = {
'blocked.temporary': 'Ваша учётная запись заблокирована до {until}.',
'blocked.reason': 'Причина:',
'boot.errorTitle': 'Не удалось загрузить игру',
'boot.errorBody': 'Попробуйте ещё раз или зайдите позже.',
'common.back': 'Назад',
'common.cancel': 'Отмена',
'common.ok': 'ОК',
+4 -1
View File
@@ -136,7 +136,10 @@ export class MockGateway implements GatewayClient {
}
// --- auth ---
async authTelegram(): Promise<Session> {
async authTelegram(initData: string): Promise<Session> {
// e2e hook: an initData carrying this sentinel simulates a backend that rejects the launch,
// so the Mini App boot-failure path (silent retries → boot-error screen) can be exercised.
if (initData.includes('bootfail')) throw new GatewayError('unavailable');
return { ...SESSION, isGuest: false };
}
async authGuest(): Promise<Session> {
+63
View File
@@ -0,0 +1,63 @@
<script lang="ts">
// The Mini App launch failed to authenticate after its silent retries (e.g. the backend was
// briefly down during a deploy). Inside Telegram there is no web login to fall back to, so this
// terminal screen offers a manual Retry that re-runs the launch (app.svelte retryTelegramBoot).
import { retryTelegramBoot } from '../lib/app.svelte';
import { t } from '../lib/i18n/index.svelte';
let retrying = $state(false);
async function retry(): Promise<void> {
if (retrying) return;
retrying = true;
try {
await retryTelegramBoot();
} finally {
retrying = false;
}
}
</script>
<div class="boot">
<div class="card">
<h1>{t('boot.errorTitle')}</h1>
<p class="msg">{t('boot.errorBody')}</p>
<button class="retry" onclick={retry} disabled={retrying}>{t('common.retry')}</button>
</div>
</div>
<style>
.boot {
height: 100%;
display: grid;
place-items: center;
padding: 24px;
background: var(--bg);
}
.card {
max-width: 28rem;
display: flex;
flex-direction: column;
align-items: center;
gap: 1rem;
text-align: center;
color: var(--text);
}
h1 {
margin: 0;
font-size: 1.25rem;
}
.msg {
margin: 0;
color: var(--text-muted);
}
.retry {
padding: 9px 16px;
border: 1px solid var(--accent);
background: var(--accent);
color: var(--accent-text);
border-radius: var(--radius-sm);
}
.retry:disabled {
opacity: 0.5;
}
</style>
+4 -3
View File
@@ -244,9 +244,10 @@
</form>
{/if}
<!-- Linking & merge. Shown to everyone, including guests, who
upgrade by binding their first identity. -->
<section class="emailbox">
<!-- Linking & merge. Hidden for now: we target provider sign-in, and the anonymous
/app/ guest (whose upgrade path this is) comes later. Kept wired — drop `hidden`
to re-enable, together with the skipped linking specs in e2e/social.spec.ts. -->
<section class="emailbox" hidden>
<h3>{t('profile.linkAccount')}</h3>
{#if !emailSent}
<div class="addrow">