feat: harden orchestrator execution runtime

This commit is contained in:
john
2026-07-24 23:53:32 +08:00
parent 396bb78200
commit f6f2cd0933
55 changed files with 5122 additions and 194 deletions
+131
View File
@@ -0,0 +1,131 @@
# Workflow Orchestrator Phase 5 Runbook
## Scope
Phase 5 makes the LangGraph control plane operationally deployable without
making it the default executor. The default remains:
```text
memindadm executionEnabled=false
Portal MEMIND_ORCHESTRATOR_EXECUTION_HANDOFF_ENABLED=0
Orchestrator MEMIND_ORCHESTRATOR_EXECUTION_ENABLED=0
enabled executor list empty
Native Agent Run owns execution
```
An execution handoff requires all gates to agree: memindadm mode and explicit
execution switch, Portal environment gate, deterministic user/workflow rollout,
healthy Orchestrator storage, Orchestrator execution gate, enabled adapter,
authorization, workspace alias policy, quota admission, and a healthy worker.
## Local Colima canary
1. Copy `deploy/orchestrator/.env.example` to the ignored local `.env`.
2. Set distinct PostgreSQL, service, and worker secrets.
3. Keep execution disabled and start the durable control plane:
```bash
docker compose -f deploy/orchestrator/compose.yaml up -d --build
curl -fsS http://127.0.0.1:8093/ready
curl -fsS http://127.0.0.1:8093/metrics
```
4. Configure a workspace alias understood by the target Goosed instance.
5. Enable only Goosed for a single local user and start the worker profile:
```text
MEMIND_ORCHESTRATOR_EXECUTION_ENABLED=1
MEMIND_ORCHESTRATOR_ENABLED_EXECUTORS=goosed
MEMIND_ORCHESTRATOR_USER_ALLOWLIST=<local-user-id>
MEMIND_EXECUTOR_WORKSPACE_ALIASES_JSON={"canary":"/absolute/canary/path"}
```
```bash
docker compose -f deploy/orchestrator/compose.yaml --profile goosed-worker up -d
```
The Orchestrator `/ready` endpoint intentionally fails while execution is
enabled and no non-stale worker heartbeat exists.
## Rollout order
1. `off`: verify storage, metrics, backup, and worker registration.
2. `shadow`: collect Native versus LangGraph planning evidence.
3. `canary`, execution switch off: verify selection denominator and readiness.
4. `canary`, execution switch on: one explicit user, one workspace alias, zero
percentage rollout.
5. Increase percentage only after success rate, timeout, expired lease, queue
age, fallback, and Native rollback signals remain within the approved SLO.
6. `active` remains workflow-allowlisted and requires a separate release
approval.
Never enable Aider or OpenHands in the Orchestrator service until its dedicated
worker is registered and healthy. Use the example worker Compose overlay as a
template; each tool gets a separate image, resource limits, writable workspace
root, read-only container root, no Docker socket, dropped capabilities, and a
distinct worker identity.
## Immediate rollback
Use any one of these independent controls:
1. Set `MEMIND_ORCHESTRATOR_KILL_SWITCH=1` on Portal.
2. Clear the memindadm execution switch or set mode to `off`.
3. Set Portal `MEMIND_ORCHESTRATOR_EXECUTION_HANDOFF_ENABLED=0`.
4. Set Orchestrator `MEMIND_ORCHESTRATOR_EXECUTION_ENABLED=0`.
5. Remove an executor from `MEMIND_ORCHESTRATOR_ENABLED_EXECUTORS`.
6. Drain a worker by stopping it gracefully; the worker records `draining=true`.
Queued jobs stay durable. Running jobs stop receiving heartbeats, and lease
recovery moves them to `retryable` or `timed_out` according to attempt limits.
Native fallback remains a Portal decision; the Orchestrator never launches a
second Native run by itself.
## Monitoring
Scrape `/metrics` and alert on:
- `memind_orchestrator_executor_expired_leases > 0`;
- claimable jobs increasing while healthy workers are zero;
- repeated `retryable`, `failed`, or `timed_out` states;
- worker heartbeat age beyond 60 seconds;
- any execution-enabled interval without durable PostgreSQL readiness.
Executor events contain bounded metadata and artifact references. They must not
contain provider keys, absolute host paths, full stdout/stderr, or binary data.
## Backup and disaster recovery
Create and verify a PostgreSQL custom-format backup:
```bash
node scripts/orchestrator-dr.mjs backup --output /safe/path/orchestrator.dump
node scripts/orchestrator-dr.mjs verify --input /safe/path/orchestrator.dump
```
Restore only into a confirmed empty target:
```bash
node scripts/orchestrator-dr.mjs restore \
--input /safe/path/orchestrator.dump \
--database-url postgresql://... \
--confirm-empty-target
```
The restore command never uses `--clean` and refuses remote targets unless
`--allow-remote-target` is explicitly supplied. After restore, keep execution
disabled, start one worker, verify checkpoints/jobs/events/worker heartbeats,
then repeat the rollout order from `off`.
## Production release gate
This runbook does not authorize a production action. Before any server-side
deployment, follow `ENGINEERING_WORKFLOW_RULES.md`,
`PRODUCTION_RELEASE_RULES.md`, and run:
```bash
bash scripts/check-release-ready.sh
```
Production still requires a clean commit on complete `main`, passing CI, an
explicit deployment approval, a current backup, and a tested rollback.