think
16px
820px

DevOps Server Request — AHU AI Platform (non-GPU hosts)

Scope: production-grade hosts for the AHU AI platform's non-GPU zones. Excludes GPU box(es) — the current 2× H100 NVL box already exists and needs no new provisioning; the future 16× B200 cluster is a separate vendor/procurement track (data-center power/cooling, not a standard server order) and is requested separately when that phase starts.

Full rationale/topology: docs/deploy/2026-07-06-server-topology-and-specs.md (this doc is the actionable subset — the zones are explained there in depth).

Every host, regardless of tier

  • OS: Ubuntu 22.04 or 24.04 LTS
  • Time sync: NTP/chrony — mandatory, the audit hash-chain ordering and queue timers depend on consistent clocks across hosts
  • Access: SSH key-only (no passwords), reachable only via bastion/VPN — no direct SSH from the internet on any host, including DMZ
  • Runtime: Docker + Docker Compose v2
  • Firewall: default-deny, explicit allow rules only (matrix below)
  • Monitoring: disk-usage alerting at 70% (warning) / 85% (critical) — the GPU host has hit 88–94% full multiple times this project; don't let the new hosts repeat that
  • Patching: monthly patch window (not fully unattended — services here have uptime/HA expectations)
  • Network: each zone on its own private VLAN/subnet

MVP tier — 6 hosts + LB (no redundancy)

Host Zone / role vCPU RAM Disk Network
ahu-ctrl-01 Control plane: Gateway + Controller + Redis 8 16 GB 100 GB SSD Internal-only
ahu-gov-01 Governance: Observatory ingester + query API + dashboard + TimescaleDB 8–16 32–64 GB 100 GB OS + 1–2 TB NVMe, growable Internal-only + admin/VPN access for dashboard users
ahu-int-01 Internal engines: OCR + internal chatbot + RAG/vector store 16 32–64 GB 500 GB SSD Internal-only
ahu-dmz-01 DMZ: public chatbot only 4–8 16 GB 50–100 GB SSD Internet-facing (behind LB/WAF)
ahu-dwh-01 DWH compute: warehouse/lakehouse engine + ETL/CDC + Domain API + governance/audit 16–32 64–128 GB 200 GB SSD Internal-only, most isolated zone (no DMZ path)
ahu-dwh-store-01 DWH storage: operational metastore + lakehouse object storage 8–16 32–64 GB multi-TB, expandable (start ~2–4 TB) Internal-only
Load balancer + WAF (managed cloud LB/WAF preferred; one small dedicated host if on-prem-only) Public-facing, fronts ahu-dmz-01 only

Plus a backup target for the governance host's database (object storage or snapshot/PITR service) — not optional even at MVP; this is the audit-integrity store. The DWH hosts can share this same backup service (size capacity for both).

Note on the DWH specs: DWH isn't built yet — these numbers are directional (from the platform's readiness plan), not derived from a working system like the other zones. Treat as a placeholder devops can act on now; expect a refinement once DWH gets its own implementation spec (particularly the storage tier — its HA shape depends on which object-storage tech gets chosen).

HA tier — 15 hosts + LB + backup target (recommended for production)

Host Zone / role vCPU RAM Disk Notes
ahu-ctrl-01, ahu-ctrl-02 Control plane ×2 8 each 16 GB each 100 GB SSD each Active-active Gateway behind an internal VIP; Controller active-standby (Postgres advisory lock)
ahu-gov-app-01 Governance app compute: ingester + query API + dashboard 8 16–32 GB 100 GB SSD Single host is correct here — the ingester cannot yet run 2 replicas safely (known backlog item: needs a per-event advisory lock first). A 2nd app host would be wasted redundancy until that ships.
ahu-gov-db-01 (primary), ahu-gov-db-02 (standby) TimescaleDB primary/standby 8–16 each 32–64 GB each 1–2 TB NVMe, growable, each Streaming replication; primary takes the backup target
ahu-int-01, ahu-int-02 Internal engines ×2 16 each 32–64 GB each 500 GB SSD each OCR + internal chatbot + RAG, queue-distributed; add more identical hosts to scale OCR throughput independent of this baseline
ahu-dmz-01, ahu-dmz-02 DMZ ×2 behind public LB 4–8 each 16 GB each 50–100 GB SSD each Enables rolling deploys with zero downtime
ahu-dwh-meta-01 (primary), ahu-dwh-meta-02 (standby) DWH operational metastore 4–8 each 16–32 GB each 100 GB SSD each Critical — losing the metastore breaks the whole lakehouse; cheap to make redundant, worth doing even pre-build
ahu-dwh-gov-01 (primary), ahu-dwh-gov-02 (standby) DWH governance/audit (own hash-chained instance) 8–16 each 32–64 GB each growable, each Same audit pattern as the AI observatory, separate schema — DWH holds core registry data, arguably even more sensitive
ahu-dwh-compute-01 DWH warehouse/lakehouse compute + ETL/CDC + Domain API 16–32 64–128 GB 200 GB SSD Stateless/scalable pieces combined for now — add more identical hosts later to scale query or ETL throughput independently, once real load is known
ahu-dwh-store-01 DWH lakehouse object storage 8–16 32–64 GB multi-TB, expandable True storage-layer HA (replication/erasure coding) depends on the storage tech chosen at DWH build time — provision expandable capacity now, revisit redundancy shape then
LB + WAF Public-facing, fronts DMZ pair only
Backup target for TimescaleDB + DWH Object storage / PITR service, sized for both

Firewall / allowed-path matrix (give this to whoever writes the rules)

From To Port/service Notes
Internet LB/WAF 443 Only public entry point in the whole platform
LB/WAF DMZ host(s) app port
DMZ host(s) Control plane (Gateway) Gateway port (8200) Only egress path allowed from DMZ — nothing else, no exceptions
Internal-engines host(s) Control plane (Gateway) 8200 Model/OCR calls
Control plane (Gateway) GPU zone (model servers) model server ports Only inbound path to the GPU zone
Governance host(s) Control plane (Redis) Redis port Ingester consuming the audit stream (cross-zone, Redis port only)
DWH (ETL) Core AHU registry systems DB query port Outbound pull only, read-only credentials — DWH never writes back
Internal-engines (RAG) DWH (Domain API) API port "Live Data Sync from DWH" consumption
Admin/VPN Governance (dashboard) dashboard port Authorized ops/exec/auditor users only
Admin/VPN All hosts SSH Via bastion only; no direct internet SSH anywhere
DMZ / Internet DWH (any host) Explicitly denied — DWH is the platform's most isolated zone, no path from the public side, ever
Everything else Denied by default

Explicitly out of scope for this request

  • Current GPU box (2× H100 NVL) — already provisioned, no action needed.
  • Future 16× B200 cluster — separate procurement (vendor system, facilities power/cooling) — request when that phase starts.