Platform · scoped estate

Command Centre

Metrics, events, logs and traces for the selected tenant / client / environment — dig into evidence when something breaks. Read-only; no autonomous remediation.

Health

85

estate index

P1 open

2

escalated

Agents

4

2 high-risk

Read-only Agent OS
  • No shell execution
  • No cluster-admin
  • No secret reads
  • No database writes
  • No firewall changes
  • No autonomous remediation

Observability signals

Metrics · events · log pipelines · traces — unified fleet view
  • Nodes

    298

    infra

  • Clusters

    13

    infra

  • Agents

    30

    APM

  • Error events

    2

    events

  • P2

    1

    events

  • SLA risks

    3

    traces

  • Log drains

    4

    logs

  • Approvals

    3

    gates

Live telemetry

Fleet metrics & monitors

Streaming timeseries widgets and threshold monitors — Datadog-style dashboard layout, vendor-neutral simulated feed (1.5s ticks).

LIVEupdated now

CPU

44.4%

fleet avg

Latency p95

221ms

model gateway

Error rate

0.50%

production

Throughput

863rps

agent invoke

Agents busy

45%

active workers

CPU utilization

Agent hosts · last ~60s

Request latency

Gateway p95 · ms

Live event stream

Rolling ingest from collectors & agents

  • Agent sidecar scrape completed

    05:00:25

  • Agent sidecar scrape completed

    05:00:23

  • Retry budget consumed · tool invoke

    05:00:26

  • P1 signature match · kubelet not ready

    05:00:18

Error rate

Failed invokes / total

Throughput

Requests per second

Active monitors

Threshold checks on live series — monitors-as-code pattern (no vendor lock-in)

  • Fleet CPU anomaly

    ok

    avg(last_5m):cpu.utilization{scope:agents}

    44.4% · thr > 85%

  • Gateway latency p95

    ok

    avg(last_5m):gateway.latency.p95

    221ms · thr > 400ms

  • Error rate spike

    ok

    sum(last_5m):errors.rate{env:production}

    0.50% · thr > 2.5%

  • Log pipeline lag

    ok

    avg(last_5m):pipeline.lag.p95

    1.5s · thr > 3s

Telemetry pipeline

Collectors forward structured container logs from agent sidecars — ops pattern analogous to Fluent Bit DaemonSets tailing /var/log/containers/*.log.

  • Log collectors (DaemonSet)

    4/4 nodes

    healthy
  • Structured JSON parse

    cri-o · containerd

    healthy
  • Export endpoint

    EU residency

    healthy
  • Pipeline lag p95

    1.4s · live

    healthy
Open evidence / log viewer

Application latency

Model gateway request performance — live p95 overlay on seeded providers

OpenAI

187 ms live · US / EU routing

healthy

Anthropic

217 ms live · US / EU routing

healthy

Google Gemini

247 ms live · US

degraded

Azure OpenAI

276 ms live · EU (Sweden Central)

healthy

AWS Bedrock

306 ms live · EU (Frankfurt)

healthy

Ollama (on-prem)

336 ms live · On-premise

healthy

vLLM Cluster

365 ms live · On-premise (sovereign)

healthy
Open Model Gateway

Infrastructure health heatmap

Composite health by customer and environment — see everything in one place

Customerproductionstagingdevdr
FS Core Banking Platform94898468
Nordic Payments Rail99998479
Card Issuing Services99938883
Grid Telemetry Fabric88837873
SCADA Edge Estate85807570
Clinical Data Platform93888378
Imaging AI Workloads99988277
National Registry Services96918670

Error & incident timeline

Historical event volume (seeded) — live series above for last-minute fleet health

Token spend and retry waste

USD per day across all tenants