Monitoring & Observability
Operate the platform with correlated metrics, logs, and traces. The deployment supplies telemetry collection so services can use a consistent pipeline.
flowchart LR
S["Services"] --> C["OpenTelemetry Collector"]
S --> M["Prometheus metrics"]
C --> T["Tempo traces"]
C --> L["Loki logs"]
M --> G["Grafana"]
T --> G
L --> G
What to monitor
| Signal | Questions it answers | Alert examples |
|---|---|---|
| Availability | Can callers reach the service? | Failed health checks, high 5xx rate |
| Latency | Is a user flow slowing down? | Sustained high p95/p99 latency |
| Saturation | Is capacity close to a limit? | CPU, memory, connection, or queue thresholds |
| Event flow | Are events being processed and delivered? | Queue growth, sink failures, processing lag |
| Security edge | Are requests unexpectedly denied or failing validation? | Auth failures, policy denials, webhook verification failures |
Minimum dashboard set
- Gateway request rate, latency, and status codes.
- One service-health view per deployed service.
- RabbitMQ queue depth and consumer health for event-driven flows.
- Database health, connections, and backup status for each stateful service.
- Sink delivery success and lag for AuditFlow pipelines.
Use X-Request-ID and the trusted tenant context carefully to correlate an incident without exposing sensitive data in broad-access dashboards.