Produccion
Alertas Activas
3
โ†‘ 2 vs. ayer
MTTA (Ack)
4.2m
โ†“ 38% vs. mes pasado
MTTR (Resolucion)
18m
โ†“ 55% vs. mes pasado
En Guardia Ahora
Marta
Rot. hasta dom 08:59
Alertas Activas
Prioridad Alerta Etiquetas Estado Asignado Tiempo
P1
PaymentService: Error rate > 5%
Tasa actual: 8.3% โ€” endpoints /api/charge, /api/refund
production payment critical Ack
MS
Marta Solis
hace 7 min
P1
Database CPU > 90% โ€” postgres-primary-eu
Media 5 min: 93.7% โ€” conexiones activas: 342/400
production database k8s-eu-1 Abierta
MS
Marta Solis
hace 2 min
P2
Pod CrashLoopBackOff โ€” payments-worker-8f9c4
Namespace: payments-prod โ€” Reinicios: 6 en 10 min
payment kubernetes production Abierta
TR
Tomas Ruiz
hace 15 min
P4
API Latencia p99 > 2s โ€” payment-gateway
p99 actual: 2.34s (threshold: 2s)
payment latency Cerrada
DV
Diego Vega
hace 1 h
Guardia On-Call โ€” Semana 24
Rotacion semanal
Lun ยท 09-18h
MS
Marta Solis HOY
Mar ยท 09-18h
MS
Marta Solis
Mie ยท 09-18h
TR
Tomas Ruiz
Jue ยท 09-18h
NC
Nina Chen
Vie ยท 09-18h
MS
Marta Solis
Sab ยท 24h
Full-team ยท P1 escalacion directa (30 min)
Dom ยท 24h
DV
Diego Vega (Payments)
Politica de Escalacion โ€” P1 Critico
Platform Engineering
1
Alerta recibida โ€” 0 min
Notificacion al guardia on-call
Push + llamada al ingeniero de turno del schedule "Platform Primary On-Call". Si no ack en 5 min, escala.
2
Si no ack โ€” +5 min
Escalacion al Team Lead
Llamada directa a Marta Solis (TL). Regla "if-not-acked". Incluye contexto y runbook.
3
Si no ack โ€” +15 min
Escalacion al equipo completo
Todos los miembros de Platform Engineering reciben alerta simultanea. Modo "all". Repeticion cada 10 min (max 3x).
Reglas de Enrutamiento
1
Alertas Criticas Produccion
priority == P1 tags โŠƒ production
โ†’
Platform Critical Esc.
Escalacion โ€” Platform Eng
2
Pagos y Transacciones
tags โŠƒ payment priority IN [P1,P2]
โ†’
Payments On-Call
Horario โ€” Payments API
3
Kubernetes y Base de Datos
tags โŠƒ kubernetes tags โŠƒ database
โ†’
Platform Primary OC
Horario โ€” Platform Eng
โ˜…
Resto de alertas (catch-all)
match-all
โ†’
Platform Eng (equipo)
Equipo directo
Equipos Configurados
Platform Engineering
MS
Marta Solis
marta@novapay.io
Admin
TR
Tomas Ruiz
tomas@novapay.io
User
NC
Nina Chen
nina@novapay.io
User
Payments API
DV
Diego Vega
diego@novapay.io
Admin
SM
Sara Moreno
sara@novapay.io
User
Integracion Prometheus Alertmanager
Activa
# alertmanager.yml โ€” Receptor Opsgenie para NovaPay receivers: - name: 'opsgenie-critical' opsgenie_configs: - api_key: '<OG_API_KEY>' message: '{{ .CommonLabels.alertname }}: {{ .CommonAnnotations.summary }}' description: '{{ .CommonAnnotations.description }}' priority: | {{ if eq .CommonLabels.severity "critical" }}P1 {{ else if eq .CommonLabels.severity "warning" }}P3 {{ else }}P5{{ end }} tags: '{{ .CommonLabels.environment }},{{ .CommonLabels.service }}' entity: '{{ .CommonLabels.service }}' source: 'prometheus' responders: - name: 'Platform Engineering' type: 'team' details: alertname: '{{ .CommonLabels.alertname }}' cluster: '{{ .CommonLabels.cluster }}' runbook: 'https://wiki.novapay.io/runbooks/{{ .CommonLabels.alertname }}' route: group_by: ['alertname', 'service', 'environment'] group_wait: 30s group_interval: 5m repeat_interval: 4h receiver: 'opsgenie-critical' routes: - match: { severity: critical } receiver: 'opsgenie-critical' repeat_interval: 30m # re-notifica cada 30 min si sigue abierta