> For AI agents: the complete documentation index is available at /llms.txt, the full documentation bundle is available at /llms-full.txt.

# On-Call Guide — 被告警打扰时看这页

> **last-verified: 2026-07-14**,对照 [callytics-infrastructure PR #1459](https://github.com/retaintive/callytics-infrastructure/pull/1459) 之后的告警模型。**改 alarm / logger 语义的 infra PR 必须同步更新本页** — 这是防止本页再次腐烂的机制。
>
> 本页是薄路由层:告诉你"这条告警是什么、第一步做什么、详细步骤在哪"。深度调查手册按处境从 [operations 入口](/operations/index.md) 选。

## 1. 分级:什么会打扰你,什么不会

三层,按"人多久之内必须行动"分,不按代码觉得多严重分:

| 层                 | 什么信号                                                                                                              | 到哪                                                         | 你的响应                                                                 |
| ----------------- | ----------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------- | -------------------------------------------------------------------- |
| **Page(立即,可能半夜)** | **PROD** 症状型信号:DLQ 有消息、queue 积压超时限、webhook 进量归零、数据安全事件(`logger.alert`)                                            | 飞书 **PROD monitoring 群**,每张卡带 owner + 第一步动作 + runbook 链接   | 现在处理,见 §2                                                            |
| **Ticket(工作时间)**  | ① Sentry **新 issue 或复发(regression)**(不分环境);② **TEST/PRE 症状型 alarm 卡**(同一套 alarm,2026-07-15 起按环境路由,infra PR #1472) | 飞书(Sentry rule;TEST 卡进 **TEST monitoring 群**,PRE 卡按环境路由同理) | Sentry issue → §3;TEST/PRE alarm 卡 → 按 §2 对应告警的 runbook 处理,时限放宽到工作时间 |
| **Log(不通知任何人)**   | 单次可自愈失败、诊断型 alarm 状态(任何环境)                                                                                        | CloudWatch / Sentry / dashboard                            | 调查时自助查                                                               |

为什么这样分、每一级的判断例子:[Observability Guide(infra repo)](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/observability-guide.md)。

## 2. 你被 page 了 — 每条告警的含义和第一步

卡片上自带 runbook 链接,直接点它最快。下表是全量清单(按 alarm 名),供事后回看和演练:

| 告警(PROD)                                                                   | 含义                                                               | 第一步                                     | 详细步骤                                                                                                                                                                  |
| -------------------------------------------------------------------------- | ---------------------------------------------------------------- | --------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `*-transcribe-dlq-alarm`                                                   | 通话 webhook 处理重试耗尽,数据正在丢                                          | 看 DLQ 里失败消息的错误类型                        | [DLQ runbook](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/runbook-pipeline-monitoring.md#dlq)                                 |
| `*-ai-analysis-dlq-alarm`                                                  | AI 分析重试耗尽(SQS 5 次),该通话没有分析结果                                     | 同上;修因后 redrive                          | 同上                                                                                                                                                                    |
| `*-reconciliation-dlq-alarm`                                               | 补数 worker 失败堆积(阈值 5)                                             | 同上                                      | 同上                                                                                                                                                                    |
| `*-lead-processor-dlq-alarm`                                               | lead 处理失败,lead 会丢                                                | 同上                                      | 同上                                                                                                                                                                    |
| `*-contacts-batch-dlq-alarm`                                               | daily batch 分析失败堆积(注意:alarm 名叫 contacts-batch,盯的队列是 daily-batch) | 同上                                      | 同上                                                                                                                                                                    |
| `MessageSystem ComputeDLQDepthAlarm`                                       | SMS/消息管道 DLQ 有消息                                                 | 同上                                      | 同上                                                                                                                                                                    |
| `*-contacts-batch-queue-age-alarm`                                         | daily batch 队列最老消息等了 2h+,worker 没在消化                             | 查 consumer 并发/throttle 和 AI provider 健康 | [Queue age runbook](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/runbook-pipeline-monitoring.md#queue-age)                     |
| `*-reconciliation-queue-age-alarm`                                         | 补数队列积压 2h+(worker 卡住但没报错)                                        | 同上                                      | 同上                                                                                                                                                                    |
| `*-ai-analysis-queue-age-alarm`                                            | 实时分析管道积压 1h+(provider 挂/限流/worker 故障)                            | 同上                                      | 同上                                                                                                                                                                    |
| `*-webhook-ingestion-stopped`                                              | **连续 14 小时零通话进入系统** — RC subscription 失效或接收端挂了,上游全静默             | 查 RingCentral subscription 是否过期,立即续订    | [Ingestion runbook](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/runbook-pipeline-monitoring.md#webhook-ingestion)             |
| `storeId NULL rate / orphan store_id / contact_phone NULL`(`logger.alert`) | 数据身份解析出洞:写入了没归属的数据;orphan = 跨租户风险                                | 跑卡片 runbook 里的画像 SQL                    | [Store identity coverage](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/runbook-pipeline-monitoring.md#store-identity-coverage) |
| `DNC cascade rejected`(`logger.alert`)                                     | **合规风险**:客户说了别联系,但他的待办 outreach task 没关掉                         | 立即手动关闭该 contact 的 open tasks            | [DNC cascade rejected](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/runbook-pipeline-monitoring.md#dnc-cascade-rejected)       |

处理完两件收尾:

1. **把 DLQ 清回 0,让 alarm 回 OK** — 常驻 ALARM 的 DLQ 不会再为新消息报警,等于告警失效。
2. 够严重的(客户可见 / 丢数据 / 做过人工干预)写复盘,门槛和格式见 §6。

## 3. Ticket 级 — Sentry 新 issue / 复发

`logger.error` 全部进 Sentry 分组去重,只有**新问题**或**已解决又复发**才通知飞书一次 — 收到这类卡不用放下手头的事,但也别让它烂在频道里:

1. 点进 Sentry issue,看 stack trace + 哪次 release 引入。
2. 当场能修就修;不能就 assign 给 owner 并转成 issue(bd / GitHub),**在飞书 thread 里回一句去向** — 没人回的卡 = 没人认领。
3. Sentry 工作台:[https://personal-b6c.sentry.io](https://personal-b6c.sentry.io) (org 显示名 retaintive;主要项目:callytics-infrastructure / studio-web / studio-api)。

## 4. 5 分钟快速体检(按名字查,不按 dashboard 行号)

怀疑系统有问题但没收到告警时(需先 `aws sso login`;PROD 在 us-east-1,TEST 在 us-west-2):

```bash
# 1. 哪些 alarm 在响?(actions 为空的是诊断型,不通知人)
aws cloudwatch describe-alarms --region us-east-1 --state-value ALARM \
  --query 'MetricAlarms[].{name:AlarmName,actions:AlarmActions}'

# 2. DLQ 有没有积压?
for Q in $(aws sqs list-queues --region us-east-1 --query 'QueueUrls' --output text | tr '\t' '\n' | grep -i dlq); do
  echo "$(aws sqs get-queue-attributes --queue-url $Q --region us-east-1 \
    --attribute-names ApproximateNumberOfMessages \
    --query 'Attributes.ApproximateNumberOfMessages' --output text)\t$(basename $Q)"; done

# 3. 最近 1 小时 ERROR 集中在哪个 Lambda?(CloudWatch Logs Insights,console 跑更方便)
#    filter level='ERROR' | stats count(*) by @log
```

Dashboard(prod):[https://console.aws.amazon.com/cloudwatch/home?region=us-east-1#dashboards:name=call-analytics-prod-monitoring](https://console.aws.amazon.com/cloudwatch/home?region=us-east-1#dashboards:name=call-analytics-prod-monitoring)

## 5. 覆盖边界 — 什么情况**不会**有告警

知道盲区比背熟告警清单更重要:

- **absence monitoring(该发生的没发生)只有一条**:`webhook-ingestion-stopped`(通话进量归零)。其他"代码该跑没跑"类故障 — EventBridge schedule 被禁、某 Lambda 部署坏了但没流量触发 — **不会报警**,部署后要人工验证。
- **TEST / PRE 不 page(半夜不响),但会发卡**:症状型 alarm(DLQ / queue age)的卡实时进**各自环境的 monitoring 群**(TEST 卡进 TEST monitoring 群,PRE 卡按环境路由同理;2026-07-15 起,infra PR #1472),工作时间处理;daily digest(TEST+PROD 汇总/兜底)在计划中([infra #1468](https://github.com/retaintive/callytics-infrastructure/issues/1468))。
- **诊断型 alarm 不通知**:单 Lambda error 数、duration p95、AI provider 失败次数只留状态。它们变红不代表有人知道 — 靠 queue-age / DLQ 这些症状级告警兜底。
- 告警质量的持续跟踪:[infra issue #1460](https://github.com/retaintive/callytics-infrastructure/issues/1460)(coverage monitor 每小时重复告警 + NULL 来源调查)。

## 6. 复盘什么时候写

不是每个告警都复盘。满足任一才写(参考 [Google SRE 门槛](https://sre.google/sre-book/postmortem-culture/),阈值团队可调):客户可见的宕机/降级、任何数据丢失、on-call 做了人工干预(回滚 / redrive / 改流量)、解决耗时明显超预期、监控失灵(问题靠人肉发现)。

最小格式(一条飞书 thread + 一页文档):一句话摘要 / 影响范围 / 时间线(thread 本身就是)/ 根因与 contributing factors / **action items(每条有单一 owner + issue 编号 + due date)**。Blameless — 找系统性原因(runbook 不清、告警没响、测试缺失),不追究个人。

## 7. 升级路径 Escalation

**现状(2026-07)**:没有正式 on-call 轮值。默认路径 — 在告警卡的飞书 thread 里 @ 卡上的 owner;15 分钟无人响应,升级 @ Max。正式的 primary/backup 轮值、ACK 时限、假期安排是 infra roadmap Phase 2 的待办 — 定了之后更新本节。
