> For AI agents: the complete documentation index is available at /llms.txt, the full documentation bundle is available at /llms-full.txt.

# Operations — 从这里选路

> **last-verified: 2026-08-04**(对照 callytics-infrastructure #2019/#2020/#2021 后的分层告警模型 + Sentry 看板体系)

## 每日一眼(直接点,不用找)

| 频率                  | 看什么                                                                                                                          | 直达链接                                                                                                                                                                                                                                                                                                                                                 |
| ------------------- | ---------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **每天扫一眼**           | PROD 系统健康 dashboard(DLQ / queue / Lambda errors / 业务指标 20+ 面板)                                                               | [CloudWatch PROD dashboard](https://us-east-1.console.aws.amazon.com/cloudwatch/home?region=us-east-1#dashboards:name=call-analytics-prod-monitoring-us-east-1)                                                                                                                                                                                      |
| **创始人,每天 1 分钟**     | **Morning Health 看板** — 生意有没有事:受影响用户 / 今天新出现的问题 / unhandled 三个数 + 哪个服务在冒烟;异常就点 Top 5 错误钻进去                                   | [Sentry · Retaintive Morning Health](https://personal-b6c.sentry.io/dashboard/9122425/)                                                                                                                                                                                                                                                              |
| **值班工程师,每天扫一眼**     | **Cross-Service Health 看板** — 三个核心服务健康度 + 按 release 错误量(发版有没有搞坏东西);默认 prod,顶部下拉切 test                                        | [Sentry · Cross-Service Health](https://personal-b6c.sentry.io/dashboard/489704/)                                                                                                                                                                                                                                                                    |
| 出事时                 | **Call Processing Pipeline 看板** — 管道哪一环坏了,按 Lambda 函数定位                                                                      | [Sentry · Call Processing Pipeline](https://personal-b6c.sentry.io/dashboard/489703/)                                                                                                                                                                                                                                                                |
| 前端负责人,每周            | **Studio User Experience 看板** — 前端报错 / 按 URL 分布 / API 错误                                                                     | [Sentry · Studio User Experience](https://personal-b6c.sentry.io/dashboard/489705/)                                                                                                                                                                                                                                                                  |
| **每天扫一眼**           | 用户反馈(Studio 前端 feedback)                                                                                                     | [Sentry User Feedback](https://personal-b6c.sentry.io/feedback/?project=4511714667003904)                                                                                                                                                                                                                                                            |
| 被动等                 | 飞书监控群(PROD 群 = page 立即行动;TEST 群 = 症状卡,工作时间;Sentry 侧 = 每项目 4 条 `[PROD]` 规则:新问题 / 回归 / 升级 / 影响用户激增,test 不进飞书)— 卡上自带 runbook 链接 | 飞书,无需主动看                                                                                                                                                                                                                                                                                                                                             |
| 查 AI 花费时            | 每次 AI 调用的成本明细                                                                                                                | [OpenRouter logs](https://openrouter.ai/logs)                                                                                                                                                                                                                                                                                                        |
| 部署验证时 / 每天顺带        | **TEST** 系统健康 dashboard + TEST 24h 新报错                                                                                       | [CloudWatch TEST dashboard](https://us-west-2.console.aws.amazon.com/cloudwatch/home?region=us-west-2#dashboards:name=call-analytics-test-monitoring-us-west-2) · [Sentry issues · TEST 24h](https://personal-b6c.sentry.io/issues/?project=4511714666938368\&project=4511714667003904\&project=4511714666938369\&environment=test\&statsPeriod=24h) |
| **每 1-2 周(告警检视仪式)** | 上周哪些告警响了、哪些没人动作 → 修剪                                                                                                         | 命令见 [On-Call Guide §4](/operations/on-call-guide.md) + [observability-guide §7(infra)](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/observability-guide.md)                                                                                                                                                   |
| **每季度**             | Sentry 配置有没有漂移(项目/规则/integration 对照清单)                                                                                       | [sentry-config.md(infra)](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/sentry-config.md)                                                                                                                                                                                                                      |

**怎么读 Morning Health**:三个大数字(受影响用户 / 今天新出现 / unhandled)任何一个异常升高,就点"Top 5 错误"或"今天新出现的问题"表进 issue 详情;"按项目分布"柱状图告诉你哪个服务在冒烟。时间范围由板顶日期选择器控制,默认 24h。已知局限:pipeline Lambda 暂未设置 user 上下文,"受影响用户"目前主要反映 studio 前后端(改进已列二期)。

## 按处境选文档

按你现在的处境走,不用记目录结构:

| 你现在的处境                             | 去哪                                                                                                                                                                |
| ---------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **被 page 了 / 收到飞书告警卡**             | [On-Call Guide](/operations/on-call-guide.md) — 每条告警的含义、第一步动作、升级路径                                                                                                |
| 客户问"我那通电话怎么了"(跨 Lambda trace 一通通话) | [Pipeline Monitoring runbook(infra repo,SoT)](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/runbook-pipeline-monitoring.md) |
| AI 分析失败 / 分析质量问题                   | [AI Analysis runbook](/operations/runbooks/ai-analysis.md)                                                                                                        |
| 查某通电话的 AI 成本 / OpenRouter 调用明细     | [OpenRouter Dashboard Lookup](/operations/runbooks/openrouter-dashboard-lookup.md)                                                                                |
| 通话数据缺漏 / 卡住,需要补数                   | [Reconciliation runbook](/operations/runbooks/reconciliation.md) · [Reprocess AI Analysis](/operations/reprocess-ai-analysis.md)                                  |
| 转录(录音 → 文字)问题                      | [Transcribe Processor runbook](/operations/runbooks/transcribe-processor.md)                                                                                      |
| 想理解"什么情况该 alert、log level 怎么选"     | [Observability Guide(infra repo)](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/observability-guide.md) — 分工比喻、7 个例子、决策树    |
| 容量 / 并发数为什么这么设                     | [Lambda Concurrency Decisions](/system-design/lambda-concurrency-decisions.md)(决策记录,不是操作手册)                                                                       |

**文档分工的一条规则**:告警的**处理步骤(alert response)只在 infra repo 维护一份** — 告警卡片上带的 runbook 链接就指向它,和代码同 repo、随 PR 原子更新、有 CI contract 强制。本站(docs.retaintive.ai)负责人读的入口(本页 + On-Call Guide)和深度调查手册(runbooks/)。发现两边内容重复时,以 infra repo 为准并回来修剪这边。
