> For AI agents: the complete documentation index is available at /llms.txt, the full documentation bundle is available at /llms-full.txt.

# AI Analysis Processor Runbook

> **last-verified: 2026-07-14**（对照 infra PR #1459;发现和代码不符以代码为准）

**服务**: `ResultProcessorTS`（Lambda `call-analytics-{env}-ai-analysis-processor-ts-{region}`），读转录结果、跑 3 段 AI pipeline、写 DynamoDB + Neon。

这本 runbook 管**深度调查**:某通电话为什么分析失败、AI 输出质量为什么差、DLQ 里的 message 怎么补。**告警响应**（DLQ 有 message、queue 积压该怎么办）不在这里 —— 那是 [Pipeline Monitoring Runbook](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/runbook-pipeline-monitoring.md) 的事,本文只在需要时链过去。

***

## Table of Contents

1. [现状速览](#现状速览)
2. [处理流程](#处理流程)
3. [3 段 AI Pipeline](#3-段-ai-pipeline)
4. [验证与重试机制](#验证与重试机制)
5. [告警来自哪里](#告警来自哪里)
6. [Common Issues 深度排障](#common-issues-深度排障)
7. [CloudWatch Logs Insights 查询](#cloudwatch-logs-insights-查询)
8. [操作手册](#操作手册)
9. [Code Reference](#code-reference)

***

## 现状速览

一句话:单模型 `deepseek/deepseek-v4-flash` 经 OpenRouter.ai 调用,**没有 model fallback**,分 3 段跑,每段独立 Zod schema 校验。以前那套 Bedrock vs Cloudflare 双 provider 已完全废弃 —— 看到旧文档写 "根据 `ai_provider` 选 Bedrock/Cloudflare" 直接忽略。

| 属性                       | 值                                                        | 说明                                                                                    |
| ------------------------ | -------------------------------------------------------- | ------------------------------------------------------------------------------------- |
| **Runtime**              | Node.js 20                                               | TypeScript,hexagonal 架构                                                               |
| **Memory**               | 512 MB                                                   | AI workload                                                                           |
| **Timeout**              | **600s（10 分钟）**                                          | PR #974 从 900s 降到 600s(fail-fast + SQS redrive);更早文档写的 15min 已过时                      |
| **Reserved Concurrency** | **8**                                                    | PR #974 从 5 提到 8(maxConcurrency 5 + buffer 3,防 throttle race)                         |
| **AI Provider**          | OpenRouter.ai                                            | Vercel AI SDK + `@openrouter/ai-sdk-provider`;共享入口 `lambda/shared/utils/ai/invoke.ts` |
| **Model**                | `deepseek/deepseek-v4-flash`                             | 单模型,无 fallback。`withRetry` 只在同一模型上重试                                                  |
| **触发**                   | SQS `ai-analysis-queue`(EventBridge S3 事件喂进来)            | batchSize 1                                                                           |
| **DLQ**                  | `call-analytics-{env}-ai-analysis-dlq-{region}`          | `maxReceiveCount=5`,14 天保留                                                            |
| **SQS visibility**       | 3600s                                                    | AWS 推荐 6× Lambda timeout                                                              |
| **成本/用量**                | [https://openrouter.ai/logs](https://openrouter.ai/logs) | **不在 Lambda 内做 cost tracking**(PR #943 移除,重复且加延迟)                                     |

**Model 迁移史**(背景,无需操作): 从 Grok 4.1 via OneRouter 换到 DeepSeek V4 Flash via OpenRouter.ai(PR #939/#941/#953),原因是成本(\~5x 便宜)+ 高 cache 命中率。OneRouter 代码已全删。DDB config 表里存量 item 还带 `ai_model: 'moonshot.kimi-k2-thinking'` / `ai_provider: 'bedrock'` 是 dead data,没人读,别信。

**Provider routing**: OpenRouter 上层 provider 顺序 `['deepseek', 'alibaba', 'atlas-cloud']`(见 `lambda/shared/utils/ai/provider-routing.ts`),`allow_fallbacks: true`,p99 latency gate 60s。这些是 OpenRouter 侧的 upstream 选择,不是我们的 model fallback。

***

## 处理流程

**入口**: `lambda/ai-analysis-processor/src/handler.ts`。SQS 每条 message = 一通电话的一次分析,失败走 `batchItemFailures` → SQS redrive。

1. **解析事件**: 从 EventBridge-wrapped SQS message 提取 S3 key,还原 `transcriptionJobName`。
2. **查 Neon**: 用 session/recording 直接查 Neon call 记录(migration 期间保留 DDB `transcriptionJobName-index` GSI 做 fallback)。
3. **去重**: 检查 Neon 行的 `s3AnalysisPath`;已有 = 跳过,避免重复 AI 成本。另有一层 claim-based defense-in-depth 防并发重复(`[SKIP] Another Lambda owns this session` → 返回成功让 SQS 删掉重复消息)。
4. **下载转录**: 从 S3 拉转录,最短 `MIN_TRANSCRIPT_LENGTH = 10` 字符。太短 → 存空分析记录(不调 AI,零成本)直接返回成功。
5. **跑 pipeline**: `runPipeline()` 顺序跑 triage → classify → coaching(详见下一节)。
6. **合并 + 校验**: `mergePipelineResult()` 把 3 段结果拼成 legacy `CallAnalysis` 形状,过一遍 legacy schema 校验。
7. **持久化**: 写 S3 分析摘要 + DynamoDB + Neon(calls + contacts + timeline 原子 batch),更新 `s3AnalysisPath`。
8. **触发下游**: enqueue Contacts Analyzer(per-call SQS)。

任何步骤 throw → `batchItemFailures` → SQS 按 `maxReceiveCount=5` redrive,5 次后进 DLQ。

***

## 3 段 AI Pipeline

**入口**: `lambda/ai-analysis-processor/src/core/stages/pipeline.ts`。核心变化:不再是"一次 AI call 返回全部字段",而是**最多 3-4 段顺序 AI call**,每段有自己的 prompt、schema、超时和重试预算。这是排障时最需要先理解的结构 —— 一通电话失败,先问是**哪一段**失败。

| 段                       | 何时跑                       | 输出                                                                   | 每段超时 × 重试                     |
| ----------------------- | ------------------------- | -------------------------------------------------------------------- | ----------------------------- |
| **Stage 0: Pre-triage** | 总跑(代码规则,无 AI)             | voicemail/no-answer 等靠规则直接判,命中就短路返回                                  | 无 AI                          |
| **Stage 1: Triage**     | 总跑                        | `call_state` / `staff_name` / `worth_analyzing`                      | `[45s, 60s]` × 3 attempts     |
| **Stage 2: Classify**   | `worth_analyzing` 为真才跑    | `customer_type` / `category` / `subcategory` / `outcome` / `summary` | **`[90s, 90s]` × 2 attempts** |
| **Stage 2.5: Verify**   | 边界(易混)subcategory 才跑      | 置信度 gate 的重分类,**observe-only,不改数据**                                  | `[45s, 60s]` × 3 attempts     |
| **Stage 3: Coaching**   | human 对话 + 值得 coaching 才跑 | `practical_coaching`                                                 | `[45s, 60s]` × 3 attempts     |

**为什么 classify 的超时最长且重试最少**(`lambda/ai-analysis-processor/src/infrastructure/ai/invoke-model.ts` 的 `STAGE_CONFIG`): 2026-05-25 PR #997 实测 classify 是结构性慢的一段(长转录输入 + 复杂 discriminated-union schema),p99 latency 在多个 provider 上 55-57s,超过默认 45s 会频繁 abort。所以 classify 拉到 90s × 2 次;真卡过 90s 的第 3 次也救不回,交给 SQS redrive 换个新 Lambda 实例(可能路由到不同 upstream)。triage / coaching 成功率 99%+,不是瓶颈,保持 `[45s, 60s]` × 3。

**Stage 2.5 Verify 是 observe-only**(2026-05-30 team 决策): verify 会对易混 subcategory 独立重推理,但**结果不写入持久化记录、不影响下游**。它只 emit 一条 `verify_decision` 日志,用来统计 adopt/disagree/confirm 率,决定将来要不要让它真改数据。排障时看到 `stage_2_5_verify_review` 日志,那是观测数据不是错误。

**Worst-case 单次调用耗时**(对照 Lambda 600s timeout):

- triage: 45+60+60 + 2×2s backoff = 169s
- classify: 90+90 + 1×2s = 182s
- coaching: 45+60+60 + 2×2s = 169s
- 合计 ≈ 520s,留 80s 给 DB 写入 + SQS publish + Sentry flush。

***

## 验证与重试机制

**每段 AI 输出用 Zod schema 校验**,校验失败会把具体错误反馈给下一次 attempt 让模型自我修正(不是盲重试)。这是保留下来仍准确的核心机制,只是实现从旧的 `CallAnalysisSchema.parse()` + 手搓 retry loop 换成了 Vercel AI SDK + `withRetry`。

**为什么用 `mode: 'json'` 而不是严格 `json_schema`**(`lambda/shared/utils/ai/invoke.ts`): Vercel AI SDK 的 `mode: 'json'` 发 `response_format: { type: 'json_object' }`,schema 作为描述注入 prompt,由 Zod 在 client 侧校验。严格 `json_schema` 模式表达不了我们用的 `oneOf` / `discriminatedUnion` 形状(triage/classify/coaching 输出都靠它)。改成严格模式会直接 break 这些 schema。

**Validation-feedback 重试**(issue #1081,`invoke.ts` 的 `extractValidationFeedback` + `buildRetryUserMessage`): 校验失败时,把失败的字段和原因(最多 10 条 issue / 1500 字符)追加到下一次 attempt 的 user message 末尾(system prompt 保持 byte-identical 以维持 cache),让 attempt 2 精准自我修正。日志里 `validationFeedbackInjected: true` 标记这类重试。

**为什么需要**: schema 从 `@retaintive/common` 组合后收窄了几个 AI 可发的 enum(如 create 的 `typeCategory` 限于 `AI_ASSIGNABLE_TASK_TYPE_CATEGORIES`)。一个越界 enum 值就让整个响应在 Zod parse 时失败。低温度下盲重试大概率重复同样的错;把具体失败字段告诉模型才能自我纠正,而不是烧完重试预算掉进 SQS redrive/DLQ。

**哪些错误会重试**(`isRetryableError`,`invoke.ts`):

| 错误类型                                     | 重试?                               | 说明                                                                       |
| ---------------------------------------- | --------------------------------- | ------------------------------------------------------------------------ |
| `AbortError`                             | ✅                                 | 我们的 90s timeout 触发的 abort,天然是 transient                                  |
| `APICallError` 408/429/5xx               | ✅                                 | provider 侧 transient;4xx 非 429 = client 错(auth/quota/bad prompt),**不重试** |
| `NoObjectGeneratedError`                 | ✅(除非 `length` / `content-filter`) | 截断会复现、安全拒绝是确定性的,这两种不重试                                                   |
| `JSONParseError` / `TypeValidationError` | ✅                                 | 模型输出结构性坏,重试给个新样本                                                         |
| `ZodError`                               | ✅                                 | raw Zod 兜底(zod 4 里 `ZodError` 不再 extends `Error`,单独提前判断)                 |

**重试预算**: 默认 `retries: 1`(= 2 次 attempt)。per-stage 覆盖见 `STAGE_CONFIG`。SQS `maxReceiveCount=5` 负责 provider 级恢复(换新 Lambda 实例)。

***

## 告警来自哪里

**先记住一条,能省一半误判**: `logger.error` **不会** page 人。它只写 CloudWatch + Sentry。日志 severity 和打扰人的紧急度是两回事(见 `lambda/shared/utils/logger.ts`)。

| 方法                            | 去哪                             | 何时用                                 |
| ----------------------------- | ------------------------------ | ----------------------------------- |
| `logger.info` / `logger.warn` | CloudWatch only                | 常规 / 可自愈                            |
| `logger.error`                | CloudWatch + Sentry            | 代码错误,要留证据但**不打扰人**                  |
| `logger.alert`                | CloudWatch + Sentry **+ Lark** | 有 owner、影响、立即动作、runbook 的真 incident |

旧文档里"`logger.error` 会发 Discord/Lark/page 人"的说法**全错** —— Discord 路径已死,忽略所有 Discord 提法。

**告警卡从哪来(所有环境同构,按环境分群 — PROD 群 = 立即,TEST 群 = 工作时间;infra PR #1472)**: 不是靠 `logger.error`,而是靠 **CloudWatch 告警 → SNS → alarm-dispatcher Lambda → Lark + Sentry**(`lib/stacks/monitoring-stack.ts`)。跟本 Lambda 相关的 pageable 告警:

| 告警                                | 阈值                              | 触发含义                                       | 响应                                                                                                                                                                   |
| --------------------------------- | ------------------------------- | ------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **AiAnalysisDLQAlarm**            | DLQ ≥ 1 条 message               | AI 处理失败 5 次后进 DLQ                          | [Pipeline Runbook #dlq](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/runbook-pipeline-monitoring.md#dlq) → 定位失败类型后回本文对应 Issue |
| **AiAnalysisQueueAgeAlarm**       | 输入队列 oldest > 1h                | pipeline 卡住(provider 宕、throttle、worker 失败) | [Pipeline Runbook #queue-age](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/runbook-pipeline-monitoring.md#queue-age)          |
| **AiProviderFailureAlarm**        | `AIRequestFailure` ≥ 10 / 15min | OpenRouter 或模型本身降级                         | 看 [Issue #4: Provider 降级](#issue-4-provider-降级或-openrouter-故障)                                                                                                       |
| **AiAnalysisProcessorErrorAlarm** | Lambda errors ≥ 5 / 5min        | Lambda 层错误(诊断性,不直接 page)                   | 查 CloudWatch Logs 找 error pattern                                                                                                                                    |

`AiAnalysisProcessorErrorAlarm` 是诊断性的 —— PR #1459 明确把 ai-analysis 的 error/provider-failure 告警"de-action"了,理由是**最终队列健康才是客户影响的 page 信号**。所以 DLQ 和 queue-age 才是真正 pageable 的症状,error count 只是 dashboard/ticket 信号。

***

## Common Issues 深度排障

告警响应的通用步骤(先看 queue、修永久原因、redrive 后确认回 0)在 [Pipeline Runbook](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/runbook-pipeline-monitoring.md#dlq),这里讲**具体某类失败为什么发生、怎么定位根因**。

### Issue #1: DLQ 里堆积 message

**先做**: 按 [Pipeline Runbook #dlq](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/runbook-pipeline-monitoring.md#dlq) 确认失败类型和影响范围,再回来对号入座根因。

**看最近一小时的 error 分布**:

```
fields @timestamp, @message, franchiseId, storeId, telephonySessionId, stage
| filter @logGroup like /ai-analysis-processor/
| filter level = "ERROR"
| sort @timestamp desc
| limit 50
```

**看 error 类型 × 哪段失败**:

```
fields errorName as errorType, stage, analyzer
| filter @logGroup like /ai-analysis-processor/
| filter @message like /AIRequestFailure/ or level = "ERROR"
| stats count() by errorType, stage
| sort count desc
```

关注:同一 error 反复出现(10+ 次)、特定 store 持续失败、集中在某一 stage(尤其 classify)。

#### 根因 A: AI 调用超时(classify 最常见)

**日志特征**: `ai_attempt_aborted` 事件,`stage: "classify"`,`latencyMs` ≈ `timeoutMs`。

```
fields @timestamp, stage, timeoutMs, latencyMs, promptChars
| filter @logGroup like /ai-analysis-processor/
| filter @message like /ai_attempt_aborted/
| stats count() as aborts, avg(promptChars) as avgPromptChars by stage
| sort aborts desc
```

**为什么**: classify 段长转录 + 复杂 schema 结构性慢。`ai_attempt_aborted` 日志会 dump `inputPreview`(前 8KB)供 on-call 复现具体卡住的 payload。abort 是设计内的(我们自己的 timer 触发),单次 abort 不是 incident;**大量 abort** 才是 provider TTFB 降级信号。

**怎么办**:

1. 单通失败 → SQS 已自动 redrive,大概率换个 upstream 就过。不用手动干预。
2. 大面积 classify abort → 看 §Issue #4 判断是不是 provider 降级。
3. 若 `promptChars` 异常大(超长转录)→ 见 §Issue #3 转录长度分布。

#### 根因 B: Schema 校验失败(重试用尽)

**错误特征**: `NoObjectGeneratedError` / `TypeValidationError` / `schema_validation` failureType,重试后仍失败。

```
fields @timestamp, stage, validationFeedbackInjected, attempt
| filter @logGroup like /ai-analysis-processor/
| filter @message like /ai_attempt_start/ and validationFeedbackInjected = 1
| stats count() as retriesWithFeedback by stage, attempt
```

**为什么**: 模型反复吐同一个越界 enum / 错类型,即使 validation-feedback 注入也没纠正过来。常见触发:

- **越界 enum**: AI 发了个 taxonomy 里没有的 subcategory / typeCategory。`@retaintive/common` 收窄 AI-emittable enum 后(PR #1077),一个越界值就整条 fail。
- **类型错**: 该给 number 给了 string。
- **outcome 条件规则违反**: `membership_cancel` 只能 `cancelled`/`retained`/`pending_follow_up`;其他 subcategory 只能 `success`/`attempted`/`na`。

**注意 classify 的 subcategory 有降级机制**(`pipeline.ts` `deriveClassifyTopicFields`): classify 的 subcategory 是 plain string,遇到未知值**降级为 `other`** 而不是烧一次 schema-repair 重试。所以 subcategory 未知不一定进 DLQ —— 看日志里 `Classify emitted unknown subcategory, falling back to "other"` 的 warn。

**怎么办**:

1. 若是最近改了 schema/taxonomy → 检查 `@retaintive/common` 的 enum 是否和 prompt 里描述的一致。
2. 若某 store 转录有特殊 pattern 持续触发 → 收集样本转录 + 错误日志开 issue。
3. 校验规则本身要改 → 走 `@retaintive/common`,消费方同 PR 适配(高风险,过 review)。

#### 根因 C: S3 转录不存在

**错误特征**: `NoSuchKey` / 找不到转录 key。

```
fields @timestamp, telephonySessionId, error
| filter @logGroup like /ai-analysis-processor/
| filter error like /NoSuchKey/ or @message like /transcript from S3/
| stats count() by telephonySessionId
| sort count desc
```

**根因**:

1. TranscribeProcessor 成功但 S3 上传失败 → 查上游 [TranscribeProcessor Runbook](/operations/runbooks/transcribe-processor.md) 同 `telephonySessionId`。
2. Neon 行里的转录 S3 key 与实际路径不符。
3. S3 对象被删(retention / 手动)→ 查 S3 versioning + CloudTrail `DeleteObject`。

**怎么办**: 若录音还在,重跑转录(见 TranscribeProcessor runbook);若有备份转录,重新上传到对应 S3 路径会触发 EventBridge 重新分析。

#### 根因 D: Neon / DDB 写入失败

**错误特征**: 写 `s3AnalysisPath` / calls / contacts / timeline batch 时 throw。

**为什么**: Neon 连接失败 / DDB 限流 / batch 里某字段约束违反。AI 分析已完成但持久化挂了 → 整条重试(浪费已完成的 AI call,所以这类要尽快修)。

**怎么办**:

1. 查是不是 Neon 侧问题(连接数 / migration drift)。
2. 短暂故障 → SQS redrive 会自愈。
3. 系统性写入失败 → 见 [Pipeline Runbook](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/runbook-pipeline-monitoring.md) 场景 B。

***

### Issue #2: 校验重试率飙升(输出质量差)

**症状**: 大量 `validationFeedbackInjected: true` 的重试,即使能成功也拖慢处理(每次重试加一整段 AI 调用)。这是**输出质量**问题,不一定进 DLQ,但吃成本和延迟。

**定位失败字段 pattern**:

```
fields stage, validationFeedbackInjected
| filter @logGroup like /ai-analysis-processor/
| filter @message like /ai_attempt_start/
| stats count() as attempts, sum(validationFeedbackInjected) as feedbackRetries by stage
```

**看哪一段质量差**:

- 集中在 classify → subcategory/outcome/enum 类问题最常见。
- 集中在 coaching → 输出结构(`detailed_feedback_item` 嵌套)问题。

**常见根因**:

1. **Schema 改动后模型没跟上**: 最近改了 `@retaintive/common` taxonomy 或某段 prompt,模型还在按旧约束吐。查 prompt version(`ai_attempt_start` 日志的 `promptVersion`)是否和 schema 一致。
2. **Prompt 里 enum 描述和 schema 不一致**: `mode: 'json'` 靠 prompt 描述引导模型,描述过时就会持续越界。
3. **模型 drift**: OpenRouter 上游 provider 换了模型 quant 变体导致行为漂移 —— 看 `ai_usage.provider` 是不是变了。

**怎么办**:

1. 先确认 validation-feedback 机制在工作: 应看到 attempt 2 有 `validationFeedbackInjected: true`,且大部分随后 `ai_usage attemptOutcome: success`。
2. 若 feedback 注入了但 attempt 2 仍失败 → prompt/schema 有硬矛盾,不是模型能自我修正的,查 prompt ↔ schema 一致性。
3. 持续 >baseline → 收集样本开 issue,考虑改 prompt(prompt 改动至少更新 snapshot,涉及 task/lifecycle 补 golden eval)。

***

### Issue #3: 处理慢(Duration 飙升)

**症状**: Lambda duration 逼近 600s timeout;`ai-analysis-queue` 积压增长。

**先做**: queue 积压响应看 [Pipeline Runbook #queue-age](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/runbook-pipeline-monitoring.md#queue-age)。这里定位**为什么单通慢**。

**看每段 latency**:

```
fields stage, latencyMs
| filter @logGroup like /ai-analysis-processor/
| filter @message like /ai_usage/
| stats avg(latencyMs) as avgMs, pct(latencyMs, 99) as p99Ms, count() as calls by stage
| sort p99Ms desc
```

**正常基线**(2026-05 实测):

- triage / coaching: p99 2-14s
- classify: p99 40-57s(结构性慢,已是最慢一段)

**异常判读**:

- classify p99 逼近 90s 且 abort 增多 → provider TTFB 降级,见 §Issue #4。
- 多段都慢 → OpenRouter 整体降级或某 upstream 出问题。
- 单通极慢 + 转录超长 → 见下面转录长度。

**看转录长度分布**:

```
fields promptChars
| filter @logGroup like /ai-analysis-processor/
| filter @message like /ai_attempt_start/ and stage = "classify"
| stats avg(promptChars) as avgChars, pct(promptChars, 95) as p95Chars, max(promptChars) as maxChars
```

超长转录(>30 分钟通话)会拉高 classify TTFB。这是已知 trade-off,单通会 redrive 自愈;若某 store 系统性超长,收集样本讨论是否要截断策略。

***

### Issue #4: Provider 降级或 OpenRouter 故障

**症状**: `AiProviderFailureAlarm` 触发(`AIRequestFailure` ≥ 10 / 15min),或某类 failureType 突增。这不是我们的 model fallback 问题(没有 fallback)—— 是 OpenRouter 或模型本身降级。

**看 failure 类型分布**:

```
fields analyzer, @message
| filter @logGroup like /ai-analysis-processor/
| filter @message like /AIRequestFailure/
| stats count() as failures by analyzer
| sort failures desc
```

`AIRequestFailure` 的维度是 `FailureType`(`timeout` / `http_429` / `http_5xx` / `http_4xx` / `no_object` / `json_parse` / `schema_validation` / `unknown`)× `Analyzer`(pipeline stage)。判读:

- **`http_429` 突增** → OpenRouter 限流。等待自愈或联系 OpenRouter。
- **`http_5xx` / `timeout` 突增** → 上游 provider 宕。OpenRouter 的 `allow_fallbacks: true` 会自动切别的 upstream,通常几分钟自愈。
- **`schema_validation` 突增** → 见 §Issue #2(是输出质量,不是 provider)。

**看当前路由到哪个 upstream**:

```
fields provider, stage
| filter @logGroup like /ai-analysis-processor/
| filter @message like /ai_usage/
| stats count() as calls by provider, stage
| sort calls desc
```

`ai_usage.provider` 显示实际路由的上游(`DeepSeek` / `Alibaba` / `AtlasCloud` 等)。若某平时常用的 upstream 突然 0 calls,说明它被 p99 latency gate 挤下去了,routing 已自动降级到下一个 —— 这是**设计内的自愈**,不用干预。

**怎么办**:

1. 大部分 provider 降级靠 OpenRouter fallback + SQS redrive 自愈,**先观察不手动干预**。
2. 若 OpenRouter 整体故障(所有 upstream 都失败)→ pipeline 会积压,DLQ 会堆,走 [Pipeline Runbook](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/runbook-pipeline-monitoring.md#queue-age);根因确认(OpenRouter 恢复)后 redrive DLQ。
3. **没有手动切 provider 的操作** —— 旧文档里 `aws dynamodb update-item ... ai_provider` 那套已完全废弃,`ai_provider` 是 dead field。

***

## CloudWatch Logs Insights 查询

**Log Group**: `/aws/lambda/call-analytics-{env}-ai-analysis-processor-ts-{region}`

日志字段来自 `invoke.ts` 的结构化 log,关键事件:`ai_attempt_start` / `ai_attempt_aborted` / `ai_usage` / `verify_decision`。租户字段是 `franchiseId` / `accountId` / `storeId`(旧 `client_id` 已 deprecated)。

### Query 1: 某通电话完整 trace

```
fields @timestamp, stage, @message
| filter @logGroup like /ai-analysis-processor/
| filter telephonySessionId = "SESSION_ID"
| sort @timestamp asc
```

跨 stage 看这通电话走到哪一段、哪段失败。

### Query 2: 各段 latency 对比(24h)

```
fields stage, latencyMs
| filter @message like /ai_usage/
| stats avg(latencyMs) as avgMs, pct(latencyMs, 50) as p50, pct(latencyMs, 99) as p99, count() as calls by stage
```

### Query 3: 各段 abort 率(定位 timeout 段)

```
fields stage
| filter @message like /ai_attempt_aborted/
| stats count() as aborts by stage
| sort aborts desc
```

### Query 4: Provider 分布 + cache 命中(成本相关)

```
fields provider, stage, cacheReadRatio, costUsd
| filter @message like /ai_usage/
| stats avg(cacheReadRatio) as avgCacheHit, sum(costUsd) as totalCost, count() as calls by provider
```

cache 命中率高 = 成本低(cache read 比 fresh input 便宜 \~5x)。`costUsd` 是 OpenRouter 已算好的账单值,不是本地估算。

### Query 5: Validation-feedback 重试有效性(#1081)

```
fields stage, validationFeedbackInjected
| filter @message like /ai_attempt_start/
| stats count() as attempts, sum(validationFeedbackInjected) as feedbackRetries by stage
```

看有多少 attempt 靠 feedback 重试;配合 Query 1 看这些重试是否随后成功。

### Query 6: 未知 subcategory 降级(质量信号)

```
fields @timestamp, telephonySessionId, subcategory, outcome
| filter @message like /Classify emitted unknown subcategory/
| stats count() by subcategory
| sort count desc
```

高频未知 subcategory = taxonomy 或 prompt 该补;这些不进 DLQ(降级为 `other`),但是分类质量流失。

### Query 7: Verify 观测(stage 2.5,observe-only)

```
fields decisionType, classifySubcategory, verifySubcategory, confidence
| filter @message like /verify_decision/
| stats count() by decisionType
```

看 verify 的 adopt/disagree/confirm 率。`decisionType` = `corrected_classify` 表示 verify 若上线会改这次分类 —— 现阶段只观测不改。

### Query 8: S3 转录缺失

```
fields @timestamp, telephonySessionId, error
| filter error like /NoSuchKey/ or error like /transcript from S3/
| stats count() by telephonySessionId
| sort count desc
```

### Query 9: 短转录处理(voicemail / 挂断量)

```
fields @timestamp, telephonySessionId
| filter @message like /transcript.*short/ or @message like /empty.*transcript/
| stats count() as shortCalls by bin(1d)
```

短转录(\<10 字符)存空记录不调 AI;高短通话率反映 voicemail/挂断多。

***

## 操作手册

### 重新分析一通电话（Reprocess）

触发条件: 分析结果错了要重跑、schema 改动后要回填、单通调试。

1. 从 Neon `call-analysis` 删掉这通电话的 `s3AnalysisPath`(去重靠它,不删会被 skip)。
2. 重新上传转录到原 S3 路径 → 触发 EventBridge → 重新分析。
3. 详细步骤见 [reprocess-ai-analysis runbook](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/callytics-operations/reprocess-ai-analysis.md)(infra repo)或 `/reprocess` skill。

### Redrive DLQ（根因修复后）

**先确认根因已修**(跑 Query 1 看最近无同类 error),再 redrive,否则 message 会立刻回 DLQ。

```bash
# 1. 看 DLQ message 数
aws sqs get-queue-attributes \
  --queue-url https://sqs.{region}.amazonaws.com/ACCOUNT/call-analytics-{env}-ai-analysis-dlq-{region} \
  --attribute-names ApproximateNumberOfMessages

# 2. redrive 回主队列
aws sqs start-message-move-task \
  --source-arn arn:aws:sqs:{region}:ACCOUNT:call-analytics-{env}-ai-analysis-dlq-{region} \
  --destination-arn arn:aws:sqs:{region}:ACCOUNT:call-analytics-{env}-ai-analysis-queue-{region} \
  --max-number-of-messages-per-second 10

# 3. 盯 redrive 进度
aws sqs list-message-move-tasks \
  --source-arn arn:aws:sqs:{region}:ACCOUNT:call-analytics-{env}-ai-analysis-dlq-{region}
```

**安全提示**: redrive 会对每条 message 重新调 Lambda(有成本 + AI 调用);message 走正常流程(最多 5 次 redrive);根因没修则回 DLQ。redrive 后盯 DLQ 掉回 0,让 `AiAnalysisDLQAlarm` 恢复 OK —— DLQ 长期留在 ALARM 会掩盖后续新增 message。

Region 按环境: test=us-west-2, pre=us-east-2, prod=us-east-1。AWS profile 用 `aws configure list-profiles` 找有账号 `237206024479` 访问权的(通常 `AdminAccess-237206024479`)。

### 本地验证 schema 改动

改 `@retaintive/common` taxonomy 或某段 schema 后,部署前先跑 Lambda 单测:

```bash
cd lambda/ai-analysis-processor && npm test   # bun:test runner
```

schema/taxonomy 改动是高风险: 走 `@retaintive/common`,消费方同 PR 适配,过 review。prompt 改动至少更新 prompt snapshot;涉及 task/lifecycle/worldview 补 golden eval case。

***

## Code Reference

**入口**:

- `lambda/ai-analysis-processor/src/handler.ts` — SQS batch 处理 + `processSingleRecord()` 编排

**Pipeline**:

- `lambda/ai-analysis-processor/src/core/stages/pipeline.ts` — 3-4 段编排 + verify 决策
- `lambda/ai-analysis-processor/src/core/stages/{triage,classify,verify,coaching}.ts` — 各段 prompt + schema
- `lambda/ai-analysis-processor/src/infrastructure/ai/invoke-model.ts` — per-stage timeout/retry 配置(`STAGE_CONFIG`)

**共享 AI 层**:

- `lambda/shared/utils/ai/invoke.ts` — Vercel AI SDK + OpenRouter 核心调用、重试分类、validation-feedback、usage 日志
- `lambda/shared/utils/ai/provider-routing.ts` — OpenRouter upstream provider 顺序
- `lambda/shared/utils/ai/model-config.ts` — 默认 model ID(`deepseek/deepseek-v4-flash`)

**日志 / 告警**:

- `lambda/shared/utils/logger.ts` — `error` (CloudWatch+Sentry) vs `alert` (+Lark)
- `lambda/ai-analysis-processor/src/utils/metrics.ts` — `AIRequestFailure` / `ProcessingDuration` / `NeonWriteDuration` EMF metric

**基础设施**:

- `lib/stacks/lambda-stack.ts` — Lambda 配置(600s / 512MB / reserved 8)
- `lib/stacks/monitoring-stack.ts` — DLQ / queue-age / provider-failure 告警 + alarm-dispatcher
- `lib/stacks/sqs-stack.ts` — `ai-analysis-queue` / `ai-analysis-dlq`(maxReceiveCount 5)

**规则**:

- `.claude/rules/ai-analysis.md`(infra repo)— AI 层的 Why + 隐式约束

***

## Related Documentation

- [Pipeline Monitoring Runbook](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/observability/runbook-pipeline-monitoring.md) — 告警响应 + 跨 Lambda trace(告警来这里先看)
- [TranscribeProcessor Runbook](/operations/runbooks/transcribe-processor.md) — 上游转录 Lambda
- [Reprocess AI Analysis](https://github.com/retaintive/callytics-infrastructure/blob/main/docs/callytics-operations/reprocess-ai-analysis.md) — 重跑单通分析
- [OpenRouter Activity](https://openrouter.ai/logs) — AI 用量 + 成本(不在 Lambda 内 track)
- [Zod](https://zod.dev/) — schema 校验库
