> For AI agents: the complete documentation index is available at /llms.txt, the full documentation bundle is available at /llms-full.txt.

# RingCentral Call Log、录音就绪与 Rate Limit 调查

> **验证日期：** 2026-08-04
>
> **代码基线：** `callytics-infrastructure` PR [#2035](https://github.com/retaintive/callytics-infrastructure/pull/2035) 的最终 head `53d0f694b`，已通过 merge commit `07e126e5f` 合入 `main`。本文记录 merged code behavior，不代表任何 environment 已完成 deployment。
>
> **调查范围：** RingCentral webhook → SQS → `transcribe-processor` → Call Log / recording / Message Store → S3 / transcription，以及 404/429 recovery。

这份文档记录一次具体事故调查，也固定几条以后 review 不能再猜的边界：

1. 主路径应该只调用一次 `fetchCallLog(telephonySessionId, accountId)`，然后复用这次 response；
2. account-level 与 extension-level Call Log response 的 coverage 确实不同，但它们是同一次 helper 内的 403 fallback，不是旧代码第二次 helper call 的理由；
3. recording metadata 可能延迟出现。正确恢复来源是后续 webhook/message 的新 snapshot，而不是同一 invocation 内约 1 秒后的即时重复 fetch；
4. 当前 reconciliation schedule 在所有 environment 都是 `DISABLED`，而且现有 detection 不会抓到本地 `recordingAvailable=false`、provider 后来变成 true 的情况；
5. PR #2035 降低 quota burn、终止无限 429 clone，并补齐诊断信号，但不等于已经有 account-aware pacing。

## 一、每个数据源回答什么

| Source                         | 回答的问题                                                       | Authority / limitation                                                              |
| ------------------------------ | ----------------------------------------------------------- | ----------------------------------------------------------------------------------- |
| Telephony Session Notification | 刚刚发生了哪个 party-level status？                                 | 快、event-driven；可能乱序、重复，`Disconnected` 不一定代表整通 multi-party session 已结束               |
| Account-level Call Log         | 这通 completed call 的 metadata、legs、recording/contentUri 是什么？ | coverage 最完整；需要相应权限；Call Log 是 Heavy API，completed record 与 recording metadata 可能延迟 |
| Extension-level Call Log       | 当前 authenticated extension 能看到什么？                           | account endpoint 返回 403 时的 fallback；只包含该 extension 可见的 partial data                 |
| Recording content endpoint     | 用已选出的 `contentUri` 下载 audio                                 | 单独的 HTTP request 和 rate-limit surface                                               |
| Message Store                  | voicemail detail、transcript、audio attachment                | voicemail fallback；部分失败是 non-blocking，不一定触发 SQS retry                               |
| DynamoDB / Neon                | 我们已经接受并持久化了什么 processing state？                             | 本地 system of record，不会自动知道 provider 在上次 fetch 后才出现的新 recording                      |

RingCentral 官方建议也支持这个分工：

- [Call Log overview](https://developers.ringcentral.com/guide/voice/call-log)
- [Call recordings](https://developers.ringcentral.com/guide/voice/call-log/recordings)
- [Telephony session notifications](https://developers.ringcentral.com/guide/voice/telephony-session-notifications)
- [Rate limits](https://developers.ringcentral.com/guide/basics/rate-limits)

不要把文档中的 default rate 当成每个 credential/account 的硬编码真相。运行时应读取 `X-Rate-Limit-*` 和 `Retry-After`；limits 可以定制，也可能受到额外 protection policy 影响。

## 二、当前处理 flow

```text
RingCentral telephony webhook
        |
        v
apps/api webhook receiver
        |
        +-- non-Disconnected --> SQS DelaySeconds=0
        |
        `-- Disconnected ----> SQS DelaySeconds=90
                                  |
                                  v
                           transcribe-processor
                                  |
                  +---------------+----------------+
                  |                                |
           non-Disconnected                    Disconnected
                  |                                |
       lifecycle/timeline only              fetch Call Log
                                                   |
                                      account-level endpoint
                                                   |
                                      403 only ----+--> extension-level endpoint
                                                   |     (partial coverage)
                                                   v
                                          select session record
                                                   |
                              +--------------------+-------------------+
                              |                                        |
                       no contentUri                             contentUri ready
                              |                                        |
                persist availability=false                  reuse same snapshot
                end this invocation                                |
                                                            download recording
                                                                    |
                                                              S3 + transcription
```

SQS queue 本身仍有 5 分钟 default delivery delay；ingress 对每条 telephony message 显式传 `DelaySeconds`，所以实际 telephony policy 是 Disconnected 90 秒、其他 status 0 秒。不能只看 queue default 推断实际等待时间。

完整 lifecycle 与 multi-party caveat 见 [Call Lifecycle Tracking](/system-design/call-lifecycle-tracking.md)。

## 三、为什么旧代码会 fetch 两次

旧 TypeScript 路径在两个层级分别调用相同 helper：

```text
BEFORE

processSingleRecord
      |
      +--> fetchCallLog #1 (same session + account)
      |       |
      |       +--> account Call Log
      |               `-- 403 only --> extension Call Log
      |
      +--> snapshot A has selected recording/contentUri?
              |
              +-- no  --> END；fetch #2 根本不会发生
              |
              `-- yes --> processRecording
                              |
                              +--> fetchCallLog #2 immediately
                              |    (same helper + same arguments)
                              |
                              `--> download recording

正常 account access：2 个 Call Log HTTP requests / message
403 fallback：        4 个 Call Log HTTP requests / message
```

```text
AFTER PR #2035

processSingleRecord
      |
      +--> fetchCallLog #1
      |
      +--> snapshot A has selected recording/contentUri?
              |
              +-- no  --> END
              |
              `-- yes --> pass snapshot A to processRecording
                              |
                              `--> download recording

正常 account access：1 个 Call Log HTTP request / message
403 fallback：        2 个 Call Log HTTP requests / message
```

历史追溯结果：

- pre-TypeScript Python 路径是 single fetch + reuse；
- TypeScript rewrite 引入了两个 call site；
- 当时架构说明仍描述 single fetch，没有找到需要即时第二次 fetch 的设计理由；
- 第一次 snapshot 没有 selected recording/contentUri 时，代码会在进入 `processRecording` 前 return。因此旧的第二次 fetch 从来不是“等 recording 变 ready”的 recovery 机制。

结论：删除的是重复的整个 helper call；account 403 → extension fallback 必须保留。

## 四、录音为什么会晚到

Webhook 与 Call Log/recording pipeline 是两条不同时间线：

```text
party Disconnected webhook ───────────────► 先到

Call Log completed record       ──────────► 可能稍后可查

recording metadata/contentUri      ───────► 可能再晚几十秒到数分钟
```

AWS test log 的具体证据：

- 同一 session 在 `2026-07-31 12:58:46.787` 观察到 `apiHasRecording=false`；
- 后续 message 在 `12:59:58.838/.907` 观察到 `true`，约 72 秒后才 ready；
- 另一个旧路径 invocation 中，第二次 Call Log fetch 只比第一次晚约 1.281 秒，而且第一次已经是 `true`。

30 天聚合调查中，`apiHasRecording=false` 共 6,269 条 log / 1,885 个 sessions；`true` 共 14,168 条 / 5,901 个 sessions。样本中存在大量 false → true，常见间隔约 1–7 分钟，也有更长 outlier。这个统计用于证明“延迟 snapshot 确实存在”，不能直接当作 recording delay 的概率分布，因为一个 session 可能产生多条 webhook 与 retry log。

长通话、transfer 与 multi-party 场景要额外谨慎：`Disconnected` 是 party-level terminal status，不保证整个 session 已完全结束；同一 session 也可能收到多个 Disconnected webhook。现有 14-scenario tests 覆盖了 1/2/3 个 webhook 以及第二或第三次才出现 recording 的组合。

## 五、后续 webhook 与 reconciliation 的边界

正常自愈路径：

```text
第一次 Disconnected message
  -> fresh Call Log snapshot
  -> recording unavailable
  -> persist false and return

RingCentral 后续再发 webhook/message
  -> 再取一个 fresh snapshot
  -> recording available
  -> update false -> true
  -> download/transcribe
```

但是不能把 reconciliation 当作当前 safety net：

- `lib/stacks/eventbridge-stack.ts` 把 reconciliation schedule 在所有 environments 都设为 `DISABLED`；
- 即使手工 enable，当前 worker 的 `recordingNotDownloaded` case 要求本地 state 已是 `recordingAvailable=true`；
- 所以“本地 false，但 RingCentral 后来 true”不会被这个 case 自动发现。

如果 provider 不再发后续 webhook，这通录音可能一直停在 unavailable。这个 gap 早于 PR #2035；旧的即时第二次 fetch 也覆盖不了分钟级 delay。修复方向应该是明确的 delayed-readiness retry/reconciliation design，而不是恢复重复 fetch。

## 六、404 与 429 recovery

### 404

- error 距 call event 小于 15 分钟：交给 SQS retry；
- 达到 15 分钟：标记 Call Log fetch failed，结束该 enrichment path。

### 429

PR #2035 后：

```text
RingCentral 429
      |
      v
read Retry-After + add 2–10 minute spread (cap at SQS 900s)
      |
      v
send cloned SQS message with redriveCount + 1
      |
      +-- redriveCount < 5 --> original acknowledged; clone continues
      |
      `-- redriveCount >= 5 --> stop cloning
                                  |
                                  v
                         normal batchItemFailure
                                  |
                                  v
                    SQS maxReceiveCount=3 -> DLQ
```

clone 是一条新 SQS message，所以 SQS retention age 会重新开始；终止 clone loop 靠 `redriveCount`，不是 14 天 retention。14 天 retention 保护普通 backlog，避免 main queue 在长时间无法 drain 时先于 DLQ 丢消息。

PR #2035 仍没有跨 invocation、按 account 协调的 token bucket/pacing。大批量 replay 前必须做 capacity check，不能把 bounded retry 当成吞吐控制。

## 七、Observability contract

`transcribe-processor` 使用现有 shared logger（当前 interim `EnhancedLogger` contract），没有新建第二套 logging system。

### Stable events

| Event                                         | 关键字段                                                             | 回答的问题                                             |
| --------------------------------------------- | ---------------------------------------------------------------- | ------------------------------------------------- |
| `ringcentral.request_started`                 | `api_surface`, `account_id`, `telephony_session_id`              | 实际发出了多少次 outbound request？                        |
| `ringcentral.request_completed`               | 上述字段 + `http_status`, `duration_ms`, `rate_limit_*`              | 哪个 endpoint 慢/失败？runtime quota 还剩多少？              |
| `ringcentral.rate_limited`                    | `retry_after_seconds`, `rate_limit_group`, `client_error_action` | 哪个 bucket 429？client 是 throw 还是 return undefined？ |
| `ringcentral.recording_availability_observed` | `api_has_recording`, `api_recording_id`                          | 同一 session 是否出现 false → true？                     |
| `ringcentral.recording_processing_started`    | `call_log_input`                                                 | 使用 `prefetched` 还是兼容性的 `self_fetch`？              |
| `ringcentral.redrive_scheduled`               | `redrive_count`, `delay_seconds`                                 | clone 是否被安排？                                      |
| `ringcentral.redrive_acknowledged`            | `previous_redrive_count`, `retry_outcome`                        | 原 message 是否由 clone 接管？                           |
| `ringcentral.redrive_exhausted`               | `redrive_count`, `retry_outcome`                                 | 是否已经回到普通 SQS/DLQ path？                            |

`api_surface` 固定为：

- `account_call_log`
- `extension_call_log`
- `recording_content`
- `message_store_search`
- `voicemail_message_detail`
- `voicemail_transcript_content`

### 为什么不能猜 header

只有 `X-Rate-Limit-Remaining` 和 `X-Rate-Limit-Limit` 都存在且为合法整数时，才 emit capacity metrics。header 缺失必须保留为 unknown；填成 `remaining=0 / limit=10` 会制造不存在的容量告警，也会掩盖 customized limit。

metric dimension contract：

- 旧 series：`Source`，保留给现有 dashboard；
- 新 series：`AccountId + Source`，用于定位单个 provider account；
- 不使用 `ClientId` 装 RingCentral account ID。

### CloudWatch Logs Insights 示例

按 API surface 看请求量、状态与延迟：

```sql
fields @timestamp, account_id, telephony_session_id, api_surface,
       http_status, duration_ms, rate_limit_group,
       rate_limit_remaining, rate_limit_limit
| filter message = "ringcentral.request_completed"
| stats count() as requests,
        pct(duration_ms, 50) as p50_ms,
        pct(duration_ms, 95) as p95_ms,
        max(duration_ms) as max_ms
  by account_id, api_surface, http_status
| sort requests desc
```

追一条 session 的 recording readiness：

```sql
fields @timestamp, message, telephony_session_id,
       api_has_recording, api_recording_id, call_log_input
| filter telephony_session_id = "<SESSION_ID>"
| filter message = "ringcentral.recording_availability_observed"
    or message = "ringcentral.recording_processing_started"
| sort @timestamp asc
```

追一条 429 是否最终 redrive 或 exhausted：

```sql
fields @timestamp, message, account_id, telephony_session_id,
       api_surface, rate_limit_group, retry_after_seconds,
       client_error_action, redrive_count, delay_seconds, retry_outcome
| filter message = "ringcentral.rate_limited"
    or message = "ringcentral.redrive_scheduled"
    or message = "ringcentral.redrive_acknowledged"
    or message = "ringcentral.redrive_exhausted"
| sort @timestamp asc
```

`client_error_action=return_undefined` 表示 voicemail client 把该 HTTP failure 作为 unavailable 返回；`throw` 只表示 client 抛给 caller，是否形成 clone 要继续查同一 session 的 `ringcentral.redrive_*`。不要仅凭一条 client log 推断端到端 retry outcome。

## 八、Review 与操作 checklist

涉及 RingCentral request 数量或 retry 的 PR，至少回答：

- 是逻辑 helper call 变了，还是 helper 内 actual HTTP surface 变了？
- account 403 → extension fallback 是否保留？
- 新 fetch 是否真的提供了更新鲜 snapshot，还是同一 invocation 的即时重复？
- 没有 recording 时，靠后续 webhook、SQS retry、reconciliation，还是会静默结束？
- 429 是 retry、non-blocking fallback，还是最终 DLQ？
- rate-limit headers 缺失时是否保持 unknown？
- log 能否按 `account_id + telephony_session_id + api_surface` 串起来？
- unit、integration 与 e2e 是否覆盖 delayed recording 和 redrive termination？

相关运维入口见 [转录处理器 Runbook](/operations/runbooks/transcribe-processor.md) 和 [数据对账 Runbook](/operations/runbooks/reconciliation.md)。
