ABR package-list drift
on a mid-stream CDN failover.
A working postmortem on the apac/teal weekday primetime broadcast that, mid-stream, pushed a regional slice of viewers from CDN-A to CDN-B through a mid-stream failover after CDN-A's regional POP saturated — CDN-A kept publishing the live-edge tail #EXT-X-MEDIA-SEQUENCE:N while CDN-B anchored a fresh replay-origin timeline #EXT-X-MEDIA-SEQUENCE:N+k. The two CDNs' ladders diverged by k=6 segments (~36 s of ABR package-list drift); the cohort's cdn.pop.ladder_segment_index_delta reads fail, the cohort's ladder-rendition re-aggregate probe fails across 240p / 540p / 720p / 1080p, while cache.availability stayed pass and the multi-CDN health probe multicdn.winner_RTT stayed flat on cdn-a / cdn-b / cdn-c. The Streamwake agentic ops layer classified it as cdn_abr_ladder_drift · dominant at 81% confidence with the cdn_cache_segment_miss lane ruled out by name on the cache leg + segment-leg + license-rollout posture, and remediated with an authorization-tiered policy: Tier 0 autonomous under gate (reissue ABR ladder via replay-origin reanchor), Tier 1 surfaced to humans (CDN-B operator side ladder resync), Tier 2 surfaced for operator-team approval forward (codify the new probe into the cohort fan-in) — recovery verified cohort-side + ladder-side across T+30 s → T+15 m, NOT cache-side.
Book a technical demo for ABR ladder drift on mid-stream CDN failover
Read the postmortem — then bring your own incident to Streamwake.
Two ways to engage on this exact failure pattern: book a 30-minute technical demo where we walk through the probe cascade on your source, or hand us an archived incident and watch the agent diagnose it end-to-end.
Both routes land on the scoping intake form — no SDR gate.
What the cohort saw
The first things to read on any real mid-stream CDN failover package-list drift incident are the cohort-level numbers on the affected apac/teal cohort — and crucially, the contrast against the cache leg that stayed green across both CDNs. Three numbers did the heavy lifting here: the ladder-segment-index delta, the cohort join-stall ratio, and the contrast between the cross-CDN ladder alignment probe (fail) and the cache.availability / segment-leg / license-rollout posture probes (all pass on the affected edge POPs across both CDNs). The figures below are simulated telemetry — the disclosure above applies to every figure on this page.
Roughly 14% of the apac/teal weekday primetime cohort — session-level ABR ladder crossed on the failover slice from cdn-A/sin02 to cdn-B/hkg01, where cdn.pop.ladder_segment_index_delta lifted from 0 baseline to k=6 (~36 s of timeline divergence) on both CDNs and cohort.ladder_rendition_reaggregate_probe failed across renditions 240p / 540p / 720p / 1080p. The cache leg was green — cache.availability stayed X-Cache: HIT and cache.segment_leg_cache_hit stayed pass on both CDNs through the entire window — but the cohort's rebuffer ratio lifted to 0.082 behind the mid-stream ladder PDAT drift.
~41 minutes between the first cross-CDN ladder alignment probe variance at T+0 m and the apac/teal cohort's cohort_join_stall_ratio + cohort_rebuffer_ratio + ladder-rendition re-aggregate probe + cohort startup-time settling within tolerance at T+41 m. The window goes longer than the natural ladder-side reanchor settle because the cdn-B operator-side ladder resync (Tier 1) and the operator-team sign-off on the new probe (Tier 2) ship after the close-out probes land — so the postmortem window is two-tier: T+0 m → T+18 m on the cohort-side + ladder-side close-out, and shaped forward by the cdn-B resync + Tier 2 sign-off that bound forward to the next weekday primetime broadcast.
cdn.pop.ladder_segment_index_delta: k=6 (~36.04 s PDAT drift) on cdn-A/sin02 vs cdn-B/hkg01 (apac/teal cohort) — CDN-A continuing the live-edge tail at #EXT-X-MEDIA-SEQUENCE:84217 /#EXT-X-PROGRAM-DATE-TIME:2026-08-20T18:42:11.000Z while CDN-B serves the fresh anchor at #EXT-X-MEDIA-SEQUENCE:84223 / #EXT-X-PROGRAM-DATE-TIME:2026-08-20T18:42:47.040Z. The discriminator is that the cache leg stayed green — cache.availability = X-Cache: HIT, cache.segment_leg_cache_hit = pass, cdn.license_rollout_posture_check = pass, and the multi-CDN health probe multicdn.winner_RTT = 40 / 41 / 42 ms (cdn-a / cdn-b / cdn-c flat at baseline). The failure is on the ladder manifest timeline alignment, NOT a cache-miss posture on either CDN leg.
How Streamwake classified this incident
Three ranked hypotheses: the top one filing the timeline as cdn_abr_ladder_drift · dominant, the second explicitly tagged cdn_cache_segment_miss · ruled out by so the recovery message lands (the cache leg stayed clean on both CDNs through the entire window), and the third filed as an alternate: cdn_pop_route_bouncing_under_failover with a low confidence that captures the rtt-side stampede shape before the multi-CDN health probe removes it.
cdn_abr_ladder_drift — a mid-stream CDN failover from CDN-A to CDN-B on a regional slice of the apac/teal weekday primetime cohort pushed the cohort's breadcrumb CDN-A to keep publishing the live-edge tail #EXT-X-MEDIA-SEQUENCE:N while the failover slice on CDN-B got a freshly-anchored timeline #EXT-X-MEDIA-SEQUENCE:N+6; the two CDNs' ladders diverged by k=6 segments, #EXT-X-PROGRAM-DATE-TIME drifted past the cohort's expected join window on the affected slice, and the cohort's ladder-rendition re-aggregate probe failed on 240p / 540p / 720p / 1080p. Four signals line up: ladder_segment_index_delta k=6 vs 0 baseline (cross-CDN), cross_cdn_alignment_probe reads fail, cohort.ladder_rendition_reaggregate_probe fails across monitored renditions, cohort_join_stall_ratio 0.124 vs 0.018 baseline.
cdn_cache_segment_miss was the secondary signal ranked at 24% — but it's tagged ruled out by so the recovery message lands. cache.availability reads X-Cache: HIT on both cdn-A/sin02 and cdn-B/hkg01, cache.segment_leg_cache_hit reads pass on the affected cohort on both CDNs, and cdn.license_rollout_posture_check reads pass on the affected ladder — the failure is on the ladder alignment, NOT on the cache-locale. The dismissal rule was "rank the cause on the cache.availability + segment-leg cache-hit + license-rollout posture pattern, not on the player-visible symptom alone"; the ladder-alignment nature of the failure is the decisive signal pattern.
- Region: apac/teal weekday primetime broadcast (cdn-A/sin02 → cdn-B/hkg01 failover slice)
- Status: resolved (window closed; Tier 1 cdn-B resync configured; Tier 2 new probe staged for operator-team sign-off forward)
- Opened: 2026-08-20 18:42 UTC
- Spread: contained to the apac/teal weekday primetime broadcast — na-east and eu-west cohorts on the same broadcast unaffected; adjacent multi-CDN cohorts on the same vendor unaffected across all three CDNs; the multi-CDN health probe settled at zero egress change on the apac/teal edge legs.
Above the 80% threshold the agent treats as a confident top-hypothesis filing. cdn_cache_segment_miss · 0.24 was cleared explicitly because cache.availability, cache.segment_leg_cache_hit, and cdn.license_rollout_posture_check together prove the cache leg was healthy at zero egress change across both CDNs — the discriminator for the ladder-alignment-not-cache-locale shape of the failure. The multicdn.winner_RTT probe reading pass on cdn-a|cdn-b|cdn-c (40/41/42ms flat at baseline) clears cdn_pop_route_bouncing_under_failover as the alternate.
Who clears the gate
The governed-action policy on this incident is split across three authorization tiers — the agent may act on its own under an explicit gate (Tier 0), surfaces the work to the cdn-B operator team (Tier 1), or hands the new probe codification to the operator-team sign-off forward (Tier 2). Each tier has an explicit gate; each gate names the threshold before the action lands, and the lane the action belongs to is what makes the split distinct from a generic postmortem.
reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill — autonomous under gate: top_confidence >= 0.80 AND cdn.pop.ladder_segment_index_delta >= 3 AND multicdn.winner_RTT.pass == true on the affected cohort. The agent clears the gate and reissues the ABR ladder via replay-origin reanchor at T_mid plus a CDN-side segment_index backfill so the cohort resolves on the refreshed ladder — the cache leg is NOT touched, the multi-CDN health probe proves the CDN leg is healthy at zero egress change.
surface_cdn_failover_evidence_to_cdn_b_for_ladder_resync_via_origin_clip — surfaced to humans: cdn_abr_ladder_drift >= 3 segments AND top_confidence >= 0.75. The operator at the cdn-B side configures the ladder anchor / segment_index backfill before cdn-B's manifest tail fully publishes — a misanchor that surfaces on a future primetime isn't masked by a "we reconciled" close-out.
surface_cdn.pop.ladder_segment_index_delta_probe_to_cohort_fanin — surfaced for operator-team approval forward: incident_resolved AND cohort_side close-out signals cleared. This is the learn-loop act — once the lane is closed, the operator team signs off on pinning the new probe (cdn.pop.ladder_segment_index_delta) into the cohort probe fan-in for future mid-stream failover windows; cohort.ladder_rendition_reaggregate_probe is pinned for the next four weekday primetime windows so the false-positive rate can be measured before promotion.
Incident timeline
Eleven events: detection on the apac/teal cohort, classification across three ranked hypotheses (with the cdn_cache_segment_miss lane tagged ruled out by so the recovery message lands), five acts the agent took under the authorization-tiered governed-action gates (Tier 0 autonomous under gate + a status-change handoff at T+5 m + a T+18 m re-probe + a T+25 m status change), four acts it surfaced to humans (Tier 1 cdn-B surface, Tier 2 operator-team sign-off, cdn-B operator ack, reliability-team assignment), and the resolution. The right-hand "act" + "tier" tags are what makes this postmortem distinct from a generic write-up — they pin the split between autonomous agentic ops, the cdn-B side coordination, and the operator-team sign-off forward. All times below are simulated telemetry — the disclosure at the top of this page applies to every minute offset on the timeline.
Today
- T+0mDetectionby cohort agent · apac/teal · cdn-A/sin02 + cdn-B/hkg01act · autonomous
Cross-CDN ladder alignment probe fan-in reads fail at k=6
cdn.pop.ladder_segment_index_delta crossed 0 → 6 within a 30 s window on cdn-A/sin02 AND cdn-B/hkg01; cross_cdn_alignment_probe reads fail (CDN-A live-edge tail at #EXT-X-MEDIA-SEQUENCE:84217 vs CDN-B replay-origin anchor at #EXT-X-MEDIA-SEQUENCE:84223 — ~36.04 s PDAT drift); cohort.ladder_rendition_reaggregate_probe reads fail across 240p / 540p / 720p / 1080p; cohort.cohort_join_stall_ratio lifted to 0.124 (vs 0.018 baseline) and cohort.cohort_rebuffer_ratio lifted to 0.082 (vs 0.012 baseline). cache.availability stays pass on both CDNs; segment-leg cache-hit stays pass; multicdn.winner_RTT stays flat on cdn-a / cdn-b / cdn-c.
Aug 20, 06:42:11 PM - T+1mClassificationby Streamwake reliability agentact · autonomous
Ranked: cdn_abr_ladder_drift · dominant (0.81) · cdn_cache_segment_miss · ruled out by (0.24) · cdn_pop_route_bouncing_under_failover · alternate (0.21)
Top hypothesis reads 81% confidence. cdn_cache_segment_miss is ruled out by name on cache.availability + segment-leg cache-hit + license-rollout posture — the cache leg was green the entire window on both CDNs and the failure is on the ladder manifest timeline alignment. cdn_pop_route_bouncing_under_failover is the alternate (the multi-CDN health probe reads pass on cdn-a|cdn-b|cdn-c — flat at baseline; the discriminator is on-segment-index delta, not on rtt_ms signature).
Aug 20, 06:43:11 PM - T+2mAutomated action · autonomousby Streamwake reliability agentact · autonomousTier 0
Governed · Tier 0 autonomous under gate: reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill
Top_confidence 0.81 >= 0.80 gate cleared; cdn.pop.ladder_segment_index_delta k=6 >= 3 gate cleared; multicdn.winner_RTT.pass == true gate cleared on cdn-a | cdn-b | cdn-c. Reissues the ABR ladder via replay-origin reanchor at T_mid plus a CDN-side segment_index backfill so the cohort resolves on the refreshed ladder — the cache leg is NOT touched (it was green the entire window), the multi-CDN health probe proves the CDN leg is healthy at zero egress change.
Aug 20, 06:44:11 PM - T+3mSurfaced to humanby agent → cdn-B operator teamact · surfaced to humansTier 1
Governed · Tier 1 surfaced to humans: surface_cdn_failover_evidence_to_cdn_b_for_ladder_resync_via_origin_clip
CDN-side coordination required. Cdn_abr_ladder_drift >= 3 segments (k=6) AND top_confidence 0.81 >= 0.75 gate cleared. Surfaces the mid-stream failover evidence (CDN-B manifest timeline PDAT drift, k=6 segment skew) to the cdn-B operator team so they configure the ladder anchor / segment_index backfill before cdn-B's manifest tail fully publishes — a misanchor that surfaces on a future primetime isn't masked by a "we reconciled" close-out.
Aug 20, 06:45:11 PM - T+5mStatus changeby Streamwake reliability agentact · autonomous
T+5 m handoff: ladder probe re-reads dropping; join-stall settle milestone
cohort.cohort_join_stall_ratio settled to within 0.05 tolerance by T+5 m (per the staged T+5 m join-stall settle gate — NOT a single-probe reanchor); ladder_segment_index_delta re-reads dropping from k=6 toward k=2 within the first same-length ladder window — ladder-side settle tracking ahead of cohort-side settle, as expected when the reanchor lands first.
Aug 20, 06:47:11 PM - T+7mSurfaced to humanby agent → reliability teamact · surfaced to humansTier 2
Governed · Tier 2 surfaced for operator-team sign-off forward: surface_cdn.pop.ladder_segment_index_delta_probe_to_cohort_fanin
Operator-team sign-off forward required. This is the learn-loop act — once the lane is closed, the operator team signs off on pinning the new probe (cdn.pop.ladder_segment_index_delta) into the cohort probe fan-in for future mid-stream failover windows; cohort.ladder_rendition_reaggregate_probe is pinned for the next four weekday primetime windows so the false-positive rate can be measured before promotion.
Aug 20, 06:49:11 PM - T+10mSurfaced to humanby cdn-B operator teamact · surfaced to humans
cdn-B operator acked; ladder anchor / segment_index backfill configured
Acknowledged within 7 m of the Tier 1 surface; ladder anchor / segment_index backfill configured on the cdn-B side so the next apac/teal weekday primetime broadcast window opens with a reissued ABR ladder — the operator confirmed manifest timeline PDAT alignment to the live-edge tail on cdn-A live-edge tail cadence.
Aug 20, 06:52:11 PM - T+15mSurfaced to humanby reliability teamact · surfaced to humans
Postmortem write-up assigned (this page)
Reliability team assigned the public postmortem; this page is the resulting write-up, with the three ranked hypotheses, the authorization-tiered remediation policy (Tier 0 / Tier 1 / Tier 2), the staged T+30 s → T+15 m viewer-level recovery verification window, and the learn-loop note pinning the new probe into the cohort fan-in.
Aug 20, 06:57:11 PM - T+18mAutomated action · autonomousby Streamwake reliability agentact · autonomous
T+18 m re-probe: first ladder-segment-index delta re-read
cdn.pop.ladder_segment_index_delta dropped from k=6 → k=2 within the first same-length ladder window (per the T+30 s staged gate — ladder-side settle ahead of cohort-side settle, as expected after the reanchor); cohort.cohort_rebuffer_ratio within 0.05 tolerance (per the T+2 m staged gate); cache.availability / segment-leg cache-hit stayed pass throughout — those probes are NOT a close-out signal (the cache leg was green the entire window).
Aug 20, 07:00:11 PM - T+25mStatus changeby Streamwake reliability agentact · autonomous
T+25 m status change: full cohort re-anchor cadence holds
cohort.cohort_join_stall_ratio remains within 0.05 tolerance (per the T+5 m staged gate); cohort.ladder_rendition_reaggregate_probe reads pass on every monitored rendition (240p / 540p / 720p / 1080p); cohort.cohort_startup_time_p95_ms reads 1920 ms (within startup-time baseline, ahead of the next apac/teal weekday primetime broadcast window opens); cdn.pop.ladder_segment_index_delta sits at k=0 within the same window.
Aug 20, 07:07:11 PM - T+41mResolutionby Operator + agentact · surfaced to humans
T+41 m incident resolved; Tier 0 ladder reissue landed autonomously; Tier 1 cdn-B resync configured; Tier 2 new probe staged for operator-team sign-off forward
Cohort-side + ladder-side close-out signals cleared on the apac/teal cohort for the current window (cohort_join_stall_ratio + cohort_rebuffer_ratio + cohort_startup_time_p95_ms + cohort.ladder_rendition_reaggregate_probe + cdn.pop.ladder_segment_index_delta within tolerance) — NOT on cache.availability / segment-leg cache-hit / multicdn.winner_RTT returning to pass (the cache leg was green the entire window, the multi-CDN health probe was healthy the entire window). The audit step codifies: "verify ladder-segment-index delta + cohort ladder-rendition re-aggregate settle within tolerance over the next same-length cohort window" — NOT cache-side close-out, NOT CDN-cache re-anchor.
Aug 20, 07:23:11 PM
- classify · ranked three hypotheses with confidence in 90 s; cdn_cache_segment_miss · ruled out by tagged (cache leg stayed green)
- Tier 0 ladder reissue · reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill under the top_confidence ≥ 0.80 + ladder_segment_index_delta ≥ 3 + multicdn.winner_RTT.pass gate
- T+5 m status · observed ladder-side settle ahead of cohort-side settle; cohort_join_stall_ratio tolerance holding
- T+18 m re-probe · ladder-segment-index delta k=6 → k=2 within the first same-length ladder window (T+30 s staged gate)
- T+25 m status · cohort ladder-rendition re-aggregate pass on every monitored rendition; cohort_startup_time_p95_ms ahead of the next apac/teal weekday primetime window
- cdn-B operator team · Tier 1 surface — configured the ladder anchor / segment_index backfill before cdn-B's manifest tail fully publishes
- reliability team · Tier 2 surface — sign-off forward on pinning cdn.pop.ladder_segment_index_delta + cohort.ladder_rendition_reaggregate_probe into the cohort probe fan-in
- reliability team · assigned the public postmortem write-up (this page) — recovery criteria on the audit step is the cohort-side + ladder-side close-out signal, NOT the cache-side re-anchoring
What recovery looks like across the ladder-side + cohort-side probes
Recovery on this incident is verified across the staged T+30 s → T+15 m window on the apac/teal cohort + on the cross-CDN ladder alignment probe — NOT on the cache layer turning green again. The cache leg stayed green the entire window; if recovery were the cache turning green, the postmortem would land on a CDN-side lane that was never the failure. The audit step writes the close-out signal into the playbook as a four-probe verification — ladder probe re-read at T+30 s → joined-cohort rebuffer settle at T+2 m → join-stall settle at T+5 m → full cohort re-anchor at T+15 m.
cdn.pop.ladder_segment_index_delta drops from k=6 → k=2 within the first same-length ladder window after the Tier 0 reissue lands (cadence 30 s). Ladder-side settle leads the cohort-side settle, as expected when the replay-origin reanchor lands first.
cohort.cohort_rebuffer_ratio settles ≤ 0.05 within ten consecutive 30 s windows on the apac/teal affected cohort — NOT ladder-side settling; cohort-side settle reads the per-segment PDAT retune on the affected cohort.
cohort.cohort_join_stall_ratio settles ≤ 0.05 within ten consecutive 30 s windows on the apac/teal affected cohort — NOT ladder-side settling; the mid-stream rejoin (PDAT retune on viewers already joined at the point of the failover) validates on the same cadence.
cohort.ladder_rendition_reaggregate_probe reads pass on every monitored rendition (240p / 540p / 720p / 1080p); cdn.pop.ladder_segment_index_delta sits at ≤ 0 within the same window; cohort.cohort_startup_time_p95_ms reads ahead of the next apac/teal weekday primetime broadcast window open.
cache.availability returning to X-Cache: HIT at zero egress change is NOT a close-out signal — the cache leg was green the entire window on both CDNs. cache.segment_leg_cache_hit staying pass is NOT a close-out signal — the segment leg was green the entire window. multicdn.winner_RTT reading pass on cdn-a / cdn-b / cdn-c (40/41/42ms flat at baseline) is NOT a close-out signal — the multi-CDN health probe proves the CDN leg is healthy, but does not prove the cohort delivers green playback when the cross-CDN ladder manifest is misanchored.
Anatomy of the evidence packet
The two packets on the failing source — a multi-CDN cycle 0 (CDN-A live-edge tail) + cycle 1 (CDN-B replay-origin anchor) ABR ladder probe packet on the apac/teal affected cohort (with the cross-CDN ladder alignment probe failing on the cohort while cache.availability + segment-leg cache-hit + license-rollout posture stay pass on both CDNs), and the agent timeline response with the three ranked hypotheses, the authorization-tiered remediation policy, and the staged T+30 s → T+15 m viewer-level recovery verification window. The probe packet is what the agent decided on; the timeline response is what the agent emitted.
GET /live/event/stream.m3u8 HTTP/1.1
host: cdn.example.com
accept: application/vnd.apple.mpegurl
----- cycle 0 (T+0m, mid-stream · CDN-A live-edge tail · apac/teal cohort) -----
# CDN-A continues the live-edge tail — the breadcrumb viewers already joined against
HTTP/2 200
content-type: application/vnd.apple.mpegurl
server-timing: manifest-fetch;dur=124
x-cdn: cdn-A/sin02
x-pop: cdn-A/sin02
x-cache: HIT
x-multicdn-route: cdn-A
#EXTM3U
#EXT-X-VERSION:9
#EXT-X-TARGETDURATION:6
#EXT-X-MEDIA-SEQUENCE:84217
#EXT-X-PROGRAM-DATE-TIME:2026-08-20T18:42:11.000Z
#EXTINF:6.000,
084217.ts
# ladder_segment_index_delta: fail (cdn-A continuing live-edge tail at #EXT-X-MEDIA-SEQUENCE:84217)
----- cycle 1 (T+0m, mid-stream · CDN-B replay-origin anchor · apac/teal slice) -----
# CDN-B serves a fresh replay-origin anchor — the new timeline for the failover slice
HTTP/2 200
content-type: application/vnd.apple.mpegurl
server-timing: manifest-fetch;dur=131
x-cdn: cdn-B/hkg01
x-pop: cdn-B/hkg01
x-cache: HIT
x-multicdn-route: cdn-B
x-failover-mode: replay-origin-anchor-at-t-mid
#EXTM3U
#EXT-X-VERSION:9
#EXT-X-TARGETDURATION:6
#EXT-X-MEDIA-SEQUENCE:84223
#EXT-X-PROGRAM-DATE-TIME:2026-08-20T18:42:47.040Z ← 36.04 s ahead of CDN-A's PDAT
#EXTINF:6.000,
084223.ts
# cdn.pop.ladder_segment_index_delta: fail (k=6 across all 4 renditions on cdn-A/sin02 AND cdn-B/hkg01 —
# mid-stream manifest timeline divergence;
# CDN-A live-edge tail vs CDN-B replay-origin anchor)
# cdn.pop.ladder_segment_index_delta.k: 6 (vs 0 baseline — ~36 s drift)
# cdn.pop.affected_renditions: [240p, 540p, 720p, 1080p]
# cdn.pop.cross_cdn_alignment_probe: fail (CDN-A continuing against CDN-B anchor — ladders diverged)
# cdn.pop.affected_pops: [cdn-A/sin02, cdn-B/hkg01]
# cohort.ladder_rendition_reaggregate_probe: fail (240p / 540p / 720p / 1080p re-aggregate probe fan-in read fail
# across renditions; rendition-level segment_index drift != 0)
# cohort.cohort_join_stall_ratio: 0.124 (vs 0.018 baseline, 90s cohort window — late-join stalled behind
# the ladder mismatch; viewer PDAT drift past expected join window)
# cohort.cohort_rebuffer_ratio: 0.082 (vs 0.012 baseline — playback stalled on the affected cohort)
# cohort.cohort_startup_time_p95_ms: 6210 (vs 1840 baseline — TTFF lifted above the 4 s threshold)
# cohort.affected_cohort_id: apac-teal-primetime-broadcast
# cache.availability on the affected edge POP: pass (X-Cache: HIT on cdn-A/sin02 AND cdn-B/hkg01 — segment leg green)
# cache.segment_leg_cache_hit: pass (segment leg green on the affected cohort on both CDNs)
# cdn.license_rollout_posture_check: pass (license-rollout posture on the affected ladder is green —
# ladder rendered correctly across all renditions)
# multicdn.winner_RTT on cdn-a | cdn-b | cdn-c: pass (cdn-a 40ms, cdn-b 41ms, cdn-c 42ms — flat at baseline;
# no rtt-side failover stampede; the failure is on-segment-index delta,
# not on rtt_ms signature)
# edge.egress_kbps: 6421 (at-or-above expected — cohort IS delivering traffic; the failure
# is on the ladder manifest, not on edge egress shape)
# cdn_pop.partial_segment_warmup_state: primed (segment leg green — not a partial-segment warm-up failure)
----- cycle 2 (T+~12 m, after reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill) -----
HTTP/2 200
content-type: application/vnd.apple.mpegurl
server-timing: manifest-fetch;dur=120
x-cdn: cdn-B/hkg01
x-pop: cdn-B/hkg01
x-cache: HIT
x-multicdn-route: cdn-B
x-failover-mode: replay-origin-anchor-at-t-mid (re-issued)
#EXTM3U
#EXT-X-VERSION:9
#EXT-X-TARGETDURATION:6
#EXT-X-MEDIA-SEQUENCE:84228
#EXT-X-PROGRAM-DATE-TIME:2026-08-20T18:55:24.040Z
#EXTINF:6.000,
084228.ts
# cdn.pop.ladder_segment_index_delta: fail→recover (k dropping toward 0 within the same-length cohort window)
# cohort.ladder_rendition_reaggregate_probe: pass (240p / 540p / 720p / 1080p ladder-rendition re-aggregate fan-in pass
# after the segment_index backfill)
# cohort.cohort_join_stall_ratio: 0.024 (within cohort baseline)
# cohort.cohort_rebuffer_ratio: 0.013 (within cohort baseline)
# cohort.cohort_startup_time_p95_ms: 1920 (within startup-time baseline)
# cache.availability: pass (X-Cache: HIT — cache leg still clean through the entire window)
# cache.segment_leg_cache_hit: pass (segment leg still clean — NOT a segment-leg cache miss)
# multicdn.winner_RTT: pass (cdn-a / cdn-b / cdn-c flat — multi-CDN health probe healthy)
# governed_action_emitted: reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill,
# surface_cdn_failover_evidence_to_cdn_b_for_ladder_resync_via_origin_clip,
# surface_cdn.pop.ladder_segment_index_delta_probe_to_cohort_fanin- cdn.pop.ladder_segment_index_delta →
k=6 (vs 0 baseline — mid-stream manifest timeline divergence, ~36.04 s PDAT drift on the apac/teal affected cohort across 4 renditions on cdn-A/sin02 AND cdn-B/hkg01) - cdn.pop.cross_cdn_alignment_probe →
fail (CDN-A live-edge tail at #EXT-X-MEDIA-SEQUENCE:84217 vs CDN-B replay-origin anchor at #EXT-X-MEDIA-SEQUENCE:84223) - cohort.ladder_rendition_reaggregate_probe →
fail across 240p / 540p / 720p / 1080p on the affected renditions - cohort.cohort_join_stall_ratio →
0.124 vs 0.018 baseline (90 s cohort window — viewer PDAT drift past expected join window on the affected cohort) - cohort.cohort_rebuffer_ratio →
0.082 vs 0.012 baseline (90 s cohort window) - cache.availability → pass; cache.segment_leg_cache_hit → pass; cdn.license_rollout_posture_check → pass; multicdn.winner_RTT → pass (40 / 41 / 42 ms flat on cdn-a / cdn-b / cdn-c). The cache leg is healthy at zero egress change — NOT a cache-miss, NOT a segment-leg hit-rate miss, NOT a multi-CDN routing flip.
{
"stream_id": "ckabrpackagelistdrift5189",
"source": "https://cdn.example.com/live/event/stream.m3u8",
"protocol": "HLS / CMAF / cdn-A/sin02 | cdn-B/hkg01 / multi-CDN cdn-a|cdn-b|cdn-c / ABR ladder / mid-stream failover",
"checked_at": "2026-08-20T18:42:11Z",
"ranked_hypotheses": [
{
"rank": 1,
"hypothesis": "cdn_abr_ladder_drift",
"tag": "dominant",
"confidence": 0.81,
"evidence_signals": [
"cdn.pop.ladder_segment_index_delta → fail (k=6 vs 0 baseline across 4 renditions on cdn-A/sin02 AND cdn-B/hkg01)",
"cdn.pop.cross_cdn_alignment_probe → fail — mid-stream manifest timeline divergence (CDN-A live-edge tail at #EXT-X-MEDIA-SEQUENCE:84217 vs CDN-B replay-origin anchor at #EXT-X-MEDIA-SEQUENCE:84223; #EXT-X-PROGRAM-DATE-TIME drift ~36.04 s)",
"cohort.ladder_rendition_reaggregate_probe → fail (240p / 540p / 720p / 1080p rendition-level segment_index drift != 0 across the affected renditions on cdn-A/sin02 and cdn-B/hkg01)",
"cohort.cohort_join_stall_ratio → fail (0.124 vs 0.018 baseline, 90s cohort window — late-join stalled behind the ladder mismatch; viewer PDAT drift past the expected join window)",
"cohort.cohort_rebuffer_ratio → fail (0.082 vs 0.012 baseline — playback stalled on the affected cohort)",
"cohort.cohort_startup_time_p95_ms → fail (6210ms vs 1840 baseline — TTFF past the 4 s threshold)",
"edge.egress_kbps → pass (6421 Kbps at-or-above expected — cohort IS delivering traffic; the failure is on the ladder manifest, not on edge egress shape)",
"cache.availability → pass on the affected edge POP — X-Cache: HIT on cdn-A/sin02 AND cdn-B/hkg01 under rebuffer load; no edge miss posture",
"cache.segment_leg_cache_hit → pass on the affected cohort across both CDNs",
"cdn.license_rollout_posture_check → pass on the affected ladder; the failure is NOT on the segment-leg cache-side",
"multicdn.winner_RTT → pass on cdn-a|cdn-b|cdn-c (40ms / 41ms / 42ms — flat at baseline; no rtt-side failover stampede; the multi-CDN health probe proves the CDN leg is healthy)"
]
},
{
"rank": 2,
"hypothesis": "cdn_cache_segment_miss",
"tag": "ruled_out_by",
"confidence": 0.24,
"evidence_signals": [
"cache.availability probe reads pass on the affected edge POP across both CDNs (cdn-A/sin02 AND cdn-B/hkg01); X-Cache: HIT under rebuffer load — no edge miss posture; the cache leg is healthy",
"cache.segment_leg_cache_hit probe reads pass on the affected cohort across both CDNs; if a segment-leg cache miss were stateful, the cohort resolve path across both CDNs and the same breadcrumb ladder would all misfire at the cache leg — they do not",
"cdn.license_rollout_posture_check reads pass on the affected ladder; the ladder rendered correctly across all renditions — the failure is on ladder alignment, not on the cache-locale",
"edge.egress_kbps reads at-or-above expected — the cohort is delivering traffic, the failure is on the ladder manifest timeline, not on segment availability",
"if the cache-locale had a stateful miss, the warm breadcrumb cohort on cdn-A/sin02 would misfire on the same rung — it does not; the failure is on the cdn_ABR package-list drift between CDN-A live-edge tail AND CDN-B replay-origin anchor"
]
},
{
"rank": 3,
"hypothesis": "cdn_pop_route_bouncing_under_failover",
"tag": "alternate",
"confidence": 0.21,
"evidence_signals": [
"multicdn.winner_RTT probe reads pass on cdn-a | cdn-b | cdn-c (40ms / 41ms / 42ms — flat at baseline); no rtt-side failover stampede",
"the multi-CDN health probe proves the CDN leg is healthy at zero egress change; the probe-side shape is on-segment-index delta, not on rtt_ms signature — the discriminator for the cdn_ABR package-list drift failure lane",
"if rtt-side route bouncing were stateful, three AS populations on adjacent multicast-cdn neighborhoods would flip together under the failover — they did not; the failure is on the ladder manifest timeline alignment, not on rtt routing"
]
}
],
"governed_actions": [
{
"action": "reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill",
"type": "governed",
"decision_lane": "autonomous",
"authorization_tier": "Tier 0 (autonomous under gate)",
"gating": "top_confidence >= 0.80 AND cdn.pop.ladder_segment_index_delta >= 3 AND multicdn.winner_RTT.pass == true on the affected cohort",
"evidence": "top-hypothesis confidence 0.81; ladder_segment_index_delta k=6 >= 3; multicdn.winner_RTT pass (40/41/42ms flat)",
"expected_effect": "ladder_segment_index_delta collapses toward 0 within the same-length cohort window; cohort.ladder_rendition_reaggregate_probe settles to pass; cohort_join_stall_ratio / cohort_rebuffer_ratio clears within tolerance"
},
{
"action": "surface_cdn_failover_evidence_to_cdn_b_for_ladder_resync_via_origin_clip",
"type": "governed",
"decision_lane": "surfaced_to_humans",
"authorization_tier": "Tier 1 (surfaced to humans)",
"gating": "cdn_abr_ladder_drift >= 3 segments AND top_confidence >= 0.75",
"evidence": "ladder_segment_index_delta k=6 >= 3; top-hypothesis confidence 0.81 >= 0.75; the operator at the cdn-b side configures the ladder anchor and segment_index backfill before cdn-b's manifest tail fully publishes, so a misanchor that surfaces on a future primetime isn't masked by a 'we reconciled' close-out",
"expected_effect": "cdn-B operator team configures the ladder anchor / segment_index backfill before the next apac/teal weekday primetime broadcast window opens"
},
{
"action": "surface_cdn.pop.ladder_segment_index_delta_probe_to_cohort_fanin",
"type": "governed",
"decision_lane": "surfaced_to_humans_for_approval_forward",
"authorization_tier": "Tier 2 (operator-team sign-off forward)",
"gating": "incident_resolved AND cohort_side close-out signals cleared",
"evidence": "incident_resolved on T+41 m; cohort-side close-out signals (cohort_join_stall_ratio + cohort_rebuffer_ratio + cohort_startup_time_p95_ms + ladder_rendition_reaggregate_probe) cleared within tolerance",
"expected_effect": "operator-team signs off on pinning the new probe into the cohort fan-in for future mid-stream failover windows; learn-loop codification"
}
],
"verification_window": {
"close_out_signal": "cohort_side_plus_ladder_side",
"probes_t_plus_30s": [
"cdn.pop.ladder_segment_index_delta drops from k=6 → k=2 within the first same-length ladder window (30s cadence)"
],
"probes_t_plus_2m": [
"cohort.cohort_rebuffer_ratio clears <= 0.05 within ten consecutive 30s windows (NOT ladder-side settling)"
],
"probes_t_plus_5m": [
"cohort.cohort_join_stall_ratio settles <= 0.05 within ten consecutive 30s windows on the affected cohort (NOT ladder-side settling); mid-stream rejoin validates"
],
"probes_t_plus_15m": [
"cohort.ladder_rendition_reaggregate_probe pass on every monitored rendition (240p / 540p / 720p / 1080p)",
"cdn.pop.ladder_segment_index_delta <= 0 within the same window",
"cohort.cohort_startup_time_p95_ms ahead of the next weekday primetime broadcast window opens"
],
"NOT_close_out_signal": [
"cache.availability returning to X-Cache: HIT at zero egress change — geometric (cache leg was green the entire window)",
"cache.segment_leg_cache_hit staying pass — geometric, the cache leg was green",
"multicdn.winner_RTT reading pass on cdn-a|cdn-b|cdn-c — geometric (proves CDN leg is healthy, but does not on its own prove the cohort delivers green playback when the ladder is misanchored)"
],
"explicit_note": "'ladder aligned' != 'cohort delivers green playback'. Recovery is verified cohort-side + ladder-side."
},
"surfaced_to_humans": [
{"owner": "cdn-B operator team", "task": "configure the ladder anchor / segment_index backfill on the cdn-B side before the next apac/teal weekday primetime broadcast window opens"},
{"owner": "reliability team", "task": "approve the cdn.pop.ladder_segment_index_delta probe + the cohort.ladder_rendition_reaggregate_probe codification into the cohort fan-in"},
{"owner": "viewer-platform team", "task": "approve the apac/teal cohort-side profile post-mortem from this entry — ladder-side settle + cohort-side settle cadence pinned from the timeline"},
{"owner": "on-call", "task": "page on the cdn_ABR package-list drift root-cause review — mid-stream failover + live-edge tail + replay-origin anchor"}
]
}- Tier 0 →
reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill(autonomous under gate: top_confidence ≥ 0.80 + ladder_segment_index_delta ≥ 3 + multicdn.winner_RTT.pass) - Tier 1 →
surface_cdn_failover_evidence_to_cdn_b_for_ladder_resync_via_origin_clip(surfaced: cdn_abr_ladder_drift ≥ 3 segments + top_confidence ≥ 0.75 + cdn-B operator configures the ladder anchor / segment_index backfill) - Tier 2 →
surface_cdn.pop.ladder_segment_index_delta_probe_to_cohort_fanin(surfaced for operator-team approval forward: incident_resolved + cohort-side close-out signals cleared; the learn-loop act) - surfaced → paged the cdn-B operator team + the reliability team on the cdn_ABR package-list drift root-cause review
- close-out signal → cohort-side + ladder-side:
cohort_join_stall_ratio,cohort_rebuffer_ratio,cohort.ladder_rendition_reaggregate_probe,cdn.pop.ladder_segment_index_deltawithin tolerance over the next same-length cohort window — NOTcache.availability/cache.segment_leg_cache_hit/multicdn.winner_RTTreturning to pass
Detection, classify, mitigate, recover (simulated telemetry)
Four timing windows on the postmortem timeline, each read off the cohort probe cadence. The figures are simulated telemetry — the disclosure near the top of this page applies to every figure on this list. Note that the close-out window is verified cohort-side + ladder-side on the apac/teal cohort for the affected window (NOT cache-side, which was green the entire window).
~5 s
cdn.pop.ladder_segment_index_delta crossed 0→6 within a 30 s window on the apac/teal affected cohort across both CDNs; the agent surfaced the detector from the cross-CDN alignment probe at T+5 s on the cohort.
~1 m
cdn_abr_ladder_drift · dominant ranked at 0.81 confidence with three ranked hypotheses at T+1 m — discriminator is cache.availability + cache.segment_leg_cache_hit + license-rollout posture reading pass on the affected cohort.
~7 m
Tier 0 reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill queued at T+2 m; cdn-B operator-side ladder resync configured by T+10 m; cohort re-anchored by T+18 m.
~22 m
cohort.cohort_join_stall_ratio + cohort.cohort_rebuffer_ratio + cohort.ladder_rendition_reaggregate_probe + cdn.pop.ladder_segment_index_delta cleared within tolerance at T+25 m — cohort-side + ladder-side close-out, NOT cache-side re-anchoring.
The loop on this incident
Three steps close the lane on an apac/teal mid-stream CDN failover ABR package-list drift incident. The fix is split explicitly into the authorization-tiered governed-action branches — Tier 0 acts on its own under the replay-origin reanchor gate, Tier 1 surfaces to the cdn-B operator team, Tier 2 surfaces for operator-team sign-off forward on the new probe.
The probe set fans in across the apac/teal affected cohort on both CDNs. cdn.pop.ladder_segment_index_delta reads k=6 (vs 0 baseline); cdn.pop.cross_cdn_alignment_probe reads fail (CDN-A live-edge tail at #EXT-X-MEDIA-SEQUENCE:84217 vs CDN-B replay-origin anchor at #EXT-X-MEDIA-SEQUENCE:84223 — ~36.04 s PDAT drift); cohort.ladder_rendition_reaggregate_probe reads fail across 240p / 540p / 720p / 1080p; the cache leg probes read pass on both CDNs; multicdn.winner_RTT reads 40 / 41 / 42 ms on cdn-a / cdn-b / cdn-c — all three CDNs flat at baseline.
cdn_cache_segment_miss · ruled out by (0.24) and cdn_pop_route_bouncing_under_failover (0.21) alternate.The discriminator against the cache-side lane is the cache.availability + segment-leg cache-hit + license-rollout posture pattern — the cache leg reads pass on the affected cohort across both CDNs, and the multi-CDN health probe proves the CDN leg is healthy at zero egress change; the failure is on the cross-CDN ladder manifest timeline alignment, not cache-locale-shaped and not rtt-side-shaped. The discriminator against the rtt-side lane is the multi-CDN health probe — flat at baseline on cdn-a / cdn-b / cdn-c.
The Tier 0 autonomous branch reissues the ABR ladder via replay-origin reanchor at T_mid plus a CDN-side segment_index backfill — clears the ladder probe within T+30 s and the cohort side within T+18 m. The Tier 1 surface branch configures the cdn-B side ladder anchor / segment_index backfill before cdn-B's manifest tail fully publishes. The Tier 2 forward branch codifies the new probe (cdn.pop.ladder_segment_index_delta + cohort.ladder_ rendition_reaggregate_probe) into the cohort probe fan-in for future mid- stream failover windows. Recovery is verified cohort-side + ladder-side across T+30 s → T+15 m on the apac/teal affected cohort.
Reissue the ABR ladder — or sync the cdn-B anchor.
On an apac/teal mid-stream CDN failover ABR package-list drift incident, the governed fix is a two-arm branch — an agent-side reissue arm that clears the cross-CDN ladder with replay-origin reanchor (Tier 0 autonomous under gate); a cdn-B operator-side ladder resync arm that configures the cdn-B side anchor + segment_index backfill (Tier 1 surfaced to humans); and a forward branch that codifies the new probe into the cohort probe fan-in (Tier 2 surfaced for operator-team sign-off forward). The three arms close the lane in the same incident window and forward.
Emit reissue_abr_ladder_via_replay_origin_reanchor_at_t_mid_plus_cdn_segment_index_backfill — replay-origin reanchors the affected ladder at T_mid and the cdn-side segment_index backfill resets the affected renditions' index against the anchor. The fix is Tier 0 autonomous under gate (top_confidence >= 0.80 AND cdn.pop.ladder_segment_index_delta >= 3 AND multicdn.winner_RTT.pass == true on the affected cohort) and surfaced for operator-team sign-off forward outside that gate so the operator team can review the false-positive rate before the reissue lands.
cdn.pop.ladder_segment_index_delta drops from k=6 → k=2 within the first same-length ladder window (T+30 s staged gate); cohort.cohort_rebuffer_ratio settles ≤ 0.05 within ten consecutive 30 s windows (T+2 m staged gate); cohort.cohort_join_stall_ratio settles ≤ 0.05 within ten consecutive 30 s windows (T+5 m staged gate); cohort.ladder_rendition_reaggregate_probe reads pass on every monitored rendition; cohort.cohort_startup_time_p95_ms reads ahead of the next apac/teal weekday primetime broadcast window opens.
NOT a close-out signal: cache.availability returning to X-Cache: HIT at zero egress change, cache.segment_leg_cache_hit staying pass, or multicdn.winner_RTT reading pass on cdn-a / cdn-b / cdn-c. The cache leg was green the entire window; reading the cache-side probes as the close-out signal misses the lane forward.
Configure the cdn-B side ladder anchor + segment_index backfill before cdn-B's manifest tail fully publishes — the operator team at the cdn-B side configures the ladder anchor / segment_index backfill on the affected cohort so a misanchor that surfaces on a future primetime isn't masked by a "we reconciled" close-out. The fix is surfaced to humans because it requires CDN-side coordination the agent cannot perform — the multi-CDN health probe proves the CDN leg is healthy at zero egress change, but the cdn-B side ladder anchor must be reconfigured by the cdn-B operator team.
cdn.pop.ladder_segment_index_delta settles ≤ 0 within the same window after the Tier 0 reissue + Tier 1 cdn-B resync configuration (T+15 m staged gate); cohort.ladder_rendition_reaggregate_probe reads pass on every monitored rendition (240p / 540p / 720p / 1080p); the cdn-B side ladder anchor ship configured for the next apac/teal weekday primetime broadcast window so a re-incident on the affected scaffold lands already absorbed.
NOT a close-out signal: cache.availability / cache.segment_leg_cache_hit staying pass at zero egress change — those probes are geometric and do not prove the cohort delivers green playback when the ladder is misanchored.
A fix that normalizes cache.availability + cache.segment_leg_cache_hit + multicdn.winner_RTT without clearing the affected cohort's join stall, cohort rebuffer, or cross- CDN ladder alignment probe within tolerance is a fix that didn't reach the cohort. The audit step on this incident writes the close-out signal into the playbook as "verify cdn.pop.ladder_segment_index_delta + cohort.ladder_rendition_ reaggregate_probe + cohort_join_stall_ratio + cohort_rebuffer_ratio settle within tolerance over the next same-length cohort window" — not "verify the cache leg returning to green / verify cache.availability + segment-leg cache-hit returning to pass".
Codifying the new probe into the cohort fan-in
The Tier 2 surface for operator-team sign-off forward on this incident codifies two probes into the cohort probe fan-in — cdn.pop.ladder_segment_index_delta pinned into the cohort fan-in for the next weekday primetime broadcast window (operator-approved; surfaces on every mid-stream failover cohort); cohort.ladder_rendition_reaggregate_probe pinned for the next four weekday primetime windows so the false-positive rate can be measured before promotion.
The audit step codifies that cdn.pop.ladder_segment_index_delta +cohort.ladder_rendition_reaggregate_probe settle within tolerance over the next same-length cohort window — meaning the cross-CDN ladder probe fan-in tracks across every mid-stream failover window, the operator-team sign-off goes on the new probe codification, and the cohort fan-in carries the ladder-mismatch signal lane forward. The cache.availability / segment-leg cache-hit / multi-CDN health probe readings stay pass through the entire window — those are geometric, NOT cohort-side close-out signals.
Want Streamwake to disambiguate ABR ladder drift on your cohort?
Sign up, register a multi-CDN probe, and the same cdn.pop.ladder_segment_index_delta · cdn.pop.cross_cdn_alignment_probe · cohort.ladder_rendition_reaggregate_probe · cohort.cohort_join_stall_ratio probes that produced the timeline above run on every prime-cohort refresh — and surface in a Slack channel, a webhook, or the streams dashboard.
Synthetic Incident — This scenario uses simulated telemetry constructed from documented streaming behaviors. It does not represent a Streamwake customer outage.
- Pick a recent on-call incident — manifest stall, edge miss, player-side stall, or peer congestion.
- We replay it through the same reliability-agent probe cascade used on the postmortem above.
- You walk away with a written what-could-have-been-Automated readout, not a sales deck.
Related writeups
The closest siblings cover the player-visible symptom lands: the broader ISP-vs-CDN triage guide at /troubleshooting/isp-congestion-vs-cdn-failure (the eight-symptom read / ISP-vs-CDN discriminator), the ISP-vs-CDN recovery criteria entry (cdn_failure · ruled out by name + cohort-side + last-mile close-out), the origin-shield queue saturation write-up (the shield queue p99 waiting past the cohort's expected join-window profile), the manifest fetch timeout storm at a regional edge POP (a regional edge POP returning manifest-timeouts above baseline during a quiet pre-peak window), and the live-event scale-out buffering postmortem (a marquee broadcast with a viewer-spike that pushes the cohort beyond the pre-provisioned capacity envelope). Together they cover the four failure-mode lanes Streamwake reliability agents are tuned for alongside ABR ladder drift on mid-stream CDN failover.