Dimension Labs/Live Monitor · Causal Intelligence
Refreshes weekly · Lemonade omni-channel

Omni-Channel Voice Monitor

Lemonade · week of May 12, 2026 · 4 review surfaces + 623 conversation sessions · every week-over-week move tested for significance before it earns an arrow
Tested: 0 of 6 weekly moves are real
Read this first

Last week's board flagged four metrics as "rising." Run each move against its own sampling error and none of them clears the noise band this week — the arrows are directional, not real change. That doesn't mean nothing is wrong. Two things survive the test, and neither is a weekly wiggle: a 12-month climb in AI-failure language, and a large, stable causal link that turns failed chats into one-star reviews. This board separates the two so you act on the mechanism, not the noise.

Dead-end rate
18%
○ wk change: noise
of "contained" chats were actually dead-ends — the customer gave up or never got the human they asked for
Δ wk/wk +1.5pp · 95% CI −3.0 to +6.0 (crosses 0). Level is severe; the week's move isn't.
Escalation gap
61%
○ wk change: noise
of customers who clearly needed a person never reached one, in the chat logs
Δ wk/wk +2.1pp · 95% CI −8.2 to +12.3. A 61% level is the story — not this week's drift.
Genuine resolution
70%
○ no real change
of chats were genuinely resolved — against the 88% the bot reports as "contained"
Δ wk/wk −1.0pp · 95% CI −6.1 to +4.1. Flat. The honest containment number, holding steady.
Chat → 1-star risk
82%
●● causal · stable
of failed or abandoned chats already carry a one-star review signal
Not a trend — a mechanism. Flat within ±3pp for 12 months. This is the causal parameter, below.
One-star share
31%
○ wk change: noise
of all public reviews this period (n=168)
Δ wk/wk +3.9pp · 95% CI −6.0 to +13.7. Small review base — this jump is inside the noise.
AI-failure mentions
11%
● trend: real (12-mo)
of reviews name a specific AI failure; "refused to escalate" leads
Wk/wk is noise (+1.7pp, CI −4.8 to +8.2), but +4pp since Jan '25 · trend p<0.05. Real climb.

What actually moved this week — nothing you can detect.

Each metric's week-over-week change, with its 95% confidence interval. When an interval crosses the zero line, the "move" is indistinguishable from sampling noise. Two-proportion z-tests · this week vs last · n=623 sessions, 168 reviews.
≈ SAMPLING NOISE 0 — no change −10pp +10pp Dead-end rate n.s. ○ Escalation gap n.s. ○ Genuine resolution n.s. ○ One-star share n.s. ○ AI-failure mentions n.s. ○ Every interval crosses zero. Not one weekly change is statistically detectable.
●● real · p<0.01 real · p<0.05 within noise · n.s.Test: two-proportion z · α=0.05

Alerts — reframed

Don't act on the four "rising" arrows.
Dead-end, escalation, one-star share, and AI-failure all drifted up, but each move sits inside its confidence band. Chasing them means chasing sampling error.
→ Suppress weekly-delta alarms until a move clears its CI.
The one real driver is structural, not weekly.
A failed chat is 6.9× more likely to become a one-star review (95% CI 5.3–8.9). That link has held for 12 months. It's the mechanism manufacturing your public score.
→ Manage the mechanism: cut the failed-chat inflow.
The 18-pt containment gap IS real.
88% "contained" vs 70% resolved. That gap's CI (13.6–22.4pp) excludes zero — a genuine structural defect, unlike the weekly wiggles.
→ Split containment; alarm on the gap, not the headline.
The backbone — what's real is a mechanism, not a move

A failed chat is 6.9× more likely to become a one-star review.

The causal parameter behind the whole board. Would-1-star rate for resolved vs failed/abandoned chats, this period, with 95% confidence intervals. The two intervals don't come close to touching — this separation is not noise. Risk ratio with Katz CI · n=623 sessions.
0%50%100% SHARE OF CHATS THAT CARRY A ONE-STAR SIGNAL Resolved chat 12% Failed / dead-ended chat 82%
6.9×
more likely to become a one-star review · 95% CI 5.3–8.9
p < 0.001

This is the sentence the rest of the board is footnotes to. It doesn't move week to week because it isn't a trend — it's how the system converts a broken conversation into a public review. You don't fix a 6.9× risk ratio by watching it; you fix it by shrinking the thing on the wrong side of it: the failed-chat inflow.

Real structure vs weekly noise — the containment gap

The gap between "contained" and "resolved" is the real change. The weekly line isn't.

Containment (bot self-report) vs genuine resolution (annotated), 16 months. The shaded band is the dead-end gap. Its width is stable and its 95% CI (13.6–22.4pp) excludes zero: this is a structural defect, not sampling drift. The week-to-week jitter along each line is noise — the level difference between them is signal.
60%80%100% 88% "contained" 70% resolved the 18-pt dead-end gap · 95% CI 13.6–22.4pp · excludes 0 Jan'25AprJulOctJan'26Apr
18pt
dead-end gap now · real (CI excludes 0)
±4.4pp
weekly jitter along each line · noise
<5%
target gap · needs the relational-intent fix

Read the vertical distance, not the horizontal wiggle. Containment alone says everything is fine; the honest number is the resolution line, and the gap between them is the defect. When the gap widens past its band, the one-star tail grows the next week — but a one-week bump in either line, on its own, tells you nothing.

Counterfactual — what closing the gap actually buys

Route relational intents to a person and one-star share falls to ~20%.

Modeled counterfactual: hold volume constant, send complaints, disputes, cancellations, and stalled claim-status chats to a human on first contact — removing ~72% of dead-ends. The one-star inflow that feeds the tail shrinks accordingly. Dashed = do-nothing projection; solid = with first-contact handoff. Shaded gap = one-star reviews prevented per week.
35% 25% 15% first-contact handoff now · 31% do nothing · ~32% with handoff · ~20% ≈11pp prevented Past 12 weeks (actual) · projection with 95% band ±4pp on the endpoint
~19
of this week's 52 one-star reviews trace to dead-ended chats · 95% CI 13–25
31% → 20%
projected one-star share with the handoff · CI 16–24%
~72%
of dead-ends removed by routing 4 relational intents to a person

This is the payoff of separating signal from noise: the intervention isn't "improve the metrics," it's one specific routing change on four intents, and the causal link tells you the size of the prize before you spend a dollar. The band on the projection is wide — treat ~20% as the center of a range, not a promise — but the direction and mechanism are not in doubt.

Where the mechanism fires — intent, with small-sample honesty

The bot dead-ends on the relational intents — but trust the dead-end rates by sample size.

Every conversation classified by what the customer came for; red = dead-ended. A 100% dead-end rate on 70 complaints is a firm finding; the same rate on 13 cancellations is directionally damning but sits on a wide interval. Read the reliability column, not just the bar.
WHAT CUSTOMERS BRING TO THE AI 623 sessions · bar = volume, red = dead-ended · right column = how much to trust the rate New quote File a claim (FNOL) Billing & payment Coverage question Complaint / dispute Claim status chase Renewal Cancellation n=150 n=110 n=104 n=89 n=70 n=52 n=35 n=13 DEAD-END RATE (95% CI) 0% · handled 0% · handled 0% · handled 15% (8–24) · firm 100% (95–100) · firm 35% (23–49) · firm 0% · handled 100% (77–100) · thin n Handled by the bot Dead-ended / needed a human

Quotes, first claim notices, billing, and renewals are handled end to end. Complaints and claim-status chases dead-end on firm samples — those are safe to act on now. Cancellations also dead-end 100% of the time, but on 13 sessions the interval runs 77–100%: directionally alarming, worth confirming next week before you build a metric around it. Same lens, applied to the queue: act on the firm findings, watch the thin ones.

Recommended actions this week
ACTION 01

Act on the mechanism, not the arrows.

Route complaints, disputes, cancellations, and stalled claim-status chats to a human on first contact. This is the one intervention the causal link says will move the public score — worth ~11pp of one-star share.

Moves: failed-chat inflow → one-star share
ACTION 02

Silence the weekly-delta alarms.

Four "rising" flags this week are all inside their confidence bands. Gate every alert on clearing its CI, so the team stops spending cycles chasing sampling noise and reacts only to real change.

Moves: alert precision, team focus
ACTION 03

Alarm on the gap and the 12-mo trend.

Two things are real: the 18-pt containment gap (CI excludes 0) and the AI-failure climb since Jan '25 (trend p<0.05). Put both on standing alert — they're the signals that survive the test.

Moves: containment gap, AI-failure trend

On the watchlist — signal, not noise

Dimension Labs · Omni-Channel Voice Monitor · causal layer: two-proportion z-tests (wk/wk), Katz-CI risk ratios, α=0.05 · n=623 sessions, 168 reviewsGenerated June 2026