An analysis of 56,788 post-game survey responses across 81 games, decomposing a 14-point satisfaction gap into the operational frictions that drive it.
The Mariners measure fan experience after every home game: a 0–10 rating plus open-text feedback. This data is joined with information about the ticket holder and game-day events, including weather. Across the 2025 season we pointed the AI agent at a simple question — does weather meaningfully affect how fans rate their experience at the park?
The answer came back clear. Fans rate their experience measurably and significantly lower when the roof is closed. The gap is real, it's consistent, and it survives every statistical check we ran against it. But "it rains in Seattle" isn't an actionable finding — so we asked the agent to keep going.
From here the agent went deeper. It cross-referenced every fan rating against the full game-day record — weather, attendance, day of week, promotions on the calendar, entry gate, seat section — looking for the condition that actually moved the dial. It tested and retested, ruling out what didn't matter until a pattern emerged that did.
The finding: roof closure by itself barely moves the rating. A roof-closed game without a giveaway rates about the same as a cloudy game — a little worse than sunny, but nothing fans describe as meaningfully different. It's only when a roof-closed day coincides with a giveaway that fan experience drops below the worst weather the sky can produce. Something operational is happening on those specific game days that doesn't happen on any other.
The structured data told us when and how much. To understand what and why, the agent turned to the thousands of open-text comments fans wrote about their experience. Read at scale, a clear story emerges.
That's the actionable insight: roof-closed games with giveaways aren't a weather problem to weather — they're a distinct operational mode that calls for a distinct operational playbook. The rest of this case study lays out the evidence, the mechanism, and the specific plays.
Five findings build the case. Together they move from evidence (the penalty is real and measurable) to where it lands (it spreads across every operational domain) to the operating mode (four specific frictions describe it) to the trigger (one scheduling decision amplifies it) to who absorbs it (the fans the organization most needs to protect).
Each card links to its section. Every finding is backed by the same fan-voice classification layer, anchored in fans' own words.
Before isolating the mechanism, we needed to confirm the signal. Across every weather condition, fan experience moves in a clean gradient — and roof-closed games sit clearly at the bottom. The gap from sunny to roof-closed is statistically robust, holds up across weekday and weekend subsets, and survives every sensitivity check. This is the evidence that something is happening. The next finding starts to explain what.
Each dot represents one weather condition's mean rating; the roof-closed dot sits 0.62 points below sunny.
The difference between roof-closed and sunny is the largest pairwise gap in the data and is statistically robust across every subset tested — weekday-only, weekend-only, and excluding responses with missing weather all reproduce the pattern. A separate cross-check using top-box rate (a simpler, distribution-free measure) confirms the finding at −14.1 pp.
At the scale of a full season, this means hundreds of additional fans per roof-closed game shift out of the highest satisfaction tier. While the effect is modest for any individual fan, the cumulative impact across ~6,600 roof-closed survey responses represents a meaningful operational signal.
If the gap came from one failed system — a broken concession line, an understaffed gate — we'd see one domain drop while the rest held steady. That's not what the data shows. Every operational domain takes a hit under roof closure, which tells us the issue isn't a single failure point. It's the way the day itself runs differently when the roof is down. The pattern points us toward investigating the conditions of the day, not the performance of any one system.
Each row compares the same domain's mean rating across two conditions; the gap is the roof-closed penalty for that domain.
Beyond mean scores, the underlying fan sentiment shifts across all three annotated operational domains. Negative mentions rise while positive mentions decline, indicating that the penalty reflects an active increase in dissatisfaction — not merely a reduction in enthusiasm. In concessions, negative mentions rise from 13.3% to 16.4% of responses under roof closure, while positive mentions of food quality drop from 5.7% to 5.0%.
A diffuse penalty raises the central analytical question: if the effect is everywhere, what's actually broken? Aggregate rating comparisons don't answer that. The next section decomposes the penalty by issue aspect — the specific operational failures fans mention in the open-text fields.
Across thousands of fan comments, four operational frictions appear far more often on roof-closed days than on sunny ones. Read individually, they look like separate problems. Read together, they describe a single shift: the stadium operates in a more constrained mode under roof closure, and the same four friction points — concessions throughput, giveaway distribution, concourse flow, service recovery — are the surfaces where that constraint shows up to fans.
When the roof is closed and something goes wrong at the game, fans who raise the issue rate their overall experience 3.3 points lower than roof-closed fans with no issue flagged. Staff interactions — issue resolution, ushers, concession cashiers — all show sharply negative within-roof-closed gaps.
Operational reading: when density and stress rise, single-point-of-failure service breakdowns feel worse and cascade. Recovery capacity matters most when the roof is closed.
In-Stock / Sold-Out mentions nearly double (1.2% → 1.9% of responses), and Lines & Speed complaints follow a similar pattern. Each roof-closed game surfaces ~125 fans voicing stock-out complaints specifically — a high per-fan rate on a smaller denominator.
Operational reading: roof closure shifts dwell patterns to covered concourses and spikes simultaneous demand — exposing inventory planning gaps and throughput bottlenecks.
Giveaways & Swag mentions rise from 1.2% of responses on sunny days to 3.3% under roof closure — the largest single-aspect lift in the data. Notably, the raw complaint count is nearly identical (218 sunny vs 222 roof-closed), but those complaints came from a fan pool one-third the size.
Operational reading: roof-closed games concentrate fans indoors and in covered zones, increasing congestion at pickup points — amplifying per-fan giveaway friction even when absolute complaint volume looks flat.
Gate Lines & Wait Time is among the most prevalent flagged issues in roof-closed responses (n = 263), and Concourse Crowd Flow shows a clean negative within-roof-closed gap. These aspects describe the same thing from different angles: covered routing plus weather-driven clustering creates pinch points.
Operational reading: covered concourse routing under roof-closed conditions doesn't disperse fans the way open-air pathways do. Active crowd management and wayfinding become disproportionately valuable.
Every issue aspect that met the sample-size gate (n ≥ 30 in both roof-closed and sunny), ranked by three different lenses. Each lens tells a different part of the story.
Each bar is one issue aspect. Color denotes domain. Sorting by "within-roof diff" surfaces the aspects most strongly associated with bad ratings among roof-closed fans; sorting by "incremental penalty" surfaces the aspects whose roof-closed penalty is largest relative to the baseline gap; sorting by "prevalence" surfaces the aspects most frequently mentioned.
So far the analysis has treated all roof-closed games as one bucket. They aren't. One game-level condition — whether the day features a giveaway — concentrates almost all of the observed penalty.
Roof-closed games without a giveaway barely register a fan-experience gap — a small dip, well within tolerance. Add a giveaway to the same day and the penalty grows nearly fourfold. The combination doesn't just add the two effects together; it triggers a categorically different fan experience. And critically: the trigger is on the calendar, not in the forecast. Weather isn't actionable. Scheduling is.
Each bar is one subset of roof-closed games. The top bar restores the sunny baseline for reference. The vertical line at 8.85 marks the sunny baseline; distance from that line is the roof-closed penalty for each subset.
Within giveaway-present games under roof closure, concessions issues roughly double compared to the overall sample. In-Stock / Sold-Out complaints rise from 1.2% to 2.2% of responses (+92%), and Lines & Speed from 1.2% to 2.2% (+86%). These rates are approximately double the overall issue lifts, confirming that giveaway logistics and roof closure interact to create compounding strain.
| Issue Aspect | Rate (Sunny) | Rate (Roof Closed + Giveaway) | Absolute Lift | Relative Lift |
|---|---|---|---|---|
| In-Stock / Sold-Out | 1.17% | 2.24% | +1.08 pp | +92% |
| Lines & Speed | 1.18% | 2.20% | +1.02 pp | +86% |
| Food Quality | 1.01% | 1.74% | +0.73 pp | +73% |
| Ordering & Payment | 0.41% | 0.50% | +0.09 pp | +22% |
Rate change, not relative lift, is the correct frame here. All rates are low in absolute terms; the relative lift metric inflates small-denominator shifts. Focus on the absolute lift column for operational prioritization.
The penalty isn't distributed evenly across the fanbase. It lands hardest on three groups: fans seated in upper-deck sections, fans with the longest tenure as Mariners supporters, and families attending with children. Each of these segments matters disproportionately to long-term loyalty — long-tenure fans are the base, families are the future, and upper-deck fans are the volume that fills the park. A roof-closed giveaway day damages experience most for the fans the organization can least afford to lose.
Each bar is the roof-closed-vs-sunny gap for that fan segment. Bars are grouped by dimension; the gap is the change in mean overall rating between the two conditions within that segment.
400-level (upper-deck) fans experience the largest penalty (−0.77), followed by the 200-level (−0.65) and 300-level (−0.60). Suite and premium seating shows the smallest gap (−0.38), consistent with climate-controlled amenities and dedicated concessions buffering the weather effect. The gradient is monotonic: higher-elevation seating corresponds to a larger penalty.
Long-tenured fans (10+ years: −0.71) and mid-tenure fans (2–3 years: −0.68) are more weather-sensitive than the overall sample — the opposite of what a "loyalty insulates" hypothesis would predict. Fans attending with children show an elevated penalty (−0.71), likely reflecting the compounding friction of navigating crowded concourses with young family members.
The findings point to a single conclusion: roof-closed games with giveaways aren't bad weather days that the organization has to endure. They're a distinct operating mode that the organization can prepare for, staff for, and execute against differently. Below is the operational playbook the analysis points toward — issue by issue, mapped to the team that owns it. The recommended actions target the specific frictions fans named in their own words.
| Domain | Issue | Recommended Action | Evidence |
|---|---|---|---|
| Concessions / Merch | In-Stock / Sold-Out | Pre-stock high-demand items in covered concourse zones; add live sold-out signage plus POS/menu updates; align promotional comms with actual inventory. | Strong |
| Fan Engagement | Giveaways & Swag | Return to entry-point distribution for high-demand items; per-gate allocation; clear pickup instructions; avoid end-of-game bottlenecks. | Strong |
| Concessions / Merch | Food Quality | Increase QA checks during roof-closed games; focus on hot-hold times and replenishment cadence. | Moderate |
| Concessions / Merch | Lines & Speed | Dynamic labor reallocation to covered stands; open supplemental grab-and-go points; add wayfinding to less-crowded stands. | Moderate |
| Staff / Service | Issue Resolution | Explicitly staff "fix-it" roles during roof-closed games (guest services, supervisors); improve wayfinding and staff knowledge scripts for common questions. | Strong |
| Venue | Concourse Crowd Flow | Additional wayfinding and marshaling for covered pinch points; active line management at gate screening and high-density concession stands. | Moderate |
| Venue | Restrooms | Increase restroom staffing and cleaning cadence during roof-closed games where indoor density increases. | Directional |
Three additional action modules complement the playbook:
1. Pre-game experience risk tiering. Weather × month × day-type combinations from 2025 are classified as High, Medium, or Low risk based on observed mean satisfaction. When a roof closure is anticipated on a giveaway day, the full weather-hardened playbook activates automatically.
2. Targeted fan communications. For roof-closed + giveaway games, proactively message earlier arrival windows, giveaway pickup instructions (gate / section + cutoff times), and concessions wayfinding to reduce covered concourse congestion.
3. 2026 measurement framework. Track overall experience rating by weather condition weekly against the 2025 baseline. Monitor issue-aspect rates under roof closure (especially stock-outs and lines). Use top-box rate as a simple scorecard.
What makes this work different is the way the agent is deployed: not as a one-shot summarizer or a chatbot over the data, but as the engine of a full analytical pipeline. It reads fan text in the language fans used, joins that reading to the structured record of the game, tests what's moving together, and returns to the text to explain the movement in fans' own words.
Classify each open-text response across an eight-domain issue-aspect taxonomy, labeling what fans talked about and how they felt about it.
Link classifications to the game-level record — weather, giveaway scheduling, attendance, day of week — and to fan-level attributes like seating, tenure, and party.
Run statistical comparisons across conditions, apply sample-size gates, and surface effect sizes and directions that warrant closer investigation.
For each meaningful difference, drill back into the unstructured text to surface the specific language and representative verbatims that explain the gap.
The pipeline ran against 56,788 post-game survey responses across 81 games in the 2025 Mariners season. Primary statistical tests include the Welch two-sample t-test on overall rating (roof-closed vs sunny), one-way ANOVA across five weather buckets, and per-aspect effect-size rankings with Benjamini–Hochberg FDR correction on the full driver table.
Every chart in this case study — the weather gradient, the domain dumbbell, the driver rankings, the giveaway moderation chart, and the fan-level sensitivity chart — was produced from the same joined dataset. The same pipeline runs weekly against incoming survey data and can be re-pointed to any operational question with a similar structure.