The Edwin's guest scorecard is genuinely strong. This brief goes underneath it, with real models, two years of history, and the guests' own words, to show what actually moves the recommend score, why the highest-paying guests are the least satisfied, and the seven-figure revenue riding on four cheap fixes.
The left is the official story; the right is the operating story living inside the guests' own words.
Structured scores are the full survey sample; the signal counts are from the comment-bearing responses.
The Edwin sits in the hardest place for a survey to help an owner: the scores are already high, so the dashboard mostly confirms the property is good. It cannot tell you which high-value stay is quietly at risk, which safety issue is real, or where a guest is trying to hand you more money.
That gap is not academic at a boutique commanding luxury rates. The scorecard is not wrong. It is incomplete. The rest of this brief is what the words, the rates, and two years of history add to it.
“After posting a Google review I was charged an additional $100, on top of $930 for a 2 night stay... I plan to move my loyalty to Hyatt.”1★ survey guest·churn + billing + reputation in one stay
One model of the recommend score on all four sub-scores at once, standardized so the bars are directly comparable.
OLS, recommend score on standardized sub-scores, n=842, R²=0.67, 95% CIs shown. Observational.
This is the single most important correction to the scorecard reflex. A dashboard flags food & beverage because it is the lowest number and sends leadership to the kitchen. The data says the kitchen is a modest driver of whether a guest recommends The Edwin. The warmth and competence of the staff is the dominant one, by a wide and statistically strong margin.
That is good news: staff is the most coachable asset in the building, and The Edwin already has the raw material. Guests name their people, unprompted.
“The receptionist, Alden I think was his name, was very welcoming and kind. Great stay overall.”5★ survey guest
Left: issues by how often guests raise them. Right: the same issues by how likely each is to actually cost you a guest, weighted by what they pay. The order nearly inverts.
Retention-risk intensity, indexed to 100: the churn-signal rate among guests raising each issue, weighted by their nightly rate. Illustrative, and deliberately independent of how often an issue is mentioned.
This is the reframe at the center of the brief. A satisfaction dashboard, and a survey report, rank issues by volume, by how often a topic comes up. That points an operator at food, the valet stand, and the front desk, the loudest moments, most of which are minor or even praised. Re-ranked by how likely each issue is to actually make a guest leave, weighted for the rate that guest pays, the list nearly inverts.
The quiet issues are the expensive ones. A guest whose elite status is ignored, or who finds a charge that should not be there, or whose room is not right at a luxury rate, rarely writes a long review, but is far likelier to take their business, and their repeat revenue, elsewhere. Loudness is not cost. Fund the fixes that keep the guest, not just the ones that make noise.
“I'm a Lifetime Titanium elite who should get upgraded automatically... you assigned me one of the worst rooms.”2★ Titanium member · the rarest complaint, the highest churn
“After posting a Google review I was charged an additional $100... I plan to move my loyalty to Hyatt.”1★ survey guest · a quiet billing dispute that became churn
Each bar is the share of all guests dissatisfied that quarter; the brass line is the same for guests paying $350+ a night. The gap, and its recent peaks, is the story.
Quarterly, 2023 Q1 to 2025 Q3. The final quarter is a partial, smaller sample. High-rate line is guests paying $350+.
A scorecard reports this quarter. An owner needs the direction of travel. Across two and a half years, The Edwin's overall dissatisfaction has not trended down, and its high-rate guests move in sharper swings, worst in the busy summer quarters when the property is fullest and the service engine is most stretched. That is the operational tell: the cracks in section three open widest exactly when volume is highest.
The recent quarters are the ones to watch. This is a live pattern, not a historical one, which is why the companion monitor tracks it weekly.
“We've stayed at this hotel a few times and have decided it's just not worth the extra money.”2★ repeat guest
The share dissatisfied, grouped by rate paid; the controlled slope beneath holds the confounders constant.
Linear probability model with controls for channel, loyalty tier, trip type; n=944, robust SE. Bars are raw rates.
This is where the meaning layer and the money meet. A satisfaction score treats every stay the same. But a disappointed guest who paid $220 is a bad night; a disappointed guest who paid $650 is lost revenue and a public review with a luxury price attached. The paradox is not that the expensive rooms are worse. It is that at a luxury rate, the small failures stop being forgivable, and section six is what those failures actually are.
Each is low-volume but high-severity: when a guest raises it, the stay usually scored poorly.
Comment-bearing responses since 2024. Severity = share of each subtopic's mentions carrying a low score.
Every one of these is an own-goal at a luxury price. A guest paying four figures a night does not expect to be charged $4 to bring hot gravy, served cold biscuits, or told the free breakfast promised as compensation never materialized. Room service that never arrives, and a Lifetime Titanium member handed the worst room in the house, are not taste, they are process, and fixing them is cheaper than the reviews they produce.
“If you're paying $1,000/night it would be nice if you had a complimentary breakfast or coffee downstairs embedded with your stay.”2★ guest, Riverview suite
“Primary issue was room service, which never produced our meals.”2★ survey guest
The food & beverage score is the lowest on the card, but dining behavior is one of the strongest positive forces in the data.
Recommend-score model with dining as treatment; controls for rate, channel, tier; n=979. Dose: +0.16 per additional outlet.
This is the nuance a scorecard destroys. Read only the number, and food & beverage looks like the problem. Read the behavior, and the rooftop and restaurant are a reason people come and a reason they recommend. The outlets themselves are loved; the narrow failures are breakfast value and room-service reliability. Protect the flywheel, fix the two leaks.
“Wonderful hotel. Professional and kind staff. Excellent bar.”5★ survey guest
Directional figures meant to size the stakes, not forecast them. Every input is shown below.
Room revenue base assumes 90 rooms, ~70% occupancy, $287 ADR (from data). OTA share 70% (from data); OTA commission ~15% (industry benchmark). Not a forecast.
Owners do not manage a recommend score; they manage a P&L. Translating the two most robust findings into money makes the case concrete. The channel mix is a margin story: seven in ten stays arrive through OTAs, and shifting even ten points of that volume to direct booking is worth roughly $99K a year, work the public-review reputation directly drives. The rate paradox is a revenue-quality story: nearly half a million dollars a year in room revenue is collected from guests who leave dissatisfied at the top of the rate card, the single most expensive guest to disappoint.
These are illustrative, not audited, and deliberately conservative. The point is the order of magnitude: the fixes in this brief are cheap, and the exposure they address is measured in seven figures.
Dissatisfaction climbs with the rate, and the $500+ guest is the costliest to lose. Build a high-rate readiness standard around room condition, hot water, and a flawless arrival, and staff up in the busy quarters when the trend shows strain.
A warm staff interaction is the single strongest driver of a recommend, and it already happens at the valet stand. Make it a coached, named, replicable standard at the front desk and elite check-in too.
Disputed charges, thin breakfast, unreliable room service, and unrecognized elites are rare, severe, and cheap to fix. Each is a process, not a taste. Fix them before they become the public review.
Seven in ten stays pay an OTA commission. The public reviews are the lever that moves the next guest to book direct, so review-response and public recovery are margin work worth roughly $99K a year per ten points shifted.
The library is read through five enrichment lenses, satisfaction, critical signals, friction, delight, and a plain-language summary, and then joined to the structured survey data: the scores, the room rate, the booking channel, the loyalty tier, and the dates.
From there the process is three steps. First, enrich: every comment becomes dozens of structured fields where there was only free text. Second, classify: friction and delight are sorted into the operating categories above, so a complaint lands in the moment a manager owns. Third, model: the dimensions are analysed against the scores and rates with controls, so we report an effect and its uncertainty rather than an anecdote. That is the difference between a dashboard, which trends topics and shows verbatims, and an analysis, which tells you which signal is a churn risk, which is a safety issue, and which is revenue, and how much each actually moves the outcomes you manage.
Window. Findings use responses from the start of 2024 onward (1,513 responses: 979 in-house surveys and 534 public reviews; ~1,041 with comment text). The two-year trend extends to 2023 to show direction of travel and is labelled as such. Older responses are otherwise excluded.
Two layers, kept separate. Structured scores (recommend score, 1–10 sub-scores, star ratings, room rate, booking channel) are facts from the survey layer. Signal, friction, and subtopic reads come from the comment text. The two are never blended into one number.
The models. The driver decomposition is an OLS of the recommend score on four standardized sub-scores (n=842, R²=0.67). The rate effect is a linear probability model of dissatisfaction with controls for channel, loyalty tier, and trip type (n=944, robust SE). The dining lift is a recommend-score model with dining as treatment and rate, channel, tier as controls (n=979). Effects carry 95% confidence intervals and are observational, strong directional evidence, not a randomized trial.
Financial figures are illustrative. They combine the property's own rates and channel mix with published benchmarks (boutique occupancy; typical OTA commission of ~15%; a Cornell review-elasticity relationship) and are labelled as directional sizing, not a forecast or an audited figure. No revenue total is presented as fact. Verbatims are exact guest quotes. Public facts about the property, the Autograph brand, and Vision Hospitality Group come only from the cited external-context file and are never combined with a review figure.