Dimension Labs · Causal Brief
Multi-source customer reviews · 5,380 distinct reviews · Jan 2024–Jun 2026
The one thing to know
The stock is up a hundredfold and profit hit a record. But read what customers actually write and a second story appears: buying a car is easy and people love it, then owning it goes wrong. The part that keeps breaking is the same one that nearly sank the company in 2022.
Wall Street won
~100×
The stock climbed from about $3.55 in 2022 to above $460, and Carvana rejoined the S&P 500 in December 2025 on record revenue of $20.3 billion.
Customers didn’t
Half
More than half of the 5,380 reviews are one or two stars. The average review scores just 2.83 out of 5.
Where it breaks
9 in 10
Nine in ten people who write about actually owning the car are unhappy, versus about one in seven of the people writing about shopping for one.
Follow it left to right: the same pile of reviews splits into a happy shopping crowd and an unhappy owning crowd, and both run into the same wall right after the sale.
The reframe · what a survey can’t tell you
Feed the same 5,183 reviews to a traditional customer-experience tool and it counts mentions. Refunds and title paperwork look like the fire, because they are said the most and they draw the regulatory fines. Rank the same five after-sale problems by what actually pushes a review to one star, holding the review site, the timing, and the length fixed, and the list turns over. The two loudest complaints barely move the rating; two of the quietest do most of the damage.
Nine in ten reviews are either one or two stars or a perfect five. Only 9% land anywhere in the middle. The 2.83-star average everyone quotes is the midpoint of two crowds that never meet, so describing Carvana by that number is like describing a coin as landing on its edge. The real question is not the average. It is what flips a customer from one pole to the other.
Count how often each place turns up a one- or two-star review. On Google Play it is about one in four. On Trustpilot and the App Store, about one in three. On the Better Business Bureau it is 97 out of 100, and on Reddit 96. That gap is the whole story. The app stores mostly catch the easy part, shopping and selling. The complaint sites catch what happens when something goes wrong after the sale and the customer needs someone to help. It is the same company met at two different moments, and the moment decides how people feel.
Each bar is how often that channel’s reviews are one or two stars. The two complaint sites sit near the top of the scale; the three shopping and app channels sit far below. This is not one unhappy crowd, it is two different crowds.
| Channel | Reviews | Unhappy rate | Avg. star |
|---|---|---|---|
| Trustpilot | 1,788 | 34.3% | 3.46 |
| BBB | 1,629 | 97.4% | 1.12 |
| Google Play | 1,194 | 23.2% | 3.98 |
| App Store | 567 | 34.6% | 3.57 |
| 202 | 96.0% | 2.30 |
This is not one harsh channel. A one-star Trustpilot review and a Better Business Bureau complaint are different groups of people, caught at different moments. Hold the moment constant and the split still holds, so the 97%-versus-23% gap is the finding, not an artifact of where people happen to post.
Better Business Bureau: 97.4% unhappy (1,629 reviews). Google Play: 23.2% (1,194 reviews). A 74-point gap between the same company’s two audiences, p < 0.001 by a chi-square test.
Take the turnaround seriously before reading the complaints. Three years ago Carvana was nearly bankrupt. Today it is one of the sharpest recoveries in recent retail history. That is exactly why the reviews matter: the whole recovery is riding on the one thing customers say still breaks.
In late 2022 Carvana was close to bankruptcy: the stock had fallen to around $3.55, debt was about $7 billion, and the company lost $1.59 billion that year. Then it turned. In 2025 it sold a record 596,641 cars, revenue reached $20.3 billion, and it made a record $1.9 billion in profit. The stock climbed back above $460, roughly a hundred times its low, and Carvana rejoined the S&P 500. Investors have decided the comeback is here to stay. The only real question left is whether the machine that actually delivers and services the cars can keep up with the growth.
Annual net income from the public filings, with revenue and unit growth labeled. The turnaround is genuine and large. The callout tracks the parallel stock recovery and the return to the S&P 500.
| Year | Net income | Revenue | Retail units |
|---|---|---|---|
| FY2022 | -$1.59B | not reported | not reported |
| FY2024 | +$0.404B | $13.67B | 416,348 |
| FY2025 | +$1.895B | $20.322B | 596,641 |
The reviews split the moment money changes hands: the people still shopping are happy, and the people who now own the car are not. And among the failures that follow the sale, the loudest complaint turns out not to be the one that does the most damage.
On Trustpilot, sellers are unhappy about 1 time in 10; buyers, about 8 times in 10. Across every site, shopping reviews run unhappy roughly 1 in 6, ownership reviews about 9 in 10. What changes is not who is writing, but how far past the sale they are. A love-it-or-hate-it brand would spread the anger evenly. Carvana’s lands almost entirely after the money changes hands.
The complaints name what goes wrong, and the same handful come up again and again: a car that was late, a car that arrived broken or dirty, a charge or financing problem, a title or plates that never showed up, and a refund that never came back. Each of these turns up in hundreds of reviews, and when one of them appears, the review is unhappy nine times out of ten.
Each line connects an easy moment (left dot) with the harder moment that comes after the sale (right dot). The gap is the story: the same people, much unhappier once they own the car, are waiting on a refund, or need support. Each pair compares like with like, not one site’s average against another’s.
| Comparison | Easy-stage rate | Hard-stage rate |
|---|---|---|
| Trustpilot: selling vs. buying | 10.9% | 76.7% |
| Trustpilot: pickup vs. ownership | 8.3% | 85.1% |
| All reviews: browsing vs. ownership | 15.5% | 90.0% |
| All reviews: normal vs. refund delayed | 53.3% | 97.0% |
Each bar is how often that problem leads to an unhappy review. The dashed line is the average across all reviews, 53%. Every one of these clears it by a wide margin, and the label shows how many reviews mention it and how much more often than usual it ends badly.
| Issue | Unhappy rate | Vs. normal | Reviews |
|---|---|---|---|
| Refund refused or delayed | 97.0% | 1.82x | 234 |
| Vehicle mechanical defect | 96.5% | 1.81x | 568 |
| Financing or charge problem | 94.0% | 1.76x | 234 |
| Vehicle arrived dirty or damaged | 92.1% | 1.73x | 368 |
| Title or registration delay | 91.3% | 1.71x | 138 |
| Customer support unreachable | 88.0% | 1.65x | 292 |
| Delivery delay or reschedule | 86.7% | 1.63x | 353 |
Each effect above is measured within the same channel, year, and review length, then stress-tested three ways: a placebo (scramble which reviews had the problem, the effect must vanish), a random confounder, and an 80% subsample. The raw multiples show the pattern; the causal figures show what each failure does on its own. Every one survives all three tests. Full method in the appendix.
The raw multiples inflate the true effect because bad experiences cluster on the complaint sites. The causal column strips that out: a broken car raises one-star risk by 23 points even holding the channel constant, a refund fight by only 4. Same anger in the raw count, very different true cause.
It is tempting to read the falling volume of ownership complaints as the problem getting solved. It is not. Fewer owners are writing, but the ones who do are exactly as unhappy as they were two years ago, nine in ten of them. The improvement is in the count, not the experience, and a count is the one thing a growing company can produce by accident.
About 1,027 reviews are from people describing life with the car after delivery, and nine in ten of them are unhappy. Compare that with roughly one in six for people still shopping. And it is not getting better: owners were unhappy 93% of the time in 2024, 87% in 2025, and 89% so far in 2026. Fewer people are writing about it, which is easy to mistake for progress, but the ones who do are just as likely to be unhappy as they were two years ago. That is a problem with how the company runs, not a bad quarter.
The blue bars show how much of each year’s reviews come from people already owning the car. The red line shows how often those people are unhappy. Notice the bars shrink but the line stays high: fewer complaints, no less pain.
| Year | Share of reviews | Unhappy rate |
|---|---|---|
| 2024 | 24.3% | 93.2% |
| 2025 | 17.6% | 87.0% |
| 2026 | 15.5% | 89.0% |
The falling complaint volume is itself the control. If ownership were actually improving, the unhappy rate among owners would drop with it. It doesn’t: 93% in 2024, 87% in 2025, 89% in 2026. Fewer people writing, no less pain, which rules out the comfortable reading that this is getting fixed on its own.
9 in 10
owners who write a review are unhappy, and that number has not moved in three years of record growth.
It would be comforting if the praise were fake, because then Carvana would simply be a bad company. It is not fake. People rave about specific, repeatable things. Which means Carvana has already proven it can delight the exact customers it later loses. The talent is there; only the follow-through is missing, and that is a harder problem to admit than a broken product.
About 1 in 3 reviews give five stars, and they are not vague. The thing people praise most is how easy and painless the process is: it comes up in roughly 1,304 reviews, almost all of them five stars. A good price comes up in another 305. The happiness is all in shopping, selling, and the app. The anger is all in delivery, owning the car, warranty claims, and getting help. Carvana has proven it can win people over, and then it loses those same people right after the sale. The easy front end sold a record 596,641 cars last year; the broken back end is the risk sitting underneath that number.
Top praise drivers by the number of praise reviews, with the five-star rate for each. The praise is genuine and specific, which is precisely why the ownership-stage collapse reads as a real, stage-specific failure rather than general negativity.
| Praise driver | Reviews | Five-star rate |
|---|---|---|
| Easy, seamless process | 1,304 | 85% |
| Good price or value | 305 | 86% |
| Helpful customer service | 226 | 64% |
| Wide selection | 24 | 92% |
| Fast, convenient delivery | 16 | 94% |
The five-star reviews are the counterfactual that makes everything else a finding. A company that delighted no one would be a simpler, duller problem. Carvana clearly can delight, and it does so on specific, repeatable things, which is exactly what makes the post-sale collapse a real, stage-specific failure rather than generic negativity.
Praise is concentrated, not vague: an easy, seamless process drives 1,304 of the five-star reviews (about 85% five-star), a good price another 305. The happiness is all before and during the sale; the anger is all after it.
Here is the most counterintuitive number in the analysis. When Carvana actually resolves a problem, 5% of those reviews are still unhappy. When it does not, 98% are. The defect itself barely matters; the silence afterward is almost the entire story. Carvana could erase most of its one-star reviews without building a single better car, purely by answering the customers it already sold to.
Here is the cleanest number in the whole analysis. When the problem got fixed, only about 5 in 100 of those reviews are unhappy. When it did not, 98 in 100 are. People are not unreasonable. A single thing going wrong does not lose them. Being left with an open problem and no answer does.
The worst of it is where money and warranty meet a problem that never closes. Refunds people are owed but cannot get: 364 reviews, almost all furious. Warranty claims that were slow or flatly denied: 430 reviews, about 96 in 100 unhappy. These people have already paid, are owed something, and cannot get it back. Which points to a hopeful fix: Carvana does not need to stop making mistakes. It needs to close the loop on the ones it makes.
Unhappy-review rate by outcome. The single green bar, resolved-well at 5.1%, is the proof that the customers are winnable. Everything below it is the cost of not closing the loop, weighted toward warranty and refund disputes.
| Outcome | Reviews | Unhappy rate |
|---|---|---|
| Problem resolved well | 311 | 5.1% |
| Resolved, but only after effort | 106 | 76.4% |
| Warranty claim: slow or no help | 273 | 95.2% |
| Warranty denied or not covered | 157 | 98.1% |
| Refund owed but not paid | 364 | 97.5% |
| Problem never resolved | 1,395 | 98.1% |
Same company, opposite ending. Hold the problem constant and change only whether Carvana closed the loop: 5% of resolved cases are unhappy versus 98% of unresolved ones (311 vs 1,395 reviews, p < 0.001). Resolution, not the defect itself, is almost the entire effect. This is the cleanest cause-and-effect contrast in the analysis.
5% vs 98%
Fix the problem and almost everyone is fine. Leave it open and almost no one is. The lever is the response, not the mistake.
Every dashboard would tell you to fix titling: it has drawn three state actions and a $1.5 million settlement in the middle of the recovery. The causal model says something no dashboard could. Hold the channel constant and a title complaint has no independent effect on the star rating at all. The count screams and the cause is silent. Fix titling for the regulators and the legal exposure, not for the reviews.
In the reviews, these paperwork complaints show up every year, almost always unhappy: 99 in 2024, 52 in 2025, and 24 so far in 2026. Outside the reviews, the same problem keeps drawing fines. Illinois suspended Carvana’s license to sell cars in 2022 over exactly this. In January 2025, in the middle of the comeback, Connecticut’s attorney general made Carvana pay $1.5 million for the same failure. Maryland fined it too.
This was the weak spot that nearly sank Carvana in 2022, and it is still costing it fines in 2025. It did not get fixed on the way up; it grew right alongside the company. This is where the two stories meet. A comeback investors are betting on is resting on the one thing customers and three state governments keep saying is broken. Fix delivery, titles, warranty claims, and refunds, and the easy buying experience becomes a real advantage. Leave them broken, and every extra car sold is one more happy buyer turned into an angry owner.
The bars count title and license-plate complaints in the reviews each year. The boxes mark two government actions outside the reviews: Illinois pulling Carvana’s license in 2022, and Connecticut’s settlement in January 2025.
| Year | Complaints | Unhappy |
|---|---|---|
| 2024 | 99 | 87 |
| 2025 | 52 | 45 |
| 2026 | 24 | 19 |
Title and plate complaints look damaging by the count: they appear in 394 reviews, 62% of them unhappy versus 52% otherwise, a raw 1.2×. But hold the channel constant and the effect reverses to minus 4 points (95% range −8 to −1). The raw signal was where these complaints pile up, the Better Business Bureau, which is angry by nature, not the titling itself. A count would have aimed the fix at the wrong thing.
Titling remains a real operational and regulatory problem: three state actions and counting, including a $1.5M Connecticut settlement in the middle of the recovery. It just is not what pushes a customer’s review to one star. Fix it for the fines and the legal exposure, not the star rating.
The same defect that adds 31 points of one-star risk on Trustpilot adds essentially zero on the Better Business Bureau, because those reviewers are already almost all furious, there is no goodwill left to lose. The damage does not show up where you would look for it. And it takes only one unhandled problem to do most of it: a single named failure nearly triples the unhappy rate, while the second and third barely add to it.
Two things stand out. A mechanical defect raises the chance of a one-star review by about 23 points, and not being able to reach a human by about 17. These are the most solid results we have: when we scramble the data as a sanity check the effect disappears, and when we rerun on random slices of the reviews it barely moves. A late delivery adds about 13 points. A refund or billing fight adds a smaller but real 4.
One result flips when you look carefully, and that is worth being honest about. Complaints about the title and plates look damaging at first, but that is mostly because they pile up on the complaint sites, which are angry anyway. Once you account for that, mentioning a title problem does not by itself push a review to one star. Titles are a real problem and a real regulatory risk, as the last section showed, but inside the reviews the thing that actually makes people furious is the car breaking and no one picking up the phone.
Each bar is how much that problem raises a customer’s chance of leaving a one-star review, after removing the effect of which site they used and when. The thin line shows how sure we are; the badge marks how strong the result is after our sanity checks. A bar to the right means that problem, on its own, makes a one-star review more likely.
| Problem | Effect on the chance of one star | 95% range | Placebo test | Reviews | Confidence |
|---|---|---|---|---|---|
| Mechanical defect | +23.5 points | +19 to +28 | −0.3 (passes) | 802 | Strong |
| Unreachable support | +17.2 points | +12 to +24 | −1.1 (passes) | 231 | Strong |
| Delivery delay | +12.8 points | +8 to +18 | −0.4 (passes) | 195 | Strong |
| Refund or billing dispute | +4.5 points | +2 to +7 | +0.2 (passes) | 825 | Moderate |
| Title or registration mention | −4.4 points | −8 to −1 | +0.2 (passes) | 394 | Suggestive |
Two more checks point the same way. First, the more separate problems a single review names, the more likely it is to be one star, in a clean step up: reviews naming no problem are unhappy 34% of the time, one problem 82%, two problems 88%, three 90%. When the result climbs with the dose of the cause like that, it is strong evidence the causes are real. Second, a broken car does the most damage exactly where Carvana still has goodwill to lose: it raises the one-star chance by about 31 points on Trustpilot and 28 on the App Store, but barely moves the Better Business Bureau, where reviews are already almost all one star.
Read all of this as a ranking you can trust the direction of, not a lab measurement. These are real-world reviews, and the people who write about a defect are the ones who had a defect, so this shows which problems turn a review one-star, not a controlled experiment. Even so, the order is stable, it survives every check, and it points straight at two fixes: the quality of the cars Carvana sends out, and whether a customer can actually reach a person when something goes wrong.
| Problems named in the review | Unhappy rate | Reviews | |
|---|---|---|---|
| None | 34% | 3,280 | |
| One | 82% | 1,439 | |
| Two | 88% | 393 | |
| Three | 90% | 62 |
When the outcome climbs step by step with the dose of the cause, that is classic supporting evidence for causation. The jump from zero to one named problem, 34% to 82%, is the cliff.
Three tests per cause. Placebo: scramble which reviews carried the problem and the effect collapses toward zero (every placebo landed within 1 point of zero). Random confounder: add a random variable and the estimate holds. 80% subset: rerun on random slices and it barely moves. The full table is above; the raw output ships in the Causal Analysis folder.
A causal forest (EconML) shows the defect effect is +31 points on Trustpilot and +28 on the App Store, but near zero on the Better Business Bureau, which is already almost all one-star. The damage lands where Carvana still has goodwill to lose.
The same company, in customers’ own words. Read the two columns side by side: buying a car and owning it are two different companies. Every quote is exact, with no paraphrase.
“Extremely difficult to deal with via the only option: chat.”
BBB · 1-star · support unreachable
“They delivered a filthy and non-functional car to me that did not match the picture they listed.”
Google Play · 1-star · vehicle condition
“Terrible company. Delivered a car with tires that are falling apart and should have never passed inspection.”
Trustpilot · 1-star · car quality
“Imagine you buying a $44,000 car not knowing it have bad tires and worn brake pad to get denied on a claim!”
BBB · 1-star · warranty denied
“Horrible delays. Had to switch to two different vehicles after financing one which got delayed a whole week and the other got delayed another week.”
App Store · 1-star · delivery delay
“Please never work with them. I have spent all weekend trying to speak with a human.”
Trustpilot · 1-star · support unreachable
“They refuse to give my money back after they never approved me.”
BBB · 1-star · refund and financing
“This is how a multibillion dollar company treats its customers. Avoid.”
Trustpilot · 1-star · refund delayed
“The entire process was smooth including the delivery.”
App Store · 5-star · buying a car
“Great app! I bought my Betsey, Volvo through Carvana 2 years ago and it went very smoothly.”
App Store · 5-star · buying a car
“The easiest experience I’ve ever had!”
Trustpilot · 5-star · selling to Carvana
“Will buy my next car from carvana!!!!”
Google Play · 5-star · buying a car
“Lela made our visit very quick and painless.”
Trustpilot · 5-star · helpful service
“The best quick way to sell a car at a very great price.”
Trustpilot · 5-star · selling to Carvana
These are priorities, not dollar promises. Each one is sized by how many reviews are at stake and the lever Carvana actually controls. The order follows Section 06: because the damage comes from problems that never get fixed, the top moves are all about closing the loop, not preventing every mistake.
Trigger: Customers cannot get a human answer, get bounced between teams, or watch an open case stall.
Scale: 1,398 reviews describe unreachable support, and 96.0% of them end unhappy.
Intervention: Fast response windows, named case ownership, and one accountable queue for document, refund, and repair escalations.
Trigger: Buyers hit coverage denial, slow SilverRock claims, or a dead-end between Carvana and the repair partner.
Scale: 430 reviews describe negative or denied warranty coverage, and 96.3% of them end unhappy.
Intervention: A tighter repair-approval path, clearer coverage communication at and after the 7-day window, and fewer dead ends.
Trigger: Money owed back is delayed, disputed, or never arrives after a return, denial, or failed delivery.
Scale: 364 reviews describe money owed back, and 97.5% of them end unhappy.
Intervention: Clear refund clocks, automatic status updates, and one workflow tying finance, support, and title release together.
Trigger: A review explicitly prefers CarMax, a traditional dealership, or another option after something breaks.
Scale: 89 comparison reviews prefer the alternative, and 94.4% of them end unhappy.
Intervention: Save-team outreach for late delivery, title delay, and unresolved ownership cases before the buyer reframes the whole experience around a competitor.
Dimension Labs turns unstructured customer language into structured, analyzable signal. Companies sit on thousands of reviews, tickets, and posts that no dashboard can read. We run each one through a purpose-built extraction model that reads only the customer’s words and returns a consistent set of labels, so a pile of free text becomes a dataset you can decompose, compare, and run causal analysis on.
For Carvana we designed 25 such dimensions, validated them against a stratified sample until they were essentially error-free, then applied them to all 5,380 reviews across five channels. Every number in this brief traces to those labels; every quote is an exact substring of a real review. The dimensions are grouped below by what they capture.
Reads the emotional register and loyalty in the customer’s own words.
Locates each review in the customer’s path so friction can be read by stage.
Names what the review is about and the single most important thing that went wrong or right.
Captures the physical car and the digital product.
Tracks the post-sale operational failures behind the regulatory record.
Isolates the financial disputes where trust collapses.
Measures whether a human helped and whether the problem actually closed.
Flags defection risk and how much is at stake.
Pulls the exact language for the report, never paraphrased.
This brief uses 5,380 distinct customer reviews collected from Trustpilot, BBB, Google Play, the App Store, and Reddit between January 2024 and June 2026 (window edges 2023-12-31 to 2026-06-12). Every count, share, and rate is calculated on distinct reviews, with zero duplicates. The enrichment layer covers 5,377 of the 5,380 reviews, a dimension coverage rate of 99.94%; the three unmatched reviews are excluded from dimension-level rates only.
The structured analysis uses the Carvana review dimensions for overall sentiment, lifecycle stage, primary topic, primary problem raised, praise driver, recommendation signal, resolution status, customer-service signal, financial-impact signal, warranty signal, transaction type, issue severity, actionability ranking, and verbatim evidence fields (overall_sentiment, lifecycle_stage, primary_topic, primary_friction, praise_driver, resolution_signal, customer_service_signal, financial_impact_signal, warranty_signal, transaction_type, and the two verbatim fields). Every quotation in this brief is an exact substring of a real review; none is paraphrased or generated.
Public figures and public milestones are kept strictly separate from review-base figures. Public-company numbers come only from the cited external context; review-base numbers come only from the enriched dataset. Where a public event and a review pattern appear side by side, they are shown as parallel evidence and never blended into a single metric.
The causal section is formal causal inference, not correlation. We built a row-level table of 5,183 reviews (every rated review plus every Better Business Bureau complaint), defined the outcome as a one- or two-star rating or a BBB complaint, and derived five candidate causes from the review wording. For each, we estimated the effect on the probability of a one-star review while holding channel, year, quarter, and review length constant, using logistic standardization (G-computation) as the primary estimate with bootstrap confidence intervals, and DoWhy’s propensity-score weighting and linear regression as cross-checks. Every estimate was put through three refutation tests: a placebo (scramble which reviews carry the cause; the effect must collapse toward zero), a random common cause, and an 80% data subset. We also ran an EconML causal forest for how the effect varies by channel, and a dose-response check on the number of problems named. Libraries: DoWhy 0.14, EconML 0.16.0, scikit-learn, statsmodels.
Because this is observational review text, the estimates reach the level of a strong, refutation-checked ranking of causes, not a randomized experiment; the honest ceiling and the one result that reverses on adjustment (titling) are stated in the body. The full analysis ships with this report: the dataset (carvana_causal_dataset.csv), the code (carvana_causal_analysis.py), the raw estimates (causal_results.json), and a written evidence pack with a per-cause robustness table are in the accompanying Causal Analysis folder.
FY2025 net income of $1.895 billion includes a one-time deferred-tax benefit of roughly $685 million; operating income ($1.881B) and adjusted EBITDA ($2.237B) are the cleaner reads and are used where per-unit economics matter. The FY2025 release does not publish a full-year gross-profit-per-unit figure, so this brief does not cite one. Neither caveat changes the shape of the recovery.