← Back to home

Weighted WAR: How Much Did Greatness Actually Matter?

August 8, 2026 11 min read Baseball · Analytics

In baseball lore, when Ralph Kiner sat across from Branch Rickey to argue for a raise, he had a good case. He had led the National League in home runs seven straight years. Rickey listened, then delivered one of the great put-downs in baseball history:

We could have finished last without you.

The line is cruel, and it isn't fair, and it also isn't entirely wrong. The Pirates were terrible. Kiner put up enormous individual numbers on teams that were going nowhere. It was production in a vacuum: beautiful, prodigious production and in the only currency Rickey cared about (wins), nearly worthless.

There's one wrinkle worth mentioning: the famous Kiner story has been questioned over the years, including by Kiner himself, so I'm treating the quote here as baseball lore rather than gospel. But whether Rickey actually said it exactly that way is almost beside the point. The question underneath it is real.

And that is the subject of this post, what does value mean?

Baseball has spent the past twenty years increasingly valuing Wins Above Replacement (WAR) as the shorthand for a player's worth. And for good reason. WAR strips away the team around you and tries to isolate what you did.

But there's a problem with that, or at least an interesting limitation.

WAR deliberately treats context as noise. A win in a blowout September game for a 95-loss team and a win that drags your club into October can contribute equally to a player's context-neutral value. Watching baseball tells you they're not the same thing.

So I built a different metric.

Call it weighted WAR: a version of career value that deliberately rewards consequence. Not just how much you produced, but how often that production came when winning actually mattered.

This is, on purpose, partly unfair to the player. It makes the player answer for the context in which his value was delivered, even when that context wasn't his fault. That's not a bug I'm apologizing for. It's the entire point.

The modern Kiner, and we'll get to him, deserves the same question Rickey asked.

How the metric works

Weighted WAR keeps raw WAR as its backbone, then asks two follow-up questions: how much of that value came while your team was actually in a meaningful race, and what did you do once you got to October?

I like to call it the Derek Jeter factor.

The whole thing is built from three pieces that add up to a single number.

Replacement-anchored value (60%). Sixty percent of your raw WAR, untouched. This is the anchor, the same for everyone, no judgment calls. It keeps weighted WAR from drifting too far from the number we already trust. Greatness is still mostly greatness.

Contention value (20%). Twenty percent of your raw WAR, multiplied by a contention factor: how much of your value came while your team was actually in a race. Were you raking in September with your club a few games out of a playoff spot, or piling up stats when the standings had already made October irrelevant?

One rule matters here: if your team was already in, or comfortably leading the division, you still get full credit for keeping it up. Coasting isn't penalized. You only lose credit here when your team was genuinely out of the picture.

Postseason value (20%). Twenty percent of your raw WAR, multiplied by a postseason factor based on how you actually played once you got there. It's about your rate, not simply how many postseason games you happened to play, with a minimum-sample floor so one hot or cold series doesn't swing an entire career.

Getting there matters. But this column is about what you did with the chances you got.

In formula form, it's basically:

Weighted WAR = Raw WAR × (0.60 + 0.20 × Contention Factor + 0.20 × Postseason Factor)

So if a player's contention factor is 0.70 and his postseason factor is 0.50, he keeps 60% of his raw WAR automatically, plus 14% from contention and 10% from October. His weighted WAR would be 84% of his raw WAR.

That's the idea. The goal isn't to replace WAR. It's to ask a different question with it.

A few calls are worth stating plainly. I'm not trying to adjust individual players for steroids or PED use here. That's a separate argument, and I'd rather not bake a giant set of assumptions about who used what into a metric that's already making one controversial adjustment.

I'm also separating baseball history by playoff environment, because "in contention" means something very different when only four teams make the postseason than when twelve do. MLB added the first Wild Card in 1995, a second in 2012, and a third in 2022; 2020 was its own pandemic-era anomaly with an expanded 16-team field.

And I'm sticking to the top ten position players in each era. Pitchers are a separate problem for a separate day.

One important caveat on the data: the raw WAR figures are straightforward, but the contention and October factors below are currently educated estimates, not the output of a fully crunched game-by-game model. I still need to finish the September-leverage and postseason-rate work. The third era's raw WAR figures are especially provisional because several of those players are active and their numbers will keep moving.

So treat this as a working model, not a finished ledger.

The point of this first pass is to see whether the idea produces interesting—and defensible—results.

Era 1 — The 1980s: October was a closed door

Start in the 1980s, the cruelest era for this metric.

No Wild Card. You won your division or you went home, and "your division" meant beating out five other teams across a full season. Most great players spent at least part of their careers behind that door.

The Don Mattingly effect.

PlayerRaw WAR60% AnchorContention (20%)Postseason (20%)Weighted WAR
Rickey Henderson72.043.213.0 (×0.90)12.2 (×0.85)68.4
Mike Schmidt60.036.012.6 (×1.05)13.2 (×1.10)61.8
Wade Boggs61.036.611.0 (×0.90)9.8 (×0.80)57.4
Alan Trammell53.031.811.1 (×1.05)12.2 (×1.15)55.1
George Brett49.029.410.8 (×1.10)11.8 (×1.20)52.0
Robin Yount53.031.810.6 (×1.00)9.5 (×0.90)51.9
Cal Ripken54.032.49.7 (×0.90)9.2 (×0.85)51.3
Eddie Murray48.028.89.6 (×1.00)9.6 (×1.00)48.0
Tim Raines48.028.87.2 (×0.75)4.8 (×0.50)40.8
Dale Murphy46.027.66.4 (×0.70)4.6 (×0.50)38.6

Rickey Henderson stays on top. When you're that far ahead, the factors can't catch you, which is exactly how the anchor is supposed to work.

But the reshuffle underneath him is the story.

Mike Schmidt stays comfortably ahead of Wade Boggs, helped by a stronger October résumé. Trammell and Brett also get rewarded for producing on teams that actually reached October. Trammell's 1984 postseason is the obvious example: the model is trying to give some numerical weight to the fact that his best work came when the Tigers were playing for something.

And then there's Dale Murphy.

Back-to-back MVPs, the face of the Braves, and a weighted WAR that falls from 46.0 to 38.6—the steepest drop in the decade.

This is the 1980s Kiner.

Murphy's greatness was genuine. His Braves were also, for most of his prime, nowhere near the postseason. The model isn't saying Murphy was responsible for that. It's saying something narrower and, I think, more interesting:

A huge amount of his value came in seasons where there was very little chance for that value to affect a pennant race.

That's the distinction weighted WAR is trying to capture.

Era 2 — One Wild Card (1995–2011): the door cracks open

The first Wild Card changed the math.

Beginning in 1995, a great player on a second-place team suddenly had a path to October. Contention stretched deeper into September for more clubs, and October stopped being the exclusive property of division winners.

PlayerRaw WAR60% AnchorContention (20%)Postseason (20%)Weighted WAR
Barry Bonds115.069.021.9 (×0.95)17.3 (×0.75)108.2
Alex Rodriguez105.063.020.0 (×0.95)18.9 (×0.90)101.9
Albert Pujols88.052.818.5 (×1.05)19.4 (×1.10)90.7
Chipper Jones85.051.017.0 (×1.00)17.0 (×1.00)85.0
Derek Jeter71.042.615.0 (×1.05)15.6 (×1.10)73.2
Carlos Beltrán68.040.814.3 (×1.05)16.3 (×1.20)71.4
Jeff Bagwell75.045.014.3 (×0.95)12.0 (×0.80)71.3
Scott Rolen70.042.014.0 (×1.00)14.0 (×1.00)70.0
Jim Thome70.042.012.9 (×0.92)11.9 (×0.85)66.8
Larry Walker68.040.812.2 (×0.90)11.6 (×0.85)64.6

Bonds stays first. At 115 raw WAR, nothing is going to erase that kind of lead. Even a mediocre postseason factor barely moves him.

But the reshuffling underneath him is more interesting.

Even though Pujols had a better postseason than A-Rod, he's still behind him, but the gap shrinks, from 17 WAR to 11.2 weighted WAR.

That's the purpose of this metric to show players who came up big in the postseason or the latter part of the season to showcase how meaningful they were, they stepped up when the lights got brighter.

This weighted WAR formula lifts Derek Jeter and Carlos Beltrán into the conversation with players who had stronger regular seasons.

Meanwhile Bagwell, Thome and Walker, who were great, take some of a hit because more of their careers happened outside October.

That's not the model saying they weren't as good.

It's saying their greatness had fewer opportunities to become consequential, either they couldn't come up big in the clutch or something else.

Era 3 — The Modern Era (2015–2025): no excuses left?

Now the door is pretty wide open.

I'm starting this final sample in 2015, rather than pretending the playoff changes line up perfectly with a neat decade. The postseason expanded again in 2022, and 2020 was its own weird 16-team experiment, so this might not be perfect as I'm trying to fit things into 10 year blocks.

It's simply the modern sample I'm using for the current version of the model.

And if you are great and your team still can't manufacture a meaningful September, the metric has less sympathy, maybe the weights need to change to showcase that.

PlayerRaw WAR60% AnchorContention (20%)Postseason (20%)Weighted WAR
Mookie Betts62.037.213.0 (×1.05)13.6 (×1.10)63.8
Mike Trout70.042.08.4 (×0.60)7.0 (×0.50)57.4
Freddie Freeman52.031.210.9 (×1.05)12.5 (×1.20)54.6
José Ramírez55.033.011.0 (×1.00)9.9 (×0.90)53.9
Aaron Judge50.030.09.5 (×0.95)8.5 (×0.85)48.0
Manny Machado48.028.89.6 (×1.00)9.6 (×1.00)48.0
Bryce Harper45.027.09.5 (×1.05)10.8 (×1.20)47.3
Francisco Lindor46.027.69.4 (×1.02)8.7 (×0.95)45.7
Nolan Arenado48.028.88.6 (×0.90)7.7 (×0.80)45.1
Corey Seager42.025.28.8 (×1.05)10.1 (×1.20)44.1

The numbers are brutal.

Mike Trout has the most raw WAR in the group by eight wins, and he is not first.

Mookie Betts, trailing Trout by eight raw WAR, comes out on top because more of Betts' career value came in meaningful races and October, including championships in Boston and Los Angeles.

Trout, meanwhile, drops from 70.0 raw WAR to 57.4 weighted WAR, that is the steepest fall anywhere in the project.

This is the modern Kiner case.

And it's more complicated than the original, because Trout had every door open.

His Angels had one playoff appearance during the period covered here, and even as the postseason expanded, the basic problem remained: enormous individual production, almost no meaningful October opportunity.

That's where the ×0.60 contention factor and ×0.50 postseason factor come from in this working model.

And this is where you should object.

It isn't Trout's fault that the Angels couldn't build a roster, but I do think he could have influenced the front office a bit more. Leadership is not just on the field.

He showed up. He was one of the best players alive. For years, he was probably the best player alive.

All true.

But the metric was never measuring blame alone.

It is mostly measuring consequence.

Those aren't the same thing.

And where there is blame, it's shared. The front office carries most of it. The player carries some.

A player can carry only a small share of the blame and still have a career in which his greatness rarely affected the outcome of a pennant race.

That's the uncomfortable part.

What the number is actually for

Weighted WAR is not a replacement for WAR. It's an argument with it.

Raw WAR asks:

How much value did this player produce?

And it answers that question beautifully.

Weighted WAR asks a different question:

How much of that value was delivered when winning had meaningful consequences? How much should the player bear the blame?

Those aren't the same question.

And I think baseball sometimes gets into trouble when we pretend they are.

The important thing is that weighted WAR isn't trying to prove Mike Trout was less talented than Mookie Betts. It isn't even really saying Betts was a better player.

It's saying something narrower:

If you care about the historical consequences of a player's career, Betts has a stronger case than his raw WAR alone would suggest, while Trout has a weaker one.

That's a different argument.

And there's a legitimate objection sitting right in the middle of it: why should a player lose statistical credit because his front office failed?

I think I do have an answer now: he only loses his share, not the whole thing.

The 60% anchor never moves. Most of his greatness is banked, no matter what the team around him did.

Only the smaller slice is ever on the table. The player carries part of it, the front office carries most.

That's actually the point.

If you believe a player's historical value should be almost entirely independent of his teammates, then WAR is probably the better number.

If you believe that a player's place in baseball history is partly about what his greatness actually accomplished in the standings and in October, then some kind of consequence adjustment makes sense.

I'm in the second camp.

A career is not just a pile of value. It's value delivered into a context, and the context is part of the story whether our favorite metric chooses to see it or not.

Kiner made the Hall of Fame, eventually, on the thirteenth try.

Trout will walk in first ballot, and he should.

Greatness deserves the monument.

But there's a second question worth asking at the plaque, the one Branch Rickey supposedly asked across a desk in Pittsburgh.

Not how good were you?

We already have WAR for that.

The question is:

How much did all that greatness actually matter?