The best tennis players in the world win about 55 percent of the points they play.

That single number is the reason we think decision-making frameworks deserve more attention in tennis than in almost any sport we cover. Not 70 percent. Not 80. Somewhere in the mid-fifties, with the leaders of the last two decades sitting a shade under 56 percent and the merely very good a point or two beneath them. Two decades of dominance, expressed as a coin that lands your way roughly eleven times out of twenty.

The verdict: when a tennis decision takes months to pay off — a coaching change, a string setup, a frame, a season's training block — the outcome you eventually observe is too noisy to grade the decision that produced it, and the only defensible way to evaluate that decision is against a written record of what you expected and why, made before the outcome existed. Every alternative we looked at is a variation on grading yourself from memory, which is grading yourself with the answer key already open.

The number, and where it comes from

Career total-points-won leaderboards are published by Tennis Abstract, Jeff Sackmann's public tennis statistics project, which aggregates match and point-level data from the professional tours. The figures for the leading men of the modern era cluster in the mid-50s. This is not a contested statistic — it falls straight out of arithmetic on published match records — but it is worth flagging what it is: an aggregation of tour data, not a controlled measurement, and the exact figure shifts by a few tenths depending on which matches and which years are counted.

What makes 55 percent interesting is not the number itself but the leverage it carries. Franc Klaassen and Jan Magnus, working from point-by-point records of several hundred Wimbledon matches from the early 1990s, spent two decades formalising how point-level probability maps onto match-level results — first in the Journal of the American Statistical Association in 2001, later in their book Analyzing Wimbledon (Oxford University Press, 2014). Their central structural finding is that tennis scoring is an amplifier. A player who wins points at a rate a couple of percentage points above their opponent wins matches at a rate that looks nothing like a couple of percentage points. The hierarchical scoring system — points into games, games into sets, sets into matches — takes a nearly symmetrical edge and compounds it into something that reads, from the stands, as a gulf in class.

That is the whole architecture of luck and judgment in one sport, sitting in plain view. At the level of the single point, the gap between the best player alive and a good club professional is close to invisible. At the level of the career, it is total.

What the 55 percent actually measures

It measures a ratio over an enormous denominator. A tour career runs to hundreds of thousands of points. At that scale, the small persistent edge — the extra half-step, the marginally better second serve, the shot selection that is right slightly more often — separates cleanly from everything random happening around it. The number is trustworthy precisely because nothing in it depends on any particular point.

It also measures the result, not the reason. That distinction is going to matter shortly.

And it measures something the pros' own performance data confirms: that points are close to, but not quite, statistically independent. Klaassen and Magnus tested the assumption directly and found small but real departures from independence — winning the previous point slightly raises the chance of winning the next, and both players tend to play more conservatively on important points, which measurably compresses the server's advantage. The deviations are small enough that models treating points as independent still work well, and large enough that anyone claiming tennis is a pure sequence of unrelated coin flips is overstating the case. We flag this because it is a rare instance where the honest answer to "is this true?" is "mostly, and the exceptions are documented."

What the 55 percent does not measure

Here is where the number stops being reassuring.

It does not tell you which points. Klaassen and Magnus's own work shows that the tour's aggregate rate hides systematic variation by pressure and by scoreline. A flat 55 percent conceals a player who is 60 percent on serve at 40–0 and something much worse at 30–40.

It does not attribute causes. A point won is a point won whether it came from a rebuilt backhand, a bad read by the opponent, a net cord, or a linesperson's error. The ledger does not record why.

Most importantly, it does not tell you anything about the decision behind the shot. A player who correctly chose to go up the line and mishit it scores zero. A player who chose wrongly, went crosscourt into strength, and got a shanked reply scores one. Over 200,000 points those cases wash out. Over the four matches you play after switching strings, they do not wash out at all — they are the data.

A wide, high-vantage photorealistic photograph of a single tennis ball frozen mid-bounce directly on…

That is the hinge of this piece. The 55 percent figure is credible because it rests on a sample large enough to drown the noise. Almost no decision you will personally make in tennis gets a sample anywhere near that size. You will make a choice with months of consequences and then read its quality off three or four outcomes, each one roughly as informative as a single point.

How we evaluated

We did not run a study, chart matches, or measure a string bed. This is a synthesis, and the sources we weighed, in descending order of how much weight we gave them:

  • Peer-reviewed and academically published tennis statistics. Klaassen and Magnus (2001, JASA; 2014, Oxford University Press) are the anchor because their work is explicitly about the point-to-match mapping and because they tested their own assumptions rather than asserting them.
  • Lab-published equipment data. Tennis Warehouse University publishes machine-measured string data — stiffness, tension loss, spin potential — with its methodology stated. That is a different and better class of evidence than a manufacturer's "tension maintenance" claim made without a protocol. Rod Cross and Crawford Lindsey's Technical Tennis remains the accessible reference for the underlying physics.
  • Aggregated public point data. Tennis Abstract and its associated Match Charting Project are volunteer-built. Coverage is uneven across tours and eras, which is a real limitation on any claim drawn from them.
  • Decision-science literature from outside tennis. Philip Tetlock's forecasting work (Expert Political Judgment, 2005; Superforecasting, 2015) and Annie Duke's Thinking in Bets (2018) supply the vocabulary. Duke's term for judging a choice by its result — "resulting" — is the cleanest name anyone has given the error this piece is about.

Where these disagree, we say so. Where a figure is manufacturer-stated rather than independently measured, we say that too.

Your string change has a sample size of one

Take the most common recurring gear decision in the sport. You switch polyester. The manufacturer's page claims improved tension maintenance. Tennis Warehouse University's lab data, across its string database, shows that tension loss in polyester is heavily front-loaded — a substantial share of it happens in the hours after the frame comes off the machine, before a ball is struck. So the string you play on Tuesday is measurably not the string that was installed on Sunday, and the string you judge in week three is a different object again.

Now count the noise sources between your decision and your verdict: opponent quality, court surface and speed, ball type and ball age, temperature and humidity, string age at the moment of play, your own form, and whatever you happened to believe about the string before you played on it. The effect you are trying to detect might be worth two points per hundred. The channel you are trying to detect it through is carrying considerably more variance than that.

This is not an argument that gear does not matter. Lab-measured differences between strings are real and, in the aggregate, meaningful. It is an argument that your personal four-match verdict is not capable of resolving them, and that treating that verdict as evidence is how players end up cycling through nine setups in two years and arriving nowhere.

Decision Evidence available when you decide Feedback lag Luck share in what you'll observe What a written record buys
String or tension change Lab stiffness and tension-loss data; owner reports 2–6 weeks Very high Separates "it felt different" from "I predicted this specific difference"
Racquet change Published specs; independent spec measurements; demo impressions 2–4 months High Catches the honeymoon effect that fades by month two
Coaching change Track record; reference conversations; a stated plan 6–18 months High, and confounded by growth Distinguishes a bad hire from a good hire in a bad season
Season training block Physical baselines; periodisation literature 6–12 months Moderate — measurable inputs help Records the assumption that failed, not just the result
Club or federation selection policy Historical cohort data; comparable programs 3–8 years Very high; cohorts are tiny Converts committee argument into a checkable claim

Three ways to grade a decision

Memory-based review. You think back to why you switched and judge whether it was smart. This is the default and it is the weakest of the three on every criterion that matters. It fails on hindsight resistance, because knowing the outcome silently rewrites the reasoning you recall having. It fails on auditability, because nobody else can check it. Its only advantage is cost: it is free, and it feels like reflection.

Outcome-based review. You judge the choice by what happened. Duke's "resulting." This scores better on auditability — the result is a fact — and worse on everything else, because it systematically punishes good decisions that lost and rewards bad decisions that won. In a domain where the elite edge is five percentage points, outcome-based review over small samples is close to random with respect to decision quality. A coach who develops players well and loses a national title to a draw is graded identically to a coach who guessed and got lucky, except the second one keeps the job.

A photorealistic environmental portrait of a lone player in plain athletic wear standing at…

Pre-committed record. Before the outcome exists, you write down: the choice, the alternatives you rejected, the specific thing you expect to observe, a confidence number, the assumptions the choice depends on, and the date you will review it. At first glance that sounds like admin overhead with a nicer name. In practice it changes three things at once. It fixes your reasoning in a form hindsight cannot edit. It makes the assumption — not the result — the reviewable unit, so when circumstances shift you can identify which premise stopped holding. And it makes a confidence number falsifiable: after twenty entries you can check whether your 70 percents come in at seventy percent, which is the only way anyone has found to measure judgment in a lagging domain. Tetlock's forecasting research is built on exactly this mechanism and is the strongest evidence any of these frameworks has behind it.

The line worth screenshotting: write down what you expect, in numbers, with a date to check it, before the result exists — and grade the prediction, not the outcome. Nothing else on this list survives contact with a 55 percent world.

Where the evidence is thin, and where sources disagree

We should be straight about the limits of this synthesis.

There is, as far as we can find, no published study showing that decision records improve outcomes in tennis coaching or administration specifically. The mechanism is well-evidenced in forecasting and in engineering practice; the transfer to tennis is an argument, not a finding. Anyone telling you otherwise is selling a system.

The field also supplies a useful cautionary tale about how confidently these things get settled. Amos Tversky, Thomas Gilovich, and Robert Vallone published a famous analysis in Cognitive Psychology in 1985 concluding that basketball's "hot hand" was a misperception of randomness. It stood as textbook fact for thirty years. Then Joshua Miller and Adam Sanjurjo, in Econometrica in 2018, identified a subtle selection bias in the original method — a bias in how streak-conditional sequences are counted — and showed that correcting for it flipped the sign of the result. The evidence now supports a modest hot hand.

We raise this not to litigate basketball but because it is the article's own logic playing out at the level of a research literature. The 1985 decision to conclude "no hot hand" was, given the analysis and priors of 1985, defensible. It was also wrong. The reason we can tell the difference — the reason this is a story about a correctable method rather than a reputation — is that the reasoning was written down in full, in public, in a form somebody could later check. That is what a decision record is for. Not being right. Being checkable.

Who this is for, and who it isn't

It's for you if your decisions have a lag: coaches carrying athletes across development years, club and federation administrators whose policies produce results on a five-year horizon, and adult players spending real money on equipment they will then evaluate with a sample of four matches and a feeling. It's also for anyone who has noticed that their confident retrospective explanations of past choices arrive a little too smoothly.

It isn't for you if your feedback is fast and clean. If you can try something, see the result in an hour, and try again, iteration beats documentation and the record is pure overhead. It also isn't for you if you want a framework that improves outcomes next month. This one does not. It improves your calibration over a horizon measured in years, which is a genuinely worse pitch and a more honest one.

And it isn't for anyone hoping to be relieved of uncertainty. A well-kept record will show you, in your own handwriting, that you were reasoning carefully and still lost. That is the point, and it is not comfortable.

Evidence grade

For the central claim — that in tennis, outcomes observed over small samples are poor evidence of decision quality: Strong. The point-to-match amplification is well-established in the published statistics literature, and the arithmetic of small samples is not in dispute.

For the secondary claim — that keeping written decision records improves judgment in tennis: Moderate, and only by extension. The forecasting evidence is solid; the tennis-specific evidence does not exist. We would rather say that plainly than dress an analogy up as a result. Editor's note. There is a file on our shared drive with one line per product, each written before anyone here reads a single review of the thing. The entry I keep returning to is a soft co-poly I logged last spring: expect testers to call it comfortable and low on spin — 65 percent. The tester consensus came back comfortable and distinctly high on spin. When I went looking for that line, I was fairly sure I had predicted the spin correctly. I had not. The line was four months old, timestamped, and unhelpfully specific, and it is the only reason I know which half of that call I actually got right.