Engine & Software Reviews

Understanding Engine Evaluation Numbers

What the +0.4 or -2.1 next to an engine's analysis actually means, how to read it correctly, and the common mistakes players make when interpreting it.

Hardik RanpariyaSeptember 14, 20266 min read

Anyone who has ever run a game through an analysis engine has seen it: a small number, usually something like "+0.7" or "-1.3," attached quietly to the current position. Understanding exactly what this number actually represents — and, just as importantly, what it deliberately doesn't represent — is essential for using engine analysis genuinely productively rather than being subtly misled by it during study.

The Basic Unit: Centipawns

Engine evaluations are typically expressed in units of pawns, with a positive number favoring White and a negative number favoring Black. Under the hood, most engines actually calculate internally in centipawns — hundredths of a pawn — purely for the sake of calculation precision, but the number ultimately displayed to users is almost always converted back into whole pawns for readability, so "+0.75" means the engine judges the current position as worth roughly three-quarters of a pawn in White's favor overall. This isn't a literal, narrow claim that White is "ahead by three-quarters of a pawn" in raw material terms specifically — it's a much broader, blended assessment that factors in material, king safety, piece activity, pawn structure, and a great deal else besides, all of which the engine's evaluation function has learned to weigh together into one single number.

What Counts as a "Meaningful" Advantage

A rough, widely used interpretive scale looks something like this: evaluations within about plus or minus 0.3 are generally considered close to equal, with no genuinely decisive advantage yet established for either side. Evaluations in the range of roughly 0.5 to 1.5 suggest a real, meaningful edge for one side — often enough to matter considerably in practical play, though still far from anything resembling a guaranteed win. Evaluations beyond about 2.0 to 3.0 typically indicate a serious, often close-to-decisive advantage, and anything beyond roughly 5.0 or 6.0 usually reflects either overwhelming material superiority or a genuinely unstoppable attack already underway. These thresholds are rough, practical guides rather than precise mathematical boundaries, and they shift somewhat depending on the specific position in question and how many total pieces still remain on the board.

Mate Scores: A Categorically Different Kind of Number

When an engine detects a forced checkmate somewhere down a calculated line, it switches from a standard pawn-based evaluation to a distinct "mate in N" display instead — for instance, "M4" indicating a fully forced checkmate in exactly four moves for the side currently to move. This is categorically different from an ordinary pawn evaluation: it's not a graded judgment of relative advantage at all, but a concrete, calculated certainty, assuming genuinely perfect play from both sides, that the game ends in exactly that specific number of moves. A mate score always takes clear priority in significance over any competing pawn-based evaluation, since it represents an absolutely guaranteed win rather than merely a probabilistic assessment of who's currently doing better.

Why the Same Position Can Show Different Numbers on Different Engines

It's genuinely common for two different engines — or even two different versions of the exact same engine — to evaluate an identical position with noticeably different numbers, sometimes differing by half a pawn or considerably more, even when both engines fully agree on which side is actually better overall. This reflects genuine, legitimate differences in how each engine's evaluation function has been trained and internally tuned over time, not necessarily an outright error in either one specifically. The precise number matters considerably less than the overall trend and the actual recommended move; treating small evaluation differences between competing engines as inherently, meaningfully significant is a common but largely unproductive habit worth deliberately unlearning.

The Danger of Overreacting to Small Swings

A frequent, genuinely costly mistake among improving players reviewing their own games is treating every small evaluation swing — say, moving from +0.2 to -0.3 across a single move — as an automatically meaningful mistake deserving deep, careful study. In reality, evaluations swinging at this smaller scale often reflect the engine's own shifting, fine-grained assessment of subtle, genuinely hard-to-convert positional factors rather than any real, concrete practical error on the player's part. The evaluation swings genuinely worth careful, deliberate attention are the considerably larger ones — typically a full pawn or more shifting within a single move — since these usually correspond to an actual missed tactic, a real structural concession, or a genuine calculation error, rather than simple noise inherent in the evaluation function's own fine-grained internal judgment.

Evaluation Depth and Why It Matters

An engine's evaluation becomes measurably more reliable the deeper it's allowed to search a given position, and a shallow, quick evaluation can occasionally miss a genuinely critical resource that a deeper, more patient search would have found without difficulty — meaning the exact same position can show a meaningfully different number after just a few more seconds of additional calculation time. This is particularly relevant when relying on cloud or online analysis tools with a fixed, often fairly shallow time budget allotted per move; for genuinely critical positions worth getting right, letting the engine run considerably longer, or switching to a locally run engine without any fixed time limit at all, produces noticeably more trustworthy numbers than a quick, time-boxed online check ever will.

What the Evaluation Doesn't Tell You

Perhaps the single most important limitation to fully internalize here: an engine's evaluation reflects objective correctness assuming genuinely perfect play from both sides involved, not practical difficulty for the specific human opponent actually sitting across the board. A position evaluated at a fairly modest +0.6 can still be extremely difficult for a human player to navigate correctly if it demands precise, only-move-level accuracy across several consecutive moves in a row, while a position evaluated at a considerably higher +1.5 might turn out to be comfortably, intuitively winning even for a relatively inexperienced player facing it. Treating the raw evaluation number as the complete picture on its own, rather than as one genuinely useful input alongside your own developing judgment about practical difficulty, is the single most common way engine analysis ends up misused rather than genuinely learned from over time.

Multi-PV: Seeing More Than Just the Top Move

Most analysis engines can be configured to display several candidate lines simultaneously — commonly called multi-PV, short for multiple principal variations — rather than showing only the single top-rated move in isolation. This is genuinely valuable during serious study: comparing the evaluation gap between the engine's first and second choice tells you something the single best move alone simply doesn't, namely how forgiving or unforgiving the position actually is in practice. A large evaluation gap between the first and second candidate moves indicates a critical, genuinely precise position where only one specific move preserves the advantage; a small gap between them instead suggests several roughly comparable options exist, and the exact move chosen matters comparatively less to the overall outcome.

Positive vs Negative Numbers and Perspective Confusion

A small but genuinely common source of confusion, especially among newer players, is forgetting which side a positive number actually favors when switching between different tools or different board orientations mid-session. The near-universal convention across essentially all engines and platforms is that positive numbers favor White and negative numbers favor Black, regardless of which side happens to be displayed at the bottom of the screen at that moment — but it's worth explicitly double-checking this on any genuinely new platform or tool before trusting the number at a glance, since a single moment of reversed-sign confusion can lead to badly misjudging an entire position, particularly under the added time pressure of reviewing a game quickly after it ends.

Put it into practice

Try these ideas out against the engine or a live opponent.