top of page

Why Changing Batch Size Can Change a Prediction

8 hours ago
2 min read

Same image. Same weights. Different batch size.

You might expect exactly the same output. But mathematically equivalent computations are not always numerically identical.

The order of addition matters

Consider three numbers stored in 32-bit floating point:

a = 100,000,000
b = −100,000,000
c = 1

(a + b) + c = 1
a + (b + c) = 0

Why? At that magnitude, float32 cannot preserve the extra 1 when computing b + c. It rounds back to −100,000,000.

Nothing random happened. We changed the grouping, which changed where information was lost. This is non-associativity: floating-point addition does not always follow the algebra we learned at school. NVIDIA explains the underlying arithmetic.

This exaggerated example makes the mechanism visible; it is not a measurement of typical model error.

What batching changes

Neural networks perform large numbers of multiply-and-add operations. Batched execution can use a different computation path from processing each input separately, changing how intermediate results are accumulated.

PyTorch explicitly documents that a matrix product computed within a batch need not be bit-for-bit identical to the same product computed individually. PyTorch numerical-accuracy documentation.

That does not mean changing batch size always changes the predicted class. It means exact numerical equality is not guaranteed.

Small difference, different decision

Suppose a visual inspection system rejects a part when its anomaly score exceeds 0.5.

Two illustrative scores—0.499999 and 0.500001—are numerically close but produce opposite decisions.

The practical question is therefore not just “Are the tensors close?” It is also “Did any decisions change?”

When changing batching or precision, I’d:

  • Compare identical inputs across the intended inference configurations.

  • Measure score differences using justified numerical tolerances.

  • Check decision disagreements, especially near the operating threshold.

Large differences still deserve investigation; rounding should not become an excuse for preprocessing or model-mode bugs.

Batch size is a performance setting—but its numerical effects deserve a correctness test.

Recent Posts

See All

Let's talk industrial AI, engineering, and teams.

Thank you for reaching out!

© 2026 by Zeeshan Karamat. All rights reserved.

bottom of page