conceptEvaluación y Medición~1 min de lecturaActualizado 2026-06-07#evaluation#nondeterminism#reproducibility#testing
Esta nota todavía no está traducida, así que se muestra la fuente en inglés.

Nondeterminism & reproducibility

A property that trips up everyone coming from traditional software: the same input can produce different outputs. This isn't a bug — it's inherent to how LLMs are sampled and served. You can't eliminate it, so you design and test around it.

Where the variability comes from

  • Samplingtemperature > 0 deliberately draws randomly from the distribution. Different draw → different text.
  • Temperature 0 is not a guarantee. Even greedy decoding can vary because of floating-point non-associativity, batching, and changing hardware/kernels on the provider side — tiny numeric differences flip a token, which cascades.
  • Provider-side changes — silent model updates and infra changes shift behavior (deprecation & migration).
  • Context differences — small prompt or context changes (even whitespace) move outputs.

Reduce it where you need consistency

  • Lower temperature (≈0) for extraction, classification, and structured tasks.
  • Constrain the output — schemas/structured outputs shrink the space of valid answers.
  • Set a seed if the API supports it (helps, doesn't fully guarantee).
  • Cache results for identical inputs when you want a stable answer (caching).

Test for distributions, not single runs

Because outputs vary, one passing run proves nothing:

  • Run each eval case multiple times and report a pass rate, not a single pass/fail.
  • Assert on properties (valid JSON, contains the fact, no banned content) rather than exact string matches.
  • Track variance over time — a rising failure rate signals drift or a silent model change.

Pitfall

Writing exact-match tests against LLM output, or trusting "it worked when I ran it." The test passes once, then flakes in CI and production. Embrace the nondeterminism: pin what you can, and evaluate behavior as a distribution.

Connects to: decoding & sampling · eval-driven development · silent model changes