Uncertainty Quantification Metrics¶
A conditional flow produces a distribution \(p(u \mid f)\), not a point prediction. Relative \(L^2\) of the ensemble mean says nothing about whether that distribution is any good — these metrics ask whether it is calibrated.
Draw an ensemble with FlowEvaluator(..., ensemble_size=K) or by calling
flow.sample() \(K\) times, then score it here.
Score in physical units
Denormalize both samples and ground truth before scoring, exactly as with the deterministic metrics.
uq_metrics
¶
Uncertainty Quantification Metrics¶
A conditional flow produces a distribution \(p(u \mid f)\), not a point prediction. Relative L2 of the ensemble mean says nothing about whether that distribution is any good: a model can produce diverse, plausible-looking samples whose spread bears no relation to its actual error. These metrics ask the harder question — is the predicted distribution calibrated?
Three families are provided.
Calibration — does the stated uncertainty match observed error?
credible_interval_coverage()— fraction of truth inside the central \(\alpha\) interval. A calibrated model covers 90% of the truth with its 90% interval.reliability_curve()— coverage across many levels; the diagonal is perfect calibration.rank_histogram()— where the truth ranks among ensemble members. Flat means calibrated, U-shaped means under-dispersed (over-confident), dome-shaped means over-dispersed. Standard in ensemble weather verification and rarely used in ML-for-PDE work.spread_skill_ratio()— ensemble spread divided by the error of the ensemble mean. Should be about 1.
Proper scoring rules — single numbers that reward accuracy and calibration together, and cannot be gamed by lying about uncertainty.
crps_ensemble()— pointwise; the standard probabilistic analogue of MAE.energy_score()— the multivariate generalization. Use it: pointwise CRPS is blind to spatial correlation, and a model that gets marginals right while producing spatially incoherent fields scores well on CRPS and badly here.
Decomposition — where does the uncertainty come from?
variance_decomposition()— splits total variance into aleatoric (within-model sampling spread) and epistemic (disagreement across independently trained models) via the law of total variance.error_spread_correlation()— does predicted spread actually predict error? The practical question for using uncertainty to flag bad predictions.
Convention: samples has shape (K, B, C, *spatial) with K ensemble
members, matching ensemble_relative_l2().
A list of (B, C, *spatial) tensors is also accepted.
UQMetrics
¶
Compute the standard UQ suite in one call.
Example
Source code in flowpde/utils/uq_metrics.py
credible_interval_coverage(samples, target, level=0.9)
¶
Fraction of ground-truth values inside the central credible interval.
For a calibrated model this equals level. Below it means the model is
over-confident (the usual failure); above it means over-dispersed.
Finite-ensemble bias
Empirical quantiles from a finite ensemble under-cover, and the effect grows with the level. For a perfectly calibrated ensemble:
K |
0.5 | 0.9 | 0.99 |
|---|---|---|---|
| 50 | 0.479 | 0.871 | 0.956 |
| 200 | 0.493 | 0.893 | 0.978 |
| 1000 | 0.498 | 0.901 | 0.986 |
So a 99% interval built from 50 samples covers about 96% even when
the model is exactly right. Reading that as over-confidence is an
artefact of the estimator, not a property of the model. Use at least
~200 members for levels up to 0.9, and avoid reporting 0.99 coverage
below ~1000. When comparing models, hold K fixed — the bias
cancels.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
samples
|
SampleInput
|
|
required |
target
|
Tensor
|
Ground truth, |
required |
level
|
float
|
Nominal coverage in (0, 1), e.g. 0.9. |
0.9
|
Returns:
| Type | Description |
|---|---|
Tensor
|
Scalar tensor — empirical coverage, comparable directly to |
Source code in flowpde/utils/uq_metrics.py
reliability_curve(samples, target, levels=None)
¶
Empirical coverage across a range of nominal levels.
Plotting empirical against nominal gives the reliability diagram;
the diagonal is perfect calibration, below it is over-confidence.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
samples
|
SampleInput
|
|
required |
target
|
Tensor
|
Ground truth, |
required |
levels
|
Optional[Sequence[float]]
|
Nominal levels. Defaults to 0.1 … 0.9 in steps of 0.1. |
None
|
Returns:
| Type | Description |
|---|---|
Dict[str, List[float]]
|
|
Dict[str, List[float]]
|
where the error is the mean absolute gap from the diagonal. |
Source code in flowpde/utils/uq_metrics.py
rank_histogram(samples, target, normalize=True)
¶
Histogram of the truth's rank among ensemble members.
At each location, count how many ensemble members fall below the truth,
giving a rank in [0, K]. For a calibrated ensemble the truth is
exchangeable with the members, so ranks are uniform and the histogram is
flat. A U shape means the truth often falls outside the ensemble
(under-dispersed); a dome means the ensemble is too wide.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
samples
|
SampleInput
|
|
required |
target
|
Tensor
|
Ground truth, |
required |
normalize
|
bool
|
Return frequencies rather than counts. |
True
|
Returns:
| Type | Description |
|---|---|
Tensor
|
Tensor of length |
Source code in flowpde/utils/uq_metrics.py
spread_skill_ratio(samples, target, eps=1e-08)
¶
Ensemble spread compared against the error of the ensemble mean.
A calibrated ensemble has spread comparable to its own error, but only in
the limit. A finite ensemble of K members under-estimates the spread
of the distribution it is drawn from: when the truth is exchangeable with
the members, Fortin et al. (2014) give
so the raw ratio \(\text{spread}/\text{skill}\) sits at
\(\sqrt{K/(K+1)} < 1\) even for a perfectly calibrated ensemble.
adjusted_ratio therefore multiplies the raw ratio by
\(\sqrt{(K+1)/K}\), landing at 1 when calibrated. The correction is
negligible at K = 50 (2 %) and substantial at K = 2 (22 %), which
is why its direction has to be pinned by a small-K test.
Returns:
| Type | Description |
|---|---|
Dict[str, float]
|
|
Dict[str, float]
|
below 1 means over-confident, above 1 means over-dispersed. |
Source code in flowpde/utils/uq_metrics.py
crps_ensemble(samples, target)
¶
Continuous Ranked Probability Score, pointwise, fair estimator.
A proper scoring rule: it is minimized only by the true predictive distribution, so a model cannot improve it by misreporting its uncertainty. For a deterministic forecast it reduces exactly to MAE, which makes it directly comparable against a point-prediction baseline.
The fair (unbiased) estimator is used, dividing the second term by
K(K-1) rather than K^2; the biased version rewards small
ensembles for being artificially narrow.
Computed via the sorted formulation, which is O(K log K) and avoids
materializing the K x K pairwise differences.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
samples
|
SampleInput
|
|
required |
target
|
Tensor
|
Ground truth, |
required |
Returns:
| Type | Description |
|---|---|
Tensor
|
Scalar tensor — CRPS averaged over the batch and field, lower better. |
Source code in flowpde/utils/uq_metrics.py
energy_score(samples, target)
¶
Energy score — the multivariate generalization of CRPS.
Norms are taken over the whole field, so unlike pointwise CRPS this is sensitive to spatial structure. A model that reproduces every pointwise marginal correctly but generates spatially incoherent fields scores well on CRPS and poorly here — which is exactly the failure mode worth catching in a PDE surrogate, where solutions are smooth and correlated by construction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
samples
|
SampleInput
|
|
required |
target
|
Tensor
|
Ground truth, |
required |
Returns:
| Type | Description |
|---|---|
Tensor
|
Scalar tensor, lower is better. |
Source code in flowpde/utils/uq_metrics.py
variance_decomposition(model_samples)
¶
Split predictive variance into aleatoric and epistemic parts.
By the law of total variance, for models \(m\) and samples \(x\),
Aleatoric is the spread the flow produces for a fixed model — genuine posterior uncertainty on an ill-posed inverse problem. Epistemic is disagreement between independently trained models — reducible with more data or better training.
The distinction matters for interpretation: on a forward problem the map is deterministic, its posterior is a Dirac, and all observed spread is model error rather than physical uncertainty. Reporting it as "uncertainty" without this split is misleading.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model_samples
|
Sequence[SampleInput]
|
One ensemble per independently trained model, each
|
required |
Returns:
| Type | Description |
|---|---|
Dict[str, float]
|
|
Dict[str, float]
|
variances (not standard deviations). |
Source code in flowpde/utils/uq_metrics.py
error_spread_correlation(samples, target, eps=1e-08)
¶
Does predicted spread actually predict error?
The practical test of an uncertainty estimate: if you flag the highest-variance predictions, do you catch the worst ones? A model can be globally well-calibrated yet assign uncertainty that is uninformative per-sample, in which case it cannot be used to triage predictions.
Also reports top_decile_error_ratio: the mean error of the 10% of
samples with the largest spread, divided by the overall mean error.
Values above 1 mean high-variance predictions really are worse.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
samples
|
SampleInput
|
|
required |
target
|
Tensor
|
Ground truth, |
required |
Returns:
| Type | Description |
|---|---|
Dict[str, float]
|
|