Metrics¶
Error metrics for scoring generated solutions. Report these in physical units — denormalize predictions first.
metrics
¶
PDE Operator Evaluation Metrics¶
Standard metrics used in neural PDE operator benchmarking literature, following the conventions of the FNO paper (Kovachki et al., 2021) and subsequent works (DeepONet, UQNO, etc.).
Primary metric¶
Relative L2 error (also called "relative \(\ell_2\) error" or "nRMSE" in some papers) is the single number reported in virtually every neural operator paper:
It is computed per sample (norms taken over channels + spatial dims) and then averaged over the batch. This removes scale-dependence so errors are comparable across PDEs with very different solution magnitudes.
Secondary metrics¶
- H1 semi-norm error: penalises gradient mismatch on top of value mismatch. Meaningful for smooth PDEs (Poisson, Darcy) where solutions are in \(H^1\). Discretised via first-order finite differences.
- MAE (Mean Absolute Error): complementary to L2, less sensitive to outliers.
- Max pointwise error / relative max error: worst-case analysis.
Flow-model-specific metrics¶
- Ensemble relative L2: generate N samples per condition, average the predictions, then compute relative L2 of the ensemble mean. This separates the model's mean-prediction quality from its uncertainty.
No external dependencies beyond PyTorch are required. The neuraloperator
library provides a similar LpLoss class, but implementing these here keeps
the dependency footprint minimal and makes the definitions transparent.
Usage:
from flowpde.utils.metrics import relative_l2_error, EvalMetrics
# Single batch
err = relative_l2_error(pred, target) # scalar tensor
# All metrics at once
em = EvalMetrics()
results = em(pred, target) # dict
EvalMetrics
¶
Compute a standard suite of PDE evaluation metrics in one call.
By default computes relative L2, H1, MAE, and relative max error. You can restrict which metrics are computed by passing a list of names.
Available metric names: "rel_l2", "h1", "mse", "mae",
"rel_max".
Example
em = EvalMetrics()
results = em(pred, target)
print(results)
# {'rel_l2': tensor(0.0312), 'h1': tensor(0.0421), ...}
em_fast = EvalMetrics(metrics=["rel_l2"])
results = em_fast(pred, target)
Source code in flowpde/utils/metrics.py
__call__(pred, target)
¶
Compute all configured metrics.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pred
|
Tensor
|
Predicted tensor, shape |
required |
target
|
Tensor
|
Ground-truth tensor, same shape. |
required |
Returns:
| Type | Description |
|---|---|
Dict[str, float]
|
Dict mapping metric name → Python float. |
Source code in flowpde/utils/metrics.py
relative_l2_error(pred, target, eps=1e-08)
¶
Relative L2 error — the primary neural-operator benchmark metric.
Norms are taken over all non-batch dimensions (channels + spatial).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pred
|
Tensor
|
Predicted tensor, shape |
required |
target
|
Tensor
|
Ground-truth tensor, same shape. |
required |
eps
|
float
|
Small constant added to the denominator for numerical stability (prevents division by zero for near-zero targets). |
1e-08
|
Returns:
| Type | Description |
|---|---|
Tensor
|
Scalar tensor — the mean relative L2 error over the batch. |
Source code in flowpde/utils/metrics.py
relative_l2_error_batch(pred, target, eps=1e-08)
¶
Per-sample relative L2 errors.
Like relative_l2_error() but returns a (B,) tensor instead of
averaging, useful for inspecting the error distribution.
Source code in flowpde/utils/metrics.py
mse(pred, target)
¶
Mean Squared Error averaged over all dimensions including batch.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pred
|
Tensor
|
Predicted tensor, shape |
required |
target
|
Tensor
|
Ground-truth tensor, same shape. |
required |
Returns:
| Type | Description |
|---|---|
Tensor
|
Scalar tensor. |
Source code in flowpde/utils/metrics.py
mae(pred, target)
¶
Mean Absolute Error averaged over all dimensions including batch.
Less sensitive to large outliers than MSE.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pred
|
Tensor
|
Predicted tensor, shape |
required |
target
|
Tensor
|
Ground-truth tensor, same shape. |
required |
Returns:
| Type | Description |
|---|---|
Tensor
|
Scalar tensor. |
Source code in flowpde/utils/metrics.py
relative_max_error(pred, target, eps=1e-08)
¶
Relative maximum pointwise error (L∞ norm, normalised by target range).
Normalises by the standard deviation of each sample's target field so the result is comparable across PDEs with different solution scales.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pred
|
Tensor
|
Predicted tensor, shape |
required |
target
|
Tensor
|
Ground-truth tensor, same shape. |
required |
eps
|
float
|
Stability constant. |
1e-08
|
Returns:
| Type | Description |
|---|---|
Tensor
|
Scalar tensor — mean relative max error over the batch. |
Source code in flowpde/utils/metrics.py
h1_error(pred, target, eps=1e-08)
¶
Relative H¹ semi-norm error for 2-D spatial fields.
The H¹ semi-norm adds a gradient penalty on top of the L2 value
mismatch. It is the standard "H1 loss" used in the neuraloperator
library and is meaningful for smooth PDEs (Poisson, Darcy) whose
solutions are in the Sobolev space H¹.
Gradients are estimated with first-order finite differences along each spatial axis. For 1-D inputs the function falls back to plain relative L2 error (only one spatial axis, same as L2 loss for 1-D problems where the FD gradient information is less informative).
where
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pred
|
Tensor
|
Predicted tensor, shape |
required |
target
|
Tensor
|
Ground-truth tensor, same shape. |
required |
eps
|
float
|
Stability constant. |
1e-08
|
Returns:
| Type | Description |
|---|---|
Tensor
|
Scalar tensor — mean relative H¹ error over the batch. |
Source code in flowpde/utils/metrics.py
ensemble_relative_l2(preds, target, eps=1e-08)
¶
Evaluate a flow model's ensemble of predictions.
Flow matching models are generative — they produce a distribution p(u | f) rather than a single prediction. By sampling multiple times from the same condition we can decompose performance into:
- Mean error: relative L2 of the ensemble mean vs. ground truth. Reflects the model's average prediction quality.
- Sample spread (std): average standard deviation of the ensemble, normalised by the ground-truth norm. Reflects how much the model explores the solution space.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
preds
|
List[Tensor]
|
List of |
required |
target
|
Tensor
|
Ground-truth tensor, shape |
required |
eps
|
float
|
Stability constant. |
1e-08
|
Returns:
| Type | Description |
|---|---|
Dict[str, Tensor]
|
Dict with keys: |
Dict[str, Tensor]
|
|
Dict[str, Tensor]
|
|
Dict[str, Tensor]
|
|