Normalization¶
Flow matching transports \(\mathcal{N}(0, I)\) to the data, so raw PDE fields with non-unit scale make the velocity regression badly conditioned. Always normalize.
Statistics are keyed by raw field name (source, solution, kappa, initial,
final) rather than by role, which is why one normalizer stays correct across
problem='forward' and problem='inverse'.
Warning
Never refit normalization on validation or test data — share the same train normalizer instance.
normalization
¶
Field Normalization for PDE Datasets¶
Flow matching transports a standard Gaussian base distribution to the data
distribution. When the data has a very different scale from N(0, I) the
velocity targets x_1 - x_0 inherit the raw data magnitude, which makes the
regression problem badly conditioned. Standardizing each PDE field to roughly
zero mean and unit variance removes that mismatch.
Statistics must always be fitted on the training split and then reused verbatim for validation/test data, otherwise the evaluation leaks information about the held-out set:
train_ds = generator.generate(num_samples=1000, seed=0)
test_ds = generator.generate(num_samples=200, seed=1)
normalizer = FieldNormalizer.from_dataset(train_ds)
train_ds.set_normalizer(normalizer)
test_ds.set_normalizer(normalizer) # same statistics, not refitted
Metrics should be reported in physical units, so predictions are mapped back
with denormalize() (or
denormalize_channels() for targets that concatenate
several fields) before computing errors.
FieldNormalizer
¶
Per-field mean/std standardization keyed by PDE field name.
Fields are addressed by their raw name ('source', 'solution',
'kappa', 'initial', 'final') rather than by their role in the
learning problem, so a single normalizer stays correct when the same data
is used for both the forward and the inverse direction.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
stats
|
Optional[Dict[str, Dict[str, float]]]
|
Mapping |
None
|
eps
|
float
|
Floor applied to standard deviations to avoid division by zero. |
1e-08
|
Example
normalizer = FieldNormalizer({'solution': {'mean': 2.0, 'std': 4.0}}) z = normalizer.normalize('solution', torch.tensor([6.0])) z tensor([1.]) normalizer.denormalize('solution', z) tensor([6.])
Source code in flowpde/datasets/normalization.py
36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 | |
fields
property
¶
Names of all fields this normalizer knows about.
from_dataset(dataset, fields=None, eps=1e-08)
classmethod
¶
Build a normalizer from statistics a dataset already carries.
Generators compute per-field mean/std at construction time and store
them in metadata['stats'], so no second pass over the data is
needed.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dataset
|
'Any'
|
Dataset exposing |
required |
fields
|
Optional[Iterable[str]]
|
Restrict to these field names. Defaults to every field
that has statistics, excluding |
None
|
eps
|
float
|
Floor applied to standard deviations. |
1e-08
|
Returns:
| Type | Description |
|---|---|
'FieldNormalizer'
|
A fitted |
Source code in flowpde/datasets/normalization.py
from_tensors(tensors, eps=1e-08)
classmethod
¶
Fit directly from raw tensors, one entry per field.
Source code in flowpde/datasets/normalization.py
add_field(name, mean, std)
¶
normalize(name, tensor)
¶
Standardize a field to zero mean and unit variance.
Fields with no registered statistics pass through unchanged, so
auxiliary channels such as obs_mask stay binary.
Source code in flowpde/datasets/normalization.py
denormalize(name, tensor)
¶
Map a standardized field back to physical units.
Source code in flowpde/datasets/normalization.py
denormalize_channels(names, tensor, channel_dim=1)
¶
Denormalize a tensor whose channels concatenate several fields.
Used for targets such as Darcy's inverse_mode='both', where the
target is cat([kappa, source]).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
names
|
Sequence[str]
|
Field name per channel group, in channel order. |
required |
tensor
|
Tensor
|
Tensor with |
required |
channel_dim
|
int
|
Dimension holding the channels (default: 1). |
1
|
Returns:
| Type | Description |
|---|---|
Tensor
|
Tensor of the same shape, in physical units. |
Source code in flowpde/datasets/normalization.py
state_dict()
¶
Serializable state, suitable for storing next to model weights.
load_state_dict(state)
¶
Restore statistics saved by state_dict().