Choosing a model topic
Choosing a model
Start with the smallest plausible model, add one component at a time, and use the residuals rather than the likelihood to decide whether each one helps.
Start here
| what you have | what to use |
|---|---|
| noisy readings of something that drifts | LocalLinearTrend (the default; a smoothing spline) |
| a level with no persistent direction | LocalLevel |
| a pattern on a calendar you know: a week, a year | TrigonometricSeasonal(period: 7, harmonics: 2, ...), or period: 365.25 |
| named, dated events: a holiday, a course of medication | RegressionComponent with IndicatorRegressors |
| a covariate that changes at known instants | RegressionComponent with a StepRegressor |
| correlated wobble the trend should not be chasing | Matern.oneHalf(...) alongside the trend, with a noise floor |
| a rhythm that is approximate, or whose period you do not know | StochasticCycle |
The choice between LocalLinearTrend and LocalLevel is about whether the
series has a direction that persists. A trend carries a slope and extrapolates
it; a level does not. If "still going down" makes sense for your data, use the
trend.
Components has the detail on each. This page is about putting them together.
Add one component at a time
final d = model.diagnose(data);
d.ljungBox(lags: 14, fittedParameters: model.parameterCount);
d.autocorrelation(7);
A pattern the model is missing shows up as autocorrelation long before it shows up as a visibly bad fit. Add the component that explains the spike, which is not always the one you expected.
A lag counts observations, not days: lag one is the previous reading, whenever that was. Under the model the residuals are independent however unevenly they are spaced, but a weekly pattern appears at lag seven only when the sampling is roughly daily. On thinner data, look wherever a period's worth of readings falls.
Pass fittedParameters when the variances were estimated on the same data.
Each estimated parameter costs a degree of freedom, and leaving them out makes
the test optimistic. Pass parameterCount: it counts the searched variances
but not the concentrated-out measurement variance, which follows Box and
Jenkins in not charging for the residual scale.
Comparing models
Do not compare logMarginalLikelihood across models with different diffuse
structure.
Under exact diffuse initialisation it is a restricted likelihood: the flat
directions have been integrated out against an improper prior of unit density, so
the result carries their units. comparability_test.dart checks both of these:
- Writing a regression column in grams rather than kilograms shifts it by exactly
log 1000, while the fit, the posterior and the coefficient are unchanged. - Measuring time in half-days rather than days shifts it by exactly
log 2per diffuse direction that carries a time dimension, for the same process on the same data.
Either can reverse a comparison. The same caveat applies to REML in general.
FitResult.isComparableWith is the check, and diffuseDimension is what has to
match:
| component | flat directions |
|---|---|
LocalLevel |
1 |
LocalLinearTrend |
2 |
TrigonometricSeasonal |
2 per harmonic |
RegressionComponent |
1 per column |
Matern, StochasticCycle |
0 |
So a trend can be compared with a trend plus a Matérn, and cannot be compared with a trend plus a weekly seasonal.
These can be compared across any models:
- The fitted noise level. A model missing a real component has to explain it as noise, and reports a larger noise level than the truth. This is usually the most useful single number.
- The Ljung–Box p-value.
- Out-of-sample error. Fit on a prefix,
forecastthe rest, score it. Slower, but the only one of the three that is not computed in-sample.
On two years of daily readings built from a trend, a weekly sinusoid and noise with standard deviation 0.3, the first two give a clear answer:
| model | fitted noise sd | Ljung–Box, 14 lags |
|---|---|---|
| trend only | 0.472 | statistic 1794 on 13 df, p below 1e-300 |
| trend + weekly | 0.302 | statistic 13.9 on 12 df, p = 0.31 |
The noise level matches the truth to three decimal places, and the Ljung–Box test goes from certain rejection to unremarkable. The two likelihoods have six and two flat directions and cannot be compared.
Read the warnings
fitted.atBracketEdge; // did anything finish on a bound?
for (final warning in fitted.warnings) print(warning);
atBracketEdge is the quick check. warnings is empty when there is nothing
to report, and otherwise lists:
- a reading far from what the rest of the data predicts, more than six typical errors away; see Bad readings;
- a parameter that finished on a bound. Its estimate is a boundary rather than an interior optimum, and the width beside it is one-sided. At the bottom, the variance of a trend or seasonal means the component is fitted as fixed over time, and the variance of a stationary component means it contributes nothing;
- a stationary component's variance at the top of its bracket, which means
it has taken over the measurement noise: pass
minimumMeasurementVariance; - a parameter the data barely constrains, meaning a plateau over two decades wide;
- a width measured while another parameter of the same component sat on a
bound. Every width in
plateauDecadesByParameteris conditional on the other parameters, so it is then taken at that bound. The common case is aStochasticCycleon a series with no cycle in it; its documentation works it through.
plateauDecadesByParameter is NaN for a parameter that is not searched on a
log scale; a damping factor is a logit, and a width in logits over ln 10 is not
decades of anything. plateauWidthByParameter carries the raw number for every
parameter.
What competes with what
Components that can produce the same shape trade off against each other, and the fit does not always tell you which one took it.
- A trend and a seasonal, over less than one period. A rigid sinusoid over
half a cycle is very nearly a constant plus a slope, so an annual component on
180 days will draw a swing in a series that has none. The component's own
posterior standard deviation then comes back larger than the amplitude it
drew, so check
componentVariancealongsidecomponentMean. Shrinking the variance to zero does not remove the component; it only stops the pattern changing. - A trend and a Matérn. Both are slow. With both free the fit is measurably worse than with neither: the two compete for the same slow variation and the trend's variance ends up undetermined over five decades. If you add a Matérn deviation, consider fixing the trend's smoothing rather than fitting it.
- A Matérn and the measurement noise. A length scale below the sampling
interval is white noise, so
fitfloors it at the typical gap between visits. Above the floor the two can still trade: when readings carry correlated day-to-day variation, the likelihood can prefer a large Matérn and a noise level near zero.fitsearches again from a larger noise share when a Matérn's variance ends at the top of its bracket and keeps the better optimum, andwarningssays when the variance is still there; aminimumMeasurementVarianceat the instrument's resolution rules the corner out. - A cycle and a free level. A random walk with a free variance can follow any
wiggle, so a
StochasticCyclebeside an unconstrainedLocalLeveloften loses. It fails in two ways andwarningsreports both: on a series with a real cycle the level's variance runs to the top of its bracket and interpolates the data, and on a series without one the level shrinks away while the cycle's variance and period come back flat over eight and two decades. In both cases the fit has not determined anything.
When to stop
Stop when adding a component no longer reduces the fitted noise level, the
Ljung–Box p-value is unremarkable, and warnings is empty. A model in which
every component has a clear job is more useful than one that fits slightly
better but cannot say which part did the work.
A worked comparison on real data
tool/calibration/
in the GitHub repository fits several candidate models to a weight diary (a
real export, or four synthetic ones) and prints the fitted smoothing, how well
it is determined, the noise level and a Ljung–Box test, for the whole history
and at several earlier lengths. It is not part of the published package; run it
from a clone of the repository:
dart run tool/calibration/calibrate.dart --all
dart run tool/calibration/calibrate.dart --all my-export.txt
Everything runs locally. Its README explains how to read the tables, including why the likelihood column can only be compared between rows of equal diffuse dimension.
Classes
- FitResult Choosing a model
- What StructuralModel fitting returned, together with enough diagnostics to tell a well-determined answer from a shrug.
- InnovationDiagnostics Choosing a model
- What the one-step-ahead prediction errors say about the model that produced them.
- Penalty Choosing a model
- A term added to the log-likelihood before it is maximised.
- StructuralModel Getting started Choosing a model
- An additive structural time-series model: a list of components plus observation noise.
Functions
-
fit(
StructuralModel initial, List< Choosing a modelObservation> observations, {double lowerLogRatio = -20, double upperLogRatio = 10, int scanPoints = 25, double tolerance = 1e-4, Penalty? penalty, double? fixedMeasurementVariance, double? minimumMeasurementVariance, SearchStart start = SearchStart.bracketScan}) → FitResult -
Estimates a model's variances from
observationsby maximum marginal likelihood.