Forecast fev-bench Tasks Using Threedx¶
Using the fev-bench benchmark, it’s possible to compare one forecast method against others in a standardized fashion (see the paper and Hugging Face space). The fev-bench benchmark defines data, data splits, forecast horizons, and evaluation metrics. All that is needed to evaluate a new method is to apply it onto the data and evaluate the results. For that, we’ll need some helper functions to pass data in the expected formats, along with the fev package.
The fev benchmark is a set of “tasks” which define all parameters of an experiment needed to make an evaluation of different forecast methods comparable, avoiding different authors defining different horizon lengths or data subsets.
Define the Task¶
Below, define a fev task by choosing one of the fev data sets stored on HuggingFace. Here, pick the M5 competition dataset of Walmart product sales, aggregated to monthly observations.
A defined task provides an iterable through which the data will be downloaded from HuggingFace as the forecast method iterates through the time series.
Note also that we specify quantile levels, as fev tasks are evaluated based on both point forecasts and marginal quantiles (but do not evaluate sample paths).
[1]:
import fev
task = fev.Task(
dataset_path="autogluon/fev_datasets",
dataset_config="m5_1M",
quantile_levels=[0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9],
horizon=12,
num_windows=1,
initial_cutoff=-12,
window_step_size=12,
min_context_length=6,
seasonality=12,
eval_metric="SQL",
extra_metrics=[
'MASE',
{'name': 'WAPE', 'epsilon': 1.0},
{'name': 'WQL', 'epsilon': 1.0}
],
target="target",
task_name="m5_1M",
id_column="id",
timestamp_column="timestamp"
)
Forecast Using Seasonal Naive¶
Based on examples provided in the fev repository, it’s possible to understand both the shape in which training data is provided by fev tasks, as well as the shape of forecasts that fev expects during evaluation.
For the purpose of this tutorial, however, we’ll use some helper functions that make it straightforward to apply Threedx onto fev tasks.
Using _forecast_task_as_dataset_dict_per_window(), univariate forecast methods can be applied onto the task iterable that iterates through time series. The result will be a list of datasets.DatasetDicts, each containing the quantile forecasts for each time series at one of the task’s forecast windows. The M5 task has a single window, so the list will contain only a single DatasetDict.
_forecast_task_as_dataset_dict_per_window only handles data, moving from one data structure to another, iterating through the task and collecting forecast results. The actual forecasting is done by the function passed to its using argument. Specify it as a function that takes a univariate time series y in the form of a list of a one-dimensional numpy array, along with quantile_levels, horizon, and period_length parameters.
Here, create first a set of seasonal naive point predictions as baseline using _predict_fev_quantiles_using_latest_period(). Follow those up with Threedx forecasts using _predict_fev_quantiles_using_threedx(). The latter requires additional Threedx-specific arguments parameter_grid and loss.
[2]:
import threedx as tdx
from threedx._fev import (
_forecast_task_as_dataset_dict_per_window,
_predict_fev_quantiles_using_latest_period,
_predict_fev_quantiles_using_threedx
)
[3]:
forecasts_using_latest_period = \
_forecast_task_as_dataset_dict_per_window(
using=_predict_fev_quantiles_using_latest_period,
task=task,
)
Warning: You are sending unauthenticated requests to the HF Hub. Please set a HF_TOKEN to enable higher rate limits and faster downloads.
Dropped 1126 out of 30490 time series because they had fewer than min_context_length (6) observations before cutoff (-12) or fewer than horizon (12) observations after cutoff.
Evaluate the forecasts in standardized fashion using the Task’s .evaluation_summary() method.
Note that SQL will be identical to MASE and WQL will be identical to WQL as the above seasonal naive method produces point predictions only (i.e., all quantiles are identical).
[4]:
evaluation_of_latest_period = task.evaluation_summary(
forecasts_using_latest_period,
model_name="latest_period"
)
print(evaluation_of_latest_period)
{'model_name': 'latest_period', 'dataset_path': 'autogluon/fev_datasets', 'dataset_config': 'm5_1M', 'horizon': 12, 'num_windows': 1, 'initial_cutoff': -12, 'window_step_size': 12, 'min_context_length': 6, 'max_context_length': None, 'seasonality': 12, 'eval_metric': 'SQL', 'extra_metrics': ['MASE', {'name': 'WAPE', 'epsilon': 1.0}, {'name': 'WQL', 'epsilon': 1.0}], 'quantile_levels': [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9], 'id_column': 'id', 'timestamp_column': 'timestamp', 'target': 'target', 'generate_univariate_targets_from': None, 'known_dynamic_columns': [], 'past_dynamic_columns': [], 'static_columns': [], 'task_name': 'm5_1M', 'test_error': 1.3266370566831691, 'training_time_s': None, 'inference_time_s': None, 'num_forecasts': 29364, 'dataset_fingerprint': '4dcc997759fda726', 'trained_on_this_dataset': False, 'fev_version': '0.10.0', 'SQL': 1.3266370566831691, 'MASE': 1.3266370566831724, 'WAPE': 0.505579527864656, 'WQL': 0.5055795278646565}
[5]:
print(f"Seasonal Naive MASE: {evaluation_of_latest_period["MASE"]}")
print(f"Seasonal Naive WAPE: {evaluation_of_latest_period["WAPE"]}")
Seasonal Naive MASE: 1.3266370566831724
Seasonal Naive WAPE: 0.505579527864656
These results could now be compared against what has been published for other models (available for example on https://github.com/autogluon/fev/tree/main/benchmarks/fev_bench/results).
For example, the StatsForecast-based Seasonal Naive results report a MASE of 1.3266 and a WAPE of 0.5108 which the above point forecasts reproduce nicely.
Forecast Using Threedx¶
To forecast with Threedx, repeat the above but pass the _predict_fev_quantiles_using_threedx() helper function. It wraps the Threedx object initialization and subsequent .fit() and .predict() calls, and aggregates the sample paths into marginal quantiles required by fev.
Note: Per default, _predict_fev_quantiles_using_threedx() uses the observation-driven approach to generating sample paths.
To not create the parameter grid again and again for every time series, initialize it once upfront.
[6]:
parameter_grid = tdx.initialize_parameters_at_random(
size=1000,
seed=582,
include_edge_cases=True,
)
forecasts_using_threedx = \
_forecast_task_as_dataset_dict_per_window(
using=_predict_fev_quantiles_using_threedx,
task=task,
parameter_grid=parameter_grid,
loss=tdx.mae
)
Dropped 1126 out of 30490 time series because they had fewer than min_context_length (6) observations before cutoff (-12) or fewer than horizon (12) observations after cutoff.
[7]:
evaluation_of_threedx = task.evaluation_summary(
forecasts_using_threedx,
model_name="threedx"
)
print(evaluation_of_threedx)
{'model_name': 'threedx', 'dataset_path': 'autogluon/fev_datasets', 'dataset_config': 'm5_1M', 'horizon': 12, 'num_windows': 1, 'initial_cutoff': -12, 'window_step_size': 12, 'min_context_length': 6, 'max_context_length': None, 'seasonality': 12, 'eval_metric': 'SQL', 'extra_metrics': ['MASE', {'name': 'WAPE', 'epsilon': 1.0}, {'name': 'WQL', 'epsilon': 1.0}], 'quantile_levels': [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9], 'id_column': 'id', 'timestamp_column': 'timestamp', 'target': 'target', 'generate_univariate_targets_from': None, 'known_dynamic_columns': [], 'past_dynamic_columns': [], 'static_columns': [], 'task_name': 'm5_1M', 'test_error': 1.099423808419335, 'training_time_s': None, 'inference_time_s': None, 'num_forecasts': 29364, 'dataset_fingerprint': '4dcc997759fda726', 'trained_on_this_dataset': False, 'fev_version': '0.10.0', 'SQL': 1.099423808419335, 'MASE': 1.2132898443475035, 'WAPE': 0.45562166895488493, 'WQL': 0.40151684786803704}
[8]:
print(f"Threedx SQL: {evaluation_of_threedx["SQL"]}")
print(f"Threedx MASE: {evaluation_of_threedx["MASE"]}")
print(f"Threedx WQL: {evaluation_of_threedx["WQL"]}")
print(f"Threedx WAPE: {evaluation_of_threedx["WAPE"]}")
Threedx SQL: 1.099423808419335
Threedx MASE: 1.2132898443475035
Threedx WQL: 0.40151684786803704
Threedx WAPE: 0.45562166895488493
For comparison, a few selected models for which fev benchmark results are available, along with Threedx: