Empirical Testing Of Macro Models

The ability to make money after model implementation is a simple quantitative metric, although one might need to wait for a large enough sample to test this.

From a bond market practitioner's perspective, model testing is straightforward: does it make money? The ability to make money after model implementation (and not just "out of sample") is a simple quantitative metric -- although one might need to wait for a large enough sample to test this. Not everyone is a market participant, and they want to evaluate models on other metrics (e.g., does it help guide policy decisions?). However, the key insight of the "does is make money?" metric is that it is related to the more vague: "does it offer useful information about the future?" It is entirely possible for a model to have some statistical properties that are seen as "good" -- yet offer no useful information about the future. 

Usefulness in Forecasting and Pricing

Although one might be able to make money from a mathematical model in any number of ways, I am considering two types of financial models that are of interest.

  • Forecasting models that generate buy/sell signals.
  • Pricing models. Although this sounds unusual in a macroeconomics context, this is related to DSGE models, given their similarity to arbitrage-free pricing models (link to the previous discussion).

Forecasting models are probably what most people would think of, the structure of DSGE models implies a need for worrying about pricing concepts. The key observation one can make about financial forecasting models is that they are not evaluated based on statistical tests (r-squared, whatever), rather the profits they generate after model creation. That is, passing statistical tests does not translate into a useful model.

Since there is less to say about pricing models, I will discuss them first.

Pricing Models

For a pricing model, the question is whether we can calibrate them to available data to generate a probability distribution that is internally consistent. Being arbitrage-free is the usual criterion, but the test is more complex. In fixed income, one needs to be able to price benchmark instruments (e.g., standard spot swaps, volatility cube for vanilla swaptions) as well as more complex instruments in an arbitrage-free manner. (This is tricky because the pricing of benchmark instruments is not enough to pin down the distribution, further assumptions about behavior are needed, such as how to smooth forward rates.)

Although DSGE models resemble fixed income pricing models -- they are projecting forward probability distributions -- the issue is that there is no observable curve for many key variables (mainly wages, although one can also question the relationship between breakeven inflation and the model inflation). There is no forward market in wages and physical consumer goods baskets -- so what are we going to arbitrage?

Where things might get interesting in this context is if we used economic arguments to enforce a relationship between nominal forward rates and forward breakeven inflation rates. (This would likely be based on assumed properties of real rates.) Affine curve structures for inflation-linked bonds already are models on this forward structure. Could one create models that imply "economic arbitrage" between the markets? Since I have not run out to set up a hedge fund based on this concept, it is clear that I am skeptical. (One of the stumbling blocks is the illiquidity of inflation options.)

Otherwise, we have to not worry about the complexity of nonlinear DSGE model solutions, and just use their outputs as a forecasting model. This is presumably how most people will look at them. We turn to forecasting models next.

Judging Forecast Models: How?

Using profitability as a measure of success of models has the advantage of simplicity. For most macro model users, there is no obvious replacement. However, I think the metric of success has to be related to how you intend to use the model, and what forms of errors you are concerned about. Standard statistical tests may not take these concerns into account.

For example, a model that does a good job of tracking real GDP growth during an expansion, but cannot predict a recession, is obviously useless as a recession forecasting tool, even if the average (in some sense) forecasting error is lower than competing models.

Another concern is the time frame for forecasts. Some models may only offer a forecast for an upcoming time period. Such a short forecast horizon can be of use if it is asset prices, but it is unclear how useful it is as a macroeconomics tool -- other than for forecasters who have to submit results to surveys. However, such a model may be the result of an econometric exercise, such as fitting a linearization of a DSGE model.

If we have a model that generates forward trajectories, we presumably want to compare them to realized data. The issue is that the forecast covers multiple time periods, and so we are no longer tracking data that has a single value at a given time point (which is the case for a buy/sell signal in a trading model). There seem to be a few broad approaches one could take to develop a metric.

  • Fix a forecast horizon (e.g., six months), and compare the error between forecast and realized.
  • Generate an error measure for the forward forecast (up to some horizon limit) versus realized. This leaves open how to weigh forecast errors.
  • If we are looking for particular events (e.g., recessions), compare the realized events versus the forecast.

My concern with just looking at an average error may not be meaningful. During an expansion, economic variables tend to follow smooth trajectories. If a model always generates smooth forecast trajectories, it might be able to get reasonably low average errors during the expansion. However, it might fail miserably during the short periods of economic turbulence around recessions.

Nonlinear DSGE Model Evaluation

I will just make a few points from comments since I expect to re-visit this topic.

  • I was quite unimpressed by the methods used to evaluate the empirical effectiveness of nonlinear DSGE models. 
  • I am not considering large-scale models ("Frankenmodels") that are used by central banks to generate forecasts. Those models are a mixture of theoretically incompatible components, and forecasts are being generated by teams of analysts. It is unclear how many ad hoc adjustments are being made on the fly to get the "right" numbers.
  • It is unclear to me how a nonlinear DSGE model is supposed to be recalibrated to generate a probability distribution that evolves from period-to-period, most analyses I have seen are of analyzing shocks (used in "all else equal" analysis). It seems to me that the model would need to be calibrated against observed yield curves to have the model match current cyclical conditions. I have not seen research that has attempted such a fitting.
  • The structure of the models implies that the forward state trajectory will smoothly converge to a steady-state (as happens in forward curves). Just eyeballing historical data tells us that this might be a good fit during an expansion, but fail miserably around recessions. Is this the behavior we want in a forecasting model?

 

Comments