Debating “Superforecaster” Training: What International Affairs Professionals Should Know

Updated research indicates that when a forecaster uses probability and de-biasing training, their behavior leads to better “noise” reduction in their information.

This blog is the fourth in a series applying the academic research on superforecasting to international affairs practitioners’ daily workflow.  The first post explained the mindset of “superforecasters”—individuals who can make more accurate forecasts of near-term geopolitical events, as described in Philip Tetlock and Dan Gardner’s 2015 book.  The second post demystified what it means that forecasters “beat” the U.S. Intelligence Community.  The third post described the de-biasing training that Mellers and Tetlock in 2016 and 2024 said had increased individual’s accuracy between 6-11% across four years of a U.S.  government forecasting tournament. 

As of 2025, researchers have further examined the mechanisms by which the Good Judgment Project’s (GJP’s) training improved forecasters’ accuracy in the initial tournaments.  There has also been a new critique arguing that superforecaster behaviors—not training—played a larger role in their tournament success.  (The first post in this series also included an earlier critique of Tetlock and Gardener’s 2015 book.)

International affairs practitioners working in the international affairs knowledge industry (IAKI) must support decision-making in a political, competitive, and fast-paced work environment.  IAKI, in some ways, is similar to a geopolitical forecasting tournament, but without Brier scores and time-bound specific questions.  Understanding the latest research on geopolitical forecasting can help practitioners understand what professional development training fits their needs.  Understanding the critiques help organizations understand when using a forecasting organization or tournament might be helpful.

Quick background

In the initial 2011-2015 Intelligence Advanced Research Projects Activity’s (IARPA’s) Aggregative Contingent Estimation (ACE) tournaments, the Good Judgment Project (GJP) beat competing teams the first two years, and were the only participants thereafter.  After the first two tournament years, GJP adopted the CHAMPS KNOW training, which adjusted the political reasoning section of their short de-biasing module.  The probability section remained the same. 

Forecasting organizations and academics are still in the process of using scientific and statistical rigor to understand how forecasters actually win tournaments, and how to reliably train others for future success.  The details matter because many government and corporate employers use tournament scores to inform organizational training, recruitment, and contracting.  The ongoing GJP tournaments also provide practitioners an opportunity to establish their own forecasting Brier scores, which can likely beef up their resumes.

Updated research on GJP training

Mellers and Tetlock in their 2016 and 2024 studies reported that individual forecasting accuracy improved after forecasters received GJP’s short de-biasing training.  The best performers—the “superforecasters”—were also among those who had received training, compared to the other participants in the control group. 

The 2024 study summarized the author’s statistical work to determine the reasons why training had improved accuracy.  The authors described that, “Our working hypotheses were that training would reduce cognitive biases, such as base-rate neglect, over confidence, and the confirmation bias, by encouraging forecasters to adopt an outside view… To our surprise, the [statistical] BIN model (described below) revealed that over 50% of the increase in accuracy…was due to noise reduction.”

“Noise” is what leads to variability in judgments which should be the same.  For example, when two people of the same background and skillset, with the same information, come to different conclusions, this indicates there’s noise in the system that must be identified.  Some information signals influencing the judgments are relevant, some are irrelevant, i.e. noise. 

Note that decreasing noise is a way to de-bias, though it was not described as such in this study based on this particular forecasting tournament data and training.

In the a different study describing the bias-information-noise model (BIN), the authors posited that in “a signal universe that contains all past and future signals of positive or negative relevance to [an] event…Forecasters sample and interpret signals with varying skill and thoroughness…Forecasters may sample relevant signals (increasing partial information) or irrelevant signals (creating noise).  Forecasters may center signals incorrectly (creating bias).  We can model the accumulation of these signals with continuous variables that exhibit degrees of bias, partial information, and noise in their forecasts.”

The training taught forecasters which type of information would best support accurate tournament forecasts, and the time they’d need to retrieve such data.  The training also advised that forecasters select questions of interest, likely given the effort that is needed to de-bias information and limit irrelevant signals. 

The key insight is that reducing irrelevant information indirectly lowered bias.  The forecasters’ actions limited noise; they did not first seek to directly mitigate their bias, for example by asking themselves, “Do I have a herd mentality right now?”

Updated Critique

A new 2025 critique found that when researchers flip the perspective from looking at the training variable to looking at forecasters as variables, other statistically significant individual characteristics better predicted high performance.  Hauenstein and team’s critique used a model with three person-level variables: latent forecast ability (individual differences in general forecasting skill), latent response strategy (individual differences in a forecaster’s willingness to respond consistently early or late in the response window across questions), and latent response propensity (individual differences in one’s choices of which questions to answer). 

Hauenstein and team showed that using the first two years of tournament data (before other participants bowed out) accurate forecasters could be identified by their behavior, much more than their participation in training.  In other words, when picking out superforecasters, their behavior more accurately identifies them, more than first identifying the group of trainees.  Not everyone who received training did well.

The critique authors concluded, “Given that latent forecasting ability was the most important predictor, this analysis would seem to confirm the special status of superforecasters as being unusually accurate forecasters.  However, superforecasters differ in more ways than just latent forecast ability because the latent person-level variable “response propensity” could uniquely classify superforecasters 80% of the time by itself….[they] answered more questions, selected easier items on average, and responded sooner in the response interval.”

In other words, superforecaster’s latent abilities like active open-mindedness and high numeracy were dominant contributors to their success.  Then, after receiving GJP’s training they learned the specific tournament strategies for winning Brier scores, such as responding early on questions with less debate to have a better running daily average score. 

GJP’s training advised forecasters to select questions of interest and for which they could find base-rate information, not to select “easy” questions.  The critique’s finding that the superforecaster’s behavior reliably included selecting questions with less debate, responding early in the response period (as a hack for Brier scores), and responding to more questions is an important takeaway from judging forecasting accuracy from a tournament.  Participants incentives and the specific tournament platform’s scoring challenges researchers’ ability to draw scientifically rigorous conclusions. 

However, at least according to GJP’s website, their training continues to support better human performance on forecasts compared to many algorithms, even outside of tournaments.  This is also the experience of current international affairs practitioners.  When practitioners change behavior to focus on projects of interest, lower noise with base-rates and precise questions, and manage time to allow time for updating, they can improve their accuracy.

Behavior changes can make the difference

The 2024 and 2025 research reveals something international affairs practitioners need to internalize: training works, but not always how we think it does.  Tetlock’s team discovered training improved accuracy primarily through noise reduction—teaching forecasters which information signals to pursue and how much effort to invest in each question.  Meanwhile, Hauenstein’s critique showed that superforecasters exhibited distinct behavioral patterns: they answered more questions, chose topics strategically, and responded early to allow time for updates.  These aren’t contradictory findings.  The training didn’t just teach cognitive skills; it shaped the behaviors that enabled those skills to pay off.  Superforecasters learned to be selective, weigh-in early, and iterate—behaviors that naturally reduced noise by filtering irrelevant signals and focusing effort where it mattered most.

For international affairs professionals, this research offers a reality check about what professional development can and cannot do.  No training program will overnight transform you into a superforecaster if you don’t adopt the behaviors that make the training effective.  The good news?  These are learnable behaviors that work well with your background and interest in foreign language and culture. 

Understanding this research—including the skeptical perspectives—helps you evaluate what professional development is worth your time and your organization’s money.  Training programs promising to eliminate cognitive biases through awareness exercises alone are insufficient.  The evidence suggests effective training teaches you to structure your information-gathering process—such as in the Crossthink Process or Authoritative Source Checklist, not just recognize your biases in the abstract.  Whether you’re considering superforecasting tournaments, open-source intelligence courses, or AI forecasting tools, ask yourself, “Will this change my behavior in ways that reduce noise?”  That’s the standard the research points toward.

Shutdown

There are pros and cons to attempting to extract best practices from geopolitical forecasting tournaments.  International affairs professionals work in a competitive, fast-paced, and political environment where their value is determined by how timely, relevant, convincing, and actionable their information is.  This environment—called IAKI—is comparable to the tournament environment, but still very different. 

The questions required for forecasting tournaments are much more specific and timebound than those practitioners get on-the-job.  Furthermore, performing well in a tournament also requires forecasters provide specific answers with probability estimates, which if provided as answers to your government or corporate boss would likely result in their being confused and annoyed.

The value of understanding the details is in seeing that training on de-biasing and probabilistic thinking like those I teach in the International Affairs Professional Development Course can and should be adapted for the specific work environment. 

Explaining the updated research on what continues to be a hot topic—the ability of trained outsiders to beat the U.S. intelligence community in geopolitical forecasts with unclassified questions—will always be important for international affairs professionals.  This is especially true if your organization is funding new training resources, like the latest push for open source intelligence training, considering engaging a forecasting organization, or purchasing artificial intelligence (AI) tools.

The next and last blog in this series on superforecasting will cover how the GJP has contributed to AI and human teaming in forecasting.  Instead of potentially missing the next post, join the Practitioner’s Network and I’ll send you a link to the upcoming post and provide some backstory that I only share with the network. 

Scroll to Top