The methods to improve geopolitical forecasts aren’t new, but finding the right combination of methods to achieve more than incremental progress remains a mystery.
When the Intelligence Advanced Research Projects Activity (IARPA) began funding geopolitical forecasting tournaments in 2011, the organization’s aim was to explore ways to limit human forecasting error for better decision-making.
In the prior four blog posts, I described the outcomes from the early competitions with a focus on identifying the personal qualities and training that can improve forecaster accuracy. I also shared updated research to provide a dose of skepticism in attempting to over-apply forecasting training to international affairs work without tailored adjustments. This last post covers the silent teammate that also contributed to the Good Judgment Project’s (GJP’s) success in the IARPA tournaments—the statistical and machine learning algorithm.
As an international affairs practitioner myself, what I describe below is for fellow practitioners and organizations considering when to engage forecasters. The first step to being better informed was to understand how GJP selected and trained the forecasters. If you haven’t read the first four blogs, you can here, here, here, and here. But latent skill and additional training aren’t the only reasons forecasting organizations can achieve higher accuracy in tournaments. The next thing to understand is how individual forecasters’ judgments are averaged or algorithmically adjusted.
It’s important to understand because the choices different organizations make when training their teams and algorithms impact outcomes. Think tanks in international affairs such as The RAND Corporation’s Integrated Forecasting and Estimates of Risk (INFER) and the Washington D.C.-based Center for Security and Emerging Technology’s (CSET’s) CSET Foretell are each different. Further, this background will also help international affairs practitioners interpret the implications of media organizations like CNN partnering with a different kind of prediction market called Kalshi to integrate live prediction data into its broadcasts.
The standard methods aren’t directly applicable
According to Stefan M. Herzog and Ralph Hertwig in a 2009 report called The Wisdom of Many in One Mind, “Although individuals’ estimates may be riddled with errors, averaging them boosts accuracy because both systematic and random errors tend to cancel out across individuals…Thus, the simple prescription for making good forecasts and accurate estimates is as follows: Gather a few predictions or estimates from sources that are likely to differ in their errors and average them.”
Forecasting platforms use events where the scoring organizations will know the outcome within a reasonable amount of time and require participants to use probability judgments. These methods capture the wisdom of the crowd where quantifiable outcomes and assessments are easier to average numerically.
Prior research outcomes showing the benefits of averaging belief and using algorithms are promising, but those methods are strained with international affairs information, which is incomplete and often inaccurate. International affairs affords hard-to-quantify data and sometimes elusive reference classes that make predictive model-building difficult. As Mellers and Tetlock described in their 2023 article titled Human and Algorithmic Predictions in Geopolitical Forecasting, “The best use of algorithms is as aggregators of crowd wisdom, as frameworks to partition human forecasting variance, and as inputs to hybrid forecasting models…But questions remain: ‘How big does the “crowd” need to be?’ ‘What is the best algorithm for combining individual judgments?’” Hence, they actively seek to develop new methods.
The 2011-2015 IARPA tournaments
That search for better methods produced some creative solutions during the first IARPA competitions. As part of the IARPA-funded GJP research, Ville A. Satopää, Shane T. Jensen, Barbara A. Mellers, Philip E. Tetlock, and Lyle H. Ungar in a 2014 article titled “Probability aggregation in time-series: Dynamic hierarchical modeling of sparse expert beliefs“ used the IARPA tournament data to develop a new approach.
Participants in forecasting tournaments and international affairs practitioners are wise to update their judgments as more information becomes available. Belief updating was not captured in prior crowd research. To advance the topic, Satopää and his team introduced a time-series model that incorporated self-reported expertise (latent forecasting ability) and captured a sharp and well-calibrated estimate of the crowd belief (averaged judgment).
As an important contribution, the paper describes the machine learning algorithm they developed for the specific tournament scenario and the broader International Affairs Knowledge Industry (IAKI) where many professionals believe in false information, hide their true beliefs, or may be biased for many other reasons. Satopää attempted to build a model that could detect potential bias, separate signal from noise, and use the collective opinion to estimate.
In simple terms, the team first began with simple algorithms—means and medians—that the team selected on the basis of their ability to reduce noise via error cancellation when forecasters worked alone. Then, as the tournament progressed, the team adjusted the means to give greater weight to more recent forecasts, building on the idea that forecasters who updated their predictions were more likely to have better information, thus improving the signal-to-noise ratio. After the first tournament year, the team assigned greater weight to the forecasters with better track records (called “superforecasters”) and ultimately added an extremizing transformation to weigh superforecasters more heavily when their individual judgments were different but pointing in the same general direction.
The 2017-2019 IARPA tournaments
By the time IARPA launched its next public tournaments, the question had shifted from how to process human judgment more effectively to how humans and machines could work together. The latest IARPA tournament is called Hybrid Forecasting Competition (HFC), and their website solicited volunteers as recently as 2025. According to the website, “HFC will test the limits of geopolitical forecasting by combining the ingenuity of human analysts with cutting-edge machine systems. HFC is a multi-year, U.S. Intelligence Community-funded research competition that seeks to develop and test hybrid forecasting methods that will radically improve the accuracy and timeliness of geopolitical and geoeconomic forecasts.” The latest research I could find covered the 2017-2019 tournament period.
Daniel Benjamin and team in a 2023 paper called Hybrid Forecasting of Geopolitical Events described their Synergistic Anticipation of Geopolitical Events (SAGE) system, which they developed under IARPA’s HFC program. Instead of simply processing human judgment in ways that improve accuracy, the latest research is on how to optimize human judgment in response to artificial intelligence (AI.)
Benjamin and team developed SAGE as a hybrid forecasting platform that allows human forecasters to combine model-based forecasts with their own judgment. The SAGE system provides forecasters automated statistical predictions and freedom to choose if and how much weight to assign to model predictions when submitting their personal forecasts. The team’s system was designed to test the “hypothesis that machine model forecasts embedded in a crowdsourced forecasting platform can improve the accuracy and efficiency of established crowdsourced forecasting methods.”
According to their paper, “Two common forecasting methods, crowdsourcing and machine learning, have complementary strengths and competing weaknesses. Statistical models perform well under the right circumstances. Factors like amount, availability, and structure of data determine how these methods perform. Machine-based forecasting methods typically perform well on problems for which there is sufficient historical data, but are ill-suited to forecast rare or idiosyncratic events for which such data may not exist, or when the underlying context has changed in ways not reflected by the historical data.”
“Human analysts, on the other hand, can often accurately forecast outcomes without exclusively depending on the availability of historical data, by leveraging their domain knowledge and prior experience.”
However, they noticed that introducing algorithm tools into crowd systems came with trade-offs. The wisdom of the crowd relies on sufficient expertise and diversity of knowledge. The introduction of statistical models could diminish diversity if human participants are too trusting in the models and do not feel empowered or motivated to add their private information into the system. The authors sought to understand whether training could improve accuracy when forecasters had to balance trust in a model with their own judgment.
New and improved forecasting training?
Their paper described a training called HABIT, the acronym for which they did not spell out. The authors described the training method as combining probabilistic reasoning with hybridization concepts using a narrative cartoon format. They designed their training to extend Mellers and Tetlock’s previous vignette-based methods and also focused on the core tenets of probabilistic estimation. The aim of the HABIT training was to teach forecasters about the machine models and how to integrate the model’s forecasts with one’s personal knowledge.
I couldn’t find more information on the HABIT training, but I expect it’s forthcoming. Benjamin and others in an earlier 2020 paper titled Quantifying Machine Influence Over Human Forecasters stated that, “We divide our cognitive hypotheses into two categories, strategic and biases. From the strategic analysis, we find that users engage with machine model predictions similarly to how individuals use expert advice. They use their own information the majority of the time and only incorporate the advice in certain situations—when the task is difficult…This provides an opportunity to improve the system by providing information to the human forecasters about how well the model is expected to perform for a given question, such as based on the data source and question format, as well as highlighting the questions that seem to be most uncertain. We could also develop training based on these results to aid forecasters in how to assess the performance of a model prospectively.”
Critiques continue
Forecasting tournaments provide a volume of data to train predictive models, but the critics point out that tournament settings have built-in incentives that won’t benefit forecasting algorithm design. As Nuno Sempere and Alex Lawsen in their 2023 paper called Alignment Problems With Current Forecasting Platforms pointed out, “forecasting competitions sometimes inadvertently provide forecasters with incentives not to reveal their best forecasts…In a forecasting tournament which rewards comparative accuracy, there is a disincentive to sharing information, because other forecasters can use it to improve relative standing.”
The authors further stated that, “publishing misleading information does not seem to be an urgent problem today. However, this might only be the case because forecasting competitions are currently relatively niche and small. If they grew further, they might encounter similar problems to PredictIt, where market participants have on occasion created fake polls which confused election bettors. Two notable cases from the PredictIt community are the cases of Delphi Analytica and CSP polling.”
These critiques matter because they highlight the tension between tournament design—a useful way to generate a lot of geopolitical forecasts—and the ultimate goal to improve decision-making in real-world events—which involves different rewards and incentives compared to winning in a tournament.
Shutdown
While it’s a longstanding practice to reduce error by averaging human judgment or employing other algorithms, researchers are still searching for the right combination of methods to achieve more than incremental progress for decision-makers. In this last blog on “superforecasters,” I looked at varying statistical and algorithmic methods that forecasting organizations use to get the best of both worlds.
De-biasing and probability training boosts open-minded and critical thinkers’ forecasting accuracy. When these forecasters adjust their behaviors for the tournament environment, they individually have more accurate forecasts. Organizations are still testing which type of algorithm rules work best to synthesize a forecaster crowd judgment. Moving forward, researchers are also exploring how to use AI systems to generate judgments to which human forecasters respond.
There are many different ways to aggregate and many different decision points at which to welcome an AI conclusion. There are also many different types of prediction markets varying in domain of information, method of aggregation and AI teaming, and incentives to human forecasters. Forecasting organizations like GJP and Metaculus incentivize participation, but not on the scale of other platforms like prediction markets. Understanding the different platforms and how tournaments work can help practitioners determine the benefits of gaining forecasting skills themselves, as well as help practitioners advise their organizations on engaging forecasters.


