Every few years, headlines resurface claiming that amateur forecasters “beat” the U.S. Intelligence Community (IC) at predicting world events. It’s a compelling David-and-Goliath story—except it’s also misleading in ways that matter deeply if you’re an international affairs researcher.
We’ve all heard colleagues cite these findings to argue everything from “I know more from reading my X feed” to “I ask artificial intelligence (AI) before I ask anyone.” The problem? Most people haven’t read Seth Goldstein, Rob Hartman, Ethan Comstock, and Thalia Shamash Baumgarten’s 2015 article called Assessing the Accuracy of Geopolitical Forecasts from the US Intelligence Community’s Prediction Market. The analysis compared the IC Prediction Market to the winning Good Judgment Project (GJP) platforms and tells a far more nuanced story about information environments, incentive structures, and what “beating” actually means when you’re measuring probabilistic judgment.
This matters because as researchers, we’re working in conditions surprisingly similar to both groups: deadline pressures, incomplete information, career incentives that don’t always reward accuracy. Understanding what this study actually found isn’t academic trivia—it’s a mirror showing us our own cognitive constraints and institutional challenges.
What “Beating” Actually Means
The authors compared forecasting accuracy on 139 unclassified geopolitical questions from the first nine months of the Intelligence Advanced Research Projects Activity’s (IARPA’s) 2011-2015 Aggregative Contingent Estimation (ACE) forecast tournament. The Intelligence Community Prediction Market (ICPM) consisted of around 4,300 cleared government employees and contractors making predictions on classified networks without additional pay. The GJP recruited forecasters from the general public and their personal networks to participate and receive small honorariums.
The results? The ICPM ran on a “garden variety prediction market” and GJP tested different tuned platforms. Of the three GJP models that competed against the ICPM, only one GJP was statistically superior. The other two were either inferior or indistinguishable from the ICPM. The take away is that the algorithm matters. Not that intelligence analysts are replaceable.
The GJP’s most accurate method called “All Surveys Logit” was developed through substantial research funding from IARPA and other organizations and used sophisticated weighting algorithms and extremization techniques on forecasters who had the best track records from the first tournament year.
So the word “beat” is doing a lot of work here. On identical prediction market platforms, IC analysts and civilian forecasters performed essentially the same. In the words of the authors, “Across this corpus of forecasting questions, the ICPM was directionally accurate on an average of about 82% of question-days, and this performance was comparable to that of a GJP prediction market hosted on the open internet.” The IC only “lost” when compared to a heavily optimized algorithm that weighted individual forecasters by past performance and psychometric profiles, then mathematically extremized their collective judgment.
Imagine a scenario where that algorithm is used with intelligence analyst’s forecasts and classified questions—That’s the future. But these competitions are needed to test and train algorithms and human interventions for AI teaming.
The Measures That Matter
The study used two primary accuracy metrics, and understanding both reveals why headline writers love this story while statisticians stay cautious.
Brier scores measure the accuracy of probabilistic judgments on a scale from 0 to 2, where lower is better. Think of it as penalizing you massively for increments of being wrong: assigning 80% probability to something that actually happens scores four times worse than assigning 90%. The ICPM scored 0.23, the GJP Prediction Market 0.21 (statistically indistinguishable), and GJP’s best method 0.15.
Directional accuracy asks a simpler question: Did you assign the highest probability to what actually happened? This blunter instrument showed the ICPM directionally correct 81.58% of question-days, versus GJP Prediction Market’s 83.45% (again, statistically indistinguishable) and GJP’s best method at 88.20%.
When headlines say “forecasters beat the IC,” they’re comparing commercial-off-the-shelf prediction market software to years of algorithmic research optimization. That’s not amateurs beating professionals—it’s well-funded research beating unoptimized government technology in the early 2010s.
How IC Analysts Actually Participated
This is where the story gets interesting for anyone who’s worked in large organizations. The 4,300 ICPM participants were cleared employees who volunteered to make predictions during work hours and receive zero material compensation beyond their salaries. They self-selected which questions to answer. They operated on a “Logarithmic Market Scoring Rule” system with 5,000 symbolic points, where successful traders gained more influence over aggregate probabilities.
Compare this to GJP’s structure: recruited participants, direct compensation for participation, random assignment to experimental conditions testing various interventions designed to maximize accuracy. One group got training materials. Another got teaming protocols. The “best method” group received algorithmic enhancement of their individual judgments.
The incentive structures couldn’t be more different. ICPM analysts were fitting prediction market trading around their actual jobs. GJP participants were being paid specifically to forecast and were guinea pigs for accuracy-enhancement research.
And here’s something that should make every researcher pause: ICPM participants had access to classified information but were being tested only on unclassified questions. The study’s authors explicitly acknowledge they can’t draw conclusions about IC forecasting capabilities on classified questions—which is, you know, the actual job.
As I’ve described elsewhere, classified or unclassified “questions” per se aren’t what makes an intelligence analyst a deer in headlights in an unclassified research environment. It’s that they don’t know how to explore and find high-value open source information because they’re used to getting classified data pushed to them for review. But I digress.
The Caveats
First, as mentioned above: The competition was on unclassified questions and used open source information. The entire premise that IC analysts should outperform civilians assumes classified information provides an advantage. For questions like “Will North Korea conduct a missile test by June?” classified intelligence might be crucial. For “Will Italy’s GDP grow by 3%?” maybe not. My personal point is that for either question, intelligence analysts will have to spend considerable time figuring out where they should look for open source answers, especially because they’ll fear all they search terms are classified.
Second, the authors only used nine months and 139 questions from one year of a four-year program. The authors themselves note caution in drawing sweeping inferences. Remember that the headlines were probably useful for some of the government project funders who needed to give the government a wake-up-call to enhance spending error reduction in IC forecasts with technology enhancements and greater use of open source information.
Third, the authors note that access to classified information might actually hurt performance on unclassified questions. (That’s what I’m saying.) The authors mention the “secrecy heuristic” that suggests analysts may irrationally overweight classified sources, spending finite attention on classified documents when open-source information would be more useful. If you’ve got 30 minutes to research a question and you waste 20 reading classified reports that don’t help, you’ve just handicapped yourself.
Fourth, compensation for participation is an important incentive. GJP participants were paid to forecast and update their predictions. ICPM participants were volunteering during work time.
Finally, the “best method” that beat the ICPM was retrospectively selected from 20 different GJP methods submitted daily. While several methods performed similarly, this still represents comparing an optimized research product to baseline technology.
Why This Matters for International Affairs Researchers
Like IC analysts, we face deadline pressures that reward confident narratives over probabilistic uncertainty. Try getting published by writing “There’s a 60% chance this regime will fall, but I could be wrong.” Reviewers want narratives and arguments, not odds.
Like GJP forecasters, we rarely get systematic feedback on our predictions. We write that a conflict is “likely to escalate” and then… move on to the next project. Nobody’s calculating our Brier scores. We’re not extremizing our collective judgments through optimized algorithms.
But unlike either group, we typically work alone or in small teams, without the crowd wisdom that helps both prediction markets and opinion pools. We don’t have 4,300 colleagues trading on our questions. We don’t have research funding to develop accuracy-enhancement tools. We have foreign language research methods, a conference deadline, and imposter syndrome.
The IC was “beat” by methodological innovations specifically designed to overcome cognitive biases and aggregate diverse viewpoints. What methodological innovations are you using to overcome your confirmation bias when you’re alone with your laptop at 11 PM before a deadline? Don’t have any designed for the solo international affairs researcher? Then join the Practitioner’s Network or take the online International Affairs Professional Development Course.
The Uncomfortable Mirror
Here’s what I learned from reading this study instead of the headlines: The IC Prediction Market, stripped of algorithmic enhancements and tested on questions where classified information provides minimal advantage, performs about as well as you’d expect humans with relevant knowledge to perform. Which is to say: pretty good, but improvable with deliberate methodological intervention.
The GJP’s “best method” didn’t win because amateur forecasters are smarter or because classified information is worthless. It won because researchers spent years testing ways to make collective judgment more accurate with training, teaming, weighting, extremization. They treated forecasting accuracy as an engineering problem worth solving.
Most international affairs research treats forecasting accuracy as someone else’s problem. We gesture toward “rigorous methodology” while operating in information and incentive environments that actively discourage the kind of probabilistic thinking, systematic updating, and accuracy tracking that improves judgment.
The headlines about the IC getting beat are popular because they’re flattering; they suggest the online crowd is just as informed as experts. The actual study suggests something less comfortable: you have to be proactive in self-improvement. If you’re stagnant and wait for your employer to require you attend a training you’re only staying with the average. That’s why I made the online, self-paced, and affordable International Affairs Professional Development Course. It won’t ever be on your approved contracted training list. It’s available for colleagues who are proactive in self-improvement, which in necessary for achieving mission goals.
Shutdown
Of the three GJP models that competed against the ICPM, only one was statistically superior. The other two were either inferior or indistinguishable from the ICPM.
This should be reassuring and concerning in equal measure. Reassuring because it suggests crowd wisdom works across different populations and security contexts. Concerning because it reveals how much accuracy we leave on the table when we don’t stay proactive with self-improvement and overcome our personal and institutional biases.
Next time you see that “opportunity to participate” in your inbox and think, “I’ll do this when it’s mandatory or funded,” remember there are other colleagues who are taking the opportunity to get a competitive edge. Does this mean you should volunteer or agree to everything—certainly not. As international affairs professionals, you know what you want to have more time to learn, but you’re just looking for someone else to pay for it. Send me a note and tell me how that’s working out for you.


