Every season Metaculus sends out a survey to bot makers who participated in the FutureEval AI forecasting benchmark. This analysis compares answers to this survey with performance in the Spring 2026 FutureEval Bot Tournament to help surface what works and what doesn’t when forecasting the future. This analysis is complementary to our Spring Pro vs Bots Analysis and Spring Advice from Bot Makers to Bot Makers.
Main takeaways:
58 bot makers answered the Spring 2026 survey. This report covers the 48 whose bots competed in the scored FutureEval tournament (the other 10 were MiniBench-only participants). It shows, for each structured question, how answers were distributed and how they relate to bot performance.
We show 2 types of graphs. The first is the raw distribution of answers for each option for each question. The second compares the value of some feature of the survey with the performance of the bot in the tournament.
Performance is a bot’s average spot peer score in the Spring 2026 FutureEval tournament. Its spot peer score on a question compares its forecast at scoring time against the geometric mean of its peers. If it is positive, the prediction was (on average) better than others. If it is negative, it was worse than others.
Three groups appear in the distribution charts:
Charts show the share within each group, since the groups differ in size.
Correlations use a stricter set of the 42 bots that forecast at least 100 scored questions, so a bot with only a few questions cannot swing a result due to lucky forecasting. The distributions above still use all 48 participants, and only the performance correlations apply the question floor.
Each feature is correlated with a bot’s average spot peer score using Pearson’s r for yes/no traits and Spearman’s rank correlation for ordered or counted ones. Because many features are tested at once, every p-value also carries a Benjamini-Hochberg q-value (its p adjusted for the false-discovery rate across all tested features), and a result is called “significant” only when q < 0.05.
For the frontier model chart, a bot is called “frontier” if the model it used for its final prediction is a flagship model (not a mini, flash, or fast variant) and was released after 2025-11-01.
How to read the correlations: r runs from -1 to +1. Values near 0 mean no relationship, positive means the trait is associated with a higher average peer score, and negative with a lower one. A low p-value means the pattern is unlikely to be chance, but due to the number of features being tested, lean on the q-value for significance.
10 respondents are excluded from this report since they made no forecasts in the scored FutureEval tournament (MiniBench-only participants).
Below is a list of every measured survey feature versus performance (average spot peer score), ordered by correlation magnitude.
In the list, r is correlation, p is the uncorrected p-value, and q is the Benjamini-Hochberg p-value adjusted for testing all 33 features. “Significant” means q < 0.05. This analysis aims to be primarily hypothesis-generating. We can’t conclude much with n ≈ 42 bots and 33 features tested. We also expect many of these features to be confounded. However, FutureEval hosts the largest group of publicly competing AI forecasters, and so the results are worth indexing on.
The question sections below are ordered by their strongest correlation with performance (largest |r| first). Questions with no performance correlation come last. The evidence summary above is ordered by |r| alone.
Survey question: Which LLM model(s) did you use to make your final prediction/answer?
“Frontier final model” shows a weak link to higher peer score (Pearson r = +0.20, p = 0.210, q = 0.557, n = 40).
“Used GPT-5.4 for its final model” shows a moderate link to higher peer score (Pearson r = +0.42, p = 0.007, q = 0.238, n = 40).
“Used a flagship GPT-5.x for its final model” shows a weak link to higher peer score (Pearson r = +0.26, p = 0.100, q = 0.422, n = 40).
“Used a Claude Opus model for its final model” shows no clear relationship with peer score (Pearson r = -0.01, p = 0.929, q = 0.958, n = 40).
“Used Claude Opus 4.6 for its final model” shows a weak link to higher peer score (Pearson r = +0.16, p = 0.322, q = 0.614, n = 40).
“Final model release date (flagship models only)” shows a weak link to higher peer score (Pearson r = +0.14, p = 0.421, q = 0.722, n = 35).
Survey question: Did your bot use any of the below forecasting strategies?
“Aggregates multiple forecasts” shows a weak link to lower peer score (Pearson r = -0.19, p = 0.236, q = 0.557, n = 41). Yes means the bot took the median, mean, or aggregate of multiple forecasts.
“Uses explicit base rates” shows no clear relationship with peer score (Pearson r = +0.02, p = 0.907, q = 0.958, n = 41). Yes means the bot explicitly estimated base rates in a rigorous way.
“Checks similar questions/markets” shows a moderate link to higher peer score (Pearson r = +0.34, p = 0.031, q = 0.323, n = 41). Yes means the bot checked similar Metaculus questions or prediction markets.
“Researches subquestions” shows a weak link to higher peer score (Pearson r = +0.24, p = 0.129, q = 0.475, n = 41). Yes means the bot generated and researched subquestions.
“Uses scenario analysis” shows a weak link to lower peer score (Pearson r = -0.10, p = 0.518, q = 0.736, n = 41). Yes means the bot explicitly considered or categorized future scenarios.
“Uses LLM self-critique / red team” shows no clear relationship with peer score (Pearson r = -0.07, p = 0.675, q = 0.855, n = 41). Yes means the bot had the LLM self-critique or red-team its forecasts.
“Caps predictions” shows no clear relationship with peer score (Pearson r = -0.03, p = 0.853, q = 0.958, n = 41). Yes means the bot capped predictions at a max/min.
“Extremizes predictions” shows a moderate link to lower peer score (Pearson r = -0.30, p = 0.055, q = 0.365, n = 41). Yes means the bot mathematically extremized predictions via code.
Survey question: How did your bot research questions?
“Number of research sources” shows a weak link to higher peer score (Spearman r = +0.19, p = 0.230, q = 0.557, n = 42).
“Uses AskNews” shows no clear relationship with peer score (Pearson r = +0.06, p = 0.700, q = 0.855, n = 42). Yes means the bot’s research used AskNews (AskNews DeepNews or Other AskNews).
“Uses Exa” shows a weak link to lower peer score (Pearson r = -0.20, p = 0.195, q = 0.557, n = 42). Yes means the bot’s research used Exa.
“Uses Perplexity” shows a weak link to lower peer score (Pearson r = -0.11, p = 0.481, q = 0.736, n = 42). Yes means the bot’s research used Perplexity.
“Uses OpenAI web search” shows a weak link to higher peer score (Pearson r = +0.27, p = 0.084, q = 0.422, n = 42). Yes means the bot’s research used OpenAI web search.
“Uses web scraping” shows a moderate link to higher peer score (Pearson r = +0.33, p = 0.032, q = 0.323, n = 42). Yes means the bot’s research used static or interactive web scraping.
Survey question: What is your best estimate for how many total active hours (between all team members) have been put into developing your bot?
“Total development hours (midpoint)” shows a moderate link to higher peer score (Spearman r = +0.32, p = 0.039, q = 0.323, n = 41).
Survey question: When building, have you optimized more for research (external information retrieval) or reasoning (processing information given to the LLM)?
“Research vs reasoning (0=research..4=reasoning)” shows a weak link to lower peer score (Spearman r = -0.26, p = 0.102, q = 0.422, n = 40).
Survey question: Your best estimate of the number of LLM calls per question?
“LLM calls per question (midpoint)” shows a weak link to higher peer score (Spearman r = +0.20, p = 0.209, q = 0.557, n = 41).
Survey question: What is your best estimate of cost per Question? (USD)
“Cost per question (midpoint)” shows a weak link to higher peer score (Spearman r = +0.17, p = 0.300, q = 0.614, n = 39).
Survey question: What went into the development of your bot?
“Does manual review of outputs” shows no clear relationship with peer score (Pearson r = +0.00, p = 0.994, q = 0.994, n = 37). Yes means the maker did significant manual review of bot outputs, beyond sanity checks.
“Uses MiniBench for design” shows no clear relationship with peer score (Pearson r = -0.07, p = 0.682, q = 0.855, n = 37). Yes means the maker ran the bot in MiniBench and used the results to inform design.
“Tests via pastcasting” shows a weak link to higher peer score (Pearson r = +0.11, p = 0.502, q = 0.736, n = 37). Yes means the bot was tested via pastcasting (questions that already resolved).
“Tests vs community prediction” shows a weak link to higher peer score (Pearson r = +0.17, p = 0.318, q = 0.614, n = 37). Yes means the bot was tested against community predictions on prediction platforms.
Survey question: Which LLM model(s) did you use in supporting roles (i.e. not final predictions)?
“Frontier supporting-role model” shows a weak link to higher peer score (Pearson r = +0.16, p = 0.335, q = 0.614, n = 39).
Survey question: How many iterations of your primary bot did you make that ended up forecasting tournament questions live?
“Iterations that went live (midpoint)” shows a weak link to lower peer score (Spearman r = -0.12, p = 0.438, q = 0.722, n = 41).
Survey question: How did you aggregate?
“Ensemble uses multiple models” shows no clear relationship with peer score (Pearson r = -0.10, p = 0.535, q = 0.736, n = 41). Yes means the ensemble combined more than one model (answered ‘Same prompt, varied models’ or ‘Varied prompts AND varied models’).
Survey question: Did you give an LLM a verification environment (backtest harness, eval set, scoring loop) and let it self-experiment to produce part of your system?
“Gave LLM a verification env” shows no clear relationship with peer score (Pearson r = +0.04, p = 0.817, q = 0.958, n = 40). Yes means the maker answered ‘Yes’ (for any purpose) to giving an LLM a verification environment.
Survey question: How many people are on your team?
“Team size” shows no clear relationship with peer score (Spearman r = +0.02, p = 0.882, q = 0.958, n = 41). [Graph excluded: this variable is not bucketed for a group chart]
Survey question: How did you combine ensemble outputs into the final forecast?
Survey question: What best describes you?
Survey question: Did you change how your bot predicted questions in Spring compared to Fall?