Polls are models of the electorate. So we treat them like models.
A modern poll is not "we called 800 random people." It is a sampling frame, a set of weighting targets, a likely-voter screen, and a good deal of judgment: a pollster's model of who will vote and how. That changes how polls should be averaged, how pollsters should be graded, and above all how much uncertainty an average deserves. Everything below was chosen by testing alternatives against 2,313 past races, using only what was knowable at the time.
What this model cannot do. It cannot say which way the polls will miss this year, only how large such misses have been. Assume a repeat of 2016–24 and the Democrats' chance of Senate control is 28%; assume the mirror image and it is 81%. It has never been run through a live election night, its House error estimates rest on 4 cycles of district polling, and its current-cycle poll feed is thinner than the big aggregators'. If you are here to decide how far to trust it, start with Uncertainty and Limits.
Where a poll’s error actually comes from
18,109 general-election polls from the final weeks of races since 1998, each scored against the certified result.
A poll is a pollster's estimate of who will vote and what those people look like, fitted to a few hundred completed interviews. The sampling is the easy part. Decompose the error of nonpartisan polls in the last three weeks of a race, and random sampling noise turns out to be real, roughly the size theory predicts, and the smallest of four contributions.
Take 2022. The average nonpartisan poll missed the final margin by 1.3 points in the Democratic direction across the whole industry. On top of that sat 4.6 points of standard deviation in error shared by every poll of a given race, 3.2 points of pollster house effect that followed a firm from contest to contest, and 2.8 points of poll-specific residual. Sampling theory predicted 3.5 points for that last term, so the idiosyncratic part is if anything quieter than textbook. 2020 was worse in a specific way: the industry-wide miss reached 5.6 points, the largest in the twelve cycles measured, while race-level shared error fell to its lowest value in the record, 3.4. Pollsters were unusually unanimous and unusually wrong at the same time.
Only the poll-specific term shrinks when you average more polls or field bigger ones. Everything else is common. That is why sample size does so little work inside a race: across 12,958 polls, the rank correlation between a poll's sample size and its error, holding the race fixed, is 0.01. The raw cross-race correlation is only −0.09, and most of even that reflects which races attract which polls. Polls of the same contest disagree by about 3.1 points no matter how large they get, against a realized error near 6.3 points for a poll of typical size.
The dispersion evidence points the other way from what you would expect. In 272 races with at least eight polls from five or more pollsters, the observed spread runs at 0.94 of pure sampling variance, against 0.95 expected from textbook random samples. That looks like textbook behavior until you account for the house effects we can measure, which should push expected dispersion to 1.22, or to 1.51 with a modest weighting design effect of 1.3. Real polls cluster about 38% tighter than the realistic benchmark. Shared weighting targets, herding toward the visible average, or both would produce that pattern. The comparison cannot tell them apart, and the 1.22 and 1.51 benchmarks depend on modeling assumptions the reader should treat as assumptions.
One caveat on the sample-size table: small polls are disproportionately House district polls, which are harder races, so the realized-versus-theory gap by bin is illustrative rather than probative. The within-race test is the clean one.
The backtest across roughly 2,000 past races says the rest plainly. Weighting by the square root of sample size changes mean absolute error by 0.008 points. Removing sponsor lean is worth 0.084, ten times more. A careful average beats a simple one by 0.27 points, 5.87 to 5.61, and that is close to the ceiling. The remaining work is estimating how wrong the shared parts can be.
| Cycle | Every poll that year | Every poll of a race | The pollster, across races | The poll itself | Sampling theory for that last column |
|---|---|---|---|---|---|
| 2000 | R+1.1 | 4.3 | 1.3 | 4.0 | 3.5 |
| 2002 | D+2.0 | 5.1 | 1.6 | 4.0 | 3.9 |
| 2004 | D+1.1 | 3.8 | 1.7 | 3.5 | 3.7 |
| 2006 | Even | 4.1 | 1.5 | 5.0 | 3.9 |
| 2008 | D+1.1 | 5.0 | 2.2 | 3.4 | 3.8 |
| 2010 | D+1.0 | 5.6 | 2.6 | 4.2 | 3.6 |
| 2012 | R+3.0 | 4.5 | 2.0 | 2.8 | 3.5 |
| 2014 | D+2.7 | 3.9 | 3.1 | 3.9 | 3.5 |
| 2016 | D+3.5 | 4.9 | 2.8 | 3.8 | 3.6 |
| 2018 | R+0.5 | 4.4 | 2.6 | 3.6 | 3.9 |
| 2020 | D+5.6 | 3.4 | 3.1 | 2.8 | 3.5 |
| 2022 | D+1.3 | 4.6 | 3.2 | 2.8 | 3.5 |
A crossed random-effects model fit within each cycle. The first three columns are shared error: the whole industry's miss that year, the miss common to everyone polling a race, and the lean a pollster carries from race to race. Only the last column behaves like sampling noise, and it is about the size sampling theory says. It is also the only part that averaging more polls, or bigger polls, can shrink.
Where the polls come from
History: FiveThirtyEight's pollster-ratings file (every poll in the final two months of general elections, 1998–2023) and its full poll database for 2018 through March 2025, recovered from the Internet Archive after the site closed. This cycle: The New York Times polls database, which keeps FiveThirtyEight's format and pollster IDs, supplies most polls; poll tables from Wikipedia's race articles and the VoteHub API fill in what it lacks. The three are merged and de-duplicated, and the Times copy wins when the same poll appears twice. Results: the MIT Election Data and Science Lab, FiveThirtyEight's results archive, and the FEC. Money: FEC itemized individual contributions.
How the Times data gets here: its robots.txt asks AI agents to stay out, and this site was assembled by one, so the file is downloaded by a person and handed to the pipeline. Questions that pit a “generic Democrat” against a “generic Republican” inside a state or district are not counted as polls of the race, and when a poll asks the race several ways the version naming the most candidates is used. In ranked-choice general elections (Alaska, and Maine's federal races) only a poll's final round counts: a first-choice tally with a real third or fourth candidate in it says nothing about where their votes go, so it is dropped unless everyone beyond the top two adds up to 5 points or less. RealClearPolling, 270toWin, Decision Desk HQ, and FiftyPlusOne are not open data and are not used.
One poll, one number. When a poll publishes several versions of a race (likely and registered voters, with and without minor candidates), those are one model run several ways. We keep the most election-like population and average the versions. Hypothetical matchups are dropped, matching candidates by full name: this year has polls of two different Sununus and two different Grahams.
Track records predict less than you would hope
A pollster is scored against other pollsters polling the same race, never against a sampling-theory benchmark, and gets no credit for sample size.
We asked how well a rating built through one year predicts the next cycle, across 1,503 pollster-cycles from 2008 to 2024. For accuracy, the out-of-sample correlation is 0.04. Even shrinking each record toward the average for its method with the weight of 100 races leaves the predictions too spread out (a slope of 0.56 where 1.0 would be calibrated). Most of any track record is luck. So the ratings are shrunk hard, carry no letter grades, are not published as a ranking, and move a poll's weight in the average by only about 15% either way.
Lean is a different story in one specific way. Polls sponsored by a party or campaign sit 3.6 points (Democratic sponsors) and 3.7 points (Republican sponsors) from nonpartisan polls of the same race, cycle after cycle; leave that in and pollster lean looks very persistent (correlation 0.32). Take it out and a pollster's lean carries over from one cycle to the next only weakly (correlation 0.09). Within a cycle it is large, about 3.2 points, so the average measures each pollster against the field as the cycle goes and removes 60% of what it finds. After the sponsor shift, party-sponsored polls turn out to be as informative as anyone else's, so they keep full weight.
Correct the lean; ignore the sample size
Type-balanced mean absolute error 7, 21, and 48 days out, 2000–2024. “Leave-one-cycle-out” means any tuning was done without the cycle being scored.
| Averaging scheme | Error |
|---|---|
| Simple mean of the last 30 days | 5.87 |
| Latest poll per pollster, 30 days (RCP-style) | 5.88 |
| √n × rating × recency (538-style) | 5.65 |
| This site’s settings | 5.61 |
| This site, re-tuned for each held-out cycle | 5.63 |
Any sensible weighted, lean-corrected average beats a simple one by about a quarter point, and ours is no better than a 538-style average that weights by the square root of the sample size. Tuning beyond that buys nothing out of sample, so the settings are round numbers from the flat part of each curve. The gains come from removing sponsor lean, correcting house effects, and a long memory when the election is far off. Weighting by sample size neither helps nor hurts, which is what you would expect if sampling noise is the smallest part of the error; we leave it out.
| Setting | Value | Change in error (points of MAE) | Cycles better / worse |
|---|---|---|---|
| Sample-size exponent (0.5 = classic √n) | 0 ✓ | chosen | — |
| Sample-size exponent (0.5 = classic √n) | 0.15 | +0.001 | 8 / 5 |
| Sample-size exponent (0.5 = classic √n) | 0.3 | +0.004 | 8 / 5 |
| Sample-size exponent (0.5 = classic √n) | 0.5 | +0.008 | 8 / 5 |
| Sample-size exponent (0.5 = classic √n) | 1 | +0.034 | 5 / 8 |
| Strength of pollster-quality weight | 0 | +0.001 | 7 / 6 |
| Strength of pollster-quality weight | 0.25 | −0.000 | 7 / 6 |
| Strength of pollster-quality weight | 0.5 ✓ | chosen | — |
| Strength of pollster-quality weight | 1 | +0.002 | 6 / 7 |
| Strength of pollster-quality weight | 1.5 | +0.011 | 5 / 8 |
| Strength of pollster-quality weight | 2.5 | +0.034 | 4 / 9 |
| Weight kept by a pollster's earlier polls (0 = latest only) | 0 | +0.019 | 4 / 9 |
| Weight kept by a pollster's earlier polls (0 = latest only) | 0.15 | +0.011 | 4 / 9 |
| Weight kept by a pollster's earlier polls (0 = latest only) | 0.35 | +0.004 | 4 / 9 |
| Weight kept by a pollster's earlier polls (0 = latest only) | 0.6 | −0.001 | 8 / 5 |
| Weight kept by a pollster's earlier polls (0 = latest only) | 1 | +0.007 | 7 / 6 |
| Recency half-life on Election Day (days) | 5 | +0.007 | 5 / 8 |
| Recency half-life on Election Day (days) | 8 | +0.003 | 6 / 7 |
| Recency half-life on Election Day (days) | 11 | +0.001 | 6 / 7 |
| Recency half-life on Election Day (days) | 14 ✓ | chosen | — |
| Recency half-life on Election Day (days) | 20 | +0.001 | 6 / 7 |
| Recency half-life on Election Day (days) | 30 | −0.000 | 6 / 7 |
| Extra half-life per day before the election | 0 | +0.038 | 3 / 10 |
| Extra half-life per day before the election | 0.1 | +0.030 | 3 / 10 |
| Extra half-life per day before the election | 0.25 | +0.020 | 3 / 10 |
| Extra half-life per day before the election | 0.5 | +0.010 | 3 / 10 |
| Extra half-life per day before the election | 1 ✓ | chosen | — |
| Hard cutoff on poll age (days) | 30 | +0.082 | 3 / 10 |
| Hard cutoff on poll age (days) | 45 | +0.045 | 2 / 11 |
| Hard cutoff on poll age (days) | 60 | +0.043 | 0 / 4 |
| Hard cutoff on poll age (days) | 90 | +0.021 | 0 / 4 |
| Hard cutoff on poll age (days) | 120 | +0.006 | 0 / 4 |
| Hard cutoff on poll age (days) | 200 | +0.003 | 0 / 4 |
| Share of the sponsor-party shift removed | 0 | +0.084 | 1 / 12 |
| Share of the sponsor-party shift removed | 0.5 | +0.019 | 5 / 8 |
| Share of the sponsor-party shift removed | 0.75 | +0.003 | 6 / 7 |
| Share of the sponsor-party shift removed | 1 ✓ | chosen | — |
| Share of the sponsor-party shift removed | 1.25 | +0.011 | 5 / 8 |
| Weight on sponsor-party polls after the shift | 0.25 | +0.052 | 3 / 10 |
| Weight on sponsor-party polls after the shift | 0.5 | +0.020 | 4 / 9 |
| Weight on sponsor-party polls after the shift | 0.75 | +0.006 | 4 / 9 |
| Weight on sponsor-party polls after the shift | 1 ✓ | chosen | — |
| Share of the house effect removed | 0 | +0.010 | 5 / 8 |
| Share of the house effect removed | 0.25 | +0.004 | 6 / 7 |
| Share of the house effect removed | 0.5 | +0.001 | 6 / 7 |
| Share of the house effect removed | 0.75 | −0.000 | 7 / 6 |
| Share of the house effect removed | 1 | +0.001 | 6 / 7 |
| Share of the house effect removed | 1.25 | +0.004 | 5 / 8 |
| Strength of the prior-cycle house effect | 0 | −0.000 | 6 / 7 |
| Strength of the prior-cycle house effect | 2 | −0.000 | 6 / 7 |
| Strength of the prior-cycle house effect | 4 ✓ | chosen | — |
| Strength of the prior-cycle house effect | 8 | −0.000 | 6 / 7 |
| Strength of the prior-cycle house effect | 16 | +0.001 | 6 / 7 |
| Shrinkage of house effects toward zero | 0 | +0.001 | 6 / 7 |
| Shrinkage of house effects toward zero | 4 | −0.000 | 5 / 8 |
| Shrinkage of house effects toward zero | 8 ✓ | chosen | — |
| Shrinkage of house effects toward zero | 16 | +0.001 | 7 / 6 |
| Shrinkage of house effects toward zero | 32 | +0.003 | 7 / 6 |
| Shrinkage of house effects toward zero | 64 | +0.005 | 5 / 8 |
| Equalizing methodological families | 0 ✓ | chosen | — |
| Equalizing methodological families | 0.25 | −0.001 | 5 / 8 |
| Equalizing methodological families | 0.5 | +0.002 | 5 / 8 |
| Equalizing methodological families | 1 | +0.030 | 4 / 9 |
| Weight on registered-voter polls (likely voters = 1) | 0.5 | −0.000 | 3 / 1 |
| Weight on registered-voter polls (likely voters = 1) | 0.7 | +0.001 | 2 / 2 |
| Weight on registered-voter polls (likely voters = 1) | 0.85 | +0.002 | 2 / 2 |
| Weight on registered-voter polls (likely voters = 1) | 1 | +0.003 | 2 / 2 |
The generic ballot runs a little blue
Generic-ballot average seven weeks out, minus the national House vote that followed. Blue bars: the polls overstated Democrats.
The generic ballot has overstated Democrats in most cycles since 1998. The model subtracts the average past miss (using only earlier cycles, shrunk toward zero), which today is 1.8 points, and carries the remaining error (about 3.0 points) into every race through the shared national term. Race polls get no such correction: their misses have run toward Democrats in most recent even years but toward Republicans in the 2025 governor's races, and the long record averages close to zero.
What a race looks like before anyone polls it
Partisan lean, the national environment, incumbency, how the incumbent ran last time, and, for federal races, the fundraising gap from FEC filings knowable at the forecast date.
| Inputs | RMSE | Close seats |
|---|---|---|
| lean + environment + incumbency | 8.33 | 6.36 |
| + incumbent's prior over-performance | 7.67 | 5.53 |
| + fundraising gap through June 30 | 6.97 | 5.11 |
| + donor-count gap through June 30 | 7.20 | 5.11 |
| + fundraising gap through Sept 30 | 6.98 | 5.18 |
A tenfold fundraising edge is worth about 4.0 points. Donor counts add nothing beyond dollars.
| Inputs | RMSE | Close seats |
|---|---|---|
| lean + environment + incumbency | 6.01 | 6.13 |
| + incumbent's prior over-performance | 5.73 | 5.62 |
| + fundraising gap through June 30 | 5.54 | 5.49 |
| + donor-count gap through June 30 | 5.56 | 5.47 |
| + fundraising gap through Sept 30 | 5.48 | 5.44 |
A tenfold fundraising edge is worth about 2.4 points. Donor counts add nothing beyond dollars.
Money only shows up when it is fit jointly with lean and incumbency; regressing it on what the other terms leave behind finds nothing, because money follows lean and incumbency. Governors get no money term (state filings are not in the FEC data), and their fundamentals are weak: voters split tickets for governor, so polls carry most of the weight there.
Polls and fundamentals, weighted by how many pollsters have weighed in
Weight on polls = n / (n + k), where n is the effective number of pollsters, not polls. For the Senate k is 4.25, for the House 0.75, and for governors 0: fundamentals get exactly as much say as the record insists on (the most poll-heavy setting within 2% of the best fit), and no more.
| Seven weeks out | Polls alone | Fundamentals alone | Blend | Typical weight on polls |
|---|---|---|---|---|
| governor (163 polled races) | 5.73 | 10.12 | 5.77 | 99% |
| house (360 polled races) | 5.86 | 6.40 | 5.03 | 64% |
| senate (217 polled races) | 5.92 | 6.67 | 5.05 | 78% |
The fundamentals model behind every row was fit only on earlier cycles. In the last three Senate cycles that model, with fundraising, was more accurate than the polling averages even two days out, which is why a well-polled Senate race still gives fundamentals about a third of the weight. Fit on all cycles pooled, the Senate constant would be 0.5, nearly all polls; that is the single most consequential judgment call on this page.
Mean absolute error in points. Counting every race, polled or not: governor 8.1 (winner called 88%), house 5.0 (winner called 96%), senate 6.2 (winner called 92%).
| k | Weight on polls, six pollsters | Democrats win control | Texas | Ohio | Michigan |
|---|---|---|---|---|---|
| 0.5 | 92% | 65% | D+2.5 | D+3.2 | D+2.2 |
| 2 | 75% | 60% | D+1.9 | D+2.4 | D+2.4 |
| 4.25 ✓ | 59% | 56% | D+1.2 | D+1.5 | D+2.6 |
| 8 | 43% | 52% | D+0.4 | D+0.5 | D+2.9 |
Trusting the polls more helps Democrats this year, because in the closest Senate races the polls are friendlier to them than the fundamentals are. The checked row is the one the site uses.
| One week out, 2018–24 | Races | Model miss | Ratings alone | Both | Winners called, model / ratings |
|---|---|---|---|---|---|
| governor | 93 | 5.8 | 7.0 | 5.3 | 94% / 94% |
| governor competitive only | 41 | 4.7 | 5.0 | 5.1 | 85% / 85% |
| house | 1584 | 4.8 | 10.5 | 4.7 | 96% / 96% |
| house competitive only | 342 | 4.3 | 4.2 | 4.4 | 81% / 84% |
| senate | 129 | 4.5 | 6.2 | 4.8 | 94% / 94% |
| senate competitive only | 54 | 4.1 | 4.1 | 4.5 | 85% / 85% |
Cook, Sabato's Crystal Ball, and Inside Elections ratings, averaged and mapped to a margin by a curve fit on the other cycles. Misses are mean absolute errors in points; competitive means a projection inside 12 points. Across all races the ratings alone miss by far more than the model, because a rating cannot tell a 20-point race from a 40-point one; in competitive races the two are about even, and adding the ratings to the model makes competitive races worse, not better. Their one real edge is calling a few more winners in competitive House districts.
The error that averaging cannot remove
Mean miss of the polling averages across all races, one week out, by cycle. When the polls are off, they are off together.
Every simulation draws one national miss that moves all races in all three offices together, plus an independent miss for each race that is larger when few pollsters have weighed in. Both come from these historical errors, not from stated margins of error. Tails are Student-t with the weight the record supports (15 degrees of freedom for the Senate, 19 for the House, and 9 for governors), scaled so that nine in ten past results fall inside the nine-in-ten interval.
Every race and horizon from 2006 to 2024, grouped by the favorite's stated chance. Points above the line mean the model was less confident than it could have been, which is the side we would rather err on.
| Mean miss of this method's projections, seven weeks out | 2006–2014 | 2016–2024 |
|---|---|---|
| senate | Even | D+2.4 |
| governor | R+0.9 | D+4.1 |
| house | — | D+1.8 |
D+ means the projection was too Democratic by that many points. We tested correcting for it the honest way, using only earlier cycles, and it did not help across the full record, because the direction flipped around 2014 and a correction would have lagged the flip. In 2025 the polls missed toward Republicans. So the model assumes no direction and carries the size of these misses as shared uncertainty (3.0 points for the Senate, with a 90% interval of 2.0 to 3.6 from only 10 cycles).
| Scenario | Senate | House | Dem. governors |
|---|---|---|---|
| No direction assumed (still carries the usual uncertainty) | 56% | 80% | 26.3 |
| Polls overstate Democrats as they did, on average, in 2016–24 | 28% | 59% | 22.8 |
| Polls understate Democrats by the same amount | 81% | 92% | 29.7 |
Chance Democrats win control (Senate, House) and expected Democratic governors.
What this does not do
- It does not know which way the polls will miss. The shared error term says how much, not which way.
- No regional correlation beyond the single national factor, and no model of how races drift between now and November beyond what past seven-week errors already contain.
- Expert race ratings are shown for comparison and never used as inputs.
- Where no Democrat is running and an independent is the alternative (Nebraska, South Dakota, and Idaho Senate), the margin is independent minus Republican, the projection is the polling average alone with a wider interval (no such race is in the backtest), and a win counts for neither party. Senate control goes to the party with more seats, with the vice president breaking a tie for Republicans, so 50 Democrats, 49 Republicans, and one such independent is counted as Democratic control; how that independent would actually vote to organize the chamber is not modeled.
- Pre-2018 cycles only have polls from the final two months, so the seven-week backtests before 2018 see a thinner slice of polling than the live model does.
- Louisiana's House seats hold all-party primaries on November 3; they are projected as party margins from fundamentals alone.
- Safe-seat probabilities are deliberately a little conservative; the difference between 98% and 99.9% is not something 14 cycles of data can resolve.