20 Comments
User's avatar
Jennifer Williams's avatar

Great post. I am ever-annoyed with the NYT and their coverage. Their response does not surprise me. Any way we could send a copy of your book to everyone in political news and opinion departments and get them to read it? Just a thought . . .

Jennifer Williams's avatar

“The primary (hah) problem with primary polls . . .”

Love it. =)

John Silas "Si" Hopkins, III's avatar

Primary polls might be off by more when one of the two top candidates is exciting and the other is steady. Those drawn to the exciting candidate may tend to make up their minds early. This could result in the undecideds skewing in favor of the steady candidate when they vote. This might be something you could check by analyzing past polls.

Judy Hennessey's avatar

I think a relevant factor this time around is that voters are much more used to general-election polls, with their more reliable factors. This is an unusual year in that many people are for the first time paying attention to primary results across the nation. To their eyes, what's normal looks peculiar because it doesn't match their past experience.

Chad Horner's avatar

not sure “weird” is fair here? in addition to polls, she was at 95% in prediction markets. she was widely expected to win!

“also journalists to react to world events in weird ways — grading candidates by their performance relative to the polls, not their performance in absolute terms”

Kathy Trullender's avatar

Appreciate you calling out the ridiculous NYTimes headline.

Drey Samuelson's avatar

with all due respect, you missed the #1 reason why primaries are extremely hard to accurately poll, and it's this: if you're a Democrat or Republican about to vote in a primary, it's not hard to be convinced to vote for another member of your political party right up to when you cast your ballot--you're not switching sides, you're simply switching candidates on the same side. However, the vast majority of people--once they hit the general election--are not going to switch which TEAM they're going to support (absent some egregious scandal), thus making polling much more accurate, there is MUCH less switching.

Cayce Jones's avatar

The NYT was just silly clickbait, not responsible reporting, because it misrepresents what people can expect from polling. The only "stunning" poll failure I can remember was Ann Selzer's last poll, but even in that case, it was an outlier from most polls.

Bruce S's avatar

When the undecided in a poll make up 20 percent of the sample, no one should be surprised at a surprise especially if the leader in the poll is only leading by 10. To 13 percent. Perhaps pollesters need to look at who makes up the undecided. Are they conservative, liberal, or interest in picking a winning candidate.

gold's avatar

Well as long as things are politically stable and demographics don’t march on and the economic backdrop doesn’t change and nobody’s home situation changes and … and … ummm …. Well this should be easy.

Having a lead in a poll a couple of weeks out is not like being up 29 in the second half (too soon?) …

And if understanding statistics were easy they wouldn’t come after lies and damned lies.

An awful lot of people who write _about_ this stuff, from the outside, really need to confront what they do not understand. There are no races. There are no “fights for the lead”. There are just games. “Elections”, they are called.

Marliss Desens's avatar

I've been referring people experiencing "polls meets reality" shock at other Substacks to Strength in Numbers. Given how close the race in Michigan was, I thought there was a good chance of a close race in Wisconsin. I think that the polls did help in one respect, in that I think they encouraged the hard run of David Crowley in the past weeks. I hope he and his supporters, joined by those who supported his opponent, will continue to campaign just as hard all the way to November. I would love to see a Democratic trifecta in Wisconsin again. (For a good examination of how the Republicans took over, see the book, The Fall of Wisconsin.)

Brien's avatar

One thing about polling I've been struggling with how to articulate for a long time is how we (laypeople) view probability as applied to election polls.

Most of us would not be all that shocked to see a 90% free throw shooter miss a shot, or to see a 20 pt favorite lose a football game. Those are surprising results, but they fall within our expecations that sometimes rare events occur.

When polls indicate a 70% probability of one candidate winning, we have a different knee-jerk response when the election goes the other way. Why? I'm having a hard time thinking of other probabilistic events where we project certainty onto estimates.

John Petersen's avatar

A way to think about the difference is the number of repeated experiences we have. Someone who watches football or basketball or baseball watches games week after week. Over the course of a season experience builds up and expectations are more finely calibrated. In addition when a low probability event occurs (Patriots come back from 28-3 or this summer Red Sox go 21-4 in July), we tend to form strong memories. While Hong's loss or Flanagan's win is notable at the moment, when the next set of primaries comes around, who will remember (besides Elliott)? On the other hand, in the short run the WI governor polls and to a lesser extent the MN senate polls were in the media all the time in the last couple weeks so we became "certain" that the polling numbers were significant because they were repeated over and over again. As they say repeating a lie doesn't make it true and reporting a polling result multiple times doesn't make it more accurate.

Brien's avatar

I think you're right that's part of it, but I don't think it fully explains what I'm attempting to describe. We have plenty of experience with events that may or may not occur and are impacted by random chance. In those cases, I'm positing that our intuition accepts that there is uncertainty in the system (the small puff of wind, the bounce of the ball off the rim).

I think that our intuition around election polling doesn't work quite the same way. Some of that comes from the way we extrapolate our own behavior to that of the entire electorate. If a pollster asked me my voting plans, I would tell her, then I would go vote that same way. Sure some small number of people may change their minds in between answering and voting, but overall (our intuition says) things should be fairly predictable.

John Petersen's avatar

Here's a method to estimate the noise in the system. Let's look at population matched counties in the range of 5-20K democratic primary votes in Wisconsin and see how noisy the data are comparing counties. Note that the bottom of this range, 5K votes, is a multiple of any statewide poll number and, of course, these are actual voters not people who intend to vote.

I found three counties, Winnebago, La Crosse and Eau Claire with about 20K votes that favored Hong, 42.5%, 41.8% and 46.0%, respectively with Crowley at 35.7%, 32.4% and 32.9%. Std deviation for Hong, 1.8% and Crowley 1.5% and that's with 20K actual votes.

I found three counties with about 8K democratic primary votes, Sauk, Columbia and Chippewa that favored Crowley, 41.0%, 44.6% and 37.6%, respectively with Hong at 35.7%, 31.8% and 35.7%. With number of votes down, std deviation is up, Crowley 2.9% and Hong 1.8%.

All I did was match counties for number of democratic votes and winner. Even then there's a lot of noise in the system that polling can never beat down. No matter how often pollsters repeat margin of error is 3% or 5% (maybe they should lead with the number, not the poll result), the audience will never factor this into their thinking, they'll remember the number, not the error. And, in particular, audience can't integrate across all the polls and races to develop an intuition about distribution of polling errors.

The voting population is fundamentally hard to sample and no one can afford "n" that can really improve polling accuracy. Even if Elliott provides a fancy monte carlo simulation based on the actual voting data, only a few geeks reading SIN will be able to calibrate themselves to the expected error. Quantum uncertainty is alive and well.

Marc Robbins's avatar

Granted, polling primaries is hard and the polls tend to miss the actual outcomes by a very wide margin.

So here's a suggestion: don't release polling information that you know has a chance of being wildly wrong. And don't publish results if you're a news organization. And if you're SIN don't aggregate these unreliable numbers and present them as if they convey valid information about the electorate.

What public benefit is served by putting out highly unreliable information?

Stefan Haag's avatar

You state that the Democratic Party primary in Texas is a plurality contest, which it is not. As you know, the primaries in Texas require a candidate to receive a majority of the primary votes for that office. If no candidate wins a majority in the initial primary, the top two vote getters in the initial primary compete in a runoff primary.

Bob Fertik's avatar

Well one problem is pretty conspicuous - the polling averages added up to 109%

Michael Marino's avatar

Good point. Some kind of nonlinearity in the polling "averages," if not an outright error.

Laura Belin's avatar

Ann Selzer was way off in her Iowa Democratic primary poll of 2018. I interviewed her about the miss at the time:

https://www.bleedingheartland.com/2018/06/22/interview-ann-selzer-stands-by-sampling-method-for-primary-polls/

No doubt it's an even bigger challenge to get a representative sample now.