9 Comments
User's avatar
Farrah Bostic's avatar

Absolutely love this.

I’m increasingly interested in seeing our survey data (which is usually not in political contexts!) unweighted and weighted. The unweighted data can “merely” be descriptive (that is, “this is what this set of 1500 people said); the weights can have more predictive qualities. My bias as a qual researcher shows here - I think description is at least as valuable as prediction (but I know that media-based incentives and “getting hired by campaigns” incentives are different than those I’m dealing with).

And the thing I absolutely want is to see the assumptions that weights are based on elevated out of the methodology section, described/explained rather than just a weights table, and like you say, a couple of scenarios explored (“if the electorate looks like this, here’s what it looks like v if the electorate looks like that”).

Love love love this.

Alan's avatar

Very good useful analysis, thank you.

I don’t know anything about polling but I know a very little bit about modeling. This may be a dumb question, but do pollsters use data which correlates with voting even though it is not directly related to voting in their models? For example, if you had a high level of uncertainty about a population’s political behavior but a high level of certainty about physical location, education, and car buying preferences, is there a formal way to factor that into predicting voting behavior?

John Petersen's avatar

"polls can represent the public when politicians can’t or won’t" In bold in the article as it should be. While trying to decide who wins the horse race is entertaining, shifting public policy is far more important. Shining a bright light on what the public wants and holding whoever wins the elections accountable to the public desire is an important step toward better government.

Originally from Cleveland's avatar

Very interesting post. I work in areas related to public health where modeling is important because the available surveillance data tend to be scarce, at best, and underlying distributions of conditions and their "determinants" (i.e., correlates) often are unknown. Over time, the ability to forecast epidemics or estimate effects of interventions has not gotten better, even though some of the surveillance data have improved. It took several decades for the US to have usable HIV surveillance data, but prediction in specific populations or even post-hoc estimation of trend modifiers remains pretty awful. We have no ideas why new HIV cases among women in the US, across demographic groups has been declining for 25 years. The implementation of even straightforward interventions like universal, opt-out HIV testing has proven to underperform models, in part because heavy users of the health care system aren't necessarily the people most at risk for HIV acquisition and models often exclude controls for this even though the characteristics are pretty well known. CDC attempted to identify counties at high risk for developing HIV or HCV outbreaks among people who inject drugs and mostly didn't predict much of anything. Part of the problem was the very low base rate of HIV and the poor surveillance of HCV. Another problem was not really knowing how these diseases enter a community even among people who inject drugs. The outbreak in the areas that were sampled were mostly among homeless people and only in a few places with concentrations of them

You don't explicitly mention this here, but the universal problem with models is that they are almost always "underspecified" such that not everything of important is able to be measured (and sometimes isn't well-known enough or easily measured). Even where you have known determinants, they may be poorly measured and complex interactions among them (which are inherently unreliable) are not well known. The advantage that polling has over public health is that you are able to validate and refine models through elections, whereas public health outcomes often are low base rate events and statistical power may be lacking. You also have a history that allows looking at changes as data collection measures have evolved and you can look at process measures (like interview completion) that may be important.

My guess is that state polling is particularly difficult because there may be regional subtleties that, paradoxically cause problems that don't exert as much outsized influence in a national sample. You need to know the characteristics and underlying distributions of many small areas and to be sure that weighting isn't defeated by not, e.g., having enough people from the purple sections of the state that look like they should be more blue or red, based on demographics.

spanghew's avatar

I appreciate the efforts to make the complexities of polling apparent to people.

But even without such awareness, it should be obvious that it's harder to poll now than it once was.

I'm 64. When I was a child, if the phone rang, you answered it. Because only a few businesses and very rich people had answering machines, and it is just What People Did.

And if you did answer the phone, and it was a poll...well, that was probably the last you'd hear about it other than perhaps a followup poll.

Whereas now, we know that if we answer the phone, or a text, or an e-mail, or follow a weblink...we are most likely going to be besieged with calls, texts, e-mails, directed ads, etc.

I know many people who refuse to respond to polling questions because they're already bombarded with political junk-correspondence.

And, of course, all those different media mean the average person not only can take in but passively receives way more information than they used to. (Yet, curiously, this does not make them better-informed. Interesting, no?)

Bob Binstock's avatar

this is a truly excellent post! i was able to clearly understand how polling has radically changed without getting lost in the statistical weeds (as sometimes happens :). and now i will be able to share this information with others, in particular cautioning them to look for that qualification of uncertainty along with the hard number. thank you!

Mike Johnson's avatar

If there were a pro-progressive bias in polling, it should probably show up in polling within primaries more often, but that's not what happened in several races this summer - hell, DSA was freaking out that Claire Valdez was going to lose a week out, and she won by almost 20 points.

Stephen Clermont's avatar

"There’s no longer a “gold-standard” method."

Thank you 🙌

Joel Rosenfield's avatar

Tom Nichols is a fantastic military analyst. I read him often. But much like I don't see you write on military strategy, he should stay in his lane and keep his opinions on polling to himself.