By Kabir Khanna and Anthony Salvanto
The estimates of the House of Representatives they come from a model, not just a survey result. The difference between the two is that while a survey tries to measure the totality of something by drawing a microcosm (a sample), a model takes that information and combines it with other factors about the people in the district and the district. more generally, to try to make an even better estimate of what is happening.
We use a procedure called multilevel regression and post-stratification (MRP), which is a great way to describe how it combines the information described above and combines data on people’s choices as well as broader district and national factors. We have worked on this model in collaboration with Professors Ben Lauderdale, Jack Blumenau and Doug Rivers, as well as the YouGov data science team. from the UK in 2015 and the 2016 US election. Here’s an introduction to how it works. (For a more technical description, see Andrew Gelman’s many writings.)
The first step is to talk to a lot of people across the country. There are more than 60 congressional districts where the race between Democrats and Republicans is likely to be competitive or close. It is a generously large number, not all of these districts will change, but it is within the range of the possible. We surveyed about 6,000 voters in all of these potentially competitive districts. We also collected about 25,000 interviews across the country, even in places outside of competitive districts, to get to know voters from other places. We asked them if they plan to vote Democrat or Republican in the House election in the district where they live.
The next step is to find out how the way you vote depends on its measurable characteristics. These characteristics include age, gender, race, education, who they voted for in 2016, where they live, and so on. Each voter has a certain combination of these, which we will call their “profile” for short. For example, a voter profile is someone who is 60 years old, a white woman with a college degree, a resident of the 7th District of the New Jersey Congress, and a Republican voter in 2016. If you change any of these characteristics, you will get a different profile. . For each of the many possible profiles, we calculate how many intend to vote Democrat and how many intend to vote Republican this year.
A good feature of MRP is that it efficiently combines information about similar types of voters, regardless of where they live in the country. This is especially useful in districts where we have received fewer interviews. So if we know a lot about, say, white working class voters, we can assume that they have some things in common, because opinions rarely stop at the boundaries of the district or even the state, especially when our policy is nationalized. If we have a good estimate of this subgroup nationwide, we can improve our estimate of the subgroup for any particular district. We also add local factors, which are also important, such as the party’s last ballot in the district and whether or not there is an incumbent running this year.
The next step is to estimate how many people with each voter profile live in each Congressional district, again using census data and other ancillary data (we increase this data by imputing turnout and voting in 2016 through the Population Survey current data and YouGov data, in addition to knowing how many people voted for each party in each district and state of Congress in 2016). This helps us determine the voting share of each party in a district. In each district, we multiply the number of people with a profile determined by the proportions of voters with that profile by choosing Democrats and Republicans. When we add these figures to all the voter profiles in the district, we get an estimate of each party’s vote share.
Each of the steps in this process has a certain statistical uncertainty, which we incorporate into our final estimates. In the final step, we simulate the congressional election 1,000 times and count how many seats each party wins in each simulation. Our seat estimate comes from the average simulation and the range of results gives us an idea of what is possible. (We report a 90% confidence interval, using the interval between the 5th and 95th percentiles of the simulated results.)
There are a couple of important warnings about this method. First of all, we are estimating how the race for the House is going from now on, with the expectation that things will change in the coming months. There is still a long way to go before the mid-term elections, during which more primaries will take place and voters will become more familiar with the candidates running in their districts.
Second, we are estimating which voters are likely to go to the polls in November based on recent historical patterns (this is known as a likely voter model). We estimate the proportion of voters in each profile who will participate using a similar model based on individual and geographic variables. Later, we will consider the implications of the different participation scenarios for which the party wins the House. For example, if Democrats can reproduce or approach the patterns of participation in a presidential year, they will improve their chances of winning the House. Republicans, on the other hand, expect a turnout pattern more similar to a typical mid-term year.
Add Comment