I checked 7 public opinion journals on Thursday, September 10, 2026 using the Crossref API. For the period September 03 to September 09, I found 11 new paper(s) in 4 journal(s).

Journal of Elections, Public Opinion and Parties

Insecurity and policy preferences: how economic and physical threats shape attitudes toward social welfare and criminal justice
Kaitlin Alper, Peter Starke
Full text

Journal of Official Statistics

From News to Index: A Practitioner’s Guide to a Deployable Sentiment Pipeline for Official Statistics
Younghwan Lee
Full text
This paper presents a deployable pipeline for constructing a news-based sentiment index (NbSI) for official-statistics use. The index is designed as a timely complement to survey-based consumer confidence measures when releases are delayed, observations are missing, or survey collection is temporarily disrupted. The pipeline is implemented using large-scale Korean economic news, manually labeled sentence-level sentiment data, pretrained word embeddings, and three standard neural classifiers: a convolutional neural network (CNN), an long short-term memory network (LSTM), and a transformer. The empirical analysis yields four main findings: (i) a compact CNN provides the strongest model-level operational trade-off, requiring substantially less training and inference time than the LSTM or transformer while delivering similar classification reliability; (ii) marginal gains from additional labeled data flatten beyond roughly 40k sentences, suggesting diminishing returns to large-scale annotation in this application; (iii) the resulting aggregate index is stable across classifier choices, supporting the use of the computationally efficient CNN as the baseline production model; and (iv) the NbSI leads Korea’s Composite Consumer Sentiment Index (CCSI) by about one month and is most useful when survey information is delayed or unavailable for sustained periods. These findings highlight the importance of transparent validation, label quality, monitoring, and maintainability when text-based indicators are adapted for official-statistics production.

Journal of Survey Statistics and Methodology

A Statewide Experiment on the Use of University Branding on Survey Mail Recruitment Envelopes
Kyle Endres
Full text
University branding is often included on envelopes used for mail recruitment materials for university-sponsored or administered surveys based on the assumption that signaling the university sponsorship improves the response rate. This postulation was experimentally tested by randomizing whether or not the university name and logo were included above the return address on the envelopes used to send recruitment materials for a statewide mixed-mode survey fielded with both an address-based sample (ABS) and an address-appended random digit dial (RDD) sample. For both samples, the statewide response rate was marginally lower when the university affiliation was included on the recruitment envelopes by 1.4 (ABS) and 1.7 (RDD) percentage points. Consistent across samples, the use of the university branding had significant negative effects on response rates in counties not geographically connected to the university. Closer to the university, however, the results were mixed, with the university branding significantly improving participation in the university’s home county by 7.9 percentage points for the ABS sample and having an insignificant effect in surrounding counties for both samples and in the university’s home county for the RDD sample.
Optimizing Survey Errors and Survey Costs: A Risk-Conscious Stopping Rule and Timing Analysis
Xinyu Zhang, James Wagner, Michael R Elliott
Full text
We introduce a risk-conscious stopping rule that considers uncertainty in predictions of survey costs and errors. The rule is particularly useful for repeated surveys with a few key statistics and tight budgets. It helps meet cost constraints by stopping a subset of cases early, aiming to minimize the negative impact on the quality of key statistics without adding new design features. To implement a decision rule that stops a subset of cases in the data collection process, a survey manager not only needs to choose which set of cases to stop, but also when to stop them. Implementing the stopping rule early may help to maximize cost savings, while decisions with reduced uncertainty can be made later as more data are collected. Dynamically identifying the “optimal” timing for implementing a stopping rule that relies on predictions can be difficult since future outcomes are unknown during data collection. We use data from the Health and Retirement Study to analyze how the timing affects the performance of the risk-conscious stopping rule. The Monte Carlo method is used to quantify uncertainty in stopping decisions. Our study illustrates several scenarios in which implementing the stopping rule after 8–10 call attempts has the lowest number of call attempts per interview while maintaining the same level of data quality.
Getting consent to survey in new ways: evidence from experiments about questions by text message in the Understanding Society Innovation Panel
Jim Vine, Annette Jäckle, Jonathan Burton, Mick P Couper
Full text
Some information cannot (reliably) be collected in annual surveys of panel members. Text messages (SMS) are well-suited to gathering time-sensitive responses, if a large and representative subset of a panel consents to being surveyed in this way. We investigate: (1) What proportion of respondents consent to text message questions? How does that vary by sample cohort, mode, and whether previously asked for consent? (2) What is the effect of consent question placement within the questionnaire? (3) Is the bias of the covered sample reduced by re-asking non-consenters at a later wave? (4) What coverage of non-internet users is gained by seeking consent for text message questioning? We use data from experiments within two waves (IP13 and IP15) of the Innovation Panel of Understanding Society: the UK Household Longitudinal Study. Most respondents consented to receive questions by text message when first asked (69 percent in IP13, 74 percent in IP15). Of those who declined consent at IP13, 55 percent consented when re-asked two years later. Face-to-face respondents were more likely to consent than web respondents. Slightly more (2.4 percentage points, s.e. 1.5 percentage points) respondents consented when experimentally asked late in the questionnaire than when asked early. Re-asking consent not only reduced the proportion of people who had not provided consent, but also reduced non-consent bias, measured across a range of socio-demographic characteristics. Some panel members who never use the internet have mobile phones and consent to receive questions via text message, so this may complement web surveying.
The Influence of Translator Backgrounds and Machine Translation on Statistical Properties of Surveys: Evidence from a Survey Experiment
Chia-Jung Tsai, Clemens Lechner, Dorothée Behr, Ulrike Efu Nkong, Anke Radinger
Full text
Comparable questionnaire translation is essential for drawing valid conclusions in cross-cultural survey research. Sound translation methodology, including the use of adequate personnel, is seen as crucial for reaching this goal. The recommended methodology should be empirically backed and stay tuned to latest developments, say machine translation. Against this backdrop, to investigate the potential effect of varied translators’ backgrounds and machine translation on the statistical properties of surveys, we conducted an experiment in which an English questionnaire was translated into German by 16 professional translators and 16 social scientists. We introduced two translation conditions: translation from scratch and post-editing (machine translation corrected by a human translator). Translations were subsequently fielded in web surveys. To investigate the quality of the survey data from these 32 translation versions (approx. 250 responses each), we use standardized mean distance and Cohen’s d with the official translation as a benchmark. We have four key findings: First, the resulting statistical means of the survey items vary, sometimes substantially, across translations. Second, post-editing, with its potential to prevent unwanted creativity and guard against gross mistranslations, is associated with a reduced gap between the survey data from the experimental questionnaires and the official translation; it also lowers the variability among different translations. Third, post-editing can lead to systematic bias if translation errors made by the machine are not identified or miscorrected. This applies to both expert groups. Fourth, social scientists are slightly more likely to deviate from a benchmark and, when translating from scratch, to produce translations leading to statistical outliers of survey data. This study highlights to what extent decisions concerning the choice of translators and the integration of machine translation can impact the statistical properties of survey data. We offer evidence to implement multi-step, collaborative translation procedures to enhance data comparability in cross-cultural studies.
Who are We Missing? Using Administrative and Contextual Data to Measure and Correct Partisan Nonresponse in Opinion Surveys
Michael Jackson, Scott Clement, Emily Guskin, Arifah Hasanbasri, Jordon Peugh, Cameron McPhee, Mark J Rozell
Full text
The 2020 US pre-election polling had the highest error in national polls in 40 years and significantly underestimated Donald Trump’s support. An AAPOR task force suggested that at least some of the polling error was likely caused by differential partisan nonresponse, but that this dynamic could not be evaluated without comparing respondents and nonrespondents. Many common sampling methods used for election, opinion, and social surveys include no reliable data on party identification for all sample units, making it difficult to measure and mitigate partisan nonresponse bias. Surveys using registration-based samples have leveraged voter databases to account for partisan nonresponse and validate self-reported turnout, suggesting that such administrative data could also be effective for general population surveys that rely on address-based samples. This article assesses methods for using voter and consumer databases, along with small-area contextual data, to measure and correct for differential partisan nonresponse in a national address-based survey of US adults conducted online and by mail. A model predicting partisan leaning was developed using an existing survey panel matched to voter and consumer databases and precinct-level election results. Among addresses sampled for the survey, 86 percent matched the voter or consumer database, and predicted partisanship aligned with self-identified party leaning for 65 percent of respondents. The inclusion of predicted partisanship in survey weighting, in addition to demographic weighting, improved the accuracy of recalled-vote estimates. The voter registration database was also used to validate self-reported turnout among respondents, which further improved the accuracy of recalled-vote estimates.
Using Large Language Models to Autocode Probing Responses to the Self-Rated Health Question: Strengths And Weaknesses in a Multilingual Survey Context
Mao Li, Stephanie Morales, Kaidar Nurumov, Sunghee Lee
Full text
Web probing is an adaptation of a cognitive interview administered in a Web survey setting to understand respondents’ thought processes in formulating their answers and/or to improve survey questions. It often uses open-ended questions, which require coding the responses, an essential step for transforming qualitative data into analyzable formats. This is labor-intensive and difficult, and more so for probing on subjective and evaluative assessments, like self-rated health (SRH). Recent advancements in large language models (LLMs) offer a promising avenue for automation of this step. However, in multilingual and cross-cultural contexts, where linguistic nuances further complicate coding, their effectiveness remains insufficiently examined. This study examined LLM autocoding of open-text responses to a Web probe on the SRH question from surveys conducted across five countries (the U.S., Mexico, Germany, Great Britain, and Spain) and in four languages (English, Spanish, German, and Korean). We evaluated the performance of five major open-source LLMs by comparing their autocodes against human codes for SRH attributes (e.g., health behaviors, illness) and tone of the coded attributes (i.e., favorable, unfavorable, or neutral in relation to health). We also compared the performance of the five LLMs against traditional machine learning models (MLMs). The best-performing LLM resulted in a micro-F1 score of 0.76 for attribute classification and 0.77 for tone classification, outperforming MLMs without requiring task-specific training. Its performance was also consistent across English, Spanish, and German responses but lower for Korean, a language generally considered low-resourced for LLMs. Overall, our results demonstrate the promise of LLMs in reducing burdens for coding complex text data from SRH probing in multilingual survey research, which likely applies to a broad range of open-ended coding tasks in survey research.
It’s Complicated: Effect of Actors’ Messages, Mode and Timing on Self-administered Survey Response Rates
Glenn D Israel, Don A Dillman
Full text
The movement toward self-administered surveys includes an opportunity to apply a holistic design across multiple requests to maximize response rates. Among ways for improving the sequence of requests may be to vary from whom the survey request is made, as well as the timing of the request, the mode used, and other design features. This study explored incorporating an organizational leader in survey requests as a strategy for improving response. Four studies using a sequence of different messages from actors in leadership and manager roles at different times and by different modes were conducted. Interspersing messages from an organizational leader did not substantially improve response rates as compared to when all of the messages were from the survey manager. However, the response rate increased and more responses came in sooner when a URL was included in an emailed leader message that closely followed one from the survey manager, and mail was used for the final two reminders. The results suggest that survey researchers should give increased attention to developing a holistic communication strategy for their target population.

Social Science Computer Review

Wild, Thick, and Wicked: Situated Evidence on AI-In-Use for Decisions About Deploying AI Systems
Reva Schwartz, Gabriella Waters
Full text
Organizations are adopting generative AI faster than the evidence base needed to govern it. Existing evaluation tools such as benchmarks, alignment scores, and safety tests were built for model development, not for judging whether systems will create value, introduce friction, or shift risk in specific real-world settings. As a result, there is little systematic evidence about how AI behaves once it is embedded in everyday work. This paper proposes a real-world AI evaluation framework focused on AI-in-use: how people actually appropriate, adapt, and work around AI systems in context, and what consequences follow over time. Instead of treating variability across users, tasks, and settings as noise to be controlled away, the framework treats that variation as the central source of deployment-relevant evidence. It sets out four design principles for producing decision-ready evidence at scale and proposes a shared evaluation architecture combining a structured observation environment, a metrics hub, and reusable consortium models that summarize system behavior across contexts. Rather than replacing traditional benchmarks, this framework adds a sociotechnical evidence layer that connects model capabilities to the organizational and practitioner level outcomes where deployment decisions are actually made.
Parsing Causal Relationships in Social Science Publications
Rasoul Norouzi, Bennett Kleinberg, Jeroen Vermunt, Caspar Van Lissa
Full text
Understanding the causes of phenomena is a key goal of scientific enquiry. For historic reasons, social scientists are reluctant to explicitly discuss causal assumptions. Nevertheless, causal claims do appear in the literature. Systematically cataloging these causal claims and parsing them into structured cause-effect relationship would allow us to synthesize them and, ultimately, construct theories based on oft-repeated causal claims. However, manually extracting these claims is infeasible due to the scale of the literature and linguistic ambiguity, and existing automated methods often lack domain specificity or fail to extract specific cause-effect pairs. To overcome these limitations, we introduce a BERT-based multitask deep learning model trained on a novel, domain-specific benchmark dataset of 3,014 annotated sentences from social science papers. The model simultaneously (1) identifies sentences containing causal claims, (2) extracts cause and effect spans, and (3) links the associated causal pairs. In our study, the multi-head model outperformed a model with sequential architecture, as well as two generative LLMs (Llama 3 8B and Qwen3 8B), which were prompted to code the sentences with a few examples of the coding schema. Specifically, the multi-head model had the highest macro F1-score across all three subtasks and lower false-positive error propagation compared to the sequential model. Its agreement with human coders on the held-out test set (Krippendorff’s α = 0.81) was of similar magnitude to the interrater agreement among the human coders, computed on an interrater reliability subsample (α = 0.80).