đŸ€– Must-Read Articles đŸ€–

Experimental Feature

Even just looking backwards a week, there are a lot more articles published than most of us could hope to read. We can always skim the titles and abstracts ourselves, but I wanted to test out some automation. The articles below, all of which can be found elsewhere on this site among other new publications, were chosen by Google Gemini Flash Thinking as "must-read" articles. The proper criteria are in the eyes of the beholder and Gemini doesn't apply my criteria without error. I may continue to refine the prompting based on experience and feedback. Like the rest of the site, this will update daily!

Do algorithmic information environments mark the end of media effects research?
Annals of the International Communication Association
Isabelle Freiling, Jakob Ohme, Isabel I Villanueva, Dietram A Scheufele
Full text
Algorithmically curated information environments are beginning to challenge both long-standing notions of what media effects look like and our ability to study them directly. Algorithmic curation and microtargeting of content based on users’ behavioral and digital trace data pose conceptual as well as methodological challenges to our understanding of communication effects. These challenges increasingly introduce significant distortions to our field’s current ability to transparently and systematically iterate between theory and empirical testing. In this article, we provide a blueprint for how a systematic, empirical rethinking, and testing of communication effects theory can look like in an era of what we call preference-based information ecologies (PBIEs). Specifically, we discuss 9 current challenges to theorizing that are conceptual and methodological in nature and discuss necessary field developments to address and overcome these challenges. We identify 4 major agenda items that our field must tackle in order to position itself for a future in which all communication processes occur in some variant of PBIEs.
A systematic review of the deficit model and call for empowerment toward informed publics
Annals of the International Communication Association
Aart van Stekelenburg, Gabi Schaap, Harm Veling, Moniek Buijzen, Marieke L Fransen
Full text
The deficit model, assuming some deficit in the public relevant to science, is one of the most influential and discussed models of science communication. Despite its popularity, the literature appears to lack a unified definition and heavily criticizes the model. The first aim of the current work is to create clarity on the deficit model and as such address the controversy surrounding it. We distill a definition of the deficit model through a systematic review of all works in which the model plays a key role (k = 178). The results demonstrate that there is no such thing as the deficit model. Instead, descriptions take many forms, making the model too broad and imprecise to test empirically or apply in practice. The second aim is to investigate if we can build on the deficit model to create a more specific and testable model for scientific information dissemination. To that end, we introduce a new model that addresses criticism of the deficit model, focuses on informing publics, and is empirically testable. This empowerment model of informed publics identifies publics not as deficient, but as active agents that should be informed and empowered. Finally, we set out areas of research for science communication scholars investigating how to inform publics.
Beyond beyond standardization: studying robustness of empirical claims based on topic modeling through multiverse analysis
Communication Methods and Measures
Paul Balluff, Christina Viehmann, Maximilian Linde, Yannik Peters, Jun Sun, Chung-Hong Chan
Full text
A comparative test of the protection motivation theory: active and passive privacy protection in Brazil, Germany, the United States, and Vietnam
Journal of Computer-Mediated Communication
Yannic Meier, Laurent H Wang, Nicole C KrÀmer, Miriam J Metzger
Full text
Online companies surveil users around the world for commercial purposes. We advance protection motivation theory (PMT), by categorizing user responses to surveillance as either active or passive (i.e., self-inhibition) privacy protection behavior, integrating privacy literacy overconfidence, and exploring privacy resignation as a moderator. Finally, we adopt a comparative lens by investigating protection behavior in four national contexts: Brazil, Germany, the United States, and Vietnam. Results of an online survey (Ntotal = 2,092) show that perceiving surveillance as privacy threat was not consistently associated with active protection behavior. Rather, self-inhibition was a more likely response to surveillance which may have detrimental consequences for both individuals and societies. Across countries, perceived effectiveness of active protection positively relates to both forms of protection while active protection does not prevent self-inhibition. Interestingly, privacy resignation may not always impede protection. In summary, not all aspects of PMT are generalizable when using a culturally diverse sample.
Mapping Out the Context
Journal of Media Psychology
Bingjie Liu, Andrew Gambino, Lewen Wei
Full text
Abstract: Generative artificial intelligence (GenAI) can write humanlike, context-aware messages and thus can aid individuals in their communication with others. Existing research suggests that using GenAI may be regarded as inappropriate or undesirable in certain interpersonal contexts, but acceptable in others. Yet, few studies have articulated the reasons underlying such context-dependent preferences. Taking a functional approach, we conducted an online survey ( N = 542) to explore how GenAI use might depend on an individual’s goals in interpersonal communication and to unravel the reasons why an individual may or may not use GenAI in specific contexts. Among the participants who had used GenAI in their interpersonal communication with others ( n = 240), we found that GenAI was most often used for goals of information and resource exchange. When message production for a goal requires more of one’s subjective experience (e.g., providing emotional support, showing others who you are), individuals are less likely to use GenAI in such contexts. Other contextual factors, including whether message production for a goal is difficult and whether it requires objectivity or human knowledge, were not found to predict GenAI use. These findings contribute to the theories and research in human–AI interaction and message production, while also informing the future design of human-centered AI.
How Online Dating Motivations and Social Networks Affect Warranting Value and Interpersonal Impressions
Journal of Media Psychology
Megan A. Vendemia
Full text
Abstract: Popular online dating platforms feature self-authored profiles in which users can indicate their relational goals and interests. The ability to easily modify self-generated information raises questions about whether online daters are who they say they are offline. Informed by warranting theory, this experiment examined how online daters’ relational motivations and connections to social networks affect viewers’ authenticity judgments and interpersonal impressions. Results revealed that romantic relationship motives heightened viewers’ anticipated future interaction beliefs, which in turn, bolstered online daters’ perceived authenticity, trustworthiness, and goodwill. Results also demonstrated that connections to online daters’ broader social networks reduced relational uncertainty. Implications for warranting theory and online daters’ profile construction are discussed.
Is Information Verification Just a Matter of Intellectual Humility?
Journal of Media Psychology
Carmela Sportelli, Paolo Giovanni Cicirelli, Giuseppe Corbelli, Marinella Paciello, Francesca D’Errico
Full text
Abstract: Adolescents can be vulnerable to misinformation, facing the daily challenge of assessing false, inaccurate, or misleading news. Understanding the psychological processes contributing to this vulnerability is crucial for enhancing their information verification skills. Previous research has established a positive association between intellectual humility – the acknowledgment of the fallibility of one’s own beliefs – and the intention to verify misleading content. However, the underlying process through which it occurs remains underexplored. Intellectual humility may also be linked to moral disengagement: individuals characterized by an awareness of their own fallibility can be less inclined to rely on moral disengagement mechanisms, which serve as self-serving defensive justifications used to preserve a positive self-image when faced with moral transgressions, thereby being more likely to engage in rightful online behaviors. Therefore, our research aims to test the relationship between intellectual humility and the intention to verify online content, investigating the mediating role of moral disengagement. The study, conducted with 380 Italian adolescents ( M = 14.43, SD = 0.74), utilized a conversational web app to measure the variables of interest. The results show that moral disengagement fully mediates the relationship between intellectual humility and the intention to verify the information. This highlights the importance of fostering metacognitive abilities, such as self-reflection and the capacity to reconsider one’s own perspective, in promoting media literacy. These findings suggest that interventions aimed at enhancing intellectual humility may reduce moral disengagement toward fake news and promote more responsible online behavior. Limitations and directions for future research are also discussed.
Authoritative Yet Conditionally Cohesive: The Cohesion Dynamics of Journalistic (Non-) Authority on Chinese TikTok (Douyin)
Journalism & Mass Communication Quarterly
Peng Xiao, Wen Shi, Yangkun Huang
Full text
The rise of algorithmic curation challenges the traditional authority of journalism. Focusing on Douyin, this study employed agent-based testing and temporal network analysis to examine how recommendation systems generate network cohesion around Chinese state and non-state media. Results revealed that state media’s cohesive power stemmed from news, whereas cohesion around non-state media was sustained by non-news content. However, for both media types, the cohesion generated by news showed a significant decline over time, and this trend varied by topic. These findings suggest that, beyond mere institutional status, strategic content adaptation is fundamental to building network cohesion in algorithmic environments.
Threat and Remedy: AI’s Technological and Normative Role in Democratic Discourse and Counter Speech
Media and Communication
Diana Rieger, Mario Haim
Full text
Artificial intelligence (AI) has fundamentally transformed online discourse, serving simultaneously as a source of and a potential solution to threats to democracy. This article conceptually examines AI’s multifaceted role in digital public spheres, beginning with an analysis of how AI shapes democratic processes online, ranging from information curation and political participation to dialogue facilitation and the construction of epistemic infrastructures. We identify several key risks: AI amplifies harmful content through automated content generation, coordinated manipulation, and the algorithmic mainstreaming of borderline content that often evades traditional detection systems. Within this broader democratic context, we position counter speech as a critical intervention strategy and examine both reactive measures (e.g., automated detection and content moderation) and proactive approaches (e.g., prebunking, algorithmic downranking, friction design, and AI-mediated dialogue). We argue that effective counter speech increasingly relies on human–AI collaboration, where AI supports human counter speakers with factual resources, emotional scaffolding, and scalability, while human actors preserve the authenticity and agency that make counter speech normatively meaningful. However, realizing the potential of such collaboration requires moving beyond technological fixes toward democratically legitimated governance structures. Drawing on examples from Wikipedia and decentralized platforms, we demonstrate how transparent and participatory institutional arrangements can foster more resilient discourse environments. We conclude that societies can harness AI’s democratic potential while mitigating its risks only by proactively establishing legitimate deliberative processes grounded in broadly shared discourse norms.
Are they virtue-signalling or sincere? The impact of perceived motives on the effectiveness of online collective action
New Media & Society
Lisette Yip, Emma F. Thomas, Hawa Muhammad Farid, M. Noor Tahir
Full text
Social media is an important tool used by supporters of modern social movements to persuade others to participate. However, people who engage in online collective action often face speculation about their motives and may be perceived as acting for self-serving reasons. How do these perceptions impact the effectiveness of collective action? In three experiments ( N = 916), participants viewed posts (Study 1) or an immersive social media feed (Studies 2–3) where a user encouraged them to take action for a political cause. They stated that they acted out of genuine concern (autonomous motivation), to avoid guilt and social disapproval (controlled motivation), or did not specify (baseline). Participants who viewed controlled-motivated posts perceived the user as less autonomously motivated and were therefore less likely to socially identify as supporters or take collective action. Thus, movements will be more effective at mobilising (online) support when they are seen to arise from autonomous motives.
Cross-platformization: How U.S. right-leaning media curate their posts on Twitter and Truth Social
New Media & Society
Josephine Lukito, Yini Zhang, Bin Chen, Stephen Prochaska, Meredith L. Pruden, Yunkang Yang, Wei Zhong, Megan A. Brown, Ross Dahlke, Jason Greenfield, Jiyoun Suk, Porismita Borah
Full text
This study examines how seven right-leaning media organizations utilize cross-platformization strategies to reach audiences across multiple platforms; in this case, Twitter and Truth Social. Focusing on both content volume and story packaging, we collect right-leaning media organizations’ social media posts and content from the 2022 U.S. midterm election. Combining computational and qualitative methods, we find that outlets posted only a small share of total stories across the two platforms and were selective in their content, often posting only on Twitter or Truth Social. For content posted on both platforms, the outlets posted content that ideologically aligned with the users of that platform. More specifically, right-leaning outlets on Truth Social tailored their posts to ideologically aligned users.
Beyond Factual Correction: Integrating Meta-Cognitive Deficits Into the Risk Information Seeking and Processing Model to Address HPV Misperception
Science Communication
Yuchen Wang, Yuan Zhong
Full text
Overconfidence, explained by the Dunning–Kruger effect (DKE), hinders engagement with corrective health information. Integrating the DKE with Risk Information Seeking and Processing (RISP) model, a 2 (meta-cognitive vs. informational intervention) × 2 (simple rebuttal vs. factual elaboration) experiment examined human papillomavirus misperceptions. Results show that meta-cognitive interventions limited subjective knowledge, effectively boosting information insufficiency. Conversely, informational inoculation increased subjective knowledge, inadvertently suppressing the insufficiency gap essential for motivation. Path modeling confirmed that information insufficiency predicts seeking and preventive intentions. We demonstrate that overcoming overconfidence requires recalibrating knowledge perceptions rather than merely providing facts, establishing meta-cognition as a contingent condition for RISP.
A Causal Inference on the Effects of Attention to Media Content About Autonomous Vehicles in the Influence of Presumed Media Influence Model
Science Communication
Tong Jee Goh, Hongjie Tang, Shirley S. Ho
Full text
Applying the influence of presumed media influence model into an autonomous vehicle (AV) context, this study examines how media viewers’ attention to content about risks and benefits of AVs affects their presumptions of influence of the content on others, perceptions of trustworthiness of the technology, and intention to ride in AVs. Through cross-lagged analyses of two-wave panel data (n wave1 = 1,306; n wave2 = 650), this study found causal evidence for the projection effect and impersonal impact. Notably, this study also found causal evidence for the consonance effect: viewers perceiving trustworthiness of AVs after developing intention to ride in AVs.
Once Upon a Genome: A Content Analysis of Narratives and Narrativity in News About Genomic Research
Science Communication
Helena Bilandzic, Susanne Kinnebrock, Markus Schug, Theresa Engstler
Full text
This content analysis investigates the functions and types of narratives in German science news about genomic research as well as their degree of narrativity. The results show that nearly half of all articles contained at least one narrative. They were used predominantly to corroborate scientific findings. Narratives about the research process were most common, followed by narratives about researchers and personal fates of affected individuals. Personal-fate narratives displayed the highest levels of narrativity, research-process narratives the lowest. The findings highlight the value of a differentiated approach to studying narratives in science communication.
Measuring Partisanship and Representation in Online Congressional Communications
American Political Science Review
MICHAEL KISTNER, MICHAEL HESELTINE, ROBERT ALVAREZ, MAYA FITCH, LUCAS LOTHAMER, ELIZABETH SIMAS
Full text
Social media and the internet have created new ways for representatives to communicate. How have members of Congress responded to these opportunities? We introduce a multi-platform dataset of congressional communications extending back to the onset of the social media era. Using computational language processing, we classify approximately 4.7 million tweets, 2.4 million Facebook posts, and 184,000 email newsletters authored by members of Congress between 2009 and 2022 based on intended purpose, and scale the partisanship of each message along a continuous left–right dimension. After validation, we demonstrate how our data can be used to study partisanship and representation in the contemporary Congress. Importantly, our data show congressional rhetoric has become more partisan and negative as social media usage has increased. We identify one potential mechanism contributing to this trend: partisanship and negativity receive inflated levels of positive engagement on social media relative to other forms of messaging like credit claiming or constituency service.
Correcting for Nonignorable Nonresponse Bias in Ordinal Observational Survey Data
Political Analysis
Lukåƥ Lafférs, Jozef Michal Mintal, Ivan Sutóris
Full text
Many political surveys rely on post-stratification, raking or related weighting adjustments to align respondents with the target population. But when respondents differ from nonrespondents on the outcome itself (nonignorable nonresponse), these adjustments can fail, introducing bias even into basic descriptives. We provide a practical method that corrects for nonignorable nonresponse by leveraging response-propensity proxies (e.g., a respondent’s rating of the interview or interviewer-coded cooperativeness) observed among respondents to extrapolate toward nonrespondents, while directly integrating observable covariates and retaining the benefits of post-stratification with known population shares. The method generalizes the variable-response-propensity framework of Peress (2010) from binary to ordinal outcomes, which are widely used to measure trust, satisfaction and policy attitudes. The resulting estimator is computed by maximum likelihood and implemented in a compact R routine that handles both ordinal and binary outcomes. Using the 2024 American National Election Study, we show that accounting for nonignorable nonresponse produces substantively meaningful shifts for life satisfaction (estimated latent correlation ρ ≈ 0.47 $\rho \approx 0.47$ rho almost equals 0.47 ), while yielding only modest changes for retrospective economic evaluations ( ρ ≈ 0.14 $\rho \approx 0.14$ rho almost equals 0.14 ), highlighting when nonignorable nonresponse substantively affects survey estimates.
High accuracy with low costs: the pretrain-finetune paradigm for classification with transformer-based language models
Political Science Research and Methods
Cecilia Y. Sui
Full text
Political science increasingly uses text classification to gauge subtle concepts such as toxicity or anger. Traditional methods treat words in isolation, overlooking the contextual dynamics where meaning resides. Transformer-based language models address these limitations but remain underutilized, especially among applied scholars, partly due to misconceptions about their mechanisms and computational requirements. This article bridges this gap by offering an accessible explication of the pretrain-finetune paradigm, focusing on underlying mechanisms and their potential to improve political text analysis by harnessing models trained on extensive datasets while requiring modest labeled data and computational resources. I demonstrate the approach by identifying toxic language in conversation threads following U.S. Senators’ tweets and provide an online tutorial to support nontechnical scholars in adopting the methodology.
A Statewide Experiment on the Use of University Branding on Survey Mail Recruitment Envelopes
Journal of Survey Statistics and Methodology
Kyle Endres
Full text
University branding is often included on envelopes used for mail recruitment materials for university-sponsored or administered surveys based on the assumption that signaling the university sponsorship improves the response rate. This postulation was experimentally tested by randomizing whether or not the university name and logo were included above the return address on the envelopes used to send recruitment materials for a statewide mixed-mode survey fielded with both an address-based sample (ABS) and an address-appended random digit dial (RDD) sample. For both samples, the statewide response rate was marginally lower when the university affiliation was included on the recruitment envelopes by 1.4 (ABS) and 1.7 (RDD) percentage points. Consistent across samples, the use of the university branding had significant negative effects on response rates in counties not geographically connected to the university. Closer to the university, however, the results were mixed, with the university branding significantly improving participation in the university’s home county by 7.9 percentage points for the ABS sample and having an insignificant effect in surrounding counties for both samples and in the university’s home county for the RDD sample.
Changes in Response Quality Over Repeated Measurement in Ecological Momentary Assessment—A Three-Week Observational Study
Journal of Survey Statistics and Methodology
Minglei Wang, Shuaiying Cao, Chan Zhang, Marc S Tibber
Full text
Ecological Momentary Assessment (EMA) is an intensive longitudinal data collection method to capture in-the-moment experiences through frequent assessments. With the advent of mobile technologies, EMA’s applications have expanded across a number of disciplines. Despite its growing popularity, methodological issues, particularly regarding response quality, have yet to be explored. This study systematically evaluates how response quality changes over time in a 21-day EMA study with five daily assessments. Data were analyzed from 100 university students who completed surveys via a mobile application. Response quality was measured using a variety of satisficing behaviors, including speeding, nondifferentiation in grids, extreme rounding (i.e., reporting 0 or 60) on questions that asked the duration of a given activity during the past hour, and anchoring (i.e., providing the same answer to a question as in the previous assessment). The findings revealed a significant increase in speeding over the three weeks, suggesting a possible decrease in response quality. There were also increases in extreme rounding and anchoring responses in weeks 2 and 3. The nondifferentiation in grids mainly stayed the same across the weeks, which might be because the grids in this study only contained a few simple items to rate. The findings also showed variation between participants regarding their satisficing behaviors across the weeks. However, such variation could not be explained by participants’ demographic characteristics or motives for participation. Compared to the changes in satisficing behaviors across the weeks, the response quality differences across the five daily assessments were less systematic and inconsistent. These findings highlight the need for effective strategies to motivate participants to provide thoughtful answers as EMA data collection proceeds, and careful consideration of EMA design parameters to ensure high-quality data collection.
Don’t Look up: Evaluating the Tradeoff Between Performance and Sustainability of Text Classification Using Open LLMs
Social Science Computer Review
Sean Hamilton-Palicki, Isaac Bravo, Clint Claessen
Full text
The increasing adoption of Large Language Models (LLMs) as a text analysis method in social science presents a critical yet under-examined trade-off between model performance and environmental sustainability. This research provides a systematic evaluation comparing the performance, energy consumption, processing time, and CO 2 emissions of various computational text analysis methods (CTAM), including dictionaries, trained classifiers, and self-hosted open LLMs when performing sentiment analysis of parliamentary speeches, classification of open-ended survey responses, and named entity recognition of newspapers. The analysis is limited to self-hosted deployment in local and server environments where per-task energy consumption is directly measurable. Although self-hosted LLMs demonstrate strong performance in sentiment analysis, closely aligning with human judgment, they require significantly more energy and time than non-LLM approaches. For classification and named entity recognition, pretrained task-specific models achieve better F1 scores with a lower carbon footprint, challenging the primacy of larger models. To navigate this trade-off, we propose a CO 2 -Adjusted F1 Score that penalizes emissions while rewarding performance. Applying this metric, we show that smaller, task-specific models may be preferred over larger general-purpose LLMs for efficient text analysis. We highlight the necessity for thoughtful and responsible model selection, promoting a “right-fit” approach for CTAM.
Saturation of moral language predicts lower content engagement on social media
Nature Human Behaviour
Cristian Candia, Mohammad Atari, Nour Kteily, Brian Uzzi
Full text
Moral language often travels widely online, but does more moral content always correspond to higher engagement? We analysed 1,621,147 observations across 13 socio-political topics on Twitter (n = 530,104), Reddit (n = 1,048,653) and 8chan (n = 42,390). Using Distributed Dictionary Representations—word embeddings scored against an expert-validated moral dictionary—we measured moral loading (that is, a post’s overall moral relevance) and moral density (that is, concentration of moral content across words). Negative-binomial models showed that moral loading was positively associated with engagement (range 1.12 [0.96, 1.28] to 9.07 [8.21, 9.93], all P < 0.001). Conditional on moral loading, however, moral density was ‘negatively’ associated with engagement (range −4.71 [−5.41, −4.02] to −0.40 [−0.52, −0.27], all P < 0.001). Engagement peaked at density 0.30 [0.30048, 0.30070], P < 0.001, with lower engagement below (2.28-fold) and above (2.78-fold), consistent with an engagement advantage for moral language that is bounded by an overmoralization penalty pattern.
State-sponsored agenda setting: Measuring the impact of narrative laundering
Proceedings of the National Academy of Sciences
Patrick L. Warren, Darren L. Linvill, Camille J. Saucier
Full text
Foreign information manipulation and interference (FIMI) is an important part of the digital information ecosystem, yet its effects are poorly understood. Most research analyzing FIMI to date is descriptive, exposing efforts to erode trust, promote favorable narratives, and target specific communities. Existing work evaluating the impacts of FIMI contrasts those directly exposed to FIMI messaging (by viewing the original messages) with those who were not, finding little evidence of changes in beliefs or behavior. Rather than persuading individuals directly, we argue FIMI often operates through agenda-setting processes that shape which issues and narratives dominate public discourse. A key component of this process is narrative laundering, whereby misleading content is scrubbed of its origins and circulated into mainstream channels. This study examines the Russian-aligned Storm-1516 campaign, which uses state media, fabricated sites, and paid influencers to elevate false narratives, especially those that implicate Ukrainian President Volodymyr Zelenskyy in corrupt acts. By analyzing the dissemination of these narratives, we assess how Storm-1516 shifts the public agenda, illustrating how foreign influence can exert substantial effects on information environments even when individual-level persuasive effects may be limited.
Leveraging generative AI for causal inference with unstructured data
Proceedings of the National Academy of Sciences
Kosuke Imai, Kentaro Nakamura
Full text
We introduce GenAI-Powered Inference (GPI), a statistical framework for causal inference using unstructured data, including text and images. GPI leverages open-source pretrained Generative AI (GenAI) models—such as large language models and diffusion models—not only to generate unstructured data at scale but also to extract low-dimensional representations that are guaranteed to capture their underlying structure. Applying machine learning to these representations, GPI enables estimation of causal effects while quantifying estimation uncertainty. Unlike existing approaches to representation learning, GPI does not require fine-tuning of GenAI models, making it computationally efficient and broadly accessible. We illustrate the versatility of the GPI framework through three applications: 1) estimating the effects of Chinese social media censorship while adjusting for textual confounders, 2) isolating the impact of specific image features from that of other correlated features in the same image, and 3) assessing the persuasiveness of political rhetoric. An open-source software package is available for implementing GPI.
Digital twins are funhouse mirrors: Five systematic distortions
Science Advances
Tianyi Peng, Melanie Brucks, George Gui, Daniel J. Merlau, Grace Jiarui Fan, Malek Ben Sliman, Eric J. Johnson, Abdullah Althenayyan, Silvia Bellezza, Dante Donati, Hortense Fong, Elizabeth Friedman, Ariana Guevara, Mohamed Hussein, Kinshuk Jerath, Bruce Kogut, Akshit Kumar, Kristen Lane, Hannah Li, Vicki Morwitz, Oded Netzer, Patryk Perkowski, Olivier Toubia
Full text
Scientists and practitioners are aggressively moving to deploy digital twins—large language model (LLM)-based models of individuals—across social science and policy research. We conducted 19 preregistered studies with 164 diverse outcomes (e.g., attitudes toward hiring algorithms and intention to share misinformation) and compared human responses with those of their digital twins (trained on each person’s previous answers to more than 500 questions). We establish an empirical benchmark for digital twin performance: Digital twins’ answers are only modestly more accurate than those from the (homogeneous) base LLM and correlate weakly with human responses (average correlation coefficient of 0.20). To guide future development, we document five ways in which digital twins distort human behavior: (i) insufficient individuation, (ii) stereotyping, (iii) representation bias, (iv) ideological biases, and (v) hyper-rationality. We make our full dataset and code public as a standardized testbed. Our results caution against premature deployment while laying the groundwork for the transparent, replicable, and iterative science necessary for responsible deployment of digital twins.
Can a public information campaign increase trust in American elections?
Science Advances
Thad Kousser, Seth J. Hill, Neil Malhotra, Robert M. Stein
Full text
Confidence in the conduct and results of American elections has declined and polarized along party lines in recent years. Can a nonpartisan public information campaign featuring election officials enhance trust in elections when delivered to skeptical voters during a contentious political contest? We analyze a preregistered, randomized field experiment conducted during the 2024 US presidential election in which conservative voters were randomly assigned to a control condition or to receive a multimodal messaging campaign designed by political professionals explaining protections on election integrity. In contrast to the large positive effects of informational messages found in recent survey experiments, we observed primarily null effects with respect to turnout ( n  = 147,488), general confidence in elections ( n  = 4740 to 4868), and perceptions of the frequency of fraud ( n  = 4925). We observed one hypothesized effect: an increase in trust in the vote-by-mail process ( n  = 6438). This study provides rigorous evidence from a real public information campaign on election integrity, yielding effect sizes in the field that are much smaller than those observed in prior survey settings.