The Rise of Hostility in Podcasts
A Five-Year Analysis of Expression Targeting Select Groups
MONTREAL | NOV 2025
“When hate manifests in our communities, it is a threat to public safety, democracy, and human rights. Hate divides us [...] It silences people, it shuts down debate and hinders our democracy. [...] Without accountability, hateful behaviour that has been appallingly normalized online is now more frighteningly common in person.”
- “Canada Must Act Now to Address Hate”, Canadian Human Rights Commission (Nov 2023)
Introduction
This study provides a longitudinal assessment of how hostile rhetoric toward select groups has evolved in popular North American political and news podcasts.
Central finding: across more than 70,000 hours of audio spanning January 2020 to August 2025, hostile rhetoric targeting specific groups has risen sharply, growing more frequent, more severe, and, most recently, more persistent, no longer spiking and cooling with individual events but holding a rising baseline.
Key Takeaways
- Hostility surged after October 7, 2023, and has not receded. Before the war in Gaza, hostile remarks toward Israeli, Palestinian, Jewish, or Muslim people appeared only sporadically in our data; afterwards, both their frequency and severity increased substantially.
- Since January 2025, hostility has decoupled from the news cycle. Event-driven spikes have given way to a sustained, escalating baseline: hostile rhetoric is becoming persistent rather than reactive.
- Hostile framing blurs its targets. Hostile commentary aimed at the Israeli side of the conflict rarely targets Israelis as a national group; it far more often invokes Jewish identifiers, blurring Jewish and Israeli categories.
- Every figure in this report is a lower bound. Our detection pipeline is deliberately conservative, validated at 95.8% precision against trained human coders, at the cost of missing roughly a third of what humans identify as hostile.
The stakes of this measurement are considerable. When negative portrayals or hostile rhetoric toward particular groups are encountered repeatedly, through trusted voices, over extended listening, they can accumulate into broader dynamics of polarization, stereotyping, and intergroup tension beyond the medium itself. The mechanism is well documented. Research links dehumanizing rhetoric about out-groups to greater intergroup hostility and support for harming them (Haslam 2006; Kteily et al. 2015; Kteily and Bruneau 2017). Repeated exposure to hate speech has been associated with desensitization and heightened prejudice toward the targeted group (Soral et al. 2018), and with reduced empathy accumulating across audiences and communities (Pluta et al. 2023; Madriaza et al. 2025). All of this unfolds inside a wider climate in which hostility is increasingly an affective, identity-based phenomenon rather than mere policy disagreement, and exaggerated perceptions of the other side's contempt feed further hostility (Iyengar et al. 2019; Moore-Berg et al. 2020). It is the cumulative, ambient character of such speech, as much as its peaks, that makes measuring it at scale worthwhile.
Podcasting is distinctively suited to that mechanism, for three reasons. First, its form is exactly the repeated, trusted, long-form exposure that persuasion research identifies as most effective. The features of the medium plausibly amplify that effect, through mere-exposure and source-credibility dynamics, the parasocial trust listeners place in familiar hosts, and the sustained opportunity for message elaboration that hours of long-form talk afford (Hovland and Weiss 1951; Zajonc 1968; Petty and Cacioppo 1986; Tukachinsky et al. 2020), so hostile characterizations arrive through a familiar, credible voice and recur across episodes rather than passing once. Second, its reach is now mainstream and skews young: about half of US adults report listening to a podcast in the past year, rising to two-thirds among those under 30 (Shearer et al. 2025), and news podcasting draws a disproportionately young, politically engaged audience (Newman et al. 2025) for whom the medium is increasingly a primary source of political information, so the norms set here may shape a cohort's sense of which groups it is acceptable to disparage. Third, the medium sits largely outside scrutiny: popular shows fall outside the editorial standards and codes of practice that govern broadcasters, and largely outside the platform-accountability regulation taking shape across democracies, whose obligations attach to platforms hosting user content rather than to host-produced audio itself.
Despite that prominence, podcasts remain comparatively understudied, an escape from scrutiny that observers of the medium flagged early (Wirtschafter 2021, 2023), and concern about this discourse has grown while the evidence has lagged. A small but growing computational literature has begun to map the podcast ecosystem at scale and to measure toxicity within it (Litterer et al. 2025; Rizwan et al. 2025), but that work is largely point-in-time and concerns general toxicity rather than the longitudinal, group-specific question this report addresses. Part of the reason for the lag is technical: ingestion and transcription at scale, and the labeling of unstructured conversational content, are substantial obstacles. But part of it is deeper. Hostility toward a group is most often carried implicitly, by insinuation, coded reference, and framing rather than by any flagged slur (ElSherief et al. 2021), which is precisely what keyword and lexicon methods miss, and what has kept the field's clearest evidence anchored in short-form text (Siegel 2020). Hearing it requires reading each passage in context, at a scale no manual effort can reach, a task that has only recently become tractable.
To make this assessment possible, we introduce a structured framework and supporting analytical systems that enable large-scale, longitudinal evaluation of how podcasts characterize specific groups. We apply this framework to five years of leading North American political and news programs, revealing how the tone toward select groups has evolved over time. Throughout, we measure language, not truth: the framework assesses how groups are characterized, never whether the underlying claims are true, justified, or fair.
A note on AI in this research: the detection pipeline at the core of this study uses large language models to classify transcript segments. The segment-level hostility labels are model outputs, validated against trained human coders (see Validation), but not individually human-adjudicated.
Podcast Sample
Our sample includes 140 leading North American podcasts, identified from Apple and Spotify rankings in June 2025 and selected for their strong focus on political and news content. Both choices are deliberate. Because this study is about hostile speech that reaches large audiences, reach is part of the population definition, and the platform rankings are a public, externally defined record of what listeners actually choose. And because hostility toward identity groups is overwhelmingly a feature of political, news, and commentary discourse, we narrow to the shows where that discourse lives. We collected all available episodes released from January 2020 to August 2025.
The timeline on the right illustrates the dataset, with each row corresponding to a podcast and columns reflecting time periods.
Darker gray squares represent periods where we collected and analyzed podcast episodes.
Transcript Segmentation
Episodes were collected directly from official RSS feeds and transcribed using a speech-to-text model.
Transcripts were divided into overlapping segments of approximately 200 words, corresponding to roughly 75 seconds of audio, about a paragraph of talk. The segment is the unit of analysis for everything that follows: small enough to be about one thing, long enough to carry the context that decides whether speech is hostile.
This approach comes at the cost of losing some long-range connections. A speaker might, for instance, introduce “those people” early in an episode, clarify the referent much later, and make hostile claims near the end, with each element falling into separate segments. Disclaimers (“I’m just asking questions”), callbacks, gradual escalation, and coded language may also occur too far apart to appear together. We accept this trade-off because it enables automated analysis on a scale that would be impossible through manual coding.
Although the method may miss long-range references, it does so consistently across groups and time periods. Later analysis shows that more than 80% of the hostile speech acts identified in this study occur entirely within a single 75-second window.
Podcast Hosts
Who says something matters as much as what is said: the parasocial trust that makes podcasts persuasive attaches to the familiar voice of the host, not to a passing guest. To distinguish whether a statement was made by the host or another participant, we created a database of podcast hosts and their characteristic voice patterns. These profiles enable us to automatically perform speaker attribution across all our podcast segments.
Defining Target Groups
Hostile expression requires two components: specific linguistic patterns and a target. In this study, targets are defined as identifiable groups: entire national, ethnic, or religious populations defined by their shared identity.
The design is equally disciplined about who the hostility targets: we count only speech aimed at groups of people as such, excluding hostility routed through states, governments, militaries, or armed factions, so that comparison across groups remains meaningful. Criticism of governments or military forces does not constitute hostile expression toward our defined target groups. We analyze only instances where speakers make claims about the collective characteristics, behaviors, or nature of groups as a whole, including major subsets within those groups (e.g. Sephardic Jews, Catholic Christians, etc.).
This people-only bar is intentionally conservative, and its costs are not evenly distributed. In particular, it understates hostility toward groups that are discussed primarily through geopolitics, because negative sentiment expressed as criticism of a state or regime is excluded by a construct that deliberately sets state-directed speech aside. We accept this trade-off because it keeps the construct narrow and internally consistent. The result is a measure of hostility directed at peoples, not criticism directed at governments.
Our target groups were selected according to three criteria: frequent mentions in the podcasts we analyzed, direct relevance to Canada, and prominence in major events during the study period.
- COVID-19’s origins in China and rising US–China competition
- The war in Gaza
- Rhetoric on Canadian annexation and US-Canada tariff disputes
- The creation of the US Task Force to Eradicate Anti-Christian Bias
Through this process, we identified seven groups for analysis: Chinese, Jewish, Israeli, Palestinian, Muslim, Christian, and Canadian people. Each of these groups received a custom lexicon, validation process, and longitudinal tracking.
These groups were not chosen at random: each is frequently named across the corpus and salient to the events that defined the study window, so that every target draws enough discussion to support a stable, comparable measurement. We acknowledge that hostile rhetoric in these podcasts extends beyond the groups analyzed here; our study focuses on this defined subset.
Framework
Our study begins with compiling instances of hostile expression directed at our target groups. This database is designed to answer the question: Did Podcast X express hostility toward group Y during a given month?
To do this, we divide the five-year study into monthly time periods. We then assess each podcast's general attitude toward a group by identifying clear and representative statements made about that group. For a statement to qualify for analysis, we apply a strict bar: one of our target groups must be unambiguously the subject of discussion, hostile speech must be clearly expressed about that group, and the speaker must voice the hostility as their own view: endorsed, not merely quoted, reported, or presented for critique.
=
(Qualifying statements about the target group in a calendar month)
All qualifying statements are then grouped by podcast, month, and target group to form what we refer to as a hostility snapshot. Each snapshot represents the collection of detected hostile remarks directed at a group during that period. To ensure that these snapshots reflect how a podcast repeatedly spoke about a group over the course of a month, rather than a one-off comment, we require multiple qualifying statements per snapshot.
Darker shades indicate greater hostility.
By collecting hostility snapshots month by month and organizing them into a timeline, we can clearly see how each podcast’s attitudes toward different target groups have changed and, in some cases, intensified during the study period.
Characterizing Hostile Expression
We assess hostile expression through a set of hallmarks: recognizable hostile rhetorical moves, defined by how a group is characterized. We treat hostility as a property of language, not of truth. Each hallmark is identified from the language of the segment alone, without ever judging whether the underlying claims are true, justified, or fair. Coding rhetorical structure rather than factual accuracy keeps judgments consistent across thousands of contested political topics, and it means a hostile move does not become less hostile because it is delivered as a joke: a humorous, sarcastic, or self-deprecating tone does not negate hostility.
Not all hostile speech is the same. We code hallmarks such as dehumanization, portraying a group as a threat, or accusing it of the very harms one fears from it. Established frameworks treat these patterns as early, measurable indicators of escalation toward harm rather than ordinary incivility (Dangerous Speech Project 2021; OHCHR 2013). Our coding scheme was developed through a bottom-up process informed by the Dangerous Speech Project (Dangerous Speech Project 2021), commonly used indicators of genocide risk (Stanton 1996), the IHRA working definition of antisemitism (IHRA 2016), and the UN Rabat Plan of Action on incitement to hatred (OHCHR 2013).
We iteratively reviewed statements expressing negativity toward our target groups and tested which indicators best captured observed patterns. This process yielded a set of hallmarks that cover the vast majority of hostile expression in our corpus. The schema also aligns with the message features that research links to greater intergroup hostility: rhetoric that makes group identities more salient, casts an out-group as threatening, contemptible, or morally inferior, and reinforces existing ideological and partisan beliefs, activating social identity, perceptions of threat, motivated reasoning, and negative affect toward out-groups (Lord et al. 1979; Kunda 1990; Iyengar et al. 2012).
Colors indicate which hostility type was most prevalent during each snapshot.
Severity Assessment
Each hallmark is coded on a four-point severity scale (0–3), where higher levels correspond to more extreme forms of hostile rhetoric. For example, Threat Construction ranges from (1) vague unease about a group’s influence, to (2) claims that the group poses an ongoing danger, to (3) depictions of the group as an existential threat. A single segment may exhibit multiple hallmarks, each at different severity levels.
The complete severity scale definitions are available here.
Once coded, we combine a segment’s hallmark–severity codes into an overall severity score. Each hallmark at each severity level is assigned a weight that captures how strongly that pattern pushes the overall hostility score upward when it is present.
We then aggregate all segment-level severity scores for each podcast–group–month combination and take the median value. The median captures the general tenor of rhetoric in that period rather than isolated extreme outbursts. Throughout this report, severity is a property of how the language was coded under this rubric, not an adjudication of real-world harm.
Colors indicate the most severe level of hostility expressed toward any group during each monthly snapshot, with darker shades representing higher severity.
Qualitative Synthesis
All segments identified as containing hostility within a podcast–month–group snapshot then undergo a processing stage in which the hostile speech acts are isolated, stripped of surrounding commentary, and restated in a clear, standardized form. These extractions are subsequently merged to remove redundant or overlapping themes. The resulting summary provides a concise record of hostility for that period, making it easier to identify recurring themes and track patterns across the dataset. These grounded summaries also serve as an audit trail: a defensible coding should expose why a passage was flagged (Mathew et al. 2021), and every snapshot in this report carries a record of what was actually said.
Hovering over a snapshot displays its qualitative summary.
Detecting Hostile Expression at Scale
Analyzing thousands of podcast episodes requires an automated approach. We employ a two-phase system to detect hostile statements across our corpus: a first, keyword-based sweep that surfaces explicit hostility, and a second, meaning-based sweep that starts from the first sweep's confirmed detections to reach hostility expressed without any of the initial keywords. Both sweeps apply the same coding bar: a statement counts only if it clears every check that follows, whichever path surfaced it.
In the first phase, we surface explicit hostility: statements where speakers directly name a target group and express clear negativity toward them. Custom lexicons identify segments mentioning these groups, and a trained classifier then evaluates whether the content is genuinely hostile or simply a neutral reference.
The second phase expands our detection to capture implicit hostility: rhetoric that doesn't name groups directly but conveys hostility through coded terms, euphemisms, or indirect references. By building on the explicit statements found in phase one, we can identify similar patterns of language even when groups aren't explicitly mentioned.
Together, these phases produce the qualifying statements that characterize each podcast's stance toward different groups over time, enabling the longitudinal analysis described in our framework. The design sits within an established computational literature that distinguishes hate speech from merely offensive language and maps the design space for detecting it at scale (Davidson et al. 2017; Fortuna and Nunes 2018), and it follows established large-scale practice in asking who the hostility targets, not only whether it is present (Silva et al. 2016).
Identifying Group References Through Custom Lexicons
We identify moments when speakers explicitly refer to specific communities using custom lexicons developed for each target group. To create these lexicons, we began with neutral group names and incorporated additional pejorative and politically charged terms drawn from resources such as Hatebase and Weaponized Word. A study that sets out to measure hostility cannot omit the most hostile vocabulary, so such terms are included for every group, not a subset. These initial term lists were used to locate potential podcast segments, which we then reviewed to identify additional relevant vocabulary. The lexicons are built for parallel coverage, so that cross-group comparisons are not distorted by one group simply being easier to name than another.
This discovery process was necessary because most available resources were developed for written online text. The distinct patterns of spoken discourse, combined with artifacts introduced by speech-to-text conversion, revealed unique lexical patterns and variants not captured in text-based sources. Through iterative manual review and filtering, we refined each lexicon to reliably capture how groups are referenced in spoken discourse.
Throughout this process, we deliberately constructed lexicons that capture references to people, rather than to states, governments, militaries, or prominent individuals. Terms such as "Israel" or "China" were generally excluded, since they more commonly describe state actions.
It is important to note that the lexicon's only job is recall of explicit mentions: it deliberately over-includes, matching any plausible reference to a group, and leaves every judgment about hostility, endorsement, and people-versus-state referents to the classification stages that follow. A keyword scan alone cannot tell hostile speech from neutral or supportive speech, and it cannot see coded, euphemistic, or implicit references at all. The first limitation is handled by the screening described below, and the second is the explicit reason for the semantic expansion phase addressed later in this report. While our lexicons capture only overt language, their consistent application across the corpus provides a reliable basis for comparing explicit references.
Colors show which group was most discussed during that time period.
Screening for Hostile Content
Among all transcript segments that reference the target groups, the majority are non-hostile. These segments may include neutral descriptions, positive portrayals, or discussions of current events that do not employ hostile framing.
To identify segments that contain hostile expression, we used a text classification model trained on our hostility framework. The model evaluates each segment strictly on the basis of its lexical content, without reference to the podcast, speaker, or broader episode context. This approach reduces the risk of bias linked to particular hosts or shows, but it also means the model is optimized for detecting explicit textual hostility and may overlook more context-dependent forms of expression. At this first pass, the system is deliberately inclusive: when genuinely unsure, it errs toward flagging, and the stricter analyses that follow do the winnowing.
When a segment is identified as containing hostile expression, we revisit the transcript and widen the surrounding text window to include more of what came immediately before and after. This added context can disambiguate language and referents, clarify speaker intent, and reduce the risk of misclassification.
Rather than relying on a single classification, each potentially hostile segment undergoes a series of analyses that apply the hostility framework from complementary perspectives. These analyses assess whether hostility hallmarks are present, identify the exact portion of the transcript in which the expression occurs, flag ambiguities that may affect judgment, assign severity scores to the hallmarks identified, evaluate how hostile claims are attributed, assess whether those claims are endorsed by the speaker, and generate a summary of the hostile content grounded in specific excerpts. At any point in this process, the model may determine that no hostility is present. Only segments that register as hostile across all components of the workflow are carried forward.
The endorsement check matters more than it might appear. Functional testing of hate-speech classifiers shows that quoted hate, counter-speech, and reclaimed or mocking use are precisely where detection systems most often fail (Röttger et al. 2021), which is why speaker stance receives its own dedicated assessment rather than being folded into the hostility judgment: quotation and critique must not be counted as the speaker's own hostility.
Furthermore, hostile language meets our inclusion threshold only when it appears more than once, either at separate points within an episode or across multiple episodes. Hostile remarks confined to a single brief exchange do not qualify. This requirement ensures that our analysis captures repeated hostility rather than an isolated instance.
A summary of the inclusion criteria is available here.
These confirmed hostile statements serve two functions: they contribute to the evidence base for our longitudinal analysis, and they provide the foundation for discovering implicit forms of hostile expression in the next phase.
Colors indicate which group was most discussed in hostile terms during that period.
Detecting Implicit Group References
The lexicon-based approach reliably identifies explicit references to target groups, but hostile rhetoric often operates more subtly: a passage that names a group obliquely (by epithet, allusion, or circumlocution) never enters a keyword funnel at all. Speakers may use euphemisms, coded language, or rhetorical framing that conveys hostility without naming groups directly. A segment discussing “our neighbors to the north” may target Canadians without using any terms from our lexicons. Similarly, invoking “those from across the pond” may refer to British communities without naming them directly.
To surface these implicit cases, we use the confirmed hostile statements from phase one as templates for finding similar rhetoric elsewhere. We extract the core hostile expressions from these statements and use a large language model to generate linguistic variations: paraphrases, different phrasings, and alternative framings that convey the same hostile idea. For instance, the concept "Muslims don't belong in Western countries" might appear as "Islam is incompatible with Western civilization," "we need to keep Muslims out," or "Islamic values threaten our way of life." While these statements use different vocabulary, they express the same underlying hostile idea.
We then use semantic similarity to search the entire corpus for segments that resemble these hostile patterns, even when they don't mention groups explicitly. Segments with high similarity scores are retrieved and evaluated by the same classification model used in phase one to confirm they actually express hostility rather than just superficial resemblance. Implicit detections are reached by a different retrieval, not held to a different bar.
This approach has limitations. Most importantly, it recovers like-for-like patterns. Because retrieval is seeded with confirmed explicit examples, it improves recall for hostility that resembles what was already captured, but it is less effective at identifying qualitatively different forms of hostility. Despite these constraints, semantic similarity detection substantially expands coverage beyond keyword matching by capturing hostile rhetorical patterns that would otherwise remain undetected.
Colors indicate which group was most discussed in hostile terms through implicit language during that time period.
Validation
Using language models as text annotators is now established practice in computational social science (Alizadeh et al. 2024), though a practice with documented risks: a single change of model or prompt can shift annotation-based conclusions (Baumann et al. 2025). Published validations are therefore no substitute for testing one's own pipeline. To assess the reliability of our automated hostility detection system, we recruited three research assistants and trained them on the hostility coding rubric. Each analyst independently reviewed a stratified sample of 223 segments covering all time periods, target groups, and podcasts in the corpus. For each segment, coders applied the framework, indicating which hostility hallmarks were present and assigning each detected hallmark a severity score on a four-point scale.
The hallmark framework provides a structured basis for identifying and assessing hostility. Each hallmark captures a distinct rhetorical pattern, but in this analysis they are treated collectively rather than as separate outcomes. The goal is to determine whether a segment expresses hostility at all. For every human-coded segment, we assign a binary label of hostility present, marked as 1 if any hallmark is rated at level 1 or higher and 0 otherwise. The automated classifier’s outputs are processed in the same way, enabling direct comparison between human and model evaluations.
Human Coding Reliability
We began by examining how consistently human coders agreed with each other when applying the hostility rubric. Agreement was measured using Cohen’s κ, a statistic that accounts for the amount of agreement that could occur by chance. Across all segments that had been reviewed by more than one person, κ values ranged from 0.33 to 0.42, which corresponds to low to moderate consistency. Because only a small number of segments were coded by all three assistants, we rely on these pairwise comparisons as our main indicator of baseline human reliability.
After the initial coding was completed, cases where coders disagreed were reviewed by experts for closer examination. The goal was to identify the sources of these differences in judgment, such as unclear definitions, ambiguous phrasing in the transcripts, or inconsistent use of the rubric. Most disagreements stemmed from variation in how coders interpreted and applied the rubric rather than issues with the framework itself. When the relevant hallmark definitions were revisited to assess whether the criteria for hostility were met, consensus was typically reached.
Coding natural speech takes practice, especially when speakers use indirect phrasing, emotional appeals, or deliberate misdirection. A more structured training process that included guided exercises, feedback on coding decisions, examples of correct and incorrect applications of the rubric, and discussion of borderline cases would likely have strengthened agreement among coders without requiring any changes to the rubric itself.
Evaluation on Consensus-Labeled Segments
To evaluate how accurately the automated classifier detects hostility, we focused on segments where human coders independently reached the same conclusion. These consensus cases likely represent instances where the rubric was applied more consistently and the determination of presence or absence of hostility was relatively unambiguous. Although consensus does not guarantee correctness, it provides a reasonable, higher-confidence reference set for comparing the model’s classifications with human judgments.
To construct the reference set, we examined all segments that had been independently reviewed by at least two human coders. For each segment, we assigned a label based on the majority decision, meaning whichever judgment, hostile or non-hostile, was chosen by most coders. This process yielded 155 segments (out of 223) with clear agreement among coders, including 69 labeled as hostile and 86 as non-hostile.
Using the human consensus labels as a reference, the classifier correctly identified 46 hostile segments (true positives) and 84 non-hostile segments (true negatives), while producing 2 false positives and 23 false negatives. These results correspond to a precision of 95.8%, and a recall of 66.7%. The agreement between the classifier and the human consensus, measured using Cohen’s κ, was 0.663.
| Sample Size | Precision | Recall | |
|---|---|---|---|
| Classifier | 155 | 95.8% | 66.7% |
While the Cohen’s κ score provides a broad indication of agreement, it does not fully capture the model’s performance profile. The classifier’s precision is very high, meaning it rarely labels a segment as hostile when human coders do not. Most disagreements result from missed detections, reflecting a more cautious approach than that of the human coders. In practice, the system is highly reliable when it identifies hostility but fails to detect roughly one third of the segments that humans classify as hostile.
This conservative configuration is intentional. It reduces the likelihood of incorrectly labeling non-hostile content as hostile, which is essential for producing dependable hostility snapshots. At the same time, it means that our corpus-level analyses likely understate the overall prevalence of hostile expression in the podcast corpus.
When false positives do occur, they typically involve podcast hosts quoting, paraphrasing, or otherwise invoking hostile statements without expressing endorsement. In these instances, the surface language may meet the formal criteria of the hostility rubric even though the speaker’s intent is descriptive, critical, or contextual rather than expressive. Distinguishing between endorsement and quotation is challenging for automated systems, particularly given limited contextual signals and the conversational formats in which speakers frequently recount others’ views, react to prior statements, or simulate opposing arguments. Although infrequent overall, such cases account for a substantial share of the classifier’s false positives and reflect a fundamental limitation of text-based hostility detection. The hardest residual case is the speaker who voices hostile lines in an adversary’s mouth, ventriloquizing what they claim their enemies believe, where the surface language can read as endorsed even though the speaker frames it as the other side’s view. We flag this as a precision caveat, but also as a substantive point: even where a speaker does not personally endorse the hostility, repeatedly voicing and amplifying it on a large platform circulates the hostile frame, and a method tuned to count only first-person endorsement will understate that diffuse contribution rather than overstate it.
Ambiguous and Disagreement Segments
Human coders disagreed on approximately one third of the sample (68 out of 223 segments) when determining whether a passage contained hostility. It is informative to examine how the automated classifier handled these disputed cases. Among the 68 segments, the classifier labeled 50 as non-hostile and 18 as hostile, again reflecting a generally cautious orientation. In most instances, subsequent expert review concluded that the classifier’s judgments were consistent with a reasonable interpretation of the rubric.
These results further indicate that, while the human-generated labels provide a valuable benchmark, they should be regarded as a noisy reference rather than a definitive ground truth. The same humility applies in the other direction: we present our detection counts as indicative of a phenomenon, a defensible and uniformly applied measure of hostile expression in which individual borderline segments remain open to reasonable reinterpretation, not as a fully adjudicated ground truth.
Year-by-Year Performance
To test whether the classifier’s performance remained stable over time, we evaluated results separately by publication year using the consensus-labeled segments. For each year, we report precision (the share of model-identified hostile segments confirmed by human coders) and recall (the share of human-identified hostile segments detected by the model). These measures capture how reliably the classifier detects hostility and how cautious it is in assigning the label.
| Year | Sample Size | Precision | Recall |
|---|---|---|---|
| 2025 | 45 | 100.0% | 66.7% |
| 2024 | 44 | 92.9% | 72.2% |
| 2023 | 30 | 90.0% | 60.0% |
| 2022 | 24 | 100.0% | 66.7% |
| 2021 | 10 | 100.0% | 57.1% |
| 2020 | 2 | 100.0% | 100.0% |
Across all years, precision remains consistently high, typically above 90 percent, indicating that the classifier rarely labels non-hostile content as hostile. Recall fluctuates between roughly 57 and 72 percent, suggesting that while the system sometimes misses hostility identified by human coders, it does so without any systematic trend over time. Although earlier years contain relatively few samples, the overall pattern suggests stable calibration and no evidence of temporal drift that would bias longitudinal patterns.
Podcast-by-Podcast Performance
We also examined performance across podcasts using the same evaluation set. Because most podcasts contributed only a small number of consensus-labeled segments, these comparisons should be interpreted cautiously. The results below report precision and recall for podcasts with the largest sample sizes.
| Podcast | Sample Size | Precision | Recall |
|---|---|---|---|
| Mark Levin Podcast | 21 | 100.0% | 71.4% |
| The Young Turks | 18 | 88.9% | 100.0% |
| The Ben Shapiro Show | 11 | 100.0% | 80.0% |
| The Alex Jones Show – Infowars.com | 9 | 100.0% | 50.0% |
| Piers Morgan Uncensored | 8 | 100.0% | 100.0% |
| The Victor Davis Hanson Show | 7 | 100.0% | 100.0% |
| ... | ... | ... | ... |
Given these limited sample sizes, podcast-level differences are best viewed as descriptive checks for gross anomalies rather than as precise estimates of podcast-specific performance.
Target Group Performance
Finally, we evaluated classifier performance by the group targeted in each segment. This comparison tests whether detection accuracy differs across groups commonly referenced in hostile speech. Because the number of annotated excerpts per group is limited, these results are exploratory.
| Target Group | Sample Size | Precision | Recall |
|---|---|---|---|
| Palestinians | 20 | 95.0% | 100.0% |
| Israelis | 14 | 92.9% | 100.0% |
| Jews | 6 | 100.0% | 100.0% |
| ... | ... | ... | ... |
These consistent results provide no indication that classifier performance differs systematically across groups, although small sample sizes limit the strength of the inferences that can be drawn from these subgroup comparisons.
Database
The database below presents the cases of hostile expression that were surfaced by our detection pipeline across more than 70,000 hours of podcast content. Only segments that registered as hostile across every component of our deliberately conservative workflow are displayed, and each case can be examined down to the underlying audio and transcript. These results should be interpreted as a lower bound; the underlying source material contains substantially more hostile rhetoric than what appears in this visualization, and the absence of a detection should not be taken as evidence that no hostility occurred. What this database records is a measurement of language at scale, not an adjudication of any claim, any show, or any host, and not a measure of real-world harm.
Use the filters below to explore hostility by target group, and click any snapshot to view the underlying source material.
Detections in this figure are generated by an automated classifier with an approximate 5% false positive rate.
Results
Is hostility rising, and if so toward whom, how severely, and does it recede when the news cycle moves on? The figures below put those questions to the full evidence base. Earlier sections provided podcast-specific snapshots; here we pool all detected cases across every show to trace how attitudes shifted over time.
How often a group is targeted and how severe that rhetoric is are distinct phenomena, and the figure encodes both. Each point's vertical position reflects the typical severity of hostile remarks detected in that month, derived from the levels assigned to the underlying hostility hallmarks: higher points indicate that rhetoric in those periods was more extreme on average, lower points that the hostile language that did appear was generally milder. The size of each marker reflects how many hostile statements were detected: larger markers indicate a greater volume of hostility, smaller markers fewer instances. A group can be discussed in hostile terms often but mildly, rarely but severely, or both at once, which is the pattern that most warrants attention. Hostility toward different peoples is also voiced in different registers and surfaces at very different rates (Obermaier et al. 2023); the asymmetries visible below are a feature of the discourse itself, not of the method, which applies one bar to every group.
Breaks in the trend lines appear when too few cases are detected to calculate a reliable severity score. These gaps do not mean hostility did not occur, only that there was not enough qualifying evidence to assign a score in those months. Because all groups are evaluated using the same detection and scoring process, an unscored month indicates there were likely fewer hostile remarks toward those groups in that period under the framework used.
The following figure shifts from a pooled overview to a podcast level view. For a selected podcast, it shows which groups were targeted by hostile rhetoric and how the typical intensity changed over the study period. Because this view analyzes one podcast at a time, fewer hostile statements are available, and fewer cases meet the evidence threshold required for scoring, leading to more gaps in the trend lines. When a month is unscored, it does not mean hostility was absent, only that not enough qualifying statements were detected to generate a reliable severity score under the study framework.
Whereas the previous figures show month-by-month trends in the typical severity of hostile content for each group, the next figure aggregates all hostile remarks across the entire study period into a single severity score for each podcast. Podcasts are ordered along the horizontal axis from least to most severe, allowing direct comparison of how intensely different shows expressed hostility toward the selected group. A podcast’s absence from the figure should not be interpreted as evidence that hostility never occurred. Rather, it indicates there were not enough qualifying hostile remarks to generate a reliable score under the study’s criteria.
Case Study | War in Gaza
Around a real-world crisis, two things tend to rise at once: how much a group is talked about, and how much of that talk is hostile. That coupling is well established beyond podcasting: a body of natural-experiment work finds that a crisis sharply raises both how much a group is discussed and the hostility aimed at it, whether measured in survey attitudes after terror attacks (Legewie 2013), in hostility toward arriving refugees (Hangartner et al. 2019), or in online hate speech in the days after an attack (Czymara et al. 2023; Müller and Schwarz 2020). The war in Gaza offers the clearest such case in our window.
We draw on our curated database of hostility incidents to present a case study examining patterns of hostile rhetoric linked to the war in Gaza. We focus on the groups most directly affected by the conflict: Palestinians, Israelis, and their major religious affiliations. This case study analyzes overall trends in hostile rhetoric, key targets, and primary contributors, as well as salient inflection points in the data and their interaction with external developments.
Post-October 7 Surge
Hostile rhetoric targeting specific groups has risen sharply since October 2023, as of analysis conducted through July 2025. The data shows a clear break in continuity: before October 2023, hostile remarks toward Israeli, Palestinian, Jewish, or Muslim people appeared only sporadically. Following October 7, the frequency and severity increased substantially, closely tracking the conflict’s timeline and signaling heightened public tensions around Gaza-related issues. Attention and hostility rose together at the outbreak, the pattern the literature predicts, with hostile speech climbing alongside the surge in discussion rather than arriving as a separate, later wave.
Focusing on Muslim and Jewish communities, the pattern of hostile rhetoric shows a clear shift after January 2025 toward a sustained, linear increase, rather than episodic spikes. Earlier periods in the data showed sharp increases followed by cooling phases, reflecting a reactive, event-driven pattern tied to specific conflict incidents. After January 2025, the trajectory becomes steadier, with extended stretches of escalating severity. This change indicates a rising baseline of hostility, suggesting that rhetoric is becoming more persistent, rather than fluctuating in response to individual events. The two signals move together on the way up but come apart on the way down: news attention recedes as the cycle moves on, while hostility lags behind it and, increasingly, does not return to its pre-crisis level at all.
We describe this only as the observed shape of the two series, but it sits at the junction of known dynamics. The fast decay of attention is the classic issue-attention cycle (Petersen 2009), accelerating as media multiply (Lorenz-Spreen et al. 2019). Whether hostility decays as fast is genuinely contested: some work finds post-event attitude shifts fade within days (Mancosu and Ferrín Pereira 2021), while other work finds the hostility durable well after the triggering salience has passed (Hangartner et al. 2019; Hobbs et al. 2021). The mechanisms the literature offers, none of them testable here, include the idea that crises can erode the norms that restrain hostile speech and embolden those already inclined to it (Alvarez-Benjumea and Winter 2020; Newman et al. 2021), with repeated exposure itself associated with desensitization over time (Soral et al. 2018).
Who the hostility names is itself revealing. In North American podcast vernacular, hostile commentary directed at the side of the conflict associated with Israel rarely targets Israelis as a national group. Instead, such remarks far more often invoke Jewish identifiers, reflecting a tendency in public discourse to blur Jewish and Israeli categories, collapsing a people into a state and a state into a people. When the focus shifts to geopolitical developments, antagonistic comments typically reference the state of Israel or the Israel Defense Forces rather than Israelis themselves, and our people-only construct sets such state-directed speech aside. That hostility surrounding an asymmetric conflict runs in both directions is consistent with experimental work showing that opposing sides blatantly dehumanize one another (Bruneau and Kteily 2017), and that being dehumanized in turn fuels reciprocal hostility (Kteily et al. 2016; Landry et al. 2022). This is a spiral rather than a one-way attack.
Conclusion
Across five years of popular North American political and news podcasting, hostile rhetoric toward the groups we track has grown more frequent, more severe, and more persistent. The rise is uneven in where it lands and how it is voiced: it moves in step with real-world events yet increasingly outlasts them, and since January 2025 it has settled into a sustained, escalating baseline rather than the reactive, event-driven spikes of earlier years.
What we report is a measurement of language at scale, not an adjudication of any claim, any show, or any host, and not a measure of real-world harm. The pipeline behind it is deliberately conservative at every stage and validated against trained human coders; the patterns it surfaces are a lower bound on the hostile rhetoric the medium carries.
This matters because podcasting combines the conditions that research ties to durable intergroup hostility (repeated exposure, trusted hosts, long-form communication) at a reach that is now mainstream, while sitting largely outside the editorial standards that govern broadcasters and outside the platform-accountability frameworks built for social media. It is the host-produced audio itself that no editorial code or transparency obligation currently reaches. Whether, and how, that gap should be addressed is a question for legislators, platforms, and civil society, not one this report answers.
What this report contributes is the measurement itself: a transparent framework applied uniformly across seven groups and five years of audio, with its evidence browsable down to the individual clip. The tools to measure hostile rhetoric at this scale now exist, and the measurement can be re-run, audited, and extended as the ecosystem changes. How hosts, platforms, and audiences respond to what podcasts now carry will shape whether a medium this trusted and this intimate remains outside the frameworks built to govern speech at scale.
References
Alizadeh, Meysam, Maël Kubli, Zeynab Samei, et al. 2024. “Open-Source LLMs for Text Annotation: A Practical Guide for Model Setting and Fine-Tuning.” Journal of Computational Social Science 8 (1): 17. https://doi.org/10.1007/s42001-024-00345-9.
Alvarez-Benjumea, Amalia, and Fabian Winter. 2020. The Breakdown of Anti-Racist Norms: A Natural Experiment on Normative Uncertainty After Terrorist Attacks. SSRN Scholarly Paper No. 3537597. Social Science Research Network. https://doi.org/10.2139/ssrn.3537597.
Baumann, Joachim, Paul Röttger, Aleksandra Urman, et al. 2025. Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation. arXiv:2509.08825. https://doi.org/10.48550/arXiv.2509.08825.
Bruneau, Emile, and Nour Kteily. 2017. “The Enemy as Animal: Symmetric Dehumanization During Asymmetric Warfare.” PLOS ONE 12 (7): e0181422. https://doi.org/10.1371/journal.pone.0181422.
Czymara, Christian S., Stephan Dochow-Sondershaus, Lucas G. Drouhot, Müge Simsek, and Christoph Spörlein. 2023. “Catalyst of Hate? Ethnic Insulting on YouTube in the Aftermath of Terror Attacks in France, Germany and the United Kingdom 2014–2017.” Journal of Ethnic and Migration Studies 49 (2): 535–53. https://doi.org/10.1080/1369183X.2022.2100552.
Dangerous Speech Project. 2021. Dangerous Speech: A Practical Guide. https://www.dangerousspeech.org/libraries/guide.
Davidson, Thomas, Dana Warmsley, Michael Macy, and Ingmar Weber. 2017. “Automated Hate Speech Detection and the Problem of Offensive Language.” Proceedings of the International AAAI Conference on Web and Social Media 11 (1): 512–15. https://doi.org/10.1609/icwsm.v11i1.14955.
ElSherief, Mai, Caleb Ziems, David Muchlinski, et al. 2021. “Latent Hatred: A Benchmark for Understanding Implicit Hate Speech.” In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.emnlp-main.29.
Fortuna, Paula, and Sérgio Nunes. 2018. “A Survey on Automatic Detection of Hate Speech in Text.” ACM Computing Surveys 51 (4): 85:1–30. https://doi.org/10.1145/3232676.
Hangartner, Dominik, Elias Dinas, Moritz Marbach, Konstantinos Matakos, and Dimitrios Xefteris. 2019. “Does Exposure to the Refugee Crisis Make Natives More Hostile?” American Political Science Review 113 (2): 442–55. https://doi.org/10.1017/S0003055418000813.
Haslam, Nick. 2006. “Dehumanization: An Integrative Review.” Personality and Social Psychology Review 10 (3): 252–64. https://doi.org/10.1207/s15327957pspr1003_4.
Hobbs, William, Nazita Lajevardi, Xinyi Li, and Caleb Lucas. 2021. Group Salience, Inflammatory Rhetoric, and the Persistence of Hate Against Religious Minorities. APSA Preprints. https://doi.org/10.33774/apsa-2021-wnwqw.
Hovland, Carl I., and Walter Weiss. 1951. “The Influence of Source Credibility on Communication Effectiveness.” Public Opinion Quarterly 15 (4): 635–50. https://doi.org/10.1086/266350.
International Holocaust Remembrance Alliance. 2016. “What Is Antisemitism? IHRA Working Definition.” https://holocaustremembrance.com/resources/working-definition-antisemitism.
Iyengar, Shanto, Gaurav Sood, and Yphtach Lelkes. 2012. “Affect, Not Ideology: A Social Identity Perspective on Polarization.” Public Opinion Quarterly 76 (3): 405–31. https://doi.org/10.1093/poq/nfs038.
Iyengar, Shanto, Yphtach Lelkes, Matthew Levendusky, Neil Malhotra, and Sean J. Westwood. 2019. “The Origins and Consequences of Affective Polarization in the United States.” Annual Review of Political Science 22: 129–46. https://doi.org/10.1146/annurev-polisci-051117-073034.
Kteily, Nour, and Emile Bruneau. 2017. “Backlash: The Politics and Real-World Consequences of Minority Group Dehumanization.” Personality and Social Psychology Bulletin 43 (1): 87–104. https://doi.org/10.1177/0146167216675334.
Kteily, Nour, Emile Bruneau, Adam Waytz, and Sarah Cotterill. 2015. “The Ascent of Man: Theoretical and Empirical Evidence for Blatant Dehumanization.” Journal of Personality and Social Psychology 109 (5): 901–31. https://doi.org/10.1037/pspp0000048.
Kteily, Nour, Gordon Hodson, and Emile Bruneau. 2016. “They See Us as Less Than Human: Metadehumanization Predicts Intergroup Conflict via Reciprocal Dehumanization.” Journal of Personality and Social Psychology 110 (3): 343–70. https://doi.org/10.1037/pspa0000044.
Kunda, Ziva. 1990. “The Case for Motivated Reasoning.” Psychological Bulletin 108 (3): 480–98. https://doi.org/10.1037/0033-2909.108.3.480.
Landry, Alexander P., Elliott Ihm, and Jonathan W. Schooler. 2022. “Hated but Still Human: Metadehumanization Leads to Greater Hostility Than Metaprejudice.” Group Processes & Intergroup Relations 25 (2): 315–34. https://doi.org/10.1177/1368430220979035.
Legewie, Joscha. 2013. “Terrorist Events and Attitudes Toward Immigrants: A Natural Experiment.” American Journal of Sociology 118 (5): 1199–245. https://doi.org/10.1086/669605.
Litterer, Benjamin, David Jurgens, and Dallas Card. 2025. “Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus.” Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 25132–54. https://doi.org/10.18653/v1/2025.acl-long.1222.
Lord, Charles G., Lee Ross, and Mark R. Lepper. 1979. “Biased Assimilation and Attitude Polarization: The Effects of Prior Theories on Subsequently Considered Evidence.” Journal of Personality and Social Psychology 37 (11): 2098–109. https://doi.org/10.1037/0022-3514.37.11.2098.
Lorenz-Spreen, Philipp, Bjarke Mørch Mønsted, Philipp Hövel, and Sune Lehmann. 2019. “Accelerating Dynamics of Collective Attention.” Nature Communications 10 (1): 1759. https://doi.org/10.1038/s41467-019-09311-w.
Madriaza, Pablo, Ghayda Hassan, Sébastien Brouillette-Alarie, et al. 2025. “Exposure to Hate in Online and Traditional Media: A Systematic Review and Meta-Analysis of the Impact of This Exposure on Individuals and Communities.” Campbell Systematic Reviews 21 (1): cl2.70018. https://doi.org/10.1002/cl2.70018.
Mancosu, Moreno, and Mònica Ferrín Pereira. 2021. “Terrorist Attacks, Stereotyping, and Attitudes Toward Immigrants: The Case of the Manchester Bombing.” Social Science Quarterly 102 (1): 420–32. https://doi.org/10.1111/ssqu.12907.
Mathew, Binny, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021. “HateXplain: A Benchmark Dataset for Explainable Hate Speech Detection.” Proceedings of the AAAI Conference on Artificial Intelligence 35 (17): 14867–75. https://doi.org/10.1609/aaai.v35i17.17745.
Moore-Berg, Samantha L., Lee-Or Ankori-Karlinsky, Boaz Hameiri, and Emile Bruneau. 2020. “Exaggerated Meta-Perceptions Predict Intergroup Hostility Between American Political Partisans.” Proceedings of the National Academy of Sciences 117 (26): 14864–72. https://doi.org/10.1073/pnas.2001263117.
Müller, Karsten, and Carlo Schwarz. 2020. Fanning the Flames of Hate: Social Media and Hate Crime. SSRN Scholarly Paper No. 3082972. Social Science Research Network. https://doi.org/10.2139/ssrn.3082972.
Newman, Benjamin, Jennifer L. Merolla, Sono Shah, Danielle Casarez Lemi, Loren Collingwood, and S. Karthick Ramakrishnan. 2021. “The Trump Effect: An Experimental Investigation of the Emboldening Effect of Racially Inflammatory Elite Communication.” British Journal of Political Science 51 (3): 1138–59. https://doi.org/10.1017/S0007123419000590.
Newman, Nic, Richard Fletcher, Craig T. Robertson, Amy Ross Arguedas, and Rasmus Kleis Nielsen. 2025. Reuters Institute Digital News Report 2025. Reuters Institute for the Study of Journalism, University of Oxford. https://reutersinstitute.politics.ox.ac.uk/digital-news-report/2025.
Obermaier, Magdalena, Ursula Kristin Schmid, and Diana Rieger. 2023. “Too Civil to Care? How Online Hate Speech Against Different Social Groups Affects Bystander Intervention.” European Journal of Criminology 20 (3): 817–33. https://doi.org/10.1177/14773708231156328.
OHCHR. 2013. “Rabat Plan of Action on the Prohibition of Advocacy of National, Racial or Religious Hatred That Constitutes Incitement to Discrimination, Hostility or Violence.” https://www.ohchr.org/en/documents/outcome-documents/rabat-plan-action.
Petersen, Karen. 2009. “Revisiting Downs’ Issue-Attention Cycle: International Terrorism and U.S. Public Opinion.” Journal of Strategic Security 2 (4). https://doi.org/10.5038/1944-0472.2.4.1.
Petty, Richard E., and John T. Cacioppo. 1986. “The Elaboration Likelihood Model of Persuasion.” In Advances in Experimental Social Psychology, vol. 19. Academic Press. https://doi.org/10.1016/S0065-2601(08)60214-2.
Pluta, Agnieszka, Joanna Mazurek, Jakub Wojciechowski, Tomasz Wolak, Wiktor Soral, and Michał Bilewicz. 2023. “Exposure to Hate Speech Deteriorates Neurocognitive Mechanisms of the Ability to Understand Others’ Pain.” Scientific Reports 13 (1): 4127. https://doi.org/10.1038/s41598-023-31146-1.
Rizwan, Naquee, Nayandeep Deb, Sarthak Roy, Vishwajeet Singh Solanki, Kiran Garimella, and Animesh Mukherjee. 2025. “Toxicity Begets Toxicity: Unraveling Conversational Chains in Political Podcasts.” Proceedings of the 33rd ACM International Conference on Multimedia, 11776–84. https://doi.org/10.1145/3746027.3754553.
Röttger, Paul, Bertie Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet Pierrehumbert. 2021. “HateCheck: Functional Tests for Hate Speech Detection Models.” In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.acl-long.4.
Shearer, Elisa, Emily Tomasik, and Mary Randolph. 2025. “Podcasts and News Fact Sheet.” Pew Research Center. https://www.pewresearch.org/journalism/fact-sheet/podcasts-and-news-fact-sheet/.
Siegel, Alexandra A. 2020. “Online Hate Speech.” In Social Media and Democracy, edited by Joshua A. Tucker and Nathaniel Persily. Cambridge University Press. Cambridge Core.
Silva, Leandro, Mainack Mondal, Denzil Correa, Fabrício Benevenuto, and Ingmar Weber. 2016. “Analyzing the Targets of Hate in Online Social Media.” Proceedings of the International AAAI Conference on Web and Social Media 10 (1): 687–90. https://doi.org/10.1609/icwsm.v10i1.14811.
Soral, Wiktor, Michał Bilewicz, and Mikołaj Winiewski. 2018. “Exposure to Hate Speech Increases Prejudice Through Desensitization.” Aggressive Behavior 44 (2): 136–46. https://doi.org/10.1002/ab.21737.
Stanton, Gregory H. 1996. “Ten Stages of Genocide.” Genocide Watch. https://www.genocidewatch.com/tenstages.
Tukachinsky, Riva, Nathan Walter, and Camille J. Saucier. 2020. “Antecedents and Effects of Parasocial Relationships: A Meta-Analysis.” Journal of Communication 70 (6): 868–94. https://doi.org/10.1093/joc/jqaa034.
Wirtschafter, Valerie. 2021. The Challenge of Detecting Misinformation in Podcasting. Brookings Institution. https://www.brookings.edu/articles/the-challenge-of-detecting-misinformation-in-podcasting/.
Wirtschafter, Valerie. 2023. Audible Reckoning: How Top Political Podcasters Spread Unsubstantiated and False Claims. Brookings Institution. https://www.brookings.edu/articles/audible-reckoning-how-top-political-podcasters-spread-unsubstantiated-and-false-claims/.
Zajonc, Robert B. 1968. “Attitudinal Effects of Mere Exposure.” Journal of Personality and Social Psychology 9 (2): 1–27. https://doi.org/10.1037/h0025848.