Decoding Hate
Three Shifts in the Study of Hate, Discourse, and Democracy
AddressHate Research Scholar at NYU’s Center for the Study of Antisemitism | Lead, Decoding Antisemitism | PI, Decoding Hate | Research Advisor, AddressHate | Editor-in-Chief, Digital Hate Review
In a nutshell: Some of the most difficult online hate to identify is coded rather than explicit — legible to its audiences, largely invisible to automated filters — so research and moderation built on detection capture only the surface of the phenomenon. And the field’s most widely used public instrument, Google Jigsaw’s Perspective API, retires at the end of 2026, leaving even that surface-level measurement without its baseline. Decoding Hate studies hate instead as a system of ideological communication — how hateful concepts occur, spread, and mutate — comparatively across antisemitism, anti-Black racism, anti-Asian racism, and misogyny, with antisemitism as the anchor and its conspiracy forms examined for what they may reveal about broader democratic erosion. The program is an instrument of understanding, not of enforcement: it moderates nothing, regulates nothing, and maintains no records on individuals. It combines fifteen years of expert interpretation with the scale of current AI to build standing observation of the digital public sphere — an observatory rather than a one-off study — so that the normalization of hate becomes visible while it is still underway. The findings feed into concrete instruments for education, journalism, law, policy, security, and the platforms themselves, beginning with the Decoding Hate Glossary, now being finalized. The essay sets out the three shifts on which all of this rests.
Much of contemporary digital hate does not announce itself. In the comment sections of mainstream news channels, under the posts of entertainers and influencers, and in reply threads and group chats, it arrives in forms built to pass unnoticed: a set of triple parentheses around a name; a “just asking questions” post about who owns the media; four innocuous-looking words that activate an entire conspiracy tradition. There is often no slur anywhere, no threat, nothing a conventional automated filter would reliably flag, and an uninitiated reader passes over the material without registering it. Audiences familiar with the codes read them fluently. A substantial part of online hate — and institutionally the most difficult part — circulates like this: legible to its audiences, largely invisible to the systems built to detect it.
This essay describes Decoding Hate, the research program I lead at NYU’s Center for the Study of Antisemitism in partnership with AddressHate, and the three shifts it rests on. The first concerns the object of study: alongside the individual hate traditions, the recurring mechanisms through which hostile narratives circulate, take hold, and become normalized — and through which, I will argue, aspects of a democracy’s condition become readable. The second concerns observational capacity: recent advances in AI make it increasingly feasible to study those mechanisms at a scale, and with a degree of contextual sensitivity, that were out of reach even five years ago. The third concerns where expertise resides: the interpretive capability this research produces has to move out of specialist communities and into the institutions that maintain democratic life.
The premise beneath all three can be stated in one sentence: digital hate should no longer be studied primarily as a collection of expressions to be detected, but as a dynamic system of ideological communication whose mechanisms become observable only through historically grounded, comparative, and longitudinal discourse analysis.
One objection — that any program which classifies speech as hate is, whatever its intentions, building an instrument of censorship — is serious enough to receive its own answer later in this essay.
There is a personal arc behind the program: fifteen years on one tradition, antisemitism, mostly in European discourse — the last five of them leading a consortium that read hundreds of thousands of comments at close range. The starting point was an observation of something that should not have been possible where it occurred. In Germany — a country with decades of post-war denazification and reeducation behind it — antisemitism was visibly being normalized again in mainstream culture long before the summer of 2014 and long before October 7, 2023, and as a researcher of racism and nationalism I found myself watching that shift in real time. It redirected my work toward Jew-hatred as one of the most adaptable and persevering hate ideologies in existence — an ideology whose capacity to survive even the most deliberate societal countermeasures is precisely what makes it worth reading closely. Decoding Hate extends that work from Europe to the United States, a political debate culture that is vast, consequential beyond its own borders, and at this moment unusually exposed — and it is the wager that an interpretive standard built on one of the most historically layered cases, together with the research architecture built around it, is what this wider discourse environment requires.
Why it matters can also be stated briefly. Healthy democracies depend on being able to distinguish sharp political disagreement, which is legitimate, from the normalization of hate, which is not; Decoding Hate aims to provide the evidence needed to make that distinction carefully. The aim is to understand how hateful concepts occur, spread, and mutate through public discourse — because a society cannot respond to what it cannot see. What follows from that understanding is a democratic decision, not a research output.
Table of Contents
Shift 1: From hate expressions to discursive systems
The first shift changes the object of study: from hate as a collection of expressions to be detected to hate as a discursive system whose mechanisms recur across traditions and whose normalization can erode the conditions of democratic life.
What detection cannot deliver
What is loosely called digital hate studies is not yet an integrated field but an assemblage: hate-speech detection in computer science, ideology-specific scholarship, platform and political-communication research, online-harm studies, extremism research, civil-society monitoring — communities studying overlapping phenomena with different definitions, different units of analysis, and different evidentiary standards. The landscape did not choose this shape; it grew into it, around platform interfaces and their closures, moderation demands, benchmark competitions, funding cycles, and whatever data happened to be collectable — which is why the critique here targets structures, not scholars. Across the assemblage one bifurcation recurs: interpretive depth without scale on one side, scale without interpretive depth on the other, and little between them that cumulates. By interpretive depth I mean the reading of an utterance against its co-text, its conversational context, its community’s codes, and its historical repertoire, rather than against its surface features alone.
Within the computational and operational branches, the dominant paradigm remains detection: keyword lists, toxicity scores, and classifier benchmarks that perform respectably on explicit material and remain least reliable on the coded, context-dependent, and strategically ambiguous forms that carry much of the content most consequential for normalization. This is not a verdict on the seriousness of the work — the computational measurement school has produced systematic, large-scale, in places genuinely longitudinal studies. The constraint sits in the operationalization, since hate measured through slur lexicons and toxicity scores is hate reduced to its explicit surface; in the unit of analysis, since the isolated post remains the dominant computational unit even though meaning routinely depends on prior turns, quoted material, imagery, and what the audience knows; and in the ground truth, since many widely used benchmark datasets were annotated by crowdworkers asked to judge interpretively complex material without the context, training, or domain knowledge such judgments require.
The failure is two-sided and documented: in published studies, roughly three-quarters of antisemitic comments scored below the toxicity thresholds of Google Jigsaw’s Perspective API, the field’s most widely used public toxicity-scoring baseline — while the same paradigm overreaches in the opposite direction, inflating scores for comments containing identity-related terms regardless of stance, so that legitimate speech by and about targeted groups is flagged while coded hostility passes. And the instrument at the center of that paradigm is going away: Perspective retires at the end of 2026, removing the field’s most widely used public baseline while the volume and velocity of online discourse continue to grow.
Beneath these operational problems sits a more basic one — the question of what, exactly, is being measured. A high benchmark score shows that a model reproduced a particular labeling scheme, not that it measured antisemitism, racism, or misogyny as social phenomena — and the labels themselves often collapse hate, prejudice, toxicity, incivility, and extremism into one undifferentiated object. Algorithmic improvements alone cannot resolve this conceptual deficit; the next advance has to be epistemological as well as computational, and it begins by settling what the object of measurement actually is. What the assemblage needs is not a better classifier but a different object of study.
A new object of study
For most of its modern history, the study of antisemitism has been organized as a specialized field, with good reasons: antisemitism is among the most historically layered and extensively documented traditions of group hatred in the Western world, and material of that depth requires specialists. But fifteen years of empirical work on its online forms made visible something that specialization can obscure. Certain mechanisms recur across otherwise distinct hate traditions — coded vocabularies that preserve deniability while delivering their message to competent audiences; conspiracy structures that route explanation through hidden malevolent agents; dehumanizing framings that move a group from opponent to enemy; the gradual normalization by which transgressive speech becomes familiar and familiar speech becomes expected — and a comparative design reveals their recurrence and variation in ways single-domain studies cannot. The histories remain distinct, and the differences in prevalence, institutional support, and lethality are real; recognizing recurrent mechanisms is not a claim of symmetry.
The wider analytical object is the discursive system in which the traditions operate: not a closed totality, but the connected processes through which hostile narratives emerge, spread, stabilize, and acquire social force — meaning, what an utterance does in its context; diffusion, how it travels between communities and platforms; normalization, how yesterday’s transgression becomes today’s opinion. Studying the system does not replace studying the traditions; the deep historical knowledge remains indispensable. The systems level supplies the analytical vantage at which findings become diagnostics of the environment a society shares, rather than reports on the troubles of one community. The interpretation of discourse has a long disciplinary history — critical discourse analysis, pragmatics, argumentation theory, and conversation analysis have read meaning in context for decades — and this program builds on their work; what remains rare — and, to our knowledge, has not existed in this integrated form — is standing, comparative, computationally extended infrastructure organized around that mode of analysis.
Discourse and democratic life
Democratic life stands on that discourse environment. Democracy depends not only on institutions but on the circumstances that allow opponents to remain members of a shared political community — conditions under which disagreement takes the form of argument rather than enemy-identification, in Mouffe’s terms, the difference between the adversary and the enemy. Hate speech erodes those conditions in identifiable ways: sustained harassment drives targeted communities out of public participation; demonizing narratives move groups from the field of contention into the register of existential threat; and the conspiracy grammar carried in much of this discourse explains visible institutions through a hidden coordinated agent, under which elections, courts, and the press become increasingly liable to be read as products of manipulation. Normalization gives each of these operations its reach: narratives that circulate unchallenged move, step by step, from the sayable to the ordinary, and what a society treats as ordinary sets the limits of what it will accept. This is not a claim that public discourse can ever be politically neutral; it is the narrower claim that demonization, conspiratorial explanation, and exclusion diminish the possibility of disagreement for everyone. Nor does the erosion reliably stay online: several of the deadliest hate-motivated attacks of the past decade in Western democracies involved perpetrators whose radicalization, ideological formation, or manifesto production was substantially documented online — though the pathways from exposure to violence are individual and varied. Which is also why the capacity to read these registers early is, among other things, a security capability — a point Shift 3 returns to.
Why discourse, rather than institutions, parties, or economic conditions? Because it is the layer in which the capacities democratic life cannot do without — mutual understanding across difference, empathy, perspective-taking — are formed and eroded. Political language is formative infrastructure: it shapes which people count as legitimate participants in the political community, and which disagreements register as disagreements rather than as threats. A democracy’s condition therefore leaves discursive traces, and those traces can be read — and the survey literature supports the premise from the other direction: two research traditions, working independently with different instruments, find that hostility toward groups and antidemocratic orientation travel together. The German prejudice-syndrome series measure antisemitism, racism, and sexism as a correlated syndrome linked to authoritarian disposition (Heitmeyer’s Group-Focused Enmity program; the FES Mitte-Studien; the Leipzig Authoritarianism Studies), and in the United States, Bartels (2020) identifies ethnic antagonism as the strongest predictor of antidemocratic attitudes in his data. Discursive traces function as indicators rather than comprehensive measures, standing alongside institutional, behavioral, and economic evidence.
And why hate, rather than discourse in general? Because hate offers an especially revealing point of entry: the operations that dismantle the conditions of shared political life often become especially visible, and appear in concentrated form, in discourse that targets groups — the opponent recast not as wrong but as illegitimate, contaminating, or dangerous. Reading it is not a narrowing of the democratic question but a point of entry at which that question becomes tractable.
Decoding, not detection
Why not simply call this digital hate studies? Because the name would misdescribe the method. The commitment marked by the word decoding is hate as meaning to be interpreted, in context, against the historical repertoires and community codes that give an utterance its force. The phenomenon has a scholarly name that predates the internet by decades: communication latency — the public indirectness of socially sanctioned prejudice — described by Bergmann and Erb in the 1980s as the displacement of prejudice into indirect, publicly deniable forms that remain intelligible to competent audiences. The digital environment did not invent this phenomenon, but it has industrialized it, making the codes rapidly mutable, networked, multimodal, and audience-specific.
Those four innocuous words from the opening, for instance: “Dan is not suicidal,” left again and again under posts by the influencer Dan Bilzerian, who has spent years attributing world events to Israel and to hidden Jewish power. No slur, no mention of Jews — and yet, within the discourse environments in which the formula recurs, its surrounding co-text supports a reading in which a future death would be understood as an assassination disguised as suicide, the fate reserved in this worldview for supposed truth-tellers who cross the imagined hidden power. The formula echoes the conspiracy repertoire surrounding Jeffrey Epstein. Little on the surface from which a detection system could reliably infer the meaning; nearly everything in the shared knowledge of the audience.
The analytical question throughout is what a statement means in its context — what concept it expresses, what it does among its readers — not what its speaker privately intends, which for anonymous online discourse is usually unknowable anyway. The codes change constantly while the structure underneath stays comparatively stable, so static lists decay quickly and interpretive frameworks can be revised as the codes change. And through everything that follows runs the practical discipline at the center of the work: distinguishing, carefully and case by case, sharp political disagreement, which is legitimate, from the normalization of hate, which is not. That concept documentation is being turned into public-facing form: the Decoding Hate Glossary, the program’s first instrument, now in the process of being finalized — Shift 3 returns to it.
Four traditions and their points of contact
The interpretive standard comes from the Decoding Antisemitism project, launched in 2020 and carried out by a consortium of twenty-five to thirty scholars in linguistics, history, sociology, and computational science across TU Berlin, King’s College London, and HTW Berlin: more than 300,000 expert-annotated comments across English, German, and French web discourse, condensed into the Decoding Antisemitism Lexicon and its 40 documented concepts — published in October 2024 and downloaded nearly half a million times since. The findings reached institutional practice: the project documented empirically how “Zionist” functions as a coded proxy for “Jew” in real discourse — “Zionists run the media,” “Zionists control the banks.” The evidence was shared in sustained working meetings with Meta’s policy team; its July 2024 policy update subsequently brought such proxy uses under the policy’s most severe enforcement tier across Facebook and Instagram. The sequence — specialist interpretation producing documented evidence, documented evidence informing an institutional decision — is the one Shift 3 generalizes.
Decoding Hate transfers the research architecture rather than the antisemitism taxonomy itself: expert annotation, documented reliability, conservative attribution, concept-level rather than keyword-level description. The substantive taxonomies for anti-Black racism, anti-Asian racism, and misogyny are built from their own intellectual traditions, with scholars grounded in those histories, and the program reads the forms together: not as a catalogue of separate prejudices, but as traditions that can converge within broader forms of antidemocratic discourse. They also span markedly different structural forms of group hatred — from hatreds that construct their targets as inferior to those that construct them as dangerously powerful, the latter marked by fear, conspiracy, and self-victimization, with antisemitism drawing on both poles — which is what makes the comparison informative: patterns that recur across such different structures are unlikely to be artifacts of any single tradition. For each domain the program is building a concept-level taxonomy on the model of the Lexicon — a scholarly reference, a coding manual, and an operational annotation guide — and the tracks are deliberately staged: one tested standard transferred domain by domain instead of four standards improvised at once. Fifteen years of work on antisemitism mean that the documentary architecture, the reliability protocols, the calibration procedures, and the computational pipeline do not have to be reinvented for every domain; each new track still requires its own specialists, literature, taxonomy, dataset, and validation, and each is built as a complete product. The antisemitism head start functions as methodological efficiency, not substantive hierarchy.
The anti-Black racism taxonomy, currently the most advanced, documents more than forty concepts across five clusters, from the stereotype repertoire through historical revisionism to discursive reversals such as colorblindness, white-victimhood claims, and statistical weaponization — the selective deployment of decontextualized statistics to naturalize group characteristics. Its architecture reflects a finding carried over from the antisemitism work and now being tested comparatively: that in mainstream spaces much of the analytically demanding material lies not in explicit abuse but in deflection, reversal, and pseudo-objective framing. First case studies are in development, including an annotated corpus around a white-nationalist march in Washington on July 4, 2026.
The anti-Asian racism taxonomy follows the same architecture and draws on a tradition whose depth is routinely underestimated, from the perpetual-foreigner frame, which withholds full membership regardless of citizenship or generation, and the disease-and-contagion attribution reactivated during the COVID-19 pandemic, through yellow-peril and techno-economic threat narratives, to the model-minority frame — complimentary on its surface, often functioning to minimize or deny discrimination — and the fetishizing framings directed above all at Asian women, where the tradition crosses into misogyny.
The misogyny track addresses a discourse environment with properties the other domains show less strongly: a coded vocabulary that mutates at high speed across the communities loosely grouped as the manosphere, organized networked harassment whose documented effect is to drive women disproportionately out of public participation online, and backlash narratives carried far beyond their coining communities by irony and humor. It begins from a documented point of contact with the anchor domain: our study of the discourse around the influencer known as Clavicular found misogyny and antisemitism fused within one community’s discourse, each hierarchy of contempt reinforcing the other.
Antisemitism remains the anchor. It is the tradition we know most deeply, and it is diagnostically important beyond itself: its conspiracy structures attach readily to wider narratives of institutional illegitimacy and democratic betrayal, so antisemitic conspiracy discourse tends — in the historical scholarship as in the corpora we have analyzed (on the conspiracy link specifically: Imhoff and Bruder 2014; Kofta, Soral, and Bilewicz 2020) — to accompany a broader degradation of a society’s capacity for argument. For readers outside the field, the point can be put simply: in its conspiracy form, antisemitism is never only antisemitism. Where the hidden-orchestrator explanation of the world gains ground, it often coexists with declining institutional trust and a growing receptivity to authoritarian answers — which is why reading this tradition closely may yield early discursive indicators bearing on a democracy’s condition. Whether such indicators have genuine leading-indicator properties is a proposition only a maintained series can test, and one the observatory is built to answer.
The four traditions meet, and the meeting points are among the most consequential objects the comparative design makes visible. The theoretical vocabulary for these junctions is well developed — Crenshaw’s intersectionality, Collins’s matrix of domination, on which the anti-Black racism taxonomy already draws — and the comparative design can be read as operationalizing that intersectional tradition at the level of communicative mechanism, where it becomes empirically testable. The replacement narrative assigns the roles of a single conspiracy across traditions, with Jews cast as hidden orchestrators and variously constructed target categories — Black, immigrant, Muslim — cast as demographic instruments, a structure documented in the manifestos of several of the deadliest attacks of the past decade. The model-minority frame operates as a wedge, invoking one minority to delegitimize the claims of another. Misogyny can function, in documented radicalization pathways, as an entry environment whose emotional grammar — humiliation, conspiracy, restoration — later admits racialized enemy figures. And the targeting compounds: women who are Black or Jewish, particularly journalists and public figures, receive hostility in which the traditions merge into a single stream. The junctions are visible at the level of single mechanisms as well: dehumanization and objectification recur across at least three traditions with domain-specific imagery — the parasite repertoire of antisemitism, the animalizing frames of anti-Black racism, the reduction of women to bodies and functions — and conspiracy structure spans antisemitism and anti-Asian racism, where espionage suspicion and pandemic origin narratives attribute covert collective agency. The comparison protocol is fixed: comparison operates at the level of communicative mechanism, never by equating histories, prevalence, social positions, or consequences.
A further convergence connects all four traditions to generalized anti-elitism. Across the corpora, hostility toward named groups can merge with conspiracism about “the elites,” in which institutional distrust hardens into the conviction that public life is managed from behind a facade. Where these structures fuse, they can reinforce authoritarian friend-and-foe constructions that differ from one national culture to another. The same grammar is cast differently in Berlin, London, Paris, New York, and Toronto, which is one reason the program is designed for comparative, multilingual observation. The four-domain design reflects the program’s current stage, not a claim of exclusivity: anti-Muslim racism is the next planned domain and is indispensable to the comparative account, above all for replacement ideology, migration discourse, and civilizational threat narratives.
Shift 2: From snapshots to standing observation
The second shift changes the mode of observation: from one-off snapshots to standing, expert-anchored, AI-extended observation that can see normalization while it unfolds.
Why snapshots are not enough
Research across the assemblage remains disproportionately organized around synchronic studies: datasets are scraped once, frozen, published, and cited for years, while the discourse they sampled has moved on, and the codes mutate faster than the benchmarks that are supposed to capture them. Standing observation already exists in civil-society monitoring, whose systems process enormous daily volumes and deliver genuinely useful trend signals — but monitoring and research are optimized for different things, monitoring for timely signal and organizational decision-making, research for construct validity, transparent inference, and explanation, and the recurring problem arises when counting is presented as though the object being counted were already conceptually settled. What remains rare is research-grade standing observation: continuous, methodologically consistent, expert-anchored reading of the same discourse environments across several hate traditions, platforms, and languages, over years — although normalization, the process at the center of the democratic question, is by definition observable only over time. A snapshot cannot show that yesterday’s transgression has become today’s opinion; only a series can.
The survey, the instrument on which societies still chiefly rely for measuring prejudice, measures a different object. Survey research is indispensable for what it delivers — representative, comparable distributions of self-reported attitudes — but the situation is artificial: fixed answer formats, decontextualized questions, priming and social-desirability pressures that weigh most heavily on the very attitudes at issue. Social media discourse inverts each limit — self-motivated statements, the speaker’s own words at self-chosen length, interactional embedding, timestamps that allow correlation with triggering events — at costs the design must carry openly: platform users are not a representative sample of any society, visible discourse is selected by participation and ranking, and public expression cannot be read straightforwardly as private belief. Neither instrument replaces the other, but a society reading only its surveys sees prejudice as a distribution of private opinions and misses it as a public, interactional process.
The research design
The method can be stated in one line: expert reading, scaled through AI trained on that expert judgment. Decoding Hate is designed against the deficits just described, and its inferential sequence is explicit, because each step adds assumptions the previous one does not carry: contextual interpretation of the single utterance, reliable coding against documented concepts, corpus-level measurement, temporal comparison, and — last and most cautiously — bounded diagnosis of a discourse environment. A judgment of meaning does not by itself establish prevalence; prevalence does not by itself establish normalization; normalization is evidence about a discourse environment, not a verdict on a democracy.
Each domain rests on a three-part documentary architecture: a scholarly reference volume grounding every concept in its historical and research literature, a coding manual translating concepts into decision rules, and an operational annotation guide applied by a standing expert team working with domain specialists, at the level of the individual unit, recording both whether a hateful concept is present and the register through which it is carried — explicit, coded, dependent on surrounding co-text, or visual. The visual register is where the method’s restraint shows most clearly: the “OK” hand gesture, in ordinary use, means approval and is annotated as exactly that; only in co-texts saturated with white-nationalist vocabulary or explicit invocation of the coded reading does it function as an in-group signal.
Two disciplines govern the process. The first is the conservative attribution principle: where an utterance is genuinely ambiguous, it is not counted as hateful, so that reported prevalence functions as a conservative lower-bound estimate within the annotated corpus. The second is documented reliability: agreement between annotators is measured and reported, disagreements are adjudicated against the manuals, and the manuals are revised as the material demands — and where disagreement persists after adjudication, it is treated as data rather than noise: concepts with structurally lower agreement are flagged as contested in the published documentation, so the interpretive debates behind the standard remain visible rather than being flattened into a single verdict. The standard claims no interpretive neutrality — no reading of discourse is a view from nowhere — but something narrower and more defensible: judgments that independently trained readers reach at documented rates of agreement, under published rules that anyone can dispute. Concepts, decision rules, and findings are published openly, continuing the open-access practice of the predecessor project.
The corpus design is event-anchored. Rather than sampling by keyword — which presupposes the vocabulary and therefore misses the coded layer — the program annotates complete discussion threads around discourse events: an influencer’s trip, a political speech, a march, an attack. Complete threads preserve the co-text on which coded meaning depends, the interactional dynamics within a community, and the engagement structure that shows what audiences visibly reward. The Clavicular study is the model: 1,379 annotated units across six discussion threads, with engagement distributions recorded as evidence of which contributions received visible audience reinforcement — not treated as a measure of private agreement — showing that reinforcement concentrated heavily on the antisemitic material. Matched pre/post designs extend the logic in time, holding a discourse community constant and measuring how a violent event reshapes what its members treat as sayable. Complete-thread designs preserve interactional context, but they are event-anchored purposive samples, and event selection, platform affordances, deletion, moderation, and engagement-driven visibility remain part of the inferential frame and are reported as such.
Annotation of this depth cannot, by itself, cover large volumes. The computational pipeline addresses the trade-off in one direction only: expert judgments are the reference standard, and only where a model reaches a predefined, reported level of performance against held-out expert annotations, for a specified concept and corpus, is it used to extend that standard across corpora too large for human reading — with performance re-evaluated under domain and temporal shift rather than assumed to transfer unchanged. Experts supply the standard, and models supply the scale: the models never define the standard but only propagate it. This corrects the workflow that has quietly become standard elsewhere, in which experts write a codebook once, machines classify at scale, and the interpretive architecture stops developing the moment the pipeline starts running: here the annotations, the manuals, and the model benchmarks are revised together, continuously. This division of labor allows concept-level analysis, long confined to small expert-annotated samples, to be extended to substantially larger and more systematically sampled corpora; claims of representativeness are reserved for the cases where the sampling design supports them.
Time as a variable
The commitment to time runs through every component of this design, and it rests on a kind of experience that cannot be improvised. Fifteen years of continuous observation, including the Decoding Antisemitism corpus built since 2020, function as a longitudinal baseline: they are what make it possible to recognize that a formula like “Dan is not suicidal” is new, that it descends from the Epstein mythology and does not appear in the corpora we have examined before 2019, and that its spread can be distinguished from background noise and investigated as a possible normalization process. The communicative patterns are the stable layer — the conspiracy grammar, the reversal structures, the in-group signals persist across decades — while their surface vocabulary turns over in months. Process patterns recur as well: across cases we have documented a three-phase cascade in which strategically ambiguous elite discourse is sharpened by digital intermediaries and collapses into explicit hate in audience participation — ambiguity, reframing, collapse, a pattern observed across several cases so far and treated as a recurring hypothesis rather than a settled law. Only a maintained series can measure normalization at all: the appearance and subsequent spread of a concept across discourse environments is a trajectory, and trajectories require time series. The value of the baseline is easiest to state with an example: our analyses of comment sections on major English-language news channels after the attacks of October 7, 2023 (Discourse Report 6; Celebrating Terror) documented a sharp and sustained surge of antisemitic discourse far above its pre-crisis level — a surge recognizable as a surge only against years of methodologically consistent prior observation.
The distinction a maintained series makes measurable can be stated as two curves. After a triggering event, discourse surges; the question that matters is what happens afterward. One curve falls back to its old baseline within weeks — an escalation episode, sharp but bounded. The other settles at a permanently elevated level: the discourse environment has absorbed the surge, and what was exceptional before the event has become ordinary after it. The two curves are identical until the event and identical through the spike; everything that matters happens afterward, and the difference is invisible in any single measurement, legible only in a maintained series. Our qualitative longitudinal observations are consistent with the second pattern: forms of open glorification that had been largely absent from politically moderate online spaces surged after October 7 and, as far as we can observe, never fully receded — they appear to have become a more persistent part of the discourse. Whether that reading holds is exactly what a maintained quantitative series can establish and a snapshot cannot; we have not identified a publicly available, independently scrutinizable instrument that measures this pattern using expert-defined categories of hate, including its coded forms.
This may be the largest difference between the program and a study, and it is not only temporal: a study produces findings and ends, while an observatory maintains a measurement standard while the object itself changes. Coded discourse is a moving object — vocabulary turns over, meanings mutate, sometimes in response to the categories and detection systems applied to them, and models trained on yesterday’s material decay. A static instrument therefore cannot measure it reliably for long. The datasets, frameworks, and documented concepts remain durable assets; the models and the measurement layer built on them hold their value only if experts annotate emerging material, the manuals are revised, and the models are revalidated against the change. The interpretive framework, the annotation, the evaluation, and the longitudinal series have to be maintained together — which is the epistemological reason, not merely the institutional one, why the infrastructure must be standing rather than episodic.
The black box and the synthetic turn
Modern societies have built communication infrastructures that shape political reality every hour of every day, yet no independent institution has broad, cross-platform, multilingual, and methodologically transparent visibility into what moves through them: governments see fragments, universities see samples, and the platforms see their own systems through measurement priorities shaped in part by operational and commercial considerations. For the institutions responsible for democratic life, the digital public sphere functions largely as a black box.
Inside it, hate ideologies display transmission dynamics I have described elsewhere in epidemiological terms: mutation, adaptation, reactivation, differential spread. Hostile narratives are often reproduced by people who do not recognize the historical repertoire they are activating — normalization is supra-individual, a property of the discourse environment rather than of any single speaker’s intent. The analogy concerns transmission rather than biology — unlike pathogens, narratives move through interpretation, contestation, and political agency — but its lesson holds: the point of understanding transmission is not to condemn the transmitters but to understand the conditions under which narratives spread and become normalized.
One assumption dominating the policy conversation needs correction: that online hate is chiefly the work of malicious actors — foreign bots, coordinated influence operations. Those actors exist and deserve scrutiny. But across the corpora we have analyzed, much of what we observe has the surface form of participatory discourse: users presenting no organizational affiliation, writing with the idiom of domestic audiences, producing and amplifying hateful narratives in seemingly spontaneous exchanges — with the caveat that coordination can never be fully excluded from discourse observation alone. Two structural conditions sustain this participation: platform architecture, where engagement-based ranking rewards conflict-oriented content — recent causal research in Science showed that ranking decisions alone measurably shifted partisan animosity — and the communication conditions the digital environment imposes: anonymity that weakens accountability, mutual reinforcement that validates extreme positions, constant exposure that normalizes extremity. A research program that watched only for coordinated campaigns would miss a substantial part of the phenomenon.
Until recently, the trade-off described at the outset — interpretive depth on small samples, or scale without depth — seemed fixed. Current large language models have begun to change its terms, and this is the wager of Decoding Hate: they can now reproduce selected expert judgments with useful accuracy on tightly defined tasks, provided the concepts, examples, decision rules, and evaluation standards are supplied by sustained domain expertise. Social media discourse has become a central marketplace of ideas of our time, and it is becoming feasible to read that arena at a scale and a level of contextual resolution that were previously difficult to combine. AI also enters on the other side of the ledger: generative systems can produce, reproduce, and repackage hateful content at scale — synthetic text, imagery, and fabricated evidence entering circulation alongside user-produced material — which makes AI an object of the observation as well as its instrument, and makes the observatory the vantage from which the growing synthetic share of the environment can be studied longitudinally.
Who decides what counts as hate?
The strongest objection to a program of this kind deserves to be stated in its strongest form, because it will be raised. It runs: any institute that classifies speech as hate is, whatever its intentions, building the intellectual infrastructure of censorship — sooner or later it will file disfavored political positions under hate, its findings will be used to pressure platforms into suppression, and the question of who decides has no good answer.
The answer is in the design. The program classifies concepts in discourse, not speakers: its analytical outputs are pattern-level findings about discourse environments — it classifies no individuals and maintains no person-level records — and it has no moderation function and seeks none. Its methodological bias runs in favor of speech: the attribution rule is deliberately calibrated to reduce false-positive overreach, even at the cost of undercounting ambiguous material. Its standard is public and contestable: concepts grounded in documented literature, decision rules published, reliability reported — anyone can check where a line was drawn and argue that it was drawn wrongly, which is more accountability than platform moderation, government regulation, or public intuition currently offers. The interpretive work behind those lines rests on established disciplines: pragmalinguistics and argumentation theory, combined with historical knowledge of exclusionary and demonizing tropes, cover a great deal of interpretive ground — a statement very often names neither its target group nor its hateful predicate, and much of that gap can be closed, controllably, from the direct co-text and the wider context; where it cannot, the utterance is not counted. And the distinction between sharp political critique and hate is not a caveat appended to the method but its central operation, performed case by case and documented. No taxonomy is intrinsically immune to repurposing; what the design provides are specific safeguards against it — the methodological ones above, paired with governance ones: limits on person-level use, on operational surveillance, and on any automated enforcement, because transparency alone does not prevent the misuse of an instrument. The reader need not take the intention on faith, because the method is checkable.
Shift 3: From specialist knowledge to institutional competence
The third shift changes where the capability lives: interpretive expertise moves out of specialist academic communities and into public institutions, such as schools, newsrooms, courtrooms, and agencies that confront the phenomenon in practice.
How the change happens
The program’s theory of change can be stated as a single pipeline, because each stage exists for the sake of the next: discourse becomes expert-labeled data; expert judgment becomes computationally extensible; longitudinal measurement becomes usable institutional knowledge. Raw discourse data — complete threads around discourse events, across platforms and languages — is read by standing expert teams against conceptual taxonomies. Those annotations serve two purposes: they are findings in their own right, and they are the reference standard against which models are trained and evaluated, extending that standard across corpora too large for human reading. The modeling stage rests on one commitment that will outlast any particular technology: expert-grounded computational extension — models trained against expert annotation to recognize the full range of hateful expression, including the coded and context-dependent forms that current systems miss, and used only where they reach published performance thresholds. In the current implementation plan, that takes the form of four models, one per hate tradition — the antisemitism model, resting on the longest baseline, leading the sequence, and the models for anti-Black racism, anti-Asian racism, and misogyny following on the pipeline it establishes, each released with published evaluations — though the technical architecture will follow the state of the art rather than remain frozen to a 2026 design. The pace at which the later models can be built is a consequence of the inheritance described in Shift 1: the pipeline precedes them.
The observatory dashboard
At the center of the observatory layer, an interactive dashboard will be built on the annotated and model-extended corpora — a standing instrument, updated at least weekly, whose organizing dimension is time: it makes the language of hate visible as it changes. A user — a researcher, an educator, a policy analyst — can ask what forms antisemitism or anti-Black racism have taken on a given platform over the past weeks, months, or years, and receive an answer at the level the method makes possible: which hateful and exclusionary concepts are circulating, and through which rhetorical devices — coded vocabulary, irony, reversal, pseudo-objective framing, imagery — they are being communicated. Its answers carry the method’s confidence structure with them: coverage, thresholds, and bounds displayed rather than buried. A surge, in particular, can be decomposed into the concepts that carry it, their distribution across platforms, and the rhetorical devices that deliver them: the dashboard is an analytical interface to the concept-level measurement architecture, not a visualization product. Standing measurement infrastructures in adjacent fields — V-Dem’s maintained democracy indices are the closest analogue in spirit — show that this kind of instrument can be sustained; what does not yet exist is one built on expert-defined categories of hate, including its coded forms. And because it runs continuously, it can distinguish what a snapshot never can — the surge that decays from the surge that holds, escalation from normalization — and, as the series lengthens, whether discourse patterns changed following educational, counter-speech, or platform interventions: descriptive evidence that can inform such assessments without by itself establishing causality. Trend reporting exists; we have not identified a standing instrument that provides concept- and device-level answers comparable across four traditions and readable as trajectories.
Access is tiered by design — aggregate trend views publicly, concept- and device-level query access for verified institutional users under use agreements — because an instrument that documents the coded layer in real time must not double as a training manual for it. Alongside the dashboard, the observatory’s outputs include openly published benchmarks, against which any detection system, ours or anyone else’s, can be measured — the underlying annotated datasets go to vetted partners under the same access architecture — and an annual report on the state of the digital discourse environment. The same corpora serve a scholarly function beyond hate research: methodologically documented, longitudinally maintained social media discourse data of this kind is scarce, and vetted researchers can work with it on the same tiered basis.
From knowledge to competence
Put simply, what cannot be seen cannot be countered. Educators do teach emerging phenomena, courts do adjudicate uncertain meanings, security services do read coded registers, and policy does regulate imperfectly measured dynamics — but each does so less reliably, less consistently, and less defensibly without a documented evidentiary infrastructure beneath it. The systematic study of digital discourse therefore belongs in the foreground of serious efforts to understand contemporary hate, polarization, and democratic strain, not at the academic margins where it largely stands today. Universities are good at producing knowledge about problems of this kind, and far less good at producing competence — the transferred, usable ability of practitioners to read the phenomenon themselves, under time pressure, in their own settings. Moving that capability outward is the third shift, and it takes the concrete form of instruments. The program’s contribution ends at competence; what institutions do with the capacity to see is their decision. The transfer will not be frictionless — institutions have professional cultures, resource limits, and political constraints of their own — and the program’s obligation is to make the instruments available, current, and good, not to assume their adoption.
The Glossary and the instruments that follow
The Decoding Hate Glossary is going to be the first instrument of this transfer, now in final preparation with interactive online publication to follow — the analytical framework turned into a usable public reference, written for the researchers, journalists, practitioners, and institutional leaders who need it under deadline, and renewable online as the codes change, at a pace no print cycle could match — and the demonstration that the pipeline ends in something a non-specialist can hold. The instruments that follow it are matched to the institutions that need them.
Teachers and journalists face versions of the same problem at different speeds: both encounter emerging forms of hate before any textbook or style guide explains them, and both need the ability to read how coded language, irony, and platform dynamics operate rather than lists of forbidden words. For education, the program is developing teaching modules and diagnostic exercises built from real annotated material, renewable as the codes change; for newsrooms, concept-level briefings allow reporting to name the mechanism at work instead of reproducing its vocabulary or dismissing it as noise.
Security services are generally better equipped to see the late stages of radicalization than the early ones, because the early stages often unfold in coded registers their tools were not built for — and, in many lone-actor cases, there may be little network or affiliation structure for conventional indicators to detect. Discourse-related situational understanding is a capability agencies could build from this work: not tracking individuals, but understanding which narratives are hardening, in which discourse environments — platforms, channels, and comment spaces as aggregates, not inferred networks of persons — and toward which targets. Such analysis may make intensifying discourse environments visible early; it cannot and should not predict individuals, and any transfer into security practice requires strict boundaries around individual-level inference, data access, and operational use.
Regulatory measurement often remains concentrated on the phenomenon’s more explicit surface; the missing instrument is periodic, independent, discourse-level reporting from infrastructure not owned by the platforms being regulated. Courts have always interpreted meaning in context, but rapidly mutating digital codes increasingly exceed the evidentiary infrastructure available to them; the transfer under development pairs concept-level trend analysis for legislative and policy contexts with expert-testimony protocols for individual communications, held to a stricter evidentiary threshold.
Platform trust and safety teams stand at the shortest distance between finding and consequence; the precedent is the Meta case described in Shift 1, and concept-level documentation and open benchmarks extend that pathway across all four traditions — evidence offered to the teams who write the policies, not enforcement performed on their behalf. Civil-society organizations, finally, monitor hostile discourse with the thinnest resources of any audience named here; annotated material and concept documentation give their monitoring a defensible evidentiary standard and their counter-speech a target — the concept being expressed rather than the surface phrasing that will have changed by next quarter.
The Glossary will be followed by the program’s first empirical outputs beyond antisemitism and, as the observatory reaches operating scale, by periodic public reporting.
The stakes
Contemporary public discourse is digitally mediated at a scale and speed that no existing configuration of institutions can adequately observe, while engagement-driven platform incentives can reward conflict, intensity, and transgression over the maintenance of common ground. The circumstances under which citizens who disagree remain members of a shared political community are not self-sustaining; they are built, eroded, and, where the erosion is noticed in time, repaired. The erosion is rarely dramatic: it proceeds through the subtle normalization of hate and exclusion, increment by increment, in exchanges too small to make news — which is also why it goes unwatched. What is at stake in those small exchanges is not only the harm done to targeted people — that harm is real, and it is itself a democratic harm — but also the operating conditions of democratic life itself.
Naming these operations accurately is not the same as suppressing them: it helps maintain an open discursive space in which understanding, empathy, and perspective-taking across differences remain thinkable. The program aims, ultimately, to contribute to a shared language in which a society can talk about politics, difference, and its own condition without sliding into demonization. Democracies will not maintain themselves on observation alone; they need a public vocabulary adequate to what the observation shows.
Democratic societies have always produced hate; that is not new, and no research program will end it. Recognizing emerging forms of hate — its mutations, its coded vocabularies, its narratives gaining traction in unnoticed or unwatched domains — before they become the background of political life opens the possibility of intervention before normalization is complete.
Whether democracies take it up is not a technical question. It is a question about whether a society intends to remain the kind of place where disagreement can still take the form of argument — and whether it acquires, while there is still time, the capacity to understand what its own public discourse is becoming.
Further reading
Selected reading
For readers who want the essential background, ten starting points: Bergmann and Erb’s “Kommunikationslatenz, Moral und öffentliche Meinung” (1986), the origin of the communication-latency concept; Mouffe’s On the Political (2005), for the adversary/enemy distinction; Alexandra Siegel’s “Online Hate Speech” chapter in Social Media and Democracy (2020), for the research landscape; Bartels (PNAS, 2020), on ethnic antagonism and antidemocratic attitudes; Decoding Antisemitism: A Guide to Identifying Antisemitism Online (Becker, Troschke, Bolton, and Chapelan, eds., Palgrave Macmillan/Springer Nature, 2024), the Lexicon on which the interpretive standard rests; Becker, Blatter, and Stanevich (2026) in Frontiers in Communication, for the current benchmark evidence on LLM-based detection; Steffen, Pustet, and Mihaljević (2024), documenting the failure of toxicity scoring on antisemitic content; Piccardi and colleagues (Science, 2025), for causal evidence on feed ranking and partisan animosity; Imhoff and Bruder (2014), on conspiracy mentality as a generalized political attitude; and The Invisible Spread, which develops the virological perspective on hateful meaning.
The full apparatus
On the theoretical background: Habermas’s The Structural Transformation of the Public Sphere (1962) and its sequel A New Structural Transformation of the Public Sphere and Deliberative Politics (2022) frame the deliberative conditions this essay treats as at risk; Mouffe’s On the Political (2005) develops the distinction between the adversary and the enemy on which the argument about disagreement rests; Arendt’s “Truth and Politics” (1967) and The Origins of Totalitarianism (1951) remain the reference points for the relationship between organized lying, conspiracy thinking, and the destruction of a shared world; Link’s Versuch über den Normalismus (1997) grounds the concept of the sayable and the analysis of normalization; Crenshaw’s “Demarginalizing the Intersection of Race and Sex” (1989) and Collins’s Black Feminist Thought (1990), the sources of intersectionality and the matrix of domination, supply the theoretical vocabulary for the junctions between hate traditions that the comparative design operationalizes at the level of communicative mechanism; and Müller’s What Is Populism? (2016) connects anti-elitism and anti-pluralism to the exclusionary logic described in the section on intersections.
On the wider field: Bergmann and Erb’s “Kommunikationslatenz, Moral und öffentliche Meinung” (Kölner Zeitschrift für Soziologie und Sozialpsychologie, 1986) is the origin of the communication-latency concept the essay builds on; Alexandra Siegel’s “Online Hate Speech” chapter in Social Media and Democracy (Persily and Tucker, eds., Cambridge University Press, 2020) surveys the research landscape and its fragmentation; Dixon and colleagues’ “Measuring and Mitigating Unintended Bias in Text Classification” (AIES, 2018) documents the identity-term inflation problem; Bach, Schmitt, and McGregor’s “Let Me Be Perfectly Unclear” (Communication Theory, 2025) develops strategic ambiguity in political communication; and boyd’s work on networked publics and context collapse (2010) underlies the platform-dynamics argument. The VOX-Pol network’s publications map the adjacent field of digital extremism research, including its own methodological self-criticism.
On prejudice, conspiracy belief, and democratic erosion: two survey traditions arrive at the same structural finding independently. In the German-language literature, Heitmeyer’s Group-Focused Enmity program (the Deutsche Zustände series, Suhrkamp, 2002–2011) documents antisemitism, racism, sexism, and further prejudices as a correlated syndrome linked to authoritarian orientation; Zick, Küpper, and Hövermann (2011) extend the design comparatively across Europe; the FES Mitte-Studien (most recently Die distanzierte Mitte, 2023) and the Leipzig Authoritarianism Studies (Decker and Brähler) carry the measurement forward in the German mainstream. In the anglophone literature, Sidanius and Pratto’s Social Dominance (1999) and Duckitt’s dual-process model account for why prejudices against different groups travel together; Altemeyer (2006) documents the authoritarianism link; Bartels (PNAS, 2020) identifies ethnic antagonism as the strongest predictor of antidemocratic attitudes in the United States; Graham and Svolik (APSR, 2020) measure voters’ willingness to trade democratic principles for partisan advantage; Mason (2018) and Kalmoe and Mason (2022) trace the enemy-construction dynamic in American partisanship; and Norris and Inglehart (2019) place the authoritarian-populist backlash in comparative perspective. On the conspiracy link specifically: Imhoff and Bruder (2014) establish conspiracy mentality as a generalized political attitude, Kofta, Soral, and Bilewicz (2020) connect antisemitic conspiracy belief to its political correlates, and Uscinski and Parent (2014) provide the American baseline. The direct longitudinal test — antisemitic conspiracy discourse as a leading indicator of democratic erosion — remains thin in this literature; it is one of the gaps the observatory is built to close.
On the method and its foundations: the Decoding Antisemitism Lexicon (Becker, Troschke, Bolton, and Chapelan, eds., Palgrave Macmillan/Springer Nature, 2024) documents the 40 concepts and the annotation framework; two peer-reviewed edited volumes, Antisemitism in Online Communication (Becker, Ascone, Placzynta, and Vincent, eds., Open Book Publishers, 2024) and Imagery of Hate Online (Becker, Scheiber, and Jensen, eds., Open Book Publishers, 2025), extend it to transdisciplinary and multimodal analysis; Becker and Troschke (2023) set out the conservative attribution principle; and Becker, Blatter, and Stanevich (2026), in Frontiers in Communication, present the four-layer typology of antisemitic discourse together with the current benchmark evidence on LLM-based detection.
On AI, platforms, and detection: Steffen, Pustet, and Mihaljević (2024) document the failure of toxicity scoring on antisemitic content; Patel, Mehta, and Blackburn (2025) evaluate large language models on antisemitism detection; Soemer and Mihaljević (2026) test how conceptual representations shape model performance; the Blue Square Alliance Command Center (2025) reports benchmark experiments with the newest model generation; and Piccardi and colleagues (2025), in Science, provide causal evidence that feed-ranking decisions shift partisan animosity.
On the discourse environment and recent cases: The Invisible Spread develops the virological perspective on hateful meaning, and At the Bottom of the Iceberg documents the Bilzerian discourse the opening draws on. Rapid case studies show the method on live material: the Clavicular Israel visit (Substack / Algemeiner), the Mamdani AIPAC speech (Substack / Algemeiner), the reactions to Kanye West’s antisemitic campaigns, the discursive afterlife of a New York Times column, and the conspiratorial aftermath of the White House Correspondents’ Dinner shooting. The three-phase mechanism the essay describes is developed in full in Becker, Beacken, and Sabra (2026), a study of the Mamdani case across 3,404 comments on YouTube and TikTok, published in NYU’s open repository. A ten-minute film, Decoding Antisemitism: Understanding Antisemitism Through Digital Discourse, introduces the approach for general audiences, with subtitles in English, German, French, Spanish, Hebrew, and Arabic. The Decoding Hate Glossary, now being finalized, will be published separately.









