Analysis

Interactive overviews and figures generated from Chapter 1 - Final results - Results.csv (the cleaned, classified finding rows). Every figure is downloadable as a CSV, SVG, or 300 ppi PNG from its own footer.

1. Overview

Table 2. Headline counts of the final corpus (Chapter 1 - Final results - Results.csv).
⬇ Download CSV Four headline counts: items in master_bibliography.csv; YES + MAYBE master items; the final-corpus item count; and how many of those have a PDF on disk.
Figure 1. Where each language looks — alluvial flow from the publication's detected Language, through every Country it discusses, into the Topic (or, in the "by CM Sub-topic" mode, the Content-moderation Sub-topic) it engages with.
⬇ Download CSV Three-stage alluvial: Language → Country → Topic (or CM Sub-topic). Each ribbon weight is the number of items that link those three values. An item that mentions several countries contributes once to each country flow; items whose Country column is blank land under "(unspecified)". The middle column is capped at the twenty most-mentioned countries; the rest are rolled into "Other (N)" so the figure fits in one viewport. The "by CM Sub-topic" mode restricts to items whose Topic is Content moderation.
Figure 2. Unique items flowing from Discipline to Topic, with a third hop to the four Sub-topics of Content moderation.
⬇ Download CSV Three-stage Sankey diagram: each ribbon is the number of unique items whose master Discipline (left) is classified by the v2 taxonomy into a given Topic (middle), and — for items whose Topic is Content moderation — into one of its four Sub-topics (right). Bands are proportional to the unique-item count and labels show the count at each stage.
Figure 3. Items per Discipline, per year, per Topic — node size = Citations. Switch to a Bump view (publications per Topic per year).
Years
⬇ Download CSV Bubble plot. X = publication year; Y = Topic; each circle is one unique item, area scaled by citation count; colour = the item's most-frequent Discipline (top-10 highlighted, the rest grey). The Bump view aggregates the same data into one line per Topic, Y = that Topic's share of all unique items published that year (lines sum to 100 % per year) — so corpus growth doesn't dominate the trend.
Figure 4. Items per Discipline, per year, per Sub-topic — Censorship or Content-moderation framings of the WHAT taxonomy. Switch to a Bump view (publications per Sub-topic per year).
Years
⬇ Download CSV Bubble plot of items grouped on Y by their WHAT Sub-category. The Censorship tab restricts to items whose dominant WHAT category is Censorship (lanes: Constitutive, Infrastructural, Regulative); the Content moderation tab restricts to items whose dominant WHAT category is Content moderation. X = publication year, size = Citations, colour = top-10 Discipline. The Bump view shows one line per Sub-topic, Y = that Sub-topic's share of items in the chosen Topic published that year (lines sum to 100 % per year).
Figure 5. Media categories per Topic / per Content-moderation Sub-topic — Sankey view, or per-year Bump view (each medium's share of that year's top-10 medium mentions).
⬇ Download CSV Sankey view: ribbons size = items associated with that medium. Bump view: one line per medium (top 10 in the chosen scope / slice), Y = items per publication year that mention it. The "by CM Sub-topic" mode restricts both views to items whose Topic is Content moderation. Source: the Media category column in Chapter 1 - Final results - Results.csv.

2. Analysis

Figure 6. Beeswarm per Topic — categories by year, coloured by Discipline. Switch to a Gantt / matrix view to see mentions per Category per year.
Years
⬇ Download CSV Each dot is one finding for one item. X = publication year; vertical jitter clusters dots by Category; colour = top-10 Discipline (rest in grey); size = Citations. Hover any dot for Mentioned item, Page, Title, Author, Discipline.
Figure 7. Beeswarm for Content-moderation Sub-topics — same encodings as Fig 7.
Years
⬇ Download CSV Same encodings as Figure 7, restricted to items whose Topic is Content moderation; vertical groups are the four Sub-topics.
Figure 8. Network — Topic → WHO → HOW → WHAT → WHY co-occurrence within items.
⬇ Edges CSV ⬇ Nodes CSV Layered network: each item links its Topic to the WHO categories it names, those WHO categories to HOW, HOW to WHAT, WHAT to WHY. Edge weight = number of unique items the link occurs in. Drag nodes; adjust the minimum edge weight to declutter.
Figure 9. Network — Content-moderation Sub-topic → WHO → HOW → WHAT → WHY.
⬇ Edges CSV ⬇ Nodes CSV Same encoding as Figure 9, restricted to items whose Topic is Content moderation; the GROUP layer holds the four Sub-topics.
Figure 10. Topic → Medium → How — three-stage alluvial mapping each Topic to the media categories it discusses, and on to the HOW mechanisms (techniques of moderation / censorship) those items describe.
⬇ Download CSV Three stages: Topic (left) → Medium (centre) → How (right). Each item contributes one unit to every (Topic × Medium × How) triple it appears in — so an item that mentions two media and annotates three HOW techniques contributes 6 units total. Source: the Type = HOW rows of Chapter 1 - Final results - Results.csv.
Figure 11. Top-10 content-moderation how techniques over time — one line per HOW category, restricted to items whose Topic is Content moderation.
Years
⬇ Download CSV Each HOW category (Filtering, Removal, Review, Enframing, …) gets one line; Y-axis is that category's share of all top-10 HOW findings published in that year — so the ten lines at any given year sum to 100 %. Rows are taken from the Type = HOW slice of the content-moderation beeswarm CSV; an item can contribute multiple HOW lines if it codes several techniques.
Figure 12. Content-moderation whohow — alluvial of the top-10 actors against the top-10 HOW techniques most used by platforms.
⬇ Download CSV Each ribbon is an (Actor × Technique) co-occurrence within an item. The HOW column is the ten techniques most associated with Platforms across the corpus; the WHO column is the ten most-mentioned actors among items that deploy any of those techniques. Width = number of items where the pair co-occurs. Source: cm_who_how.csv (derived from beeswarm_by_cm_subtopic.csv).
Figure 14. Keyword Venn — Who, How and Why keywords that are either specific to one WHAT topic or overlap across WHAT topics, linked to the six WHAT categories that organise the corpus.
⬇ Download CSV Six-set flower Venn over the six WHAT topics (Algorithmic sorting excluded). Set elements are the keyword vocabulary extracted from each publication's full text by GPT-5.4 and normalised to English via Argos Translate. Per Type (Who, How, Why), the diagram shows which keywords appear EXCLUSIVELY under one WHAT topic, which occur in EXACTLY one PAIR of adjacent topics, which span exactly three consecutive topics, and which span the core (≥4 of 6 topics). Non-adjacent overlaps are in the CSV. Closed-set taxonomy categories of the same six topics are also available in venn_categories.csv. Source: venn_keywords.csv.