Civil society organisations mentioned in academic papers and reports on censorship, moderation and AI alignment. First iteration before further validation. Faceted directory across academic literature.
About this project. This is a repository of civil society organisations mentioned in academic papers and reports on censorship, moderation and AI alignment, part of the Power over Platforms project at the University of Copenhagen.
It is currently under refinement and validation.
The corpus combines two sources.
1. Academic literature on censorship, moderation, and AI alignment. A bibliography of 7,772 records was built by scraping Google Scholar with 71 queries across English, Spanish, Portuguese, French, and Italian (e.g. censorship, moderation, modération censure, moderazione censura, moderação de conteúdo, trust and safety, content moderation, debate moderation). Records were screened for relevance to censorship/moderation as forms of public-debate or content gatekeeping; 3,563 were tagged YES/MAYBE, of which 1,785 PDFs were successfully downloaded. The bibliography was discipline-classified using a separate enrichment pass and joined to the extraction by PDF path or title.
2. AI model cards (43 PDFs from major frontier AI labs — GPT, Claude, Gemini, DeepSeek, Llama, Grok). These were treated as a parallel discipline (AI / model card) and processed alongside the academic literature, on the premise that model alignment, refusal training, RLHF/RLAIF, evaluation, and red-teaming are continuous with the longer history of moderation and censorship.
Alongside the academic literature, the directory integrates a hand-curated NetGov directory of organisations engaged in platform governance (netgov_directory.csv, 1,002 entries). For each NetGov organisation we map: Name → canonical name; Link → URL; category / specific_category → Type; headquarters_country → Country; issue_focus_relabelled (semicolon- or comma-delimited) → Issue focus; influence_evidence → Mechanism; establishment_year → When (lower bound, capped at 2025); NetGov Sources → the venue from which the entry was identified (used as the citation author). The synthetic citation under each NetGov-sourced CSO is labelled NetGov Sources rather than an academic article, with the link going to the organisation's own website.
NetGov entries are merged with academic-literature CSOs by canonical name (case-insensitive, accent-folded, leading-article stripped), so an organisation appearing in both data sources gets a single card with both kinds of citations beneath.
To address these research questions, we develop an iterative multi-source method that combines traditional purposive sampling with systematic web searches across multiple languages and countries. The data collection and processing approach is illustrated below.
Our data collection integrates two complementary approaches designed to capture both established and emerging civil society actors across different contexts of engagement.
Organizations identified from existing platform governance data sources. We began with organisations from established platform governance venues (hereinafter Source A). These sources include public databases; policy consultation submissions for major legislative processes (e.g. the EU's Digital Services Act, the UK's Online Safety Act); official lists of platform partners (e.g. Meta's Safety Partners Programme) and EU trusted flaggers; coalition and network memberships from established transnational advocacy networks; and speaker and participant lists from key field events (e.g. TrustCon, Trust and Safety Summit).
These venues represent the most visible points of civil society engagement with both platforms and policymakers. Yet each also carries biases. Policy-consultation processes favour organisations with sufficient resources and legal expertise to engage in formal regulatory processes and tend to attract those that prioritise such engagement. Conference participation skews toward organisations with travel funding and existing networks, and those that seek publicity and access. Platform partnerships may exclude organisations that maintain independence from direct platform engagement, or that take a more adversarial position. To preserve these differences as an analytical variable rather than a limitation, we recorded the sources from which each organisation was identified, allowing comparison of actors by pathway of entry into governance processes.
Since we identified both organisational and individual actors (e.g. conference speakers, individual members of certain networks), we incorporated the organisational affiliations of individual actors into the list of organisations for the current study, so that our analysis focuses on more "organised forms" of social action.
Organizations identified from Google search results. To address coverage limitations of traditional sources, we conducted Google searches (Source B). We developed 30 search queries related to content moderation and platform governance. Each query is executed in English, Danish, and German across our four main target countries (US, UK, Germany, Denmark). These countries are sampled to cover spaces where most Western platforms are headquartered in European markets both inside and outside the EU, in addition to the team's domain and language expertise. We used the Google Search SerpAPI to collect the top 100 results per query, language, and country combination. This approach yielded 16,304 deduplicated search results. Each result contained an actor's name, domain URL, and page snippet — providing scope for identifying organisations engaged in platform governance debates.
The advantage of our multi-source approach is that combining venue-based sampling with search queries from multiple languages helps us build a more inclusive picture than any one approach alone. In principle this could be expanded to more actors via additional languages for a more global mapping; for the limited purposes of this paper we focus on these four countries as preliminary starting points.
After deduplication based on domain names, the dataset was reduced to 7,257 organisational entries — 3,960 from Source A and 3,297 from Source B. Where Source A lacked links, we used the Google Search SerpAPI to obtain the three most relevant website results. We then downloaded each website in text format using a combination of Python libraries: scrapy for crawling, trafilatura and beautifulsoup for HTML parsing, and readability-lxml as a fallback. Text was downloaded in markdown format and stripped of images and other non-textual data, thus missing additional indicators of partners or funding.
Pages clearly irrelevant to organisational identity (e.g. analytics, donation forms, cookies) or inaccessible (e.g. 404/403 errors) were excluded. We also filtered pages using a whitelist of indicative keywords — mission, about us, vision, our work, projects, policy, research, training, outreach — translated into the most common non-English languages in our original datasets (Danish, German, French where relevant). The resulting corpus merged all pages per organisation into a single markdown file. Each organisation corresponds to one text file containing all available descriptive material.
Given the scale of the corpus (7,257 organisations), manual coding alone would have been prohibitively resource-intensive. Following recent work demonstrating the potential of LLMs for measuring complex social-science concepts (Laurer et al., 2025; Stolwijk et al., 2025; Weber & Reichardt, 2023; Ziems et al., 2024), we used GPT-5-mini to assist with classification, with a structured prompt derived from a codebook, and adopted an iterative approach focused on conceptual refinement.
Initial operationalisation and prompt calibration. We developed an initial coding schema grounded in our theoretical framework of civil society power in platform governance. Informed by previous work (Grover, 2022; Tjahja et al., 2021), the schema sought to assess observable structural attributes (e.g. organisational category, scope, country) based on website data, and substantive variables capturing thematic orientations (e.g. issue focus, functions, demands) and relational characteristics (e.g. platform relationships, funding, affiliations).
To assess operational clarity, three researchers first independently coded 30 random samples and discussed conceptual ambiguities, overlapping categories, and difficulties in consistently applying theoretical distinctions to website information. Following collective discussion, we revised definitions, clarified category boundaries, and selected calibrated few-shot examples to embed in the LLM prompt. The prompt instructed the model to (1) assign predefined labels, (2) extract actors' own issue framings and functional descriptions, and (3) provide short evidence excerpts from the website texts to justify each classification. The model was run in batches, with a maximum of 25,000 characters per organisation.
During this stage, we implemented a pre-classification filtering process. Entries were excluded if websites were inaccessible (n = 268), consisted solely of personal blogs or aggregated content platforms (n = 560), or represented public authorities (n = 617), private companies (n = 717), broad academic institutions or generic academic outputs (n = 1,756), or were coded as non-relevant (n = 352). We manually reviewed flagged cases to correct false negatives and then deduplicated organisations appearing across multiple sources (n = 844). At the conclusion of Round 1, the codebook remained provisional.
Conceptual refinement and iteration. The second round served two main purposes: conceptual validation and robustness testing. First, all three researchers conducted a close reading of the first-round LLM classifications for non-profit/non-state and academic-specific actors. Because several variables relied on extracting information from open-ended website texts, in many cases pre-assigned labels in the first round proved either too granular or insufficiently aligned with how organisations articulated their missions. Owing to the granularity of these labels, we applied keyword-matching to group actors' own descriptions into interpretable variables and manually verified the reclassification. We then revised the schema by incorporating new labels to better reflect patterns emerging from the data.
Second, three researchers independently coded an additional 50 organisations to stress-test both the LLM classifications and the revised coding scheme. As in Round 1, we compared coding decisions and discussed disagreements within the research team, and between human coders and LLM, qualitatively to determine their source. In many cases divergences reflected conceptual indeterminacy or overlapping definitions across variables in the original codebook. These discussions led to further simplification of the schema, clearer definitional boundaries and examples, and removal of variables that proved persistently unstable or theoretically redundant. The model was again run in batches, with a maximum of 30,000 characters per organisation.
The final codebook records the following per organisation: Relevance to platform governance (high / low / irrelevant); Establishment year; Organisational category (advocacy, coalition, network, industry/trade association, academic research centre/lab, think tank/policy institute, foundation/philanthropy, non-profit media/fact-checking, legal advocacy, tech infrastructure/standards organisation); Geographic scope (global / regional / national / sub-national); Headquarters country; Coverage focus (social_focus / tech_focus); Issue focus (human rights, free speech, censorship, privacy & data protection, algorithmic justice, mis/disinformation, child safety, counter-extremism, media/journalism, digital inclusion, consumer protection, hate or dangerous speech, minority rights, environmental impact, prosocial content moderation, middleware, challenging Big Tech, compliance with local speech regulations, others); Primary functions (knowledge production & analysis; advocacy, campaigning & agenda setting; capacity building & resource provision; policy & norm development; community building; legal & accountability actions; technical intervention & infrastructure); and Main funding category (government & intergovernmental; philanthropic foundations & private wealth; corporate & industry; membership-based / earned revenue; individual / public donation & crowdfunding; other; unknown).
After applying the revised schema and repeating classification and deduplication procedures, the final analytic dataset contains 2,143 unique organisations classified as CSO actors engaged in platform governance. Each record includes structural observable variables (establishment year, geographic scope, country) and substantive variables (coverage focus, issue focus, functions, funding types, specific funders). Where information was incomplete or ambiguous, variables were coded as "unknown".
There are at least three limitations accompanying our approach. First, the compilation of established platform-governance sources relied on the expertise and judgment of the research team. Although we consulted a diverse range of venues, this list is not exhaustive and inevitably reflects positional and language biases. While we sought to complement these sources with systematic Google search results, we recognise that Google's ranking algorithms shape which organisations are most visible, and that this visibility fluctuates over time. We also overlooked social media networks and professional sites such as LinkedIn as sources for data collection — both useful for finding international actors without a strong presence in either public-consultation fora or the Web sphere; indeed, our dataset revealed that a few actors had an exclusive presence on Facebook.
Second, our multilingual searches were limited to English, Danish, and German, while the approach can in principle be extended to other languages. Although the final dataset includes actors from more than 130 countries/regions, organisations operating primarily in other linguistic contexts are underrepresented. Further, our analysis focused on organisations rather than individuals, and aggregated personal affiliations into institutional entities; this may obscure the internal diversity of roles and positions within each organisation. Moreover, our reliance on publicly available website texts means classifications are based on organisations' self-descriptions, which are strategic representations and may emphasise some activities while omitting others, including informal or non-public forms of engagement. Finally, while LLM-based classification enables large-scale coverage, the model may infer attributes not explicitly stated or miss contextual nuance (Laurer et al., 2025; Stoll et al., 2025; Stolwijk et al., 2025).
To mitigate this risk, each record includes short evidence excerpts that facilitate manual verification. We do not claim that the dataset constitutes a representative sample of all civil society organisations engaged in platform governance. Rather, it captures a large and diverse subset of publicly visible actors across multiple venues and languages. The findings should be interpreted as identifying structural patterns, issue alignments, and means of engagement within this mapped ecosystem, rather than as distributional estimates of the broader field of platform governance.
Search. Google Scholar was queried via SerpAPI across 71 search terms covering English (e.g. censorship, moderation, content moderation, trust and safety, debate moderation), Spanish (censura, moderador censor, moderación de medios), Portuguese (moderação censura, moderação de conteúdo, moderação de mídia), French (censeur, modération censure, modérateur de débat, modération des médias), and Italian (il censore, moderazione censura, moderatore del dibattito). Each query returned up to ~100 records; results were deduplicated and merged into scholar_results_serpapi.csv and then exported to a Zotero-compatible master_bibliography.csv with all standard fields (title, author, year, abstract, DOI, URL, publication title, etc.) plus the originating Query.
PDF retrieval (scholar_2_download.py). For every record the script attempts a cascading retrieval strategy, falling through each option only on failure:
Resources / Link columns harvested by the scraper.best_oa_location).scihub.py), then via direct requests / browser fallback by DOI, article URL, and Resources URLs.citation_pdf_url meta tags, embedded PDF viewers, and PDF / download / fulltext links from the publisher's HTML.Successful downloads write to PDF downloads/ with a per-file metadata JSON. Failures log to failed_downloads.csv for later retry. The script is resume-safe and parallelisable.
Relevance screening (scholar_7_relevance.py). After the first download pass, every bibliography row is classified by Gemini 2.5 Flash-Lite as YES, NO, or MAYBE based on title, author, abstract, and URL — without reading the PDF (cheap text-only call). The classification is written back to two new columns Relevant and Relevant Note in the master bibliography. The criteria are explicit:
RELEVANT includes: censorship in any historical period or country; speech moderation in philosophy, political theory, sociology and other social sciences; moderation in conflict resolution; content moderation, platform governance, trust & safety; speech regulation, hate speech, harmful-content policies; information control, propaganda, disinformation policy; press freedom, media law, broadcasting regulation; deplatforming, shadowbanning, algorithmic demotion; index of forbidden books, publication bans, prior restraint; computer science / AI / NLP applied to moderation or censorship; internet governance, online-safety legislation; the moderation of public debates and conflicts; moderating powers and roles in political and media systems; moderation in political philosophy or philosophy in general; the moderating role of different societal actors (media, schools, religious institutions); moderation in media, political, and religious contexts.
NOT RELEVANT: medicine, clinical psychology, psychiatry, neuroscience; physics, chemistry, biology, ecology, earth sciences; engineering, materials science (unless about moderation tech); mathematics (unless applied to moderation/censorship); veterinary, agriculture, food science; items where censorship/moderation is mentioned only incidentally.
The screening is resume-safe (rows with a non-empty Relevant value are skipped) and rate-limited to fit Gemini's free-tier 15 RPM ceiling. Of the 7,772 records, 3,563 were classified YES or MAYBE.
Targeted re-download for relevant items (scholar_9_download_relevant.py). A second download pass is run on YES/MAYBE rows that don't yet have a valid local PDF. Before retrying, the script harvests every URL from all Zotero/Scholar link fields (Url, Additional Link, Scholar Link, Resources, attachment URLs) and resolves missing DOIs from embedded URLs or Crossref metadata, expanding the input each item gives to the cascading download chain. This recovered substantially more PDFs than the first pass: 1,785 of the 3,563 YES/MAYBE items now have a local PDF (~50 %); the remainder are mostly paywalled monographs, theses, or items behind institutional logins that none of the OA / shadow channels reach. Failures are logged and the run is interruptible (Ctrl-C triggers a clean save).
The PDFs that do land are the input to the CSO extraction pass described next.
Each PDF was processed by Gemini 2.5 Flash (extract_csos.py) with a prompt asking it to identify every civil society organisation substantively discussed in the document. The CSO category was defined to be historically variable: modern CSOs (NGOs, advocacy groups, watchdogs, trade associations, religious associations, trusted-flagger networks, AI-safety / evals organisations, industry consortia, external red-team auditors) and historical analogues (guilds, confraternities, learned societies, salons, reading societies, lodges, samizdat circles, abolitionist or temperance societies). State agencies, regulators, courts, for-profit platforms, and individuals were excluded; non-profit AI labs (e.g. EleutherAI) and third-party reviewers cited within commercial model cards were included.
For each CSO the model returned:
For each paper (paper-level, generated once per PDF) the model also returned Where (country, polity, platform, media type, or AI model family — e.g. Weibo, China; MENA region; Claude 3 family) and When (the temporal scope of the world described, not the world of writing — e.g. 16th century, Cold War, post-2010 platform era).
The pipeline is resume-safe (already-extracted PDFs are skipped) and parallelised across 3 Gemini workers, with a --retry-errors mode that re-queues failed extractions. Output is cso_extraction_results.csv: long format, one row per (paper, CSO).
loading…
The raw extraction contains messy free-text values that would be unusable as filters. A second Gemini pass (build_site.py) normalises them:
(start, end) year tuple, clamped to 1400–2025 to suppress hand-waved “since antiquity” ranges. Paper-level periods spanning more than 200 years are discarded as too vague.Results are cached in normalizations.json, so incremental rebuilds only normalise newly-seen values. A local migration step parses centuries, period names, and year ranges with regex/lookup before falling back to Gemini, making rebuilds essentially free.
CSOs are then merged by canonical name using a case-insensitive, accent-folded, article-stripped key. A merged CSO retains all aliases under which it appears, the most common type across mentions, the union of countries/media/norms/topics/periods, and the full list of source papers with each paper's individual mechanism and verbatim quote preserved.
The bibliography is joined in to provide each paper with its URL (from Url → Additional Link → Scholar Link → DOI) and Topic (the Query field — the literal search keyword that recalled the paper, e.g. content moderation, modération censure, debate moderator).
Article language. Each paper is assigned a Language (used as a filterable facet on cards). Inference is two-step: (1) prefer the Zotero Language field if populated (1,884 of 7,772 rows have an ISO code such as en, es, fr, pt, it, de, ca, nl), mapped to its canonical English name; (2) otherwise fall back to a hand-curated map from each of the 71 SerpAPI queries to its primary language — moderazione censura → Italian, modération de contenu → French, moderação censura → Portuguese, Moderador de medios → Spanish, content moderation → English. Queries that exist in several languages (censura in Spanish, Portuguese, Italian, Catalan) default to the most populous (Spanish) and are overridden whenever the Zotero Language field is set, so the rare-language counts (Catalan, Dutch, German, Indonesian) come purely from the Zotero metadata, never from query inference. Papers with neither a Zotero language nor a recognised query (the 708 manually-imported entries) are left unlabelled.
Across the 7,026 merged CSOs the breakdown is approximately: English 3,897, Spanish 1,205, French 948, Portuguese 818, Italian 379, then a long tail of Catalan, Dutch, German, and Indonesian. (Numbers are mention counts, so a CSO discussed in both an English and a Spanish paper is counted under both.)
loading…
This site is a single static page (index.html + data.json), published via GitHub Pages. The interface is a faceted directory of CSOs:
[when_min, when_max] overlaps the selected range.localStorage) hide Topic/Norms/etc. tags for a less cluttered view.Bibliography: the source bibliography master_bibliography_disciplines_enriched.csv is available here. It contains 7,772 Zotero-style records with Url, DOI, Additional Link, Query, Discipline, Local PDF Path, and a Relevant verdict (YES/MAYBE/NO) used to filter the extraction set.
Every run is incremental: new PDFs added to the corpus, new model cards dropped into the alignment-actors folder, or new normalised values are picked up automatically without re-doing prior work. Each rebuild stamps a BUILD_VERSION into the page so users always see fresh data.