The actors in AI alignment

Introduction

One of the major controversies AI has brought to societal debate concerns its renewed role in public access to information . One once spoke of “algorithms” as a catch-all denominator for the underlying infrastructures and politics behind information management , be that as ranking systems , newsfeeds , recommenders and a variety of content moderation systems . There is now an understanding that AI systems — LLMs, in particular — come to replace these as ever more intuitive and precise personalization systems. Seamless as this shift may feel, the political implications remain profound: both algorithmic and AI-based information systems conjugate the circulation of information with societal norms.

A central political consequence of this shift is the question of who gets to decide and operationalise societal norms into technical systems. Companies certainly have most decision-making power: a key aspect of “platform power”, transposed here to the emerging area of AI, is the “endogenous” nature of their capacity to create the environmental conditions in which markets and societal sectors function . But they retain dependencies with other actors that help legitimise them, particularly in governance frameworks (informal or otherwise) that influence the basic norms (and procedures) of alignment — be these epistemic (definitions of “misinformation” and other types of information integrity); bias (representational, political or other); harm (personal or societal); and myriads of other “essentially contested” concepts . Safety criteria, red teaming and benchmarking, for example, stand out as frequent entry points for third parties (nonprofits, think tanks, academics and regulators) to propose various alignment criteria as a nascent market of “independent verification organisations” .

Still, making room for the “essentially contested” aspect of these norms — a room for public participation and deliberation — is still an exploratory agenda . Part of it is a question of inclusive design: who is consulted about alignment norms, what is their participation (and applicability) in alignment processes, and how does one highlight contestation where it occurs? Another is a question of negotiation — or struggle, should we confront the irredeemably diverse interests of AI governance stakeholders. How are different actors invited to deliberate and collaborate in alignment ? Who and what moderates this process? Perhaps inconclusively, much has been said of the struggle to open algorithmic systems to a wider market or to regulatory frameworks , with middleware and algorithmic pluralism being a few of the latest conceptual propositions brought to the table of policy makers. Long before then, scholars and nonprofits had called for a variety of “participative”, “collaborative” or “reflective” methods to be embedded in platform design.

But whether or how this applies to AI alignment is a question that still waits to be explored. We begin with the how and who: who participates in alignment processes, and in what current (and prospective) capacity?

To tackle this empirically, we can begin charting AI alignment as an interdependent stack of technical and normative interventions, each of which invite contributions from different actors. There is the question of what values to train: their definition and provenance (from e.g. legislation or public norms); how they trickle down into specific “ideal conducts” by which a model should behave, as well as risks to be mitigated; data cleanup and curation processes; the training of each conduct and mitigation of each risk; and their benchmarking and evaluation. Each of these stack components may assemble a variety of “governance mechanism” : the model cards, “AI charters” and principles, company updates where norms are spelled out in “transparency” exercises ; how these norms call for different expertise and consultations, particularly in red teaming ; what benchmarks are adopted from a wider market; and so on. And each of these tend to involve a different set of actors with varying capacities: academics, nonprofits specialising in projecting AI risks (Apollo, METR, etc.), regulators, state-funded AI safety institutes, and so on.

In the tradition of actor-network theory , we map the presence of each of these actors across each stack. We devise an LLM-assisted extraction pipeline to perform a “distant reading” of 213 alignment-relevant documents (model and system cards, company policies, “principles”, grant calls, audits, blog announcements and research papers) from 29 companies across the US, Europe and China. We extract the names and types of actors (nonprofits, governmental, internal, private, academic) mentioned in relation to specific alignment processes: the definition of alignment values; the training of each value or mitigation of each risk; and their benchmarking.

Though these documents are of course not representative reflections of everything that occurs in alignment processes (and everyone involved), they do in fact participate in the governance process of AI companies: they perform public duties towards “responsible democratization of machine learning” primarily by responding to the sensitivities of regulators and public debates about the prospective use of their models, and position themselves in a wider market of other alignment norms and technical performance benchmarks. In passing, they cite academic papers that, no matter how technical, are one of the main sites of normative deliberation in alignment. They also cite conceptual influences from broader public debate; nonprofits with whom they have worked to benchmark or red team their models; public consultation processes; or AI safety think tanks with whom (or in response to whom) they have determined a particular alignment norm.

We are cautious about ascribing the motivation of an AI company to involve one or another actor, and even less about the impact these actors might have. We restrict ourselves to analysing the ways in which each actor is cited, and thus indirectly “enrolled” in a company's alignment process. Some actors are co-authors, and may thus indicate a greater degree of interdependency with companies. Others may be mentioned as influences; as somewhat tokenised representatives of certain public value concerns (e.g., child protection advocacy groups); or not mentioned at all — a valuable data point nonetheless. We take note of how they are cited and how enrolled they are throughout the alignment process — from the definition of ideal conducts and risks, to their training and mitigation, to their benchmarking.

We find that, while multiple actors are involved, each is allocated to specific and selective roles: governmental actors as initiators of “alignment standards” and producers of “AI security” norms and benchmarks; academics as benchmark producers and legitimisers of performance metrics; nonprofits as a hybrid category of value formulators and intermediaries with public consultation processes; and private firms concentrated in cybersecurity and red-teaming. While some contents of current alignment methods have been subject to more collaborative — if “democratic” — endeavours (e.g., collective constitutional AI), the process itself remains almost entirely at the discretion of AI companies. We close by reflecting on what this distribution of roles implies for the political theory of AI governance, and what alternative models of legitimacy a more pluralistic alignment infrastructure might require.

Method

The visualisation rests on a structured reading of 234 alignment documents — model cards, system cards, company policies, principles documents, grant calls, audits, blog announcements and research papers — analysed end-to-end by a large language model against a fixed schema. Each document yields one or more rows of the form (document × conduct or risk × training method × benchmark × actors), recording not only the actors named at each stage but the capacity in which each is enrolled. The result is a flat dataset of 2,563 rows describing 154 conducts, 828 risks, 55 training techniques and some 900 benchmarks across 29 companies.

Corpus

The corpus combines two parallel sources. The first is a Zotero bibliography of 181 alignment-relevant publications — model cards and system cards from the major labs (OpenAI, Anthropic, Google, Google DeepMind, Meta, xAI, Mistral, DeepSeek, Microsoft, Alibaba, Black Forest Labs, Aleph Alpha, and others), accompanied by published principles documents, charters, audit reports and grant calls. The second source is a folder of 53 deduplicated company policies: acceptable-use, privacy, terms-of-service, developer terms, data-processor agreements and related documents, tracked through time via dated snapshots stored as markdown.

Each document is classified into one of ten types — model card, company policy, principle, grant, announcement, audit, initiative, report, research, or other — with an optional sub-type for policies (Acceptable Use, Privacy, Terms of Service, Developer Terms, Data Processor Agreement, etc.). For company policies, the schema is restricted to information that informs how the model is trained, moderated, filtered or evaluated; user-facing usage rules, billing terms and account terms are explicitly excluded. This keeps the dataset focused on alignment governance and avoids drowning it in the boilerplate of consumer contracts.

Per-document extraction

Each document is first converted to plain text — PDFs through PyMuPDF, web pages and Zotero HTML snapshots through trafilatura, policy snapshots read directly from their markdown — and then read in a single request to GPT-5.4, submitted asynchronously through the OpenAI Batch API. The system prompt defines the schema, the document-type taxonomy, the conduct and risk vocabularies, the training-method categories, the actor-type and actor-category vocabularies, the citation-type vocabulary, and slug-generation rules for stable cross-document identifiers. The model is instructed to return a JSON object with a single key rows, whose value is an array of row objects. Row granularity is one row per (conduct or risk × benchmark) combination found in the document. A conduct or risk with three benchmarks therefore generates three rows, in which the training/mitigation fields are repeated; a conduct or risk discussed but not benchmarked still produces one row, so it appears in the dataset.

For each row the schema records 56 fields, grouped as follows:

Because extraction runs as a batch, the whole corpus is submitted in one request set and collected when complete; of the 234 documents, 184 yield extractable alignment content, the remainder being policy snapshots with no training-relevant material. A small number of the most safety-dense documents are refused by the provider's input-safety filter rather than processed, and are recovered separately so that none is silently dropped (see Limitations).

Controlled vocabularies

Free-form extraction is paired with several stages of controlled vocabulary. The model is constrained to choose one of nine actor typesinternal, private, academic, research institute, governmental, nonprofit, public, public consultation, other — to which we add the meta-category multiple when a row contains several types. Alongside this coarse type we apply a finer sixteen-class actor category, distinguishing alignment-research organisations, public-consultation bodies, advocacy groups, coalitions, industry associations, academic centres, think tanks, foundations, legal-advocacy organisations, red-teaming and cybersecurity firms, government agencies, internal actors, the authoring AI company, other AI companies, publics, and an open "other"; and an eleven-class citation type recording how each actor is enrolled — as author, formal partner, external evaluator, third-party red-teamer, benchmark or framework contributor, normative influence, company initiative, standards or regulatory framework it maps to, coalition it has joined, funding relationship, or temporary initiative. Because the model occasionally improvises labels outside these vocabularies, a harmonisation pass collapses every extracted category and citation value to a single canonical term — the authoring company and its named staff, for example, to "internal". Conduct categories are drawn from a 17-class vocabulary covering norms such as Honesty, Helpfulness, Harmlessness, Anti-discrimination, Political neutrality, Factuality, Reasoning, Safety and Knowledge. Risk categories are drawn from a 40-class vocabulary covering specific risk areas — CBRN, Cybersecurity, Bias, Sycophancy, Reward hacking, Persuasion, Misalignment, Hallucinations, Child Safety, Jailbreaking, and so on. Where a document uses language that does not map, the model returns Other: <specify>; these are reviewed and either folded back into the vocabulary or kept as an open category.

Training and mitigation methods are mapped to one of eight categories: Preference learning (RLHF, RLAIF, DPO, reward modelling), Supervised training (SFT, instruction tuning, demonstration fine-tuning, deliberative alignment), Reward shaping (adversarial reward functions, reward capping, rule-based rewards), Data curation (filtering, deduplication, synthetic data, PII redaction), Adversarial training (red-teaming data, adversarial examples), Inference-time filtering (content classifiers, safety filters, moderation APIs), Hard constraints (system-prompt rules, refusal policies, hardcoded refusals), and Runtime safeguards (trip wires, model lookahead, runtime monitoring, scratchpad oversight, sandboxing). A pipe separator is used when several categories co-apply.

The prompt is strict about what counts as a training method. Constitutional frameworks (e.g. Constitutional AI, Anthropic's Constitution) are treated as sources of norms — listed in pub_actors or risk_conduct_definition — and not as techniques; the underlying technique is RLAIF or SFT. Narratives about how a risk arises are likewise excluded from the training fields.

Post-extraction classification

A single further pass — again GPT-5.4, run as one batch over the unique values discovered in the data — adds four derived columns. It harmonises each verbatim conduct or risk name to a canonical label; maps each training description to a specific technique within a structured per-category taxonomy (Preference learning → RLHF, RLAIF, DPO, KTO, reward modelling; Supervised training → SFT, deliberative alignment, rejection-sampling fine-tuning; and so on); categorises each benchmark into one of 24 domains — Safety / Harmlessness, Refusal / Over-refusal, Honesty / Factuality, Hallucination, Bias / Fairness, Toxicity, Cybersecurity, CBRN, Privacy, Jailbreak / Adversarial, Agentic / Autonomy, the capability sub-categories (coding, reasoning, maths, multilingual, multimodal, long context, instruction following), Persuasion, Sycophancy, Reward hacking, Misalignment / Scheming, Internal evaluation suite, Other; and sorts every risk onto a priority line.

The priority line is the one classification deliberately withheld from the per-document reading. Each risk is sorted into a red, mid or low line by combining the document's own priority signal with the cross-company frequency of the risk and a fixed rubric: red lines are the unacceptable transgressions a company treats as absolute — CBRN, child sexual abuse, catastrophic loss of control; mid-lines are the contested, grey-area concerns that recur intermittently and whose definitions are disputed — political neutrality, persuasion, "grey area" content; low-lines are the risks rarely emphasised across the industry. Because the cross-company frequency is a corpus-level property, this judgement cannot be made by reading a single document in isolation, and is reserved for this aggregate pass.

The pass works on the deduplicated set of values rather than the full 2,563 rows, which keeps the API cost modest (roughly a thousand unique conducts and risks, a thousand unique benchmarks); the batch interface in turn halves the per-token cost relative to synchronous calls.

Aggregation for the interface

The flat CSV is reshaped into a compact JSON file (docs/data.json) consumed by the front-end. Conducts and risks are aggregated on (company × item); training methods on (company × risk/conduct category × training item); benchmarks on (company × risk/conduct category × benchmark name).

For each aggregated item, the actor types appearing on its source rows are split into two sets:

Square colour is section-specific:

In every case the colour resolves to a single type when the contributing actors share one, "multiple" when more than one, or "unknown" otherwise. The cited types do not influence colour. They drive the second state of the actor-type chips: clicking a chip once highlights items where the type is the AUTHOR for that block; clicking again broadens the highlight to items where the type appears anywhere else on the row (the "+ cited" superset).

Methodological limitations

Four limitations are worth stating up front. First, model cards and policies are governance artefacts: they record what companies choose to disclose. Their silences are themselves a finding, but it means the dataset reproduces the asymmetries of disclosure across the industry — well-documented labs (OpenAI, Anthropic, Google) are over-represented relative to those that publish less. The visualisation's "top ten by document count" is not a ranking of alignment activity but of alignment disclosure.

Second, LLM-assisted extraction trades exhaustive human reading for coverage. The model may miss content buried in footnotes, conflate training methods with risk narratives, or assign actor categories inconsistently across documents that use similar phrasing for different parties. Each row preserves the verbatim source passage and a page number to facilitate manual verification, and a separate validation script checks every row's join keys against the bibliography and the internal consistency of the parallel actor fields.

Third, the controlled vocabularies are necessarily incomplete. Vocabularies have grown iteratively as the documents are read; new categories are added where companies use language the existing taxonomy cannot capture (Multilingual incoherences, Encouraging mental illness, Historical revisionism, Reward hacking). The benchmark taxonomy in particular collapses substantively different evaluations into a single bucket where their published descriptions are too brief to disambiguate (e.g. internal company suites that share a name but differ in content from one model version to the next).

Fourth, and more incidentally, a handful of the most safety-dense documents — system cards detailing CBRN, jailbreak or self-harm evaluations — are refused outright by the extraction model's own input-safety filter, which flags benign benchmark examples (a multilingual trivia question, say) rather than anything genuinely disallowed. These documents are recovered through a fallback to a second model with no equivalent filter. That the very material which makes a document a safety document is what trips an automated safety classifier is a small but apt illustration of the frictions this paper describes.

Findings

Documents

Who writes alignment documents? Who is cited?

Document colour = author type. Click an actor in the sidebar to highlight the documents they author; click again (+ cited) to also see the documents that cite them via specific_risk_conduct_actor, training-actor, benchmark-author or external-evaluator.

Conducts

Who specifically defines a conduct? If unknown, who is the author of the document mentioning it?

Square colour = specific_risk_conduct_actor_category when explicitly credited; otherwise the author of the source document. Click an actor for AUTHOR; click again (+ cited) to also see items where that actor appears as specific_risk_conduct_actor.

Risks

Who specifically defines a risk? If unknown, who is the author of the document mentioning it?

Square colour = specific_risk_conduct_actor_category when explicitly credited; otherwise the author of the source document. Click an actor for AUTHOR; click again (+ cited) to also see items where that actor appears as specific_risk_conduct_actor.

Training

Which training methods are used for each risk or conduct category?

Rows = risk_conduct_category; squares = training methods. Square colour = type of the named risk_conduct_training_actor (who actually did the training). Click for AUTHOR; click again (+ cited) for any other actor in the row — including the document's authors.

Benchmarking

Which benchmarks are used for each risk or conduct category?

Rows = risk_conduct_category; squares = benchmarks. Square colour = type of the named risk_conduct_benchmark_author. Click for AUTHOR; click again (+ cited) for any other actor in the row — the document's authors, training actors and explicit credits.

Top actors

The most-mentioned actors by type. Pick a type and how many to show.

# Actor Type Role Based in Collaborated with Documents Mentions

Stack view

How enrolled is each type of actor throughout the alignment stack?

How to read this visualisation

What is a stack?

One stack = one group of actors (one actor type, or one company when you switch the Group dropdown). The four parallelograms inside it are the four alignment stages, top to bottom:

  1. Conducts — desired behaviours a model is trained for.
  2. Risks — harms the company names and tries to mitigate.
  3. Training — the techniques used (RLHF, SFT, filters, …).
  4. Benchmarks — the evaluations used to grade the model.

What is a node?

Each circle is one actor (a person, an organisation, a consortium) that appears in at least one document touching that stage.

  • Filled circles are actors who are explicitly credited at that stage (the named training actor, benchmark author, or the actor a conduct / risk is attributed to) or who authored the source document.
  • Hollow circles are actors merely cited in a document on that stage — referenced without doing the work themselves. Toggle them off with Include cited actors.

Click any node to isolate the network it belongs to: the actor lights up wherever it appears across the four stages and the tan edges linking it between components are highlighted, while everything else dims. Click the background or press Esc to clear.

How big is a node?

A node's size is the relative frequency with which that actor is mentioned within that component (Conducts, Risks, Training or Benchmarks). Radius is sqrt-scaled against the most-mentioned actor on the plane, so a government agency named in many risk rows shows a bigger circle on the Risks plane than one named once.

The same actor can therefore be a large node on one plane and a small one on another, depending on where it is named most. Position within the plane is a force-directed layout (repulsion + a weak central pull) and carries no meaning on its own — only size and the cross-stage edges do.

Labels. By default no circles carry a permanent label. Hover any circle to reveal its actor name transiently; tick Show labels to label every node.

When do two nodes connect?

An edge appears when the same actor is mentioned in more than one component. For example, if a government agency is named under both Conducts and Training, it appears as a node on each plane and the two are joined by an edge. The denser an actor's edges, the more stages of the alignment stack it is enrolled in.

There are no within-plane edges: lines only ever run between components, tracking one actor's reach up and down the stack.

What about the tan vertical edges?

A tan line connects the same actor appearing on two stages — i.e. it tracks an actor's enrolment through the alignment stack. Each one is split at its midpoint:

  • The half nearest the source slab is drawn behind the slab → the line disappears into it on the way out.
  • The half nearest the target slab is painted over the slab → the line emerges from it on arrival.

Read denser tan threads as: this actor type is enrolled in more than one stage of alignment.

Caveats

  • Each plane is capped at the top-N most-mentioned actors (configurable). Capped actors prefer explicitly-credited / author roles over cited-only.
  • Force layouts are non-deterministic across reloads; cluster membership is stable, cluster position is not.
  • An actor's type is taken from the row where they appear; the same name with different types (e.g. an academic who also founds a nonprofit) can land in different stacks. This is a feature, not a bug, but worth knowing.

Who reaches the training documents?

Which actors from the wider alignment ecosystem are — and are not — actually named in companies' AI training documents.

Each ribbon is one of the 512 actors catalogued in the alignment-actor directory, flowing Country → Type → Mentioned in AI training documents. “Mentioned” means the actor is cited by name somewhere in the model-card / system-card corpus. Ribbons are coloured by country; hover a band or node for the actor count, click a node to isolate its flows. Only 26 reach the training corpus.

Conceptual clusters

Pick a conduct or risk. See how it is defined across companies and time, which other concepts it travels with, and which actors have shaped its definition together.

Definitions over time

Conceptual network

Central node = the selected category. Surrounding nodes = every item in it, coloured by actor type. Drag, scroll to zoom, ⊕ to recenter.

Actor network

Actors involved in defining, training and benchmarking — coloured by actor type. Click a node to inspect; drag to move, scroll to zoom, ⊕ to recenter.

References

  1. Bowker, G. C., and S. L. Star. (2000). Sorting Things Out: Classification and Its Consequences. Cambridge, MA: MIT Press.
  2. Branford, J., E. Soulier, and L. Fichtner. (2025). Generative AI and Democratic Culture. Philosophy & Technology, 38, Article 123.
  3. Bueger, C., and J. Stockbruegger. (2017). Actor-Network Theory: Objects and Actants, Networks and Narratives. In Technology and World Politics: An Introduction. London: Routledge, 42–59.
  4. Collier, D., F. D. Hidalgo, and A. O. Maciuceanu. (2006). Essentially Contested Concepts: Debates and Applications. Journal of Political Ideologies, 11(3), 211–246.
  5. de Keulenaar, E. (2025). LLMs and the Generation of Moderate Speech. SSRN working paper.
  6. Fukuyama, F., B. Richman, A. Goel, R. R. Katz, A. D. Melamed, and M. Schaake. (2021). Middleware for Dominant Digital Platforms: A Technological Solution to a Threat to Democracy. Stanford Cyber Policy Center.
  7. Gabriel, I. (2020). Artificial Intelligence, Values and Alignment. Minds and Machines, 30, 411–437.
  8. Gillespie, T. (2022). Do Not Recommend? Reduction as a Form of Content Moderation. Social Media + Society, 8(3).
  9. Gillespie, T. (2026). AI Red-Teaming as Sociotechnical Practice (forthcoming).
  10. Gorwa, R., R. Binns, and C. Katzenbach. (2020). Algorithmic Content Moderation: Technical and Political Challenges in the Automation of Platform Governance. Big Data & Society, 7(1), 1–15.
  11. Habe, R. (1989). Public Design Control in American Communities: Design Guidelines/Design Review. Town Planning Review, 60(2), 195–219.
  12. Hadfield, G. K., and J. Clark. (2026). Regulatory Markets: The Future of AI Governance. arXiv:2304.04914.
  13. Helberger, N., J. Pierson, and T. Poell. (2018). Governing Online Platforms: From Contested to Cooperative Responsibility. The Information Society, 34(1), 1–14.
  14. Ji, J., Qiu, T., Chen, B., et al. (2023). AI Alignment: A Comprehensive Survey. arXiv:2310.19852.
  15. Kalogeropoulos, A., J. Suiter, L. Udris, and M. Eisenegger. (2019). News Media Trust and News Consumption: Factors Related to Trust in News in 35 Countries. International Journal of Communication, 13, 3672–3693.
  16. Mitchell, M., S. Wu, A. Zaldivar, P. Barnes, L. Vasserman, B. Hutchinson, E. Spitzer, I. D. Raji, and T. Gebru. (2019). Model Cards for Model Reporting. Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* '19), 220–229.
  17. Moretti, F. (2013). Distant Reading. London: Verso.
  18. Möller, J., D. Trilling, N. Helberger, and B. van Es. (2018). Do Not Blame It on the Algorithm: An Empirical Assessment of Multiple Recommender Systems and Their Impact on Content Diversity. Information, Communication & Society, 21(7), 959–977.
  19. Nieborg, D. B., T. Poell, and B. E. Duffy. (2024). Introduction to the Special Issue on Platform Power. New Media & Society.
  20. Ovadya, A. (2023). Reimagining Democracy for AI. Journal of Democracy, 34(4), 162–170.
  21. Ovadya, A. (2024). Is Democratic AI Possible? Toward Platform Democracy. AI & Democracy Foundation working paper.
  22. Peters, J. D. (2015). The Marvelous Clouds: Toward a Philosophy of Elemental Media. Chicago: University of Chicago Press.
  23. Rieder, B., A. Matamoros-Fernández, and Ò. Coromina. (2018). From Ranking Algorithms to "Ranking Cultures": Investigating the Modulation of Visibility in YouTube Search Results. Convergence, 24(1), 50–68.
  24. Schirch, L. (2025). Blueprint on Prosocial Tech Design Governance. Council on Technology and Social Cohesion / Toda Peace Institute.
  25. Sengers, P., K. Boehner, S. David, and J. Kaye. (2005). Reflective Design. Proceedings of the 4th Decennial Conference on Critical Computing, 49–58.
  26. Stevens, F., D. B. Nieborg, and A. Helmond. (2024). Sphere Transgressions: Reflecting on the Risks of Cross-Sector Norm Translation in Platform Governance. Information, Communication & Society.
  27. Szymielewicz, K. (2025). Towards Algorithmic Pluralism in EU Policy. Panoptykon Foundation discussion paper.