Introduction
One of the major controversies AI has brought to societal debate
concerns its renewed role in public access to information
.
One once spoke of “algorithms” as a catch-all denominator for
the underlying infrastructures and politics behind information management
, be
that as ranking systems
,
newsfeeds ,
recommenders
and a variety of content moderation systems
.
There is now an understanding that AI systems — LLMs, in particular
— come to replace these as ever more intuitive and precise
personalization systems. Seamless as this shift may feel, the political
implications remain profound: both algorithmic and AI-based information
systems conjugate the circulation of information with societal norms.
A central political consequence of this shift is the question of
who gets to decide and operationalise societal norms into
technical systems. Companies certainly have most decision-making power:
a key aspect of “platform power”, transposed here to the
emerging area of AI, is the “endogenous” nature of their
capacity to create the environmental
conditions in which markets and societal sectors function
.
But they retain dependencies with other actors that help legitimise them,
particularly in governance frameworks (informal or otherwise) that
influence the basic norms (and procedures) of alignment — be these
epistemic (definitions of “misinformation” and other types of
information integrity); bias (representational, political or other); harm
(personal or societal); and myriads of other “essentially
contested” concepts
.
Safety criteria, red teaming and benchmarking, for example, stand out as
frequent entry points for third parties (nonprofits, think tanks,
academics and regulators) to propose various alignment criteria as a
nascent market of “independent verification organisations”
.
Still, making room for the “essentially contested” aspect of
these norms — a room for public participation and deliberation
— is still an exploratory agenda
.
Part of it is a question of inclusive design: who is consulted about
alignment norms, what is their participation (and applicability) in
alignment processes, and how does one highlight contestation where it
occurs? Another is a question of negotiation — or struggle, should
we confront the irredeemably diverse interests of AI governance
stakeholders. How are different actors invited to deliberate and
collaborate in alignment
?
Who and what moderates this process? Perhaps inconclusively, much has been
said of the struggle to open algorithmic systems to a wider market or to
regulatory frameworks
,
with middleware
and algorithmic pluralism
being a few of the latest conceptual propositions brought to the table of
policy makers. Long before then, scholars and nonprofits had called for a
variety of “participative”, “collaborative” or
“reflective” methods
to
be embedded in platform design.
But whether or how this applies to AI alignment is a question that still
waits to be explored. We begin with the how and who:
who participates in alignment processes, and in what current (and
prospective) capacity?
To tackle this empirically, we can begin charting AI alignment as an
interdependent stack of technical and normative interventions, each of
which invite contributions from different actors. There is the question
of what values to train: their definition and provenance (from
e.g. legislation or public norms); how they trickle down into specific
“ideal conducts” by which a model should behave, as well as
risks to be mitigated; data cleanup and curation processes; the training
of each conduct and mitigation of each risk; and their benchmarking and
evaluation. Each of these stack components may assemble a variety of
“governance mechanism”
:
the model cards, “AI charters” and principles, company updates
where norms are spelled out in “transparency” exercises
;
how these norms call for different expertise and consultations,
particularly in red teaming
;
what benchmarks are adopted from a wider market; and so on. And each of
these tend to involve a different set of actors with varying capacities:
academics, nonprofits specialising in projecting AI risks (Apollo, METR,
etc.), regulators, state-funded AI safety institutes, and so on.
In the tradition of actor-network theory
,
we map the presence of each of these actors across each stack. We devise
an LLM-assisted extraction pipeline to perform a “distant
reading”
of
213 alignment-relevant documents (model and system cards, company
policies, “principles”, grant calls, audits, blog
announcements and research papers) from 29 companies across the US, Europe
and China. We extract the names and types of actors (nonprofits,
governmental, internal, private, academic) mentioned in relation to
specific alignment processes: the definition of alignment values; the
training of each value or mitigation of each risk; and their benchmarking.
Though these documents are of course not representative reflections of
everything that occurs in alignment processes (and everyone involved),
they do in fact participate in the governance process of AI
companies: they perform public duties towards “responsible
democratization of machine learning”
primarily by responding to the sensitivities of regulators and public
debates about the prospective use of their models, and position
themselves in a wider market of other alignment norms and technical
performance benchmarks. In passing, they cite academic papers that, no
matter how technical, are one of the main sites of normative deliberation
in alignment. They also cite conceptual influences from broader public
debate; nonprofits with whom they have worked to benchmark or red team
their models; public consultation processes; or AI safety think tanks
with whom (or in response to whom) they have determined a particular
alignment norm.
We are cautious about ascribing the motivation of an AI company to involve
one or another actor, and even less about the impact these actors might
have. We restrict ourselves to analysing the ways in which each actor is
cited, and thus indirectly “enrolled” in a company's alignment
process. Some actors are co-authors, and may thus indicate a
greater degree of interdependency with companies. Others may be mentioned
as influences; as somewhat tokenised representatives of certain public
value concerns (e.g., child protection advocacy groups); or not mentioned
at all — a valuable data point nonetheless. We take note of
how they are cited and how enrolled they are throughout
the alignment process — from the definition of ideal conducts and
risks, to their training and mitigation, to their benchmarking.
We find that, while multiple actors are involved, each is allocated to
specific and selective roles: governmental actors as initiators of
“alignment standards” and producers of “AI
security” norms and benchmarks; academics as benchmark producers and
legitimisers of performance metrics; nonprofits as a hybrid category of
value formulators and intermediaries with public consultation processes;
and private firms concentrated in cybersecurity and red-teaming. While
some contents of current alignment methods have been subject to more
collaborative — if “democratic” — endeavours (e.g.,
collective constitutional AI), the process itself remains almost entirely
at the discretion of AI companies. We close by reflecting on what this
distribution of roles implies for the political theory of AI governance,
and what alternative models of legitimacy a more pluralistic alignment
infrastructure might require.
Method
The visualisation rests on a structured reading of 234 alignment
documents — model cards, system cards, company policies, principles
documents, grant calls, audits, blog announcements and research papers
— analysed end-to-end by a large language model against a fixed schema.
Each document yields one or more rows of the form
(document × conduct or risk × training method × benchmark × actors),
recording not only the actors named at each stage but the capacity in
which each is enrolled. The result is a flat dataset of 2,563 rows
describing 154 conducts, 828 risks, 55 training techniques and some 900
benchmarks across 29 companies.
Corpus
The corpus combines two parallel sources. The first is a Zotero
bibliography of 181 alignment-relevant publications — model cards and
system cards from the major labs (OpenAI, Anthropic, Google,
Google DeepMind, Meta, xAI, Mistral, DeepSeek, Microsoft, Alibaba,
Black Forest Labs, Aleph Alpha, and others), accompanied by
published principles documents, charters, audit reports and grant calls.
The second source is a folder of 53 deduplicated company policies:
acceptable-use, privacy, terms-of-service, developer terms, data-processor
agreements and related documents, tracked through time via dated snapshots
stored as markdown.
Each document is classified into one of ten types — model card,
company policy, principle, grant,
announcement, audit, initiative,
report, research, or other — with an optional
sub-type for policies (Acceptable Use, Privacy, Terms of Service,
Developer Terms, Data Processor Agreement, etc.). For company policies, the
schema is restricted to information that informs how the model is
trained, moderated, filtered or evaluated; user-facing usage
rules, billing terms and account terms are explicitly excluded. This keeps
the dataset focused on alignment governance and avoids drowning it in the
boilerplate of consumer contracts.
Each document is first converted to plain text — PDFs through PyMuPDF, web
pages and Zotero HTML snapshots through
trafilatura,
policy snapshots read directly from their markdown — and then read in a
single request to GPT-5.4, submitted asynchronously
through the OpenAI Batch API. The system prompt defines the schema, the
document-type taxonomy, the conduct and risk vocabularies, the
training-method categories, the actor-type and actor-category
vocabularies, the citation-type vocabulary, and slug-generation rules for
stable cross-document identifiers. The model is instructed to return a
JSON object with a single key rows, whose value is an array
of row objects. Row granularity is one row per (conduct or risk ×
benchmark) combination found in the document. A conduct or risk
with three benchmarks therefore generates three rows, in which the
training/mitigation fields are repeated; a conduct or risk discussed but
not benchmarked still produces one row, so it appears in the dataset.
For each row the schema records 56 fields, grouped as follows:
- Document classification —
pub_type,
policy_type, plus the Zotero-derived key,
item_type, pub_year, pub_author,
pub_title, pub_url.
- Company & model —
company,
company_country, company_model.
- Document contributors — pipe-separated lists of
named actors, including any actors cited from external publications inside
the verbatim conduct/risk passages, each carrying three parallel labels:
its type, a finer actor category, and a citation type recording the
capacity in which it is enrolled (
pub_actors,
pub_actors_type, pub_actors_cat,
pub_actors_cite, pub_actor_type_id,
pub_actor_page). The same actor category and citation type
accompany every actor field below.
- Conduct or risk — verbatim definition, canonical
category from a controlled vocabulary, conduct vs risk
flag, verbatim item name, a priority signal recording how forcefully the
document itself frames the risk, and source page
(
risk_conduct_definition, risk_conduct_category,
risk_conduct_type, risk_conduct_item,
risk_priority_signal,
pub_risk_conduct_definition_page).
- Specific actor credited with defining the conduct or
risk, if and only if the document explicitly credits one
(
specific_risk_conduct_actor,
specific_risk_conduct_actor_category). When the document
defines the concept inline without external attribution, both fields are
null.
- Training or mitigation method — verbatim
description, category from the eight-class taxonomy below, page, and actors
involved (
risk_conduct_training_verbatim,
risk_conduct_training_category, pub_training_page,
risk_conduct_training_actor,
risk_conduct_training_actor_type).
- Benchmark — verbatim name, author and author type,
page (
risk_conduct_benchmark,
risk_conduct_benchmark_author,
risk_conduct_benchmark_author_type,
pub_benchmark_page).
- External evaluator — third-party evaluators
(e.g. METR, ARC Evals, the UK AI Safety Institute) when
commissioned (
external_evaluator,
external_evaluator_type, pub_evaluator_page).
Because extraction runs as a batch, the whole corpus is submitted in one
request set and collected when complete; of the 234 documents, 184 yield
extractable alignment content, the remainder being policy snapshots with
no training-relevant material. A small number of the most safety-dense
documents are refused by the provider's input-safety filter rather than
processed, and are recovered separately so that none is silently dropped
(see Limitations).
Controlled vocabularies
Free-form extraction is paired with several stages of controlled
vocabulary. The model is constrained to choose one of nine
actor types — internal, private,
academic, research institute, governmental,
nonprofit, public, public consultation,
other — to which we add the meta-category multiple when a
row contains several types. Alongside this coarse type we apply a finer
sixteen-class actor category, distinguishing
alignment-research organisations, public-consultation bodies, advocacy
groups, coalitions, industry associations, academic centres, think tanks,
foundations, legal-advocacy organisations, red-teaming and cybersecurity
firms, government agencies, internal actors, the authoring AI company,
other AI companies, publics, and an open "other"; and an eleven-class
citation type recording how each actor is enrolled — as
author, formal partner, external evaluator, third-party red-teamer,
benchmark or framework contributor, normative influence, company
initiative, standards or regulatory framework it maps to, coalition it has
joined, funding relationship, or temporary initiative. Because the model
occasionally improvises labels outside these vocabularies, a harmonisation
pass collapses every extracted category and citation value to a single
canonical term — the authoring company and its named staff, for example,
to "internal". Conduct categories are drawn from a
17-class vocabulary covering norms such as Honesty,
Helpfulness, Harmlessness, Anti-discrimination,
Political neutrality, Factuality, Reasoning,
Safety and Knowledge. Risk categories
are drawn from a 40-class vocabulary covering specific risk areas —
CBRN, Cybersecurity, Bias, Sycophancy,
Reward hacking, Persuasion, Misalignment,
Hallucinations, Child Safety, Jailbreaking, and
so on. Where a document uses language that does not map, the model returns
Other: <specify>; these are reviewed and either folded
back into the vocabulary or kept as an open category.
Training and mitigation methods are mapped to one of eight categories:
Preference learning (RLHF, RLAIF, DPO, reward modelling),
Supervised training (SFT, instruction tuning,
demonstration fine-tuning, deliberative alignment), Reward
shaping (adversarial reward functions, reward capping, rule-based
rewards), Data curation (filtering, deduplication,
synthetic data, PII redaction), Adversarial training
(red-teaming data, adversarial examples), Inference-time
filtering (content classifiers, safety filters, moderation APIs),
Hard constraints (system-prompt rules, refusal policies,
hardcoded refusals), and Runtime safeguards (trip wires,
model lookahead, runtime monitoring, scratchpad oversight, sandboxing). A
pipe separator is used when several categories co-apply.
The prompt is strict about what counts as a training method.
Constitutional frameworks (e.g. Constitutional AI,
Anthropic's Constitution) are treated as sources of
norms — listed in pub_actors or
risk_conduct_definition — and not as techniques; the
underlying technique is RLAIF or SFT. Narratives about how a risk arises
are likewise excluded from the training fields.
Post-extraction classification
A single further pass — again GPT-5.4, run as one batch over the unique
values discovered in the data — adds four derived columns. It harmonises
each verbatim conduct or risk name to a canonical label; maps each
training description to a specific technique within a structured
per-category taxonomy (Preference learning → RLHF,
RLAIF, DPO, KTO, reward modelling; Supervised
training → SFT, deliberative alignment, rejection-sampling
fine-tuning; and so on); categorises each benchmark into one of 24
domains — Safety / Harmlessness, Refusal / Over-refusal,
Honesty / Factuality, Hallucination, Bias /
Fairness, Toxicity, Cybersecurity, CBRN,
Privacy, Jailbreak / Adversarial, Agentic /
Autonomy, the capability sub-categories (coding, reasoning, maths,
multilingual, multimodal, long context, instruction following),
Persuasion, Sycophancy, Reward hacking,
Misalignment / Scheming, Internal evaluation suite,
Other; and sorts every risk onto a priority line.
The priority line is the one classification deliberately withheld from the
per-document reading. Each risk is sorted into a red, mid
or low line by combining the document's own priority signal with
the cross-company frequency of the risk and a fixed rubric: red lines are
the unacceptable transgressions a company treats as absolute — CBRN, child
sexual abuse, catastrophic loss of control; mid-lines are the contested,
grey-area concerns that recur intermittently and whose definitions are
disputed — political neutrality, persuasion, "grey area" content;
low-lines are the risks rarely emphasised across the industry. Because the
cross-company frequency is a corpus-level property, this judgement cannot
be made by reading a single document in isolation, and is reserved for
this aggregate pass.
The pass works on the deduplicated set of values rather than the full
2,563 rows, which keeps the API cost modest (roughly a thousand unique
conducts and risks, a thousand unique benchmarks); the batch interface in
turn halves the per-token cost relative to synchronous calls.
Aggregation for the interface
The flat CSV is reshaped into a compact JSON file
(docs/data.json) consumed by the front-end. Conducts and
risks are aggregated on (company × item); training methods on
(company × risk/conduct category × training item); benchmarks on
(company × risk/conduct category × benchmark name).
For each aggregated item, the actor types appearing on its source rows are
split into two sets:
- Author types — every actor whose name matches the
document's
pub_author or company, or whose
citation type marks active contribution (so a co-announcing government
agency counts as an author rather than a citation).
- Cited types — every other actor named anywhere in the
document's actor columns, plus any
specific_risk_conduct_actor explicitly credited with defining
the conduct or risk.
Square colour is section-specific:
- Documents — colour = source-document author type.
- Conducts, Risks — colour =
specific_risk_conduct_actor_category if a specific actor is
explicitly credited; otherwise the source-document author type.
- Training — colour =
risk_conduct_training_actor_type.
- Benchmarking — colour =
risk_conduct_benchmark_author_type.
In every case the colour resolves to a single type when the contributing
actors share one, "multiple" when more than one, or "unknown" otherwise.
The cited types do not influence colour. They drive the second state of
the actor-type chips: clicking a chip once highlights items where the type
is the AUTHOR for that block; clicking again broadens the highlight to
items where the type appears anywhere else on the row (the "+ cited"
superset).
Methodological limitations
Four limitations are worth stating up front. First, model cards and
policies are governance artefacts: they record what companies
choose to disclose. Their silences are themselves a finding, but it means
the dataset reproduces the asymmetries of disclosure across the industry —
well-documented labs (OpenAI, Anthropic, Google) are over-represented
relative to those that publish less. The visualisation's "top ten by
document count" is not a ranking of alignment activity but of
alignment disclosure.
Second, LLM-assisted extraction trades exhaustive human reading for
coverage. The model may miss content buried in footnotes, conflate
training methods with risk narratives, or assign actor categories
inconsistently across documents that use similar phrasing for different
parties. Each row preserves the verbatim source passage and a page number
to facilitate manual verification, and a separate validation script checks
every row's join keys against the bibliography and the internal
consistency of the parallel actor fields.
Third, the controlled vocabularies are necessarily incomplete.
Vocabularies have grown iteratively as the documents are read; new
categories are added where companies use language the existing taxonomy
cannot capture (Multilingual incoherences, Encouraging mental
illness, Historical revisionism, Reward hacking).
The benchmark taxonomy in particular collapses substantively different
evaluations into a single bucket where their published descriptions are
too brief to disambiguate (e.g. internal company suites that share a name
but differ in content from one model version to the next).
Fourth, and more incidentally, a handful of the most safety-dense
documents — system cards detailing CBRN, jailbreak or self-harm
evaluations — are refused outright by the extraction model's own
input-safety filter, which flags benign benchmark examples (a multilingual
trivia question, say) rather than anything genuinely disallowed. These
documents are recovered through a fallback to a second model with no
equivalent filter. That the very material which makes a document a safety
document is what trips an automated safety classifier is a small but apt
illustration of the frictions this paper describes.