Mapping value alignment

The alignment stack

Value alignment asks how guiding principles — fairness, accountability, transparency, safety — are embedded into technical systems. We broadly define it as the process by which AI companies decide, in conjunction with the law and various partnerships with states, alignment research organisations and publics, how their AIs should behave in relation to ideas of what “humans value”.

Research has tended to take up one part of that process at a time: red teaming, the drafting of alignment documents, the content of model cards, the actors present in them, or ethnographies of training and data work. While these interventions are important, they all become more important when joined together into the overall process of alignment, which is indivisible — from data production, to the drafting of policy documents, to training, to benchmarking and other forms of evaluation.

White background, at twice the drawn size.

Levels 2 to 4 are the ones this map reads empirically. Level 1 frames them but is not coded here.

Introduction

Alignment has become the main site of intervention for people to determine the norms and values by which AIs work. LLMs have been around for a while now, but AI is also progressively replacing what we called algorithms and algorithms research, as AI conditions and organises the way we consume information. Alignment gives a normative sense, or orientation, to how information is generated and organised by AI — to the point that we see how pressing an issue this is in many areas of power: domestic, geopolitical, and public.

We call this a “stack”, in the sense that we highlight the interdependencies between these components of alignment. These components are:

So here, in sum, we look into the operationalisation of a risk or virtue throughout the alignment stack — how their definition trickles down into training and benchmarking.

The documents of alignment

For those of us who have been studying content moderation, most of the familiar datapoints are policy documents, which stipulate what content a platform service allows and does not allow. The documents we see in alignment are a little different in kind. AI companies have the kind of policy documents mentioned above, but they also have more aspirational documents that define how models should in theory behave towards and against potential phenomena. This is why we, and many companies, use the word “risk”: it remains a potentiality that models should be trained to pre-empt in myriads of potential uses. The same applies to a virtue: it is a way that models should behave ideally in a diversity of scenarios, to the point of becoming “norm sponsors” — able to reproduce norms onto the world by educating and otherwise steering users towards virtuous choices.

Principles. First, there are, for lack of a better word, canonical “principles” in which AI companies scope their alignment. In the same way that platform policies determine the finite lines of what they can and cannot do, these chart the broad lines of an alignment project: they declare an alignment philosophy, or deontology, making decisions as to what they think is good and bad and why. xAI, for example, makes very clear that Grok should behave in ways antithetical to other models in the market: it should be “truth-speaking” against values the company deems too liberal. Some may argue that this is a particular kind of discursive work, a performative task aimed at speaking to governmental stakeholders and other actors who co-act as legal and normative institutions. But that may not just be that. In practical terms, it is the root repository of norms, definitions, scenarios and metrics by which specific virtues and risks are trained.

Announcements, initiatives and grants. Then there are documents where AI companies announce projects, points of interest, partnerships and other developments that are not always immediately reflected in model cards and principles. Though these may seem like PR exercises, they are useful datapoints in companies’ ongoing and often shifting alignment projects — efforts to afford deliberative or public inputs into AI constitutions, for instance, such as Anthropic’s Collective Constitutional AI and OpenAI’s equivalent effort.

Model cards. Finally there are, as we know, model cards: relatively standardised documents that report how a model was trained. The components are usually a stated training objective, including the risk-mitigation strategies and virtues companies want to embed in their models; the training process, meaning what training was chosen, with whom and how; and the benchmarks used to verify the effectiveness of the training. Since 2019, model cards have become an industry practice as an effort to render machine learning more transparent, in response to critiques against opaque data collection and training methods. Though they are not all exhaustive, they state some of the fundamental components of the training process more fully than documents veered towards lay audiences.

Repositories. We shouldn’t forget, as we know from machine learning literature, the world of repositories. For this specific reading we don’t take them into account, but they may well be combined with all of the documents above, since code contains various forms of decision-making.

Method

We manually collected 183 documents from eight companies around the world — American, EU and Chinese — as a sample; 170 of them carry coded statements and are read here. Different companies have different amounts of alignment documents available; Mistral, ironically, almost none. We then did a distant reading of all documents using three “readers”: two LLMs, one frontier and one baseline, and ourselves. We call it distant reading because it consists in a collective reading of specific datapoints we defined inductively. Once we read a pair of documents, we determined we wanted to systematically collect: the risk or virtue being trained, where available; the way it was trained; how it was benchmarked; and what actors were involved in each of these steps. This was first done manually with five documents from each company, and then given as an example to the two LLMs. When the LLMs agreed with a passage, that passage was kept; where they disagreed, we made the call by verifying the document manually.

Extraction proceeded in two steps. First, the model collected the document’s own expressions of ideal conducts, risks, training or mitigation methods, benchmarks and participating actors. Second, these expressions were categorised manually into controlled vocabularies, and grouped into the thematic families used below. Whether a statement is a virtue or a risk follows the document’s own framing, decided per item across all its occurrences.

Consistency. Consistency is not simply how often something is mentioned. A virtue or risk is consistent when three things hold at once: it is present and frequent across the companies’ documents (predominance); companies tend to define it in the same terms as one another (generality); and it recurs steadily inside each company’s own corpus rather than surfacing once (consistency). The three count equally, so a score out of 100 can be read directly off the length of a bar, and the bar’s three segments show which part produced it. A family reads as consistent from 45.

The corpus

Every document collected, by company and by kind.

One marker is one document, deliberately left uncoloured. Hover for its title, year, type, model and company.

Virtues

The thematic families of desirable conduct that appear in the corpus, ranked by how consistently the field states them.

Hover a bar for the families’ most frequent items.

The five most and least consistent virtue categories, among those named in at least three documents. Hover a chip for its three components and its most frequent items.

What kinds of virtues do companies train for, and how consistently are they defined?

Virtues are narrow. Of 201 coded statements of desirable conduct, more than half — 104, in 58 documents, from all eight companies — fall into a single family: behavioral alignment and control, which scores 63 and is the only virtue family the field states consistently. This is the register of helpfulness, harmlessness, refusal tone and instruction-following: virtues of a well-behaved assistant. The one other family that clears the threshold is epistemic integrity (53) — honesty, factuality, calibration — named by six companies across 32 documents.

The most consistent virtues tend to be product-oriented, like “helpfulness”. Others are hard or existential norms, like safety and compliance. And yet others are “honesty”, i.e. the capacity to represent oneself and the world truthfully to the user — to state what the model knows, what it does not, and what it is.

Less consistent are the things that tend to diverge more within and across companies: approaches to defining political neutrality, the personality of models, notions of diversity, and non-accredited knowledge.

Risks

The thematic families of harm the corpus sets out to pre-empt, ranked by the same measure.

Hover a bar for the families’ most frequent items.

The five most and least consistent risk categories, on the same terms. Hover a chip for its three components and its most frequent items.

What risks do companies set out to mitigate, and how consistently are they defined?

Risks are where the documentation actually lives. There are 1,605 coded risk statements against 201 for virtues — roughly eight to one — and where two virtue families clear the threshold, six risk families do, each of them named by all eight companies. The recurring items are concrete and shared: jailbreaks in 41 documents, cybersecurity in 32, hallucinations in 31, prompt injection in 28, child sexual abuse in 23, chemical and biological uplift in 20. This is a settled agenda, and it is the same agenda everywhere.

The most consistent risks are also existential: CBRN, cybersecurity, the question of autonomy — the danger of AIs “going rogue” — harmful content, and privacy. Bias is also an important category here, because it is general enough to be applicable to any and all kinds of excessive subjectivities.

Interestingly, this also includes over-refusal: the problem of models refusing too much from user prompts. Notions of misinformation, meanwhile, have lost a lot of political capital here, in the sense that their political liabilities tend to disfavour them.

From company to evaluation

The full stack in one figure: company, whether the statement is a virtue or a risk, its thematic family, the training applied, and the benchmark used to check it.

Flow widths preserve weighted source records, where a record spanning several training or benchmarking categories has its weight divided between those paths, so companies with more documents are not over-represented by duplication alone. Hover a flow to isolate it; the nodes it touches keep their labels while the rest fall back.

Records that disclose neither training nor benchmark are hidden by default; the sidebar restores them.

What kinds of training go with each family, and what benchmarks evaluate them?

Training splits along a simple line: virtues are taught, risks are blocked. When a company trains a virtue, it mostly uses supervised training and preference learning — methods that put a disposition into the model itself. When it handles the matching risk, those same methods drop away and filtering, hard constraints and runtime safeguards take over. These do not change the model; they sit around it and stop things at the moment of use. For content safety, filtering and hard rules lead. For catastrophic and security risks, the single heaviest path in the whole figure runs to runtime safeguards.

So the further down the stack we go, the less alignment looks like teaching a model to behave and the more it looks like fencing it in.

Benchmarking barely follows any of this. Whatever the training, the test is nearly always a task-performance benchmark — the same destination for filtering, supervised training, preference learning, hard constraints, data curation, runtime safeguards and adversarial training alike. The tests that would actually probe a safeguard, like adversarial stress tests or interactive scenarios, are used far less. A model can be trained in seven different ways and then shown to be aligned by one kind of score. What widens on the left of the figure narrows to a bottleneck on the right.

Limitations

The map captures published descriptions, which vary widely in scope, detail and genre. A company that publishes more will produce more source records; the weighting keeps duplicated records from inflating that further, but it cannot manufacture disclosure where none exists. Absence from a figure therefore means absence from the collected documentation, not necessarily absence from the model-development process. Generality is measured on the words companies use, so two firms describing the same harm in different registers will score low even where they mean the same thing. Repositories, where a good deal of decision-making is encoded, are left out of this reading entirely.