Industry insight · October 2026

Where AI Safety Research Is Moving: A Map of 1,842 Projects, 2021–2027

We classified 1,842 public AI safety projects from 2021 to 2027 into 23 research directions and 82 subtopics, to see which parts of the field are crowded, which are heating up, which have lasted, and which are still close to empty.

The short version

  • The overall balance has barely moved. Technical alignment has been about two thirds of research projects and governance about one third in both periods we compare, up to 2025 and 2026–27. The change is inside each half.
  • Two directions hold a quarter of all research projects. Interpretability (13.9%) and AI governance (10.3%) are the most crowded parts of the map, and both are losing share.
  • Model welfare and biosecurity are the fastest risers. Model welfare went from 0.6% to 4.4% of projects (×4.7) and biosecurity from 1.9% to 7.6% (×3.5). Both results hold when any single programme is removed.
  • The "catch it in the act" agenda is growing steadily. AI control (×1.4) and evaluations (×1.3) both gained share, and work on whether models notice they are being tested doubled.
  • Inside interpretability, the work is moving from dictionary learning to applied tools. Sparse-autoencoder projects fell from 32 to 10 between the two periods, while probing and tooling rose from 36 to 56.
  • Mentor supply and project volume do not line up. Evaluations has 1.7 recent mentor listings per recent project. Values and sycophancy has 0.35, and deception and scheming 0.44.
  • Several hot directions depend on one programme. 62% of model-welfare projects come from a single programme. Governance is the most evenly spread.
  • Some important problems are nearly empty. Sycophancy, chemical and nuclear risk, DNA-synthesis screening, AI companionship and military AI each have fewer than ten projects across seven years.

The field at a glance

Figure 1
Projects per year, by family
All 1,735 dated projects. 2027 includes only programmes that have already published next season's listings.

Project volume roughly doubled every year from 2022 to 2026, and 2026 alone holds 731 projects. Through all of that growth, the split between technical alignment and governance has stayed close to two to one.

Figure 2 · interactive
The AI safety project map
Each dot is one project, grouped by subtopic inside its direction. Bright dots are current or open; faded dots are past. Click a direction to see its profile, and switch periods to watch the map shift.

Figure 3 puts every measure behind the map into one table. Sort by any column, filter by family, or open a direction to see how each of its subtopics moved.

Figure 3
Direction explorer
Every measure in this report, one row per direction. Sort by any column; open a row for its subtopics, where the ≤2025 and 2026–27 columns show project counts.
Momentum = share of 2026–27 projects ÷ share up to 2025; the range shows the lowest and highest value when any single programme is left out. Mentors per project compares mentor listings from 2025 onward with 2026–27 projects.

Where the field is crowded

Interpretability is the largest direction on the map with 216 projects, 13.9% of all research projects. AI governance follows with 160 (10.3%). No other direction passes 6%. Below them sits a broad middle tier of about ten directions, each holding 4–6% of the field: agent foundations, deception and scheming, compute and security, AI control, evaluations, alignment training, forecasting, misalignment and biosecurity.

Both are still growing in absolute terms. New work is spreading across more directions, so their share of the field is falling.

Figure 4
Share of research projects by direction, all years
1,549 research projects. Hover a bar for its largest subtopics.

What is heating up, and what is cooling

Figure 5 compares each direction's share up to 2025 with its share in 2026–27. The bracket after each momentum score shows the lowest and highest values across the twenty leave-one-programme-out runs. Where the whole bracket is above or below ×1, the trend is robust.

Figure 5
Then vs now: share of research projects, up to 2025 → 2026–27
Sorted by momentum. Filled arrowheads mark trends that survive removing any single programme.
gaining sharelosing share▶ filled = robust · ▷ hollow = depends on one programme

Rising: welfare, biosecurity, control and evaluation

Model welfare and consciousness grew from 4 projects up to 2025 to 35 in 2026–27, from 0.6% to 4.4% of the field. Most of the new work asks whether models have morally relevant states and whether their reports about themselves can be trusted. Two years ago this was a fringe topic. It is now a standing research stream.

Biosecurity rose from 1.9% to 7.6% of research projects, the largest absolute gain on the map. The growth is split between policy (biodefence and pandemic preparedness, 6 → 31 projects) and technical work on AI-enabled biological misuse (3 → 25). Even with its largest single source removed, biosecurity's momentum stays above ×2.

AI control and evaluations are rising more slowly but just as consistently. Control went from 4.2% to 6.0% (×1.4, robust), with growth in both protocol design and monitoring. Evaluations went from 4.4% to 5.7% (×1.3, robust). Within evaluations the fastest-growing topics are evaluation methodology and agentic, AI-R&D-style evaluations. Next to these, projects on situational and evaluation awareness, whether a model notices it is being tested, doubled from 10 to 21. Together these suggest the field increasingly assumes it will need to catch problems in deployed systems, not only prevent them in training.

Field-building (×1.9, robust) and societal impacts (×1.2) are also rising. In societal impacts, almost all the growth is work on gradual disempowerment and concentration of power, up from 4 projects to 15.

Rising, but on narrow ground

Values and model behaviour (×1.6) and scalable oversight (×1.3) gained share, but neither trend survives removing every single programme. Values work is growing mostly through model personas and character training. Sycophancy, a widely discussed failure mode, has only three projects on the whole map.

Cooling: governance, multi-agent, robustness, national security

Four directions lost share robustly. AI governance fell from 14.2% to 8.2% (×0.59), driven mostly by fewer projects on safety frameworks, audits and standards (29 → 14). Part of that activity seems to have moved into compute, hardware and security, which grew to 6.2%. Its fastest subtopic is technical verification of international agreements, up from 7 projects to 21. Governance research is becoming more technical, not disappearing.

Multi-agent and cooperative AI (×0.70), adversarial robustness (×0.70) and national security (×0.67) all held roughly flat in project counts while the field around them grew. For multi-agent risk this is notable: agents that hold credentials and money and interact with each other are being deployed quickly, and research attention to how they fail together is falling as a share.

Emerging subtopics

Subtopic-level shifts are sharper than direction-level ones. Figure 6 shows the subtopics with the largest relative growth among those with at least five projects in 2026–27.

Figure 6
Fastest-growing subtopics: projects up to 2025 vs 2026–27
Counts, not shares. Ordered by growth ratio.
up to 20252026–27

Interpretability shows the clearest reshuffle inside a single direction. Sparse-autoencoder projects fell from 32 to 10 and steering projects from 20 to 8. Meanwhile interpretability methods and tooling grew from 21 to 35, and probing and representation reading from 15 to 21. The field appears to be moving from building dictionaries of features to applying cheaper techniques that are useful in deployment.

In alignment training, character and constitution work grew from 2 projects to 11. Unlearning and tamper resistance shrank from 8 to 2.

Where mentors are waiting

Mentor listings measure what senior researchers want to supervise. Comparing them with the volume of projects gives a rough picture of supply and demand. Figure 7 plots each direction's recent mentor listings (2025 onward) against its recent projects (2026–27).

Figure 7
Mentor listings vs projects, recent cohorts
Above the 1:1 line, mentors outnumber projects. Below it, projects outnumber mentors. Dot size = all-time projects.

Mentor-rich directions. Evaluations (1.7 recent mentor listings per recent project), adversarial robustness (1.6), AI control, misalignment and forecasting (about 1.45 each), and multi-agent work (1.4) all have more mentors than projects. For applicants, these are directions where supervision is relatively easy to find. Robustness and multi-agent work stand out: mentors keep offering supervision there even as the topics lose share of projects.

Mentor-thin directions. Values and model behaviour (0.35), deception and scheming (0.44), field-building (0.65), alignment training (0.71) and biosecurity (0.75) have far more projects than mentors. In values and scheming, much of the work is likely supervised by people outside the mentor lists or by peers. For experienced researchers, these are the directions where offering to mentor would add the most.

What lasts

Momentum shows change. Persistence shows commitment. Figure 8 shows each direction's share of each year's research projects. A row that stays filled across the width is a direction the field has kept coming back to.

Figure 8
Share of each year's research projects, by direction
n = dated research projects that year. 2021–22 rest on few projects. Hover a cell for counts.

Four directions appear in all seven years: AI governance, forecasting and strategy, alignment training and biosecurity. Interpretability, evaluations, societal impacts, national security, values and multi-agent work appear in six. These are the field's backbone, directions that predate the recent expansion and are still active.

AI control is the youngest large direction. It first appears in 2024, and by 2026 it is the fourth-largest research direction by project count. Model welfare also first appears in 2024. Its rise is the fastest on the map, but it has the shortest track record.

Concentration: which trends rest on one programme

A direction carried by many programmes is unlikely to vanish if one of them changes course. A direction that relies on a single programme is more fragile. Figure 9 shows, for each direction, how many programmes contribute projects and what share comes from the largest one.

Figure 9
Share of each direction's projects from its largest programme
Labels show how many programmes contribute. Programmes are not named.

Three of the hottest directions are the most concentrated. 62% of model-welfare projects, 56% of values projects and 51% of field-building projects come from one programme each. Governance (26% from its largest programme), national security (27%), scalable oversight (29%) and compute (30%) are the most evenly spread. Deception and scheming, and societal impacts, are each run by 11 of the 20 programmes, the widest reach on the map.

For funders, concentration is useful to know. A fast-rising direction that depends on one programme may need a second programme behind it before the trend can be counted on.

Figure 10
Share of projects published as papers
Only papers count as published here. Policy work often ends in reports, briefs or placements, so lower shares do not mean less output.

Bridges between directions

Many projects span two directions. When a project was labelled with a secondary direction, we counted a link between the two. Figure 11 shows the strongest links.

Figure 11
The strongest links between research directions
Line width = number of projects that span both directions. Hover a direction to isolate its links.

The strongest bridge connects misalignment and alignment training (37 projects): showing a failure and training it away are increasingly done in the same project. Governance connects to national security (34) and compute (33), and forecasting feeds into governance (26). On the technical side, evaluations link closely with deception and scheming (33), and interpretability links with agent foundations (29), alignment training (28) and scheming (22). Interpretability is the field's most connected technical direction. Its tools are increasingly used to answer other directions' questions.

Thin spots

Some subtopics have almost no projects despite clear relevance. These are the subtopics with the fewest projects across all seven years.

Figure 12
The emptiest subtopics on the map
All-time project count, with the split up to 2025 → 2026–27.

A few patterns stand out. Biosecurity is growing fast, but almost all of that growth is in policy and evaluation work. Chemical and nuclear risk, and DNA-synthesis screening, together have eight projects. Sycophancy and AI companionship are among the most visible everyday failures of deployed assistants, yet together they have eight projects. Agent security and sandboxing, adversarial attacks, steganographic reasoning and cyber-offense evaluations are each in single or low double digits, though agent security is starting to grow.

The practice layer

Not every project is research. Two programmes place people in policy institutions and newsrooms. We counted 93 policy placements, mostly in think tanks and NGOs (52), Congress (21) and the executive branch (20), and 134 journalism records, including 55 newsroom placements and coverage of the AI industry, AI policy and societal harms.

This layer runs on a different calendar, since placements are listed before their work is published, so it sits outside the momentum analysis. It is how safety expertise reaches legislation and public understanding, and one of the strongest cross-links in the data joins policy placements with biosecurity.

What this means

For people choosing a direction

The crowded directions are not closed, but competition for the best-known projects is strongest there. Mentor-rich directions with stable or falling share, such as robustness, multi-agent work, evaluations and forecasting, offer supervision with less competition. Fast-rising, mentor-thin directions such as values, welfare and scheming give more room for independent work and less guidance.

For mentors and senior researchers

The biggest gap between projects and supervision is in values and model behaviour, deception and scheming, and biosecurity. Mentoring there would cover more unmet need than in evaluations or control.

For programme designers

Several rising directions depend on one programme. A second home for model welfare, values and field-building would make those trends less fragile. The thin spots in Figure 12 are natural candidates for a targeted stream.

For funders

The strongest signal is a rising share, broad programme reach and enough mentors, all at once. AI control and evaluations meet all three. Multi-agent risk is the clearest case of a direction whose share is falling just as the real-world exposure it studies is growing.

Method

Figure 13
From public listings to a map
What we collected, how it was cleaned and coded, and what we measure.
Agreement is between two independent passes that each assigned every project a primary direction and subtopic without seeing the other's labels. The 281 disagreements were settled by a third pass. "Including secondary" counts a match when one pass's primary label was the other's secondary label.

Collection. We gathered 1,846 project records and 1,038 mentor listings from programme websites, published project lists, research libraries, and papers and posts that credit a programme. Four records were excluded during cleaning, leaving 1,842 projects.

Taxonomy. We sorted projects into 23 directions and 82 subtopics, grouped into three families. Technical alignment covers work on the models themselves. Governance and strategy covers rules, compute, security, forecasting, biosecurity and societal impact. Practice covers policy placements, journalism and field-building. The full taxonomy is in the interactive map (Figure 2).

Double coding. Two separate passes labelled every project from its title and description without seeing each other's labels. They agreed on the direction for 90.5% of projects and on the exact subtopic for 84.8%. A third pass settled every disagreement.

Time. We compare two periods: projects dated up to 2025 (642 research projects) and projects dated 2026 or 2027 (801). We measure each direction's share of its period rather than raw counts, because 2026 listings are much larger than earlier ones. Momentum is the 2026–27 share divided by the earlier share, so ×2 means a direction takes twice as much of the field as before.

Robustness. One large programme could dominate a period on its own, so we recomputed every momentum score twenty times, each time leaving out one programme. A trend is called robust only when it points the same way in all twenty runs.

Scope. Research measures exclude policy placements, journalism and untitled records. The report reproduces no project titles or names; every number is an aggregate.

Limits

Credits and sources

This report would not exist without the programmes that publish their work openly, and the researchers and mentors whose projects make up the map. We thank all of them. The map counts their work; it does not rank or judge any individual project.

Corrections. If you run one of these programmes, worked on a project, or think a category is wrong, write to info@cygnuxlabs.com. We will correct the data and publish the changes along with the reasons for them.