AI Safety and SecurityField Map
Every hex is a problem someone could work on. Gold means a little work, green active and dark forest busy, and right now every problem has someone on it. A small token says who holds it: a pink ring for frontier labs only, a blue dot for another field.
Lenses
Map of every problem in the field. Arrow keys move between problems, Enter opens one, Escape goes back up a level.
Arrow keys move · Enter opens · Esc goes back
Every problem, as a list
The Model
The thing being built
Alignment
- Scalable oversight7 organisations
- Reward hacking and specification gaming5 organisations
- Scheming and deceptive alignment6 organisations
- Goal misgeneralisation and emergent misalignment5 organisations
- Unlearning and tamper-resistant safeguards5 organisations
- Alignment for superintelligence3 organisations
- Automated alignment research2 organisations
- Conceptual foundations of alignment7 organisations
Interpretability
- Understanding model internals10 organisations
- Chain-of-thought monitorability8 organisations
- Probes and internal monitoring5 organisations
- Interpretability benchmarks3 organisations
- Interpretability applied to safety5 organisations
Control
- Control evaluations6 organisations
- Control protocols and monitoring7 organisations
- Agent sandboxing and containment4 organisations
- Formal verification and guaranteed-safe AI2 organisations
- Detecting loss-of-control incidents3 organisations
Evaluations
- Dangerous capability evaluations10 organisations
- Propensity and character evaluations7 organisations
- Science of evaluations5 organisations
- Benchmarks and evaluation tooling7 organisations
- Sandbagging and capability elicitation5 organisations
Model security
- Weights and infrastructure security7 organisations
- Agent and application security16 organisations
- Supply chain and model integrity5 organisations
- Insider threat3 organisations
- Jailbreaks and adversarial robustness15 organisations
Reliability and failures
- Agent reliability and error propagation3 organisations
- Hallucination and factual errors2 organisations
- AI failures in critical systems2 organisations
- Safety-critical AI assurance3 organisations
Developer assurance
- Frontier safety frameworks12 organisations
- Safety cases7 organisations
- Third-party auditing9 organisations
- Developer incident reporting4 organisations
- Internal safety governance3 organisations
- Automated AI R&D and internal deployment4 organisations
- Deployment safeguards against misuse3 organisations
Harmful use
People and states using it to cause harm
Cyberattacks and critical infrastructure
- Offensive cyber uplift4 organisations
- AI for cyber defence6 organisations
- Autonomous cyber operations5 organisations
- Detecting AI-driven attacks in the wild5 organisations
- Attacks on critical infrastructure7 organisations
- Cyber defence where capacity is thin6 organisations
Biological and chemical weapons
- Bioweapons knowledge uplift8 organisations
- Biological design tool misuse6 organisations
- DNA synthesis screening4 organisations
- Chemical weapons uplift4 organisations
- AI for biosecurity defence6 organisations
Military AI and autonomous weapons
- Autonomous weapons and their regulation6 organisations
- AI targeting and decision support4 organisations
- Military AI testing and assurance4 organisations
Influence operations and persuasion
- Covert influence campaigns10 organisations
- AI persuasion at scale1 organisation
- Violent extremism and terrorism4 organisations
Fraud and abuse
- Fraud and impersonation7 organisations
- AI-generated child sexual abuse material8 organisations
- Non-consensual intimate imagery6 organisations
Surveillance and repression
- Mass surveillance7 organisations
- AI-enabled censorship4 organisations
- Predictive policing and social scoring3 organisations
- Surveillance technology export controls3 organisations
Open-weight models
- Safeguard removal from open weights2 organisations
- Uncensored model supply3 organisations
- Open-weight release: marginal risk and benefit6 organisations
Society and Government
The world adapting
Structural risk
- Race dynamics4 organisations
- Concentration of power2 organisations
- Gradual disempowerment2 organisations
- Multi-agent risks5 organisations
- AI and strategic stability4 organisations
Public policy and regulation
- Domestic AI legislation10 organisations
- Liability for AI harms2 organisations
- Accountability and liability for agent actions1 organisation
- Standards and certification8 organisations
- Whistleblower protection3 organisations
- Government technical capacity5 organisations
- Agent identity, visibility and protocols2 organisations
International governance
- International agreements5 organisations
- International institutions8 organisations
- US-China and cross-bloc dialogue7 organisations
- Inclusive participation in global governance4 organisations
Compute governance
- Chip export controls and diversion8 organisations
- Verifying AI agreements6 organisations
- Cloud compute know-your-customer3 organisations
Economic transition
- Labour market disruption8 organisations
- Sharing the gains from AI3 organisations
- Adapting safety nets2 organisations
Epistemics and information
- AI and collective reasoning4 organisations
- Deliberation tools5 organisations
- Content provenance and deepfake detection5 organisations
- Decision quality in key institutions4 organisations
Ethics, fairness and welfare
- Fairness and bias8 organisations
- AI welfare and moral status5 organisations
- Machine ethics3 organisations
- Human-AI relationships and wellbeing4 organisations
Meta
What the other three stand on
Funding
- Concentration and diversity of funding22 organisations
- Funding infrastructure13 organisations
Talent and training
- Early-career pipeline capacity21 organisations
- On-ramps for experienced professionals8 organisations
- Senior and specialist capacity6 organisations
- Talent matchmaking6 organisations
Evidence and strategy
- Capability forecasting8 organisations
- Incident tracking4 organisations
- Incident investigation2 organisations
- Living reviews of AI risk3 organisations
- Capability demonstrations for decision-makers4 organisations
- Threat modelling4 organisations
- Macrostrategy and prioritisation11 organisations
Watchdogs and accountability
- Lab accountability tracking6 organisations
- Whistleblower support4 organisations
Communications and public engagement
- Public opinion and message research4 organisations
- Explaining AI risk to the public6 organisations
- Advocacy8 organisations
Convenings and coordination
- Convenings and events9 organisations
- Regional community hubs10 organisations
Research tooling and public goods
- Shared compute, data and environments5 organisations
- AI uplift for safety research2 organisations