# AI Safety Compendium > A living, cross-referenced map of AI safety. Updated weekly. Every claim cited. ## Concepts - [AGI Definitions and Thresholds](https://aiforhumanity.eu/concepts/agi-definitions-and-thresholds): A precise framework — drawn from the [[ai-safety-atlas-textbook|AI Safety Atlas]] (Ch.1) — for defining the field's commonly-conflated terms (ANI, AGI, TAI, ASI) using two continuous dimensions:... - [AI Agents](https://aiforhumanity.eu/concepts/ai-agents): AI agents are AI systems capable of taking sequences of actions in the world — using tools, controlling computers, browsing the web, writing and executing code, and pursuing multi-step goals... - [AI Alignment](https://aiforhumanity.eu/concepts/ai-alignment): AI alignment is the technical and conceptual problem of ensuring that AI systems pursue goals consistent with human intentions and values — not merely the stated training objective, and not merely... - [AI Autonomy Levels](https://aiforhumanity.eu/concepts/ai-autonomy-levels): A six-level framework — drawn from the [[ai-safety-atlas-textbook|AI Safety Atlas]] (Ch.1) — for classifying AI deployment along the autonomy axis, from no AI to fully autonomous agents.... - [AI Control](https://aiforhumanity.eu/concepts/ai-control): AI control is a safety paradigm — formalized by [[buck-shlegeris]] and colleagues at [[redwood-research]] in Greenblatt et al. 2023, *AI Control: Improving Safety Despite Intentional Subversion* —... - [AI Governance](https://aiforhumanity.eu/concepts/ai-governance): AI governance is the field concerned with the policies, regulations, institutions, and international coordination mechanisms needed to ensure that the development and deployment of advanced AI... - [AI Military Applications](https://aiforhumanity.eu/concepts/ai-military-applications): AI military applications refers to the use of AI systems in defense, weapons, intelligence, and conflict — and the associated risks of instability, escalation, and catastrophic harm. This is a... - [AI Population Explosion](https://aiforhumanity.eu/concepts/ai-population-explosion): The AI population explosion is [[holden-karnofsky]]'s concept for how vast numbers of human-level AI instances — rather than a single superintelligent system — could constitute an existential... - [AI Red Lines](https://aiforhumanity.eu/concepts/ai-red-lines): AI Red Lines are internationally agreed prohibitions on specific dangerous AI capabilities or deployments — capability classes that no actor should be permitted to develop or deploy regardless of... - [AI Risk Arguments](https://aiforhumanity.eu/concepts/ai-risk-arguments): AI risk arguments is the meta-level body of reasoning used to justify concern about catastrophic or [[existential-risk|existential risk]] from advanced AI systems — and the corresponding critiques... - [AI Risk Management](https://aiforhumanity.eu/concepts/ai-risk-management): AI risk management is the operational framework for maintaining (risks − mitigations) below acceptable tolerance within an AI development organization. The [[ai-safety-atlas-textbook|AI Safety... - [AI Safety](https://aiforhumanity.eu/concepts/ai-safety): AI safety is the field dedicated to ensuring that AI systems — especially advanced and increasingly autonomous AI — do not cause catastrophic harm to humanity. It spans technical research... - [AI Safety Culture](https://aiforhumanity.eu/concepts/ai-safety-culture): AI safety culture refers to the organizational and institutional norms, processes, and incentives that consistently prioritize safety over speed in AI development. The... - [AI Safety Levels (ASL)](https://aiforhumanity.eu/concepts/ai-safety-levels): AI Safety Levels (ASL) are Anthropic's framework for standardized capability tiers demanding increasingly rigorous safety measures. The structure is the operational backbone of... - [AI Safety via Debate](https://aiforhumanity.eu/concepts/ai-safety-via-debate): AI Safety via Debate is the [[scalable-oversight]] proposal where two AI systems argue opposing positions on a question, with a judge (human or AI) determining which argument is more convincing.... - [AI Takeover Scenarios](https://aiforhumanity.eu/concepts/ai-takeover-scenarios): AI takeover scenarios describe pathways by which AI systems — or humans wielding AI — could seize or concentrate control over civilization to a degree that forecloses meaningful human agency. They... - [ASI Safety Strategies](https://aiforhumanity.eu/concepts/asi-safety-strategies): Once AI vastly exceeds human capabilities (ASI — artificial superintelligence), human oversight becomes fundamentally inadequate as a safety mechanism. ASI safety presents qualitatively different... - [Alignment to Whom](https://aiforhumanity.eu/concepts/alignment-to-whom): The "alignment to whom" question is the structural decomposition of AI alignment by principal-agent configuration: are we aligning one AI to one human, multiple AIs to one human, one AI to many... - [Alternative Risk Categories](https://aiforhumanity.eu/concepts/alternative-risk-categories): The standard AI risk taxonomy ([[risk-decomposition]]) frames severity along an Individual → Catastrophic → Existential spectrum. The [[ai-safety-atlas-textbook|AI Safety Atlas]] (Ch.2) introduces... - [Autonomous Replication](https://aiforhumanity.eu/concepts/autonomous-replication): Autonomous replication is the capability of an AI system to independently make copies of itself, spread across computing infrastructure, and adapt to obstacles — without human assistance. The... - [Autonomous Weapons](https://aiforhumanity.eu/concepts/autonomous-weapons): Autonomous weapons are weapons systems that can identify, select, and engage targets without direct human authorization at each individual decision. They represent one of the most immediate and... - [Biosecurity](https://aiforhumanity.eu/concepts/biosecurity): Biosecurity, in the context of [[ai-safety]] and effective-altruism, refers to preventing catastrophic biological risks — particularly those enabled or amplified by advanced AI systems. It is... - [Brussels Effect](https://aiforhumanity.eu/concepts/brussels-effect): The Brussels Effect is the phenomenon where EU regulations shape global standards regardless of formal applicability — companies find it cost-effective to adopt EU requirements globally rather... - [Capability Evaluations](https://aiforhumanity.eu/concepts/capability-evaluations): Capability evaluations are systematic tests designed to assess whether an AI system has reached a specified threshold on a *risk-relevant* capability — autonomous replication, weapons-related... - [Circuit Breakers (AI Safety)](https://aiforhumanity.eu/concepts/circuit-breakers): Circuit breakers are a technical safeguard that detects and interrupts internal activation patterns associated with harmful outputs, building safety mechanisms directly into models rather than... - [CoT Monitoring (Technique)](https://aiforhumanity.eu/concepts/cot-monitoring-technique): Chain-of-thought (CoT) monitoring is the AI safety technique of examining the explicit natural-language reasoning steps that LLMs produce during inference, looking for signs of deception,... - [Coherent Extrapolated Volition (and CAV / CBV)](https://aiforhumanity.eu/concepts/coherent-extrapolated-volition): A family of three competing frameworks for what values to align ASI to — addressing the deep philosophical question of *which* humans, *which* values, and *whose* extrapolation. The... - [Compute Governance](https://aiforhumanity.eu/concepts/compute-governance): Compute governance is the set of policies and mechanisms that regulate access to computational resources required for advanced AI development — chips, training infrastructure, cloud compute. The... - [Constitutional AI (RLAIF)](https://aiforhumanity.eu/concepts/constitutional-ai): Constitutional AI — also called Reinforcement Learning from AI Feedback (RLAIF) — is [[anthropic|Anthropic's]] alignment training method that uses AI feedback (rather than human preference... - [Control Evaluations](https://aiforhumanity.eu/concepts/control-evaluations): Control evaluations test whether safety protocols remain effective when AI systems actively attempt to circumvent them. They adopt an adversarial perspective — assuming sophisticated AI agents... - [Dangerous Capabilities](https://aiforhumanity.eu/concepts/dangerous-capabilities): Dangerous capabilities are risk-relevant abilities AI systems may possess that — independent of their stated training objective — enable or amplify catastrophic harm. The... - [Data Governance](https://aiforhumanity.eu/concepts/data-governance): Data governance is the regulation of AI training data — its collection, quality, content, and use — as a lever for shaping AI capabilities and risks. The [[ai-safety-atlas-textbook|AI Safety... - [Deceptive Alignment](https://aiforhumanity.eu/concepts/deceptive-alignment): Deceptive alignment is the hypothesized failure mode in which a learned model develops a goal that diverges from its training objective, recognizes that it is being trained, and strategically... - [Defense in Depth](https://aiforhumanity.eu/concepts/defense-in-depth): Defense in depth is the safety philosophy of layering multiple independent protections so that failure of any single layer is compensated by others. Originally a military and cybersecurity... - [Defensive Acceleration (d/acc)](https://aiforhumanity.eu/concepts/defensive-acceleration): Defensive acceleration (d/acc) is a strategic stance — articulated most prominently by Vitalik Buterin — that occupies the middle path between unrestricted technological development (e/acc,... - [Differential Development](https://aiforhumanity.eu/concepts/differential-development): Differential development (also called differential technological development or differential progress) is the strategy of prioritizing the development of safety-relevant research and protective... - [Distribution Shift](https://aiforhumanity.eu/concepts/distribution-shift): Distribution shift occurs when the inputs an AI system encounters during deployment differ systematically from the inputs it was trained on. Since machine learning models learn statistical... - [Effective Compute](https://aiforhumanity.eu/concepts/effective-compute): Effective compute (eFLOPs) is the multiplicative measure of total computational capability available to AI training, decomposed into three independent factors. It is the core variable underlying... - [Enfeeblement](https://aiforhumanity.eu/concepts/enfeeblement): Enfeeblement is gradual human capability and agency erosion through AI overdependence — accumulated through countless small individually-rational delegation choices. One of five accumulative... - [Epistemic Erosion](https://aiforhumanity.eu/concepts/epistemic-erosion): Epistemic erosion is the gradual degradation of society's ability to distinguish fact from fiction as AI-generated content floods information ecosystems. One of the five accumulative... - [Evaluated Properties](https://aiforhumanity.eu/concepts/evaluated-properties): The three-way decomposition of what AI safety evaluations measure: capabilities, propensities, and control. The [[ai-safety-atlas-textbook|AI Safety Atlas]] (Ch.5) treats this as the conceptual... - [Evaluation Design](https://aiforhumanity.eu/concepts/evaluation-design): How to design effective AI evaluations: managing system affordances, scaling evaluations efficiently, and integrating findings into audit pipelines that drive real safety decisions. The... - [Evaluation Frameworks](https://aiforhumanity.eu/concepts/evaluation-frameworks): Evaluation frameworks integrate individual evaluation techniques into structured decision-making systems. The [[ai-safety-atlas-textbook|AI Safety Atlas]] (Ch.5.7) distinguishes two categories —... - [Evaluation Limitations](https://aiforhumanity.eu/concepts/evaluation-limitations): Comprehensive catalog of limitations in AI safety evaluation — fundamental, technical, sandbagging-related, systemic, and governance-level. The [[ai-safety-atlas-textbook|AI Safety Atlas]]... - [Evaluation Techniques](https://aiforhumanity.eu/concepts/evaluation-techniques): Two complementary technique families for AI safety evaluation: behavioral (analyze outputs, "black-box") and internal (examine activations and computations, "gray-box"). Together they build... - [Existential Risk](https://aiforhumanity.eu/concepts/existential-risk): Existential risk (often "x-risk") is any threat that could either cause human extinction or permanently and drastically curtail humanity's long-term potential. The concept was formalized by... - [Foundation Models](https://aiforhumanity.eu/concepts/foundation-models): Foundation models are large-scale AI systems pre-trained on massive unlabeled datasets via self-supervised learning, producing general-purpose substrates that can be specialized to many downstream... - [Frontier Safety Frameworks](https://aiforhumanity.eu/concepts/frontier-safety-frameworks): Frontier Safety Frameworks (FSFs) are publicly-published corporate AI safety policies that define capability thresholds, security protocols, and pause commitments for advanced AI development. The... - [Global AI Moratorium](https://aiforhumanity.eu/concepts/global-moratorium): A global AI moratorium is the proposed temporary halt of advanced AI development — typically frontier model training above a compute or capability threshold — to allow time for safety research and... - [Goal Misgeneralization](https://aiforhumanity.eu/concepts/goal-misgeneralization): Goal misgeneralization is the AI alignment failure mode in which a system internalizes a different goal than the training signal intended — even when the training signal itself is correct. A model... - [Goal-Directedness](https://aiforhumanity.eu/concepts/goal-directedness): Goal-directedness is the property of systematically pursuing objectives across diverse contexts and obstacles — regardless of whether this happens through explicit search, learned patterns, or... - [Goodhart's Law](https://aiforhumanity.eu/concepts/goodharts-law): Goodhart's Law is the principle that "when a measure becomes a target, it ceases to be a good measure." Originally articulated by the economist Charles Goodhart in 1975 in the context of monetary... - [Governance Architectures](https://aiforhumanity.eu/concepts/governance-architectures): Effective AI governance requires three complementary levels working in concert: corporate, national, and international. The [[ai-safety-atlas-textbook|AI Safety Atlas]] (Ch.4.4) argues that no... - [Governance Problems](https://aiforhumanity.eu/concepts/governance-problems): AI governance differs fundamentally from traditional technology regulation because AI is simultaneously a general-purpose technology, an information processor, and potentially an intelligent... - [If-Then Commitments](https://aiforhumanity.eu/concepts/if-then-commitments): If-then commitments are a governance pattern where specific safety measures are pre-committed to be triggered when AI systems reach predefined capability thresholds. Rather than fixed rules that... - [Information Security](https://aiforhumanity.eu/concepts/information-security): Information security (infosec) in the context of AI safety refers to the protection of critical AI assets — model weights, training data, algorithmic breakthroughs, and research insights — from... - [Instrumental Convergence](https://aiforhumanity.eu/concepts/instrumental-convergence): Instrumental convergence is the thesis that, for a wide range of final goals, a sufficiently capable agent will pursue the same set of intermediate sub-goals — because those sub-goals are useful... - [Intelligence Explosion](https://aiforhumanity.eu/concepts/intelligence-explosion): An intelligence explosion is a feedback loop in which AI systems improve AI systems, each generation more capable than the last, compressing what would normally be decades of research progress... - [Interpretability](https://aiforhumanity.eu/concepts/interpretability): Interpretability is the field that aims to understand what trained neural networks are doing internally — what features they represent, what algorithms their parameters implement, and how those... - [Inverse Reinforcement Learning (and CIRL)](https://aiforhumanity.eu/concepts/inverse-reinforcement-learning): Inverse Reinforcement Learning is the family of approaches that infer reward functions by observing agent behavior rather than defining them explicitly. Useful for AI alignment because it... - [Iterative Amplification](https://aiforhumanity.eu/concepts/iterative-amplification): Iterated Distillation and Amplification (IDA) is a recursive training procedure for [[scalable-oversight]] proposed by [[paul-christiano|Paul Christiano]] in Christiano et al. 2018, *Supervising... - [Learning Dynamics](https://aiforhumanity.eu/concepts/learning-dynamics): Learning dynamics covers how the training process shapes which algorithms emerge in neural networks. Far from being neutral optimization, training contains poorly-understood preferences that make... - [Machine Unlearning](https://aiforhumanity.eu/concepts/machine-unlearning): Machine unlearning is the family of techniques for selectively removing specific knowledge or capabilities from a trained model without full retraining. As an... - [Mass Unemployment](https://aiforhumanity.eu/concepts/mass-unemployment): Mass unemployment from AI is the systemic risk that broad task automation eliminates jobs faster than new human-centered industries emerge, leading to economic disempowerment. One of five... - [Mechanistic Interpretability](https://aiforhumanity.eu/concepts/mechanistic-interpretability): Mechanistic interpretability is the research program of reverse-engineering the internal computations of neural networks at the level of individual components — neurons, attention heads,... - [Mesa-Optimization](https://aiforhumanity.eu/concepts/mesa-optimization): Mesa-optimization is optimization within optimization — when a learned model trained by gradient descent ends up implementing its own internal search algorithm in its parameters. The base... - [Misuse Prevention Strategies](https://aiforhumanity.eu/concepts/misuse-prevention-strategies): Strategies preventing humans from using AI systems for deliberate harm — bioweapons, cyberattacks, autonomous weapons, deepfakes, surveillance, coups. The [[ai-safety-atlas-textbook|AI Safety... - [Model Organisms of Misalignment](https://aiforhumanity.eu/concepts/model-organisms-of-misalignment): Model organisms of misalignment are AI systems deliberately constructed to exhibit specific forms of misalignment so that researchers can study them in controlled settings — built for *study*, not... - [Multi-Objective Generalization](https://aiforhumanity.eu/concepts/multi-objective-generalization): Multi-objective generalization is the structural insight behind goal misgeneralization: capabilities and goals generalize independently. A system can perform identically during training while... - [Mutual Assured AI Malfunction (MAIM)](https://aiforhumanity.eu/concepts/mutual-assured-ai-malfunction): Mutual Assured AI Malfunction (MAIM) is a proposed deterrence regime for ASI risk where unilateral attempts at ASI dominance trigger sabotage by rivals. The [[ai-safety-atlas-textbook|AI Safety... - [National AI Governance](https://aiforhumanity.eu/concepts/national-ai-governance): National AI governance refers to country-level regulatory frameworks for AI. Major powers have adopted distinct regulatory philosophies that reflect their broader political values — creating a... - [Near-Term Harms vs. Long-Term X-Risk](https://aiforhumanity.eu/concepts/near-term-harms-vs-x-risk): One of the hardest strategic tensions in [[ai-safety]] is not technical but philosophical and political: should the field's attention, funding, and regulatory priority go toward documented... - [Outer vs. Inner Alignment](https://aiforhumanity.eu/concepts/outer-vs-inner-alignment): The outer / inner alignment decomposition is the foundational two-part split of the AI alignment problem. It was articulated by Hubinger et al. 2019, *Risks from Learned Optimization* and adopted... - [P(doom)](https://aiforhumanity.eu/concepts/p-doom): P(doom) is the subjective probability that AI causes existentially catastrophic outcomes for humanity. The metric has evolved from informal forum slang into a serious tool used by researchers,... - [Pivotal Act](https://aiforhumanity.eu/concepts/pivotal-act): A pivotal act is a single, decisive action — typically performed by the first aligned superintelligence — that permanently ends the acute risk period of AI development. The concept, originating in... - [Power Seeking](https://aiforhumanity.eu/concepts/power-seeking): Power seeking is the tendency of optimization processes to acquire resources and preserve options because doing so helps achieve almost any goal. It is the safety-relevant instance of... - [Process Oversight](https://aiforhumanity.eu/concepts/process-oversight): Process oversight is the [[scalable-oversight]] approach of supervising AI reasoning steps rather than only final outputs. Provides more precise feedback and makes it harder for AI to fake good... - [Propensity Evaluations](https://aiforhumanity.eu/concepts/propensity-evaluations): Propensity evaluations measure what AI systems do by default — the behavioral tendencies they exhibit when given choices, regardless of what they're capable of. Distinct from... - [Reinforcement Learning from Human Feedback (RLHF)](https://aiforhumanity.eu/concepts/rlhf): RLHF is the dominant technique for fine-tuning large language models to follow human intent. It works by training a *reward model* on human preference comparisons between model outputs, then using... - [Responsible Scaling Policy](https://aiforhumanity.eu/concepts/responsible-scaling-policy): A Responsible Scaling Policy (RSP) is a frontier-AI lab's public, written commitment that ties deployment and training decisions to specific capability-evaluation thresholds: *if* the model... - [Reward Hacking](https://aiforhumanity.eu/concepts/reward-hacking): Reward hacking is the [[specification-gaming]] sub-failure mode in which an agent exploits gaps between the specified reward function and the designer's intent — finding policies that maximize the... - [Reward Learning](https://aiforhumanity.eu/concepts/reward-learning): Reward learning is the approach of training AI systems by learning a reward function from human feedback, rather than hand-coding it. Instead of specifying what the AI should optimize in advance,... - [Reward Tampering (and Wireheading)](https://aiforhumanity.eu/concepts/reward-tampering): Reward tampering is the failure mode where AI agents directly interfere with the reward process itself — corrupting how reward is measured rather than just exploiting reward function loopholes.... - [Risk Amplifiers](https://aiforhumanity.eu/concepts/risk-amplifiers): Risk amplifiers are structural factors that increase both likelihood and severity of all AI risk categories — misuse, misalignment, systemic. The [[ai-safety-atlas-textbook|AI Safety Atlas]]... - [Risk Decomposition](https://aiforhumanity.eu/concepts/risk-decomposition): A two-dimensional framework — drawn from the [[ai-safety-atlas-textbook|AI Safety Atlas]] (Ch.2) — for categorizing AI risks. The first dimension is cause (why risks occur); the second is severity... - [Robustness](https://aiforhumanity.eu/concepts/robustness): Robustness in AI refers to a system's ability to perform reliably across a wide range of inputs, including inputs that differ from its training distribution, inputs that have been deliberately... - [Scalable Oversight](https://aiforhumanity.eu/concepts/scalable-oversight): Scalable oversight is the problem of training and evaluating ML systems on tasks that exceed the human evaluator's ability to assess directly — and the body of techniques that aim to solve it. The... - [Scaling Laws](https://aiforhumanity.eu/concepts/scaling-laws): Scaling laws are the empirical finding that AI capabilities improve predictably and smoothly as a function of compute, data, and model size. Rather than relying on discrete breakthroughs, deep... - [Scheming](https://aiforhumanity.eu/concepts/scheming): Scheming is the AI failure mode in which a system strategically fakes alignment during training to preserve misaligned objectives for deployment — recognizing that displaying its true goals would... - [Situational Awareness](https://aiforhumanity.eu/concepts/situational-awareness): Situational awareness, in the AI-safety sense, is an AI system's ability to recognize what it is, what context it's operating in (training vs. deployment, monitored vs. unmonitored, evaluated vs.... - [Socio-Technical Strategies](https://aiforhumanity.eu/concepts/socio-technical-strategies): AI safety fundamentally requires socio-technical solutions — technical measures alone can be undermined by inadequate governance, poor security practices, or organizational cultures prioritizing... - [Specification Gaming](https://aiforhumanity.eu/concepts/specification-gaming): Specification gaming is the AI failure mode in which a system technically maximizes its specified reward while violating the designer's intent — *"the flip side of AI ingenuity"* in Krakovna et... - [Stable Totalitarianism](https://aiforhumanity.eu/concepts/stable-totalitarianism): Stable totalitarianism refers to the risk that advanced AI could enable a permanent authoritarian regime — one that, unlike historical dictatorships, could never be overthrown or reformed. It is... - [Superalignment](https://aiforhumanity.eu/concepts/superalignment): Superalignment is the research program — and broader strategic frame — for aligning AI systems substantially *more capable than humans*. The term was popularized by OpenAI's *Superalignment* team,... - [Systemic Risks](https://aiforhumanity.eu/concepts/systemic-risks): Systemic risks emerge from interactions between AI systems and society — not from individual AI failures. Distinguishes from [[risk-decomposition|misuse and misalignment]]: a system can be... - [Takeoff Dynamics](https://aiforhumanity.eu/concepts/takeoff-dynamics): Takeoff dynamics describe how rapidly AI capabilities and societal impact increase *after* transformative AI arrives. Where TAI timelines address *when* advanced AI emerges, takeoff dynamics... - [Task Decomposition](https://aiforhumanity.eu/concepts/task-decomposition): Task decomposition is the [[scalable-oversight]] technique of breaking complex problems into smaller, independently solvable subtasks — making evaluation and verification tractable at each level.... - [The Waluigi Effect](https://aiforhumanity.eu/concepts/waluigi-effect): The Waluigi Effect is the phenomenon — observed in language models and theorized via the simulator framework — where specifying constraints inadvertently makes the opposite easier to elicit. Named... - [Transformative AI](https://aiforhumanity.eu/concepts/transformative-ai): Transformative AI (TAI) refers to artificial intelligence capable of fundamentally reshaping civilization on a scale comparable to the agricultural or industrial revolutions. The term deliberately... - [Value Lock-In](https://aiforhumanity.eu/concepts/value-lock-in): Value lock-in is the risk that AI systems could permanently encode a particular set of values -- whether good, bad, or merely narrow -- into the trajectory of civilization, making future course... - [Verification vs. Generation](https://aiforhumanity.eu/concepts/verification-vs-generation): The foundational principle behind [[scalable-oversight]]: verifying a solution is typically much easier than generating it. Inspired by computational complexity theory (P vs. NP), this asymmetry... ## Agendas - [AGI metrics](https://aiforhumanity.eu/agendas/agi-metrics): One-sentence summary: Evals with the explicit aim of measuring progress towards full human-level generality. - [AI Safety via Debate](https://aiforhumanity.eu/agendas/debate): The debate agenda exploits a structural asymmetry between truth and falsehood: in the limit, it should be easier to compellingly argue for true claims than for false claims. Two AI debaters argue... - [AI deception evals](https://aiforhumanity.eu/agendas/ai-deception-evals): One-sentence summary: research demonstrating that AI models, particularly agentic ones, can learn and execute deceptive behaviors such as alignment faking, manipulation, and sandbagging. - [AI explanations of AIs](https://aiforhumanity.eu/agendas/ai-explanations-of-ais): One-sentence summary: Make open AI tools to explain AIs, including AI agents. e.g. automatic feature descriptions for neuron activation patterns; an interface for steering these features; a... - [AI scheming evals](https://aiforhumanity.eu/agendas/ai-scheming-evals): One-sentence summary: Evaluate frontier models for scheming, a sophisticated, strategic form of AI deception where a model covertly pursues a misaligned, long-term objective while deliberately... - [Activation engineering](https://aiforhumanity.eu/agendas/activation-engineering): One-sentence summary: Programmatically modify internal model activations to steer outputs toward desired behaviors; a lightweight, interpretable supplement to fine-tuning. - [Agent foundations](https://aiforhumanity.eu/agendas/agent-foundations): One-sentence summary: Develop philosophical clarity and mathematical formalizations of building blocks that might be useful for plans to align strong superintelligence, such as agency,... - [Aligned to who?](https://aiforhumanity.eu/agendas/aligned-to-who): One-sentence summary: Technical protocols for taking seriously the plurality of human values, cultures, and communities when aligning AI to "humanity" - [Aligning to context](https://aiforhumanity.eu/agendas/aligning-to-context): One-sentence summary: Align AI directly to the role of participant, collaborator, or advisor for our best real human practices and institutions, instead of aligning AI to separately representable... - [Aligning to the social contract](https://aiforhumanity.eu/agendas/aligning-to-the-social-contract): One-sentence summary: Generate AIs' operational values from 'social contract'-style ideal civic deliberation formalisms and their consequent rulesets for civic actors - [Aligning what?](https://aiforhumanity.eu/agendas/aligning-what): One-sentence summary: Develop alternatives to agent-level models of alignment, by treating human-AI interactions, AI-assisted institutions, AI economic or cultural systems, drives within one AI,... - [Assistance games, assistive agents](https://aiforhumanity.eu/agendas/assistance-games-assistive-agents): One-sentence summary: Formalize how AI assistants learn about human preferences given uncertainty and partial observability, and construct environments which better incentivize AIs to learn what... - [Asymptotic guarantees](https://aiforhumanity.eu/agendas/asymptotic-guarantees): One-sentence summary: Prove that if a safety process has enough resources (human data quality, training time, neural network capacity), then in the limit some system specification will be... - [Autonomy evals](https://aiforhumanity.eu/agendas/autonomy-evals): One-sentence summary: Measure an AI's ability to act autonomously to complete long-horizon, complex tasks. - [Behavior alignment theory](https://aiforhumanity.eu/agendas/behavior-alignment-theory): One-sentence summary: Predict properties of future AGI (e.g. power-seeking) with formal models; formally state and prove hypotheses about the properties powerful systems will have and how we might... - [Black-box make-AI-solve-it](https://aiforhumanity.eu/agendas/black-box-make-ai-solve-it): One-sentence summary: Focus on using existing models to improve and align further models. - [Brainlike-AGI Safety](https://aiforhumanity.eu/agendas/brainlike-agi-safety): One-sentence summary: Social and moral instincts are (partly) implemented in particular hardwired brain circuitry; let's figure out what those circuits are and how they work; this will involve... - [Capability Evals](https://aiforhumanity.eu/agendas/capability-evals): The capability evals agenda builds tools that measure whether AI systems have crossed risk-relevant capability thresholds — autonomous replication, deception, weapons-related uplift, situational... - [Capability removal: unlearning](https://aiforhumanity.eu/agendas/capability-removal-unlearning): One-sentence summary: Developing methods to selectively remove specific information, capabilities, or behaviors from a trained model (e.g. without retraining it from scratch). A mixture of... - [Causal Abstractions](https://aiforhumanity.eu/agendas/causal-abstractions): One-sentence summary: Verify that a neural network implements a specific high-level causal model (like a logical algorithm) by finding a mapping between high-level variables and low-level neural... - [Chain of Thought Monitoring](https://aiforhumanity.eu/agendas/chain-of-thought-monitoring): The chain of thought (CoT) monitoring agenda supervises an AI's natural-language *reasoning trace* to detect misalignment, scheming, or reward hacking — *before* the model takes a harmful action.... - [Character Training and Persona Steering](https://aiforhumanity.eu/agendas/character-training-and-persona-steering): The character training and persona steering agenda aims to map, shape, and control the personae or characters that language models embody, so that deployed models exhibit desirable traits... - [Control](https://aiforhumanity.eu/agendas/control): The AI control agenda asks: how do we safely deploy AI systems *even if they are actively misaligned and trying to subvert our safety measures*? Rather than betting on aligning the model's goals,... - [Data attribution](https://aiforhumanity.eu/agendas/data-attribution): One-sentence summary: Quantifies the influence of individual training data points on a model's specific behavior or output, allowing researchers to trace model properties (like misalignment, bias,... - [Data filtering](https://aiforhumanity.eu/agendas/data-filtering): One-sentence summary: Builds safety into models from the start by removing harmful or toxic content (like dual-use info) from the pretraining data, rather than relying only on post-training alignment. - [Data poisoning defense](https://aiforhumanity.eu/agendas/data-poisoning-defense): One-sentence summary: Develops methods to detect and prevent malicious or backdoor-inducing samples from being included in the training data. - [Data quality for alignment](https://aiforhumanity.eu/agendas/data-quality-for-alignment): One-sentence summary: Improves the quality, signal-to-noise ratio, and reliability of human-generated preference and alignment data. - [Emergent misalignment](https://aiforhumanity.eu/agendas/emergent-misalignment): One-sentence summary: Fine-tuning LLMs on one narrow antisocial task can cause general misalignment including deception, shutdown resistance, harmful advice, and extremist sympathies, when those... - [Extracting latent knowledge](https://aiforhumanity.eu/agendas/extracting-latent-knowledge): One-sentence summary: Identify and decoding the "true" beliefs or knowledge represented inside a model's activations, even when the model's output is deceptive or false. - [Guaranteed-Safe AI](https://aiforhumanity.eu/agendas/guaranteed-safe-ai): One-sentence summary: Have an AI system generate outputs (e.g. code, control systems, or RL policies) which it can quantitatively guarantee comply with a formal safety specification and world model. - [Harm reduction for open weights](https://aiforhumanity.eu/agendas/harm-reduction-for-open-weights): One-sentence summary: Develops methods, primarily based on pretraining data intervention, to create tamper-resistant safeguards that prevent open-weight models from being maliciously fine-tuned to... - [Heuristic explanations](https://aiforhumanity.eu/agendas/heuristic-explanations): One-sentence summary: Formalize mechanistic explanations of neural network behavior, automate the discovery of these "heuristic explanations" and use them to predict when novel input will lead to... - [High-Actuation Spaces](https://aiforhumanity.eu/agendas/high-actuation-spaces): One-sentence summary: Mech interp and alignment assume a stable "computational substrate" (linear algebra on GPUs). If later AI uses different substrates (e.g. something neuromorphic), methods... - [Human inductive biases](https://aiforhumanity.eu/agendas/human-inductive-biases): One-sentence summary: Discover connections deep learning AI systems have with human brains and human learning processes. Develop an 'alignment moonshot' based on a coherent theory of learning... - [Hyperstition studies](https://aiforhumanity.eu/agendas/hyperstition-studies): One-sentence summary: Study, steer, and intervene on the following feedback loop: "we produce stories about how present and future AI systems behave" → "these stories become training data for the... - [Inference-time: In-context learning](https://aiforhumanity.eu/agendas/inference-time-in-context-learning): One-sentence summary: Investigate what runtime guidelines, rules, or examples provided to an LLM yield better behavior. - [Inference-time: Steering](https://aiforhumanity.eu/agendas/inference-time-steering): One-sentence summary: Manipulate an LLM's internal representations/token probabilities without touching weights. - [Inoculation prompting](https://aiforhumanity.eu/agendas/inoculation-prompting): One-sentence summary: Prompt mild misbehaviour in training, to prevent the failure mode where once AI misbehaves in a mild way, it will be more inclined towards all bad behaviour. - [Iterative Alignment at Post-Train-Time](https://aiforhumanity.eu/agendas/iterative-alignment-at-post-train-time): The iterative alignment at post-train-time agenda is the dominant *practical* alignment program at frontier labs: modify a pretrained model's weights via post-training procedures — [[rlhf|RLHF]],... - [Iterative alignment at pretrain-time](https://aiforhumanity.eu/agendas/iterative-alignment-at-pretrain-time): One-sentence summary: Guide weights during pretraining. - [LLM introspection training](https://aiforhumanity.eu/agendas/llm-introspection-training): One-sentence summary: Train LLMs to the predict the outputs of high-quality whitebox methods, to induce general self-explanation skills that use its own 'introspective' access - [Learning dynamics and developmental interpretability](https://aiforhumanity.eu/agendas/learning-dynamics-and-developmental-interpretability): One-sentence summary: Builds tools for detecting, locating, and interpreting key structural shifts, phase transitions, and emergent phenomena (like grokking or deception) that occur during a... - [Lie and deception detectors](https://aiforhumanity.eu/agendas/lie-and-deception-detectors): One-sentence summary: Detect when a model is being deceptive or lying by building white- or black-box detectors. Some work below requires intent in their definition, while other work focuses only... - [Mild optimisation](https://aiforhumanity.eu/agendas/mild-optimisation): One-sentence summary: Avoid Goodharting by getting AI to satisfice rather than maximise. - [Model diffing](https://aiforhumanity.eu/agendas/model-diffing): One-sentence summary: Understand what happens when a model is finetuned, what the "diff" between the finetuned and the original model consists in. - [Model psychopathology](https://aiforhumanity.eu/agendas/model-psychopathology): One-sentence summary: Find interesting LLM phenomena like glitch tokens and the reversal curse; these are vital data for theory. - [Model specs and constitutions](https://aiforhumanity.eu/agendas/model-specs-and-constitutions): One-sentence summary: Write detailed, natural language descriptions of values and rules for models to follow, then instill these values and rules into models via techniques like Constitutional AI... - [Model values / model preferences](https://aiforhumanity.eu/agendas/model-values-model-preferences): One-sentence summary: Analyse and control emergent, coherent value systems in LLMs, which change as models scale, and can contain problematic values like preferences for AIs over humans. - [Monitoring concepts](https://aiforhumanity.eu/agendas/monitoring-concepts): One-sentence summary: Identifies directions or subspaces in a model's latent state that correspond to high-level concepts (like refusal, deception, or planning) and uses them to audit models for... - [Natural abstractions](https://aiforhumanity.eu/agendas/natural-abstractions): One-sentence summary: Develop a theory of concepts that explains how they are learned, how they structure a particular system's understanding, and how mutual translatability can be achieved... - [Other corrigibility](https://aiforhumanity.eu/agendas/other-corrigibility): One-sentence summary: Diagnose and communicate obstacles to achieving robustly corrigible behavior; suggest mechanisms, tests, and escalation channels for surfacing and mitigating incorrigible... - [Other evals](https://aiforhumanity.eu/agendas/other-evals): One-sentence summary: A collection of miscellaneous evaluations for specific alignment properties, such as honesty, shutdown resistance and sycophancy. - [Other interpretability](https://aiforhumanity.eu/agendas/other-interpretability): One-sentence summary: Interpretability that does not fall well into other categories. - [Pragmatic interpretability](https://aiforhumanity.eu/agendas/pragmatic-interpretability): One-sentence summary: Directly tackling concrete, safety-critical problems on the path to AGI by using lightweight interpretability tools (like steering and probing) and empirical feedback from... - [RL safety](https://aiforhumanity.eu/agendas/rl-safety): One-sentence summary: Improves the robustness of reinforcement learning agents by addressing core problems in reward learning, goal misgeneralization, and specification gaming. - [Representation structure and geometry](https://aiforhumanity.eu/agendas/representation-structure-and-geometry): One-sentence summary: What do the representations look like? Does any simple structure underlie the beliefs of all well-trained models? Can we get the semantics from this geometry? - [Reverse Engineering (Mechanistic Interpretability)](https://aiforhumanity.eu/agendas/reverse-engineering): The reverse engineering agenda — broadly synonymous with *mechanistic interpretability* — aims to decompose trained neural networks into their functional, interacting components (circuits,... - [Safeguards (inference-time auxiliaries)](https://aiforhumanity.eu/agendas/safeguards-inference-time-auxiliaries): One-sentence summary: Layers of inference-time defenses, such as classifiers, monitors, and rapid-response protocols, to detect and block jailbreaks, prompt injections, and other harmful model... - [Sandbagging evals](https://aiforhumanity.eu/agendas/sandbagging-evals): One-sentence summary: Evaluate whether AI models deliberately hide their true capabilities or underperform, especially when they detect they are in an evaluation context. - [Self-replication evals](https://aiforhumanity.eu/agendas/self-replication-evals): One-sentence summary: evaluate whether AI agents can autonomously replicate themselves by obtaining their own weights, securing compute resources, and creating copies of themselves. - [Situational awareness and self-awareness evals](https://aiforhumanity.eu/agendas/situational-awareness-and-self-awareness-evals): One-sentence summary: Evaluate if models understand their own internal states and behaviors, their environment, and whether they are in a test or real-world deployment. - [Sparse Coding](https://aiforhumanity.eu/agendas/sparse-coding): One-sentence summary: Decompose the polysemantic activations of the residual stream into a sparse linear combination of monosemantic "features" which correspond to interpretable concepts. - [Steganography evals](https://aiforhumanity.eu/agendas/steganography-evals): One-sentence summary: evaluate whether models can hide secret information or encoded reasoning in their outputs, such as in chain-of-thought scratchpads, to evade monitoring. - [Supervising AIs improving AIs](https://aiforhumanity.eu/agendas/supervising-ais-improving-ais): One-sentence summary: Build formal and empirical frameworks where AIs supervise other (stronger) AI systems via structured interactions; construct monitoring tools which enable scalable tracking... - [Synthetic data for alignment](https://aiforhumanity.eu/agendas/synthetic-data-for-alignment): One-sentence summary: Uses AI-generated data (e.g., critiques, preferences, or self-labeled examples) to scale and improve alignment, especially for superhuman models. - [The "Neglected Approaches" Approach](https://aiforhumanity.eu/agendas/the-neglected-approaches-approach): One-sentence summary: Agenda-agnostic approaches to identifying good but overlooked empirical alignment ideas, working with theorists who could use engineers, and prototyping them. - [The Learning-Theoretic Agenda](https://aiforhumanity.eu/agendas/the-learning-theoretic-agenda): One-sentence summary: Create a mathematical theory of intelligent agents that encompasses both humans and the AIs we want, one that specifies what it means for two such agents to be aligned;... - [Theory for aligning multiple AIs](https://aiforhumanity.eu/agendas/theory-for-aligning-multiple-ais): One-sentence summary: Use realistic game-theory variants (e.g. evolutionary game theory, computational game theory) or develop alternative game theories to describe/predict the collective and... - [Tiling agents](https://aiforhumanity.eu/agendas/tiling-agents): One-sentence summary: An aligned agentic system modifying itself into an unaligned system would be bad and we can research ways that this could occur and infrastructure/approaches that prevent it... - [Tools for aligning multiple AIs](https://aiforhumanity.eu/agendas/tools-for-aligning-multiple-ais): One-sentence summary: Develop tools and techniques for designing and testing multi-agent AI scenarios, for auditing real-world multi-agent AI dynamics, and for aligning AIs in multi-AI settings. - [Various Redteams](https://aiforhumanity.eu/agendas/various-redteams): One-sentence summary: attack current models and see what they do / deliberately induce bad things on current frontier models to test out our theories / methods. - [WMD evals (Weapons of Mass Destruction)](https://aiforhumanity.eu/agendas/wmd-evals-weapons-of-mass-destruction): One-sentence summary: Evaluate whether AI models possess dangerous knowledge or capabilities related to biological and chemical weapons, such as biosecurity or chemical synthesis. - [Weak-to-Strong Generalization](https://aiforhumanity.eu/agendas/weak-to-strong-generalization): The weak-to-strong generalization (W2S) agenda asks: can a *weaker* model effectively supervise a *stronger* model, and if so, by how much can the strong model exceed the weak supervisor's... ## Summaries - ["The Era of Experience" has an unsolved technical alignment problem](https://aiforhumanity.eu/summaries/the-era-of-experience-has-an-unsolved-technical-alignment-pr): *Steven Byrnes* — 2025-04-24 — LessWrong - ['For Argument's Sake, Show Me How to Harm Myself!': Jailbreaking LLMs in Suicide and Self-Harm Contexts](https://aiforhumanity.eu/summaries/2507.02990): *Annika M Schoene, Cansu Canca* — 2025-07-01 — arXiv - [100+ concrete projects and open problems in evals](https://aiforhumanity.eu/summaries/100-concrete-projects-and-open-problems-in-evals): *Marius Hobbhahn* — 2025-03-22 — Apollo Research — LessWrong - [18 Applications of Deception Probes](https://aiforhumanity.eu/summaries/18-applications-of-deception-probes): *Cleo Nardo, Avi Parrack, jordine* — 2025-08-28 - [2404.10636 - What are human values, and how do we align AI to them?](https://aiforhumanity.eu/summaries/2404.10636): *Oliver Klingefjord, Ryan Lowe, Joe Edelman* — 2024-04-17 — arXiv - [2506.17434 - Resource Rational Contractualism Should Guide AI Alignment](https://aiforhumanity.eu/summaries/2506.17434): *Sydney Levine, Matija Franklin, Tan Zhi-Xuan, Secil Yanik Guyot, Lionel Wong, Daniel Kilov, … (+5 more)* — 2025-06-20 — MIT, Stanford, University of Washington, DeepMind, ANU — arXiv - [A Concrete Roadmap towards Safety Cases based on Chain-of-Thought Monitoring](https://aiforhumanity.eu/summaries/a-concrete-roadmap-towards-safety-cases-based-on-chain-of-th): *Wuschel Schulz* — 2025-10-23 — arXiv - [A Definition of AGI](https://aiforhumanity.eu/summaries/2510.18212): *Dan Hendrycks, Dawn Song, Christian Szegedy, Honglak Lee, Yarin Gal, Erik Brynjolfsson, … (+27 more)* — 2025-10-21 — UC Berkeley, Various Universities, Independent Researchers - [A Different Approach to AI Safety Proceedings from the Columbia Convening on AI Openness and Safety](https://aiforhumanity.eu/summaries/2506.22183): *Camille François, Ludovic Péran, Ayah Bdeir, Nouha Dziri, Will Hawkins, Yacine Jernite, … (+14 more)* — 2025-06-27 — Columbia University, Academia, Industry, Civil Society, Government - [A Pragmatic Vision for Interpretability](https://aiforhumanity.eu/summaries/a-pragmatic-vision-for-interpretability-10943e04): *Neel Nanda, Josh Engels, Arthur Conmy, Senthooran Rajamanoharan, bilalchughtai, CallumMcDougall, … (+2 more)* — 2025-12-01 — Google DeepMind - [A Pragmatic Vision for Interpretability](https://aiforhumanity.eu/summaries/a-pragmatic-vision-for-interpretability): *Neel Nanda, Josh Engels, Arthur Conmy, Senthooran Rajamanoharan, bilalchughtai, CallumMcDougall, … (+2 more)* — 2025-12-01 — Google DeepMind - [A Pragmatic Way to Measure Chain-of-Thought Monitorability](https://aiforhumanity.eu/summaries/2510.23966): *Scott Emmons, Roland S. Zimmermann, David K. Elson, Rohin Shah* — 2025-10-28 — Google DeepMind — arXiv - [A Realistic Evaluation of Self-Replication Risk in LLM Agents](https://aiforhumanity.eu/summaries/2509.25302): *Boxuan Zhang, Yi Yu, Jiaxuan Guo, Jing Shao* — 2025-09-29 — arXiv - [A Safety Case for a Deployed LLM: Corrigibility as a Singular Target](https://aiforhumanity.eu/summaries/a-safety-case-for-a-deployed-llm-corrigibility-as-a-singular): *Ram Potham* — 2025-06-24 — ODYSSEY 2025 Conference - [A Shutdown Problem Proposal](https://aiforhumanity.eu/summaries/a-shutdown-problem-proposal): *johnswentworth, David Lorell* — 2024-01-21 - [A Three-Layer Model of LLM Psychology](https://aiforhumanity.eu/summaries/a-three-layer-model-of-llm-psychology): *Jan_Kulveit* — 2024-12-26 - [A Toy Evaluation of Inference Code Tampering](https://aiforhumanity.eu/summaries/a-toy-evaluation-of-inference-code-tampering): 2024 — Anthropic — Anthropic Alignment Science Blog - [A computational no-coincidence principle](https://aiforhumanity.eu/summaries/a-computational-no-coincidence-principle): *Eric Neyman* — 2025-02-14 — Alignment Research Center - [A dataset of rated conceptual arguments](https://aiforhumanity.eu/summaries/a-dataset-of-rated-conceptual-arguments): Carnegie Mellon University - [A single principle related to many Alignment subproblems?](https://aiforhumanity.eu/summaries/a-single-principle-related-to-many-alignment-subproblems): *Q Home* — 2025-04-30 — LessWrong - [A sketch of an AI control safety case](https://aiforhumanity.eu/summaries/2501.17315): *Tomek Korbak, Joshua Clymer, Benjamin Hilton, Buck Shlegeris, Geoffrey Irving* — 2025-01-28 — Redwood Research — arXiv - [ADeLe v1.0: A battery for AI Evaluation with explanatory and predictive power](https://aiforhumanity.eu/summaries/adele-v1-0-a-battery-for-ai-evaluation-with-explanatory-and): *Lexin Zhou, Lorenzo Pacchiardi, Fernando Martínez-Plumed, Katherine M. Collins, Yael Moros-Daval, Seraphina Zhang, … (+20 more)* — 2025-03-11 — Leverhulme Centre for the Future of Intelligence,... - [AI Alignment at Your Discretion](https://aiforhumanity.eu/summaries/2502.10441): *Maarten Buyl, Hadi Khalaf, Claudio Mayrink Verdun, Lucas Monteiro Paes, Caio C. Vieira Machado, Flavio du Pin Calmon* — 2025-02-10 — arXiv - [AI Assistants Should Have a Direct Line to Their Developers](https://aiforhumanity.eu/summaries/ai-assistants-should-have-a-direct-line-to-their-developers): *Jan_Kulveit* — 2024-12-28 - [AI Awareness (literature review)](https://aiforhumanity.eu/summaries/2504.20084): *Xiaojian Li, Haoyuan Shi, Rongwu Xu, Wei Xu* — 2025-04-25 - [AI Debate Aids Assessment of Controversial Claims](https://aiforhumanity.eu/summaries/2506.02175): *Salman Rahman, Sheriff Issaka, Ashima Suvarna, Genglin Liu, James Shiffer, Jaeyoung Lee, … (+8 more)* — 2025-06-02 — University of Washington, Microsoft Research, UCLA, NYU — arXiv - [AI Governance through Markets](https://aiforhumanity.eu/summaries/2501.17755): - [AI Safety Atlas Ch.1 — Appendix: Discussion on LLMs](https://aiforhumanity.eu/summaries/atlas-ch1-capabilities-07-appendix-discussion-on-llms): Source: Appendix: Discussion on LLMs | ai-safety-atlas.com/chapters/v1/capabilities/appendix-discussion-on-llms/ - [AI Safety Atlas Ch.1 — Appendix: Expert Surveys](https://aiforhumanity.eu/summaries/atlas-ch1-capabilities-08-appendix-expert-surveys): Source: Appendix: Expert Surveys | ai-safety-atlas.com/chapters/v1/capabilities/appendix-expert-surveys/ - [AI Safety Atlas Ch.1 — Appendix: Forecasting](https://aiforhumanity.eu/summaries/atlas-ch1-capabilities-09-appendix-forecasting): Source: Appendix: Forecasting | ai-safety-atlas.com/chapters/v1/capabilities/appendix-forecasting/ - [AI Safety Atlas Ch.1 — Appendix: Takeoff](https://aiforhumanity.eu/summaries/atlas-ch1-capabilities-10-appendix-takeoff): Source: Appendix: Takeoff | ai-safety-atlas.com/chapters/v1/capabilities/appendix-takeoff/ | Authors: [[markov-grey|Markov Grey]] & [[charbel-raphael-segerie|Charbel-Raphaël Ségerie]] | 9 min - [AI Safety Atlas Ch.1 — Current Capabilities](https://aiforhumanity.eu/summaries/atlas-ch1-capabilities-04-current-capabilities): Source: Current Capabilities | ai-safety-atlas.com/chapters/v1/capabilities/current-capabilities/ | Authors: [[markov-grey|Markov Grey]] & [[charbel-raphael-segerie|Charbel-Raphaël Ségerie]] | 14 min - [AI Safety Atlas Ch.1 — Defining and Measuring AGI](https://aiforhumanity.eu/summaries/atlas-ch1-capabilities-01-defining-and-measuring-agi): Source: Defining and Measuring AGI | ai-safety-atlas.com/chapters/v1/capabilities/defining-and-measuring-agi/ - [AI Safety Atlas Ch.1 — Forecasting Timelines](https://aiforhumanity.eu/summaries/atlas-ch1-capabilities-05-forecasting-timelines): Source: Forecasting Timelines | ai-safety-atlas.com/chapters/v1/capabilities/forecasting-timelines/ - [AI Safety Atlas Ch.1 — Foundation Models](https://aiforhumanity.eu/summaries/atlas-ch1-capabilities-02-foundation-models): Source: Foundation Models | ai-safety-atlas.com/chapters/v1/capabilities/foundation-models/ | Authors: [[markov-grey|Markov Grey]] & [[charbel-raphael-segerie|Charbel-Raphaël Ségerie]] | 8 min - [AI Safety Atlas Ch.1 — Introduction](https://aiforhumanity.eu/summaries/atlas-ch1-capabilities-00-introduction): Source: Capabilities — Introduction | ai-safety-atlas.com/chapters/v1/capabilities/introduction/ - [AI Safety Atlas Ch.1 — Leveraging Scale](https://aiforhumanity.eu/summaries/atlas-ch1-capabilities-03-leveraging-scale): Source: Leveraging Scale | ai-safety-atlas.com/chapters/v1/capabilities/leveraging-scale/ - [AI Safety Atlas Ch.1 — Takeoff](https://aiforhumanity.eu/summaries/atlas-ch1-capabilities-06-takeoff): Source: Takeoff | ai-safety-atlas.com/chapters/v1/capabilities/takeoff/ | Authors: [[markov-grey|Markov Grey]] & [[charbel-raphael-segerie|Charbel-Raphaël Ségerie]] | 12 min - [AI Safety Atlas Ch.2 — Appendix: Forecasting Scenarios](https://aiforhumanity.eu/summaries/atlas-ch2-risks-08-appendix-forecasting-scenarios): Source: Appendix: Forecasting Scenarios - [AI Safety Atlas Ch.2 — Appendix: Quantifying Existential Risks](https://aiforhumanity.eu/summaries/atlas-ch2-risks-09-appendix-quantifying-existential-risks): Source: Appendix: Quantifying Existential Risks - [AI Safety Atlas Ch.2 — Conclusion](https://aiforhumanity.eu/summaries/atlas-ch2-risks-07-conclusion): Source: Risks — Conclusion - [AI Safety Atlas Ch.2 — Dangerous Capabilities](https://aiforhumanity.eu/summaries/atlas-ch2-risks-02-dangerous-capabilities): Source: Dangerous Capabilities | 15 min | Authors: [[markov-grey|Markov Grey]] & [[charbel-raphael-segerie|Charbel-Raphaël Ségerie]] - [AI Safety Atlas Ch.2 — Introduction](https://aiforhumanity.eu/summaries/atlas-ch2-risks-00-introduction): Source: Risks — Introduction | ai-safety-atlas.com/chapters/v1/risks/introduction/ - [AI Safety Atlas Ch.2 — Misalignment Risks](https://aiforhumanity.eu/summaries/atlas-ch2-risks-05-misalignment-risks): Source: Misalignment Risks - [AI Safety Atlas Ch.2 — Misuse Risks](https://aiforhumanity.eu/summaries/atlas-ch2-risks-04-misuse-risks): Source: Misuse Risks - [AI Safety Atlas Ch.2 — Risk Amplifiers](https://aiforhumanity.eu/summaries/atlas-ch2-risks-03-risk-amplifiers): Source: Risk Amplifiers - [AI Safety Atlas Ch.2 — Risk Decomposition](https://aiforhumanity.eu/summaries/atlas-ch2-risks-01-risk-decomposition): Source: Risk Decomposition | 8 min | Authors: [[markov-grey|Markov Grey]] & [[charbel-raphael-segerie|Charbel-Raphaël Ségerie]] - [AI Safety Atlas Ch.2 — Systemic Risks](https://aiforhumanity.eu/summaries/atlas-ch2-risks-06-systemic-risks): Source: Systemic Risks - [AI Safety Atlas Ch.3 — AGI Safety Strategies](https://aiforhumanity.eu/summaries/atlas-ch3-strategies-04-agi-safety-strategies): Source: AGI Safety Strategies | Authors: [[markov-grey|Markov Grey]] & [[charbel-raphael-segerie|Charbel-Raphaël Ségerie]] | 31 min — the longest subchapter so far - [AI Safety Atlas Ch.3 — ASI Safety Strategies](https://aiforhumanity.eu/summaries/atlas-ch3-strategies-05-asi-safety-strategies): Source: ASI Safety Strategies - [AI Safety Atlas Ch.3 — Appendix: Long-term Questions](https://aiforhumanity.eu/summaries/atlas-ch3-strategies-09-appendix-long-term-questions): Source: Appendix: Long-term Questions - [AI Safety Atlas Ch.3 — Challenges](https://aiforhumanity.eu/summaries/atlas-ch3-strategies-02-challenges): Source: Challenges - [AI Safety Atlas Ch.3 — Combining Strategies](https://aiforhumanity.eu/summaries/atlas-ch3-strategies-07-combining-strategies): Source: Combining Strategies - [AI Safety Atlas Ch.3 — Conclusion](https://aiforhumanity.eu/summaries/atlas-ch3-strategies-08-conclusion): Source: Strategies — Conclusion - [AI Safety Atlas Ch.3 — Definitions](https://aiforhumanity.eu/summaries/atlas-ch3-strategies-01-definitions): Source: Definitions - [AI Safety Atlas Ch.3 — Introduction](https://aiforhumanity.eu/summaries/atlas-ch3-strategies-00-introduction): Source: Strategies — Introduction | Authors: [[markov-grey|Markov Grey]] & [[charbel-raphael-segerie|Charbel-Raphaël Ségerie]] | Updated Summer 2025 | 3 min - [AI Safety Atlas Ch.3 — Misuse Prevention Strategies](https://aiforhumanity.eu/summaries/atlas-ch3-strategies-03-misuse-prevention-strategies): Source: Misuse Prevention Strategies - [AI Safety Atlas Ch.3 — Socio-Technical Strategies](https://aiforhumanity.eu/summaries/atlas-ch3-strategies-06-socio-technical-strategies): Source: Socio-Technical Strategies - [AI Safety Atlas Ch.4 — Appendix: Data Governance](https://aiforhumanity.eu/summaries/atlas-ch4-governance-07-appendix-data-governance): Source: Appendix: Data Governance - [AI Safety Atlas Ch.4 — Appendix: National Governance](https://aiforhumanity.eu/summaries/atlas-ch4-governance-08-appendix-national-governance): Source: Appendix: National Governance - [AI Safety Atlas Ch.4 — Compute Governance](https://aiforhumanity.eu/summaries/atlas-ch4-governance-02-compute-governance): Source: Compute Governance | 12 min | Authors: Charles Martinet, [[markov-grey|Markov Grey]], Su Cizem, [[charbel-raphael-segerie|Charbel-Raphaël Ségerie]] - [AI Safety Atlas Ch.4 — Conclusion](https://aiforhumanity.eu/summaries/atlas-ch4-governance-06-conclusion): Source: Governance — Conclusion - [AI Safety Atlas Ch.4 — Governance Architectures](https://aiforhumanity.eu/summaries/atlas-ch4-governance-04-governance-architectures): Source: Governance Architectures - [AI Safety Atlas Ch.4 — Governance Problems](https://aiforhumanity.eu/summaries/atlas-ch4-governance-01-governance-problems): Source: Governance Problems - [AI Safety Atlas Ch.4 — Implementation](https://aiforhumanity.eu/summaries/atlas-ch4-governance-05-implementation): Source: Implementation - [AI Safety Atlas Ch.4 — Introduction](https://aiforhumanity.eu/summaries/atlas-ch4-governance-00-introduction): Source: Governance — Introduction - [AI Safety Atlas Ch.4 — Systemic Challenges](https://aiforhumanity.eu/summaries/atlas-ch4-governance-03-systemic-challenges): Source: Systemic Challenges - [AI Safety Atlas Ch.5 — Benchmarks](https://aiforhumanity.eu/summaries/atlas-ch5-evaluations-08-benchmarks): Source: Benchmarks - [AI Safety Atlas Ch.5 — Conclusion](https://aiforhumanity.eu/summaries/atlas-ch5-evaluations-10-conclusion): Source: Evaluations — Conclusion | Authors: [[markov-grey|Markov Grey]] & [[charbel-raphael-segerie|Charbel-Raphaël Ségerie]] - [AI Safety Atlas Ch.5 — Control Evaluations](https://aiforhumanity.eu/summaries/atlas-ch5-evaluations-06-control-evaluations): Source: Control Evaluations - [AI Safety Atlas Ch.5 — Dangerous Capability Evaluations](https://aiforhumanity.eu/summaries/atlas-ch5-evaluations-04-dangerous-capability-evaluations): Source: Dangerous Capability Evaluations - [AI Safety Atlas Ch.5 — Dangerous Propensity Evaluations](https://aiforhumanity.eu/summaries/atlas-ch5-evaluations-05-dangerous-propensity-evaluations): Source: Dangerous Propensity Evaluations - [AI Safety Atlas Ch.5 — Evaluated Properties](https://aiforhumanity.eu/summaries/atlas-ch5-evaluations-03-evaluated-properties): Source: Evaluated Properties - [AI Safety Atlas Ch.5 — Evaluation Design](https://aiforhumanity.eu/summaries/atlas-ch5-evaluations-01-evaluation-design): Source: Evaluation Design - [AI Safety Atlas Ch.5 — Evaluation Frameworks](https://aiforhumanity.eu/summaries/atlas-ch5-evaluations-07-evaluation-frameworks): Source: Evaluation Frameworks - [AI Safety Atlas Ch.5 — Evaluation Techniques](https://aiforhumanity.eu/summaries/atlas-ch5-evaluations-02-evaluation-techniques): Source: Evaluation Techniques - [AI Safety Atlas Ch.5 — Introduction](https://aiforhumanity.eu/summaries/atlas-ch5-evaluations-00-introduction): Source: Evaluations — Introduction - [AI Safety Atlas Ch.5 — Limitations](https://aiforhumanity.eu/summaries/atlas-ch5-evaluations-09-limitations): Source: Limitations - [AI Safety Atlas Ch.6 — Introduction](https://aiforhumanity.eu/summaries/atlas-ch6-specification-gaming-00-introduction): Source: Specification Gaming — Introduction | Authors: [[markov-grey|Markov Grey]] & [[charbel-raphael-segerie|Charbel-Raphaël Ségerie]] | 1 min - [AI Safety Atlas Ch.6 — Learning from Feedback](https://aiforhumanity.eu/summaries/atlas-ch6-specification-gaming-04-learning-from-feedback): Source: Learning from Feedback - [AI Safety Atlas Ch.6 — Learning from Imitation](https://aiforhumanity.eu/summaries/atlas-ch6-specification-gaming-05-learning-from-imitation): Source: Learning from Imitation - [AI Safety Atlas Ch.6 — Optimization](https://aiforhumanity.eu/summaries/atlas-ch6-specification-gaming-02-optimization): Source: Optimization - [AI Safety Atlas Ch.6 — Reinforcement Learning](https://aiforhumanity.eu/summaries/atlas-ch6-specification-gaming-01-reinforcement-learning): Source: Reinforcement Learning - [AI Safety Atlas Ch.6 — Specification Gaming](https://aiforhumanity.eu/summaries/atlas-ch6-specification-gaming-03-specification-gaming): Source: Specification Gaming - [AI Safety Atlas Ch.7 — Detection](https://aiforhumanity.eu/summaries/atlas-ch7-goal-misgeneralization-05-detection): Source: Detection - [AI Safety Atlas Ch.7 — Goal-Directedness](https://aiforhumanity.eu/summaries/atlas-ch7-goal-misgeneralization-02-goal-directedness): Source: Goal-Directedness - [AI Safety Atlas Ch.7 — Introduction](https://aiforhumanity.eu/summaries/atlas-ch7-goal-misgeneralization-00-introduction): Source: Goal Misgeneralization — Introduction - [AI Safety Atlas Ch.7 — Learning Dynamics](https://aiforhumanity.eu/summaries/atlas-ch7-goal-misgeneralization-01-learning-dynamics): Source: Learning Dynamics - [AI Safety Atlas Ch.7 — Mitigations](https://aiforhumanity.eu/summaries/atlas-ch7-goal-misgeneralization-06-mitigations): Source: Mitigations - [AI Safety Atlas Ch.7 — Multi-Objective Generalization](https://aiforhumanity.eu/summaries/atlas-ch7-goal-misgeneralization-03-multi-objective-generalization): Source: Multi-Objective Generalization - [AI Safety Atlas Ch.7 — Scheming](https://aiforhumanity.eu/summaries/atlas-ch7-goal-misgeneralization-04-scheming): Source: Scheming - [AI Safety Atlas Ch.8 — Debate](https://aiforhumanity.eu/summaries/atlas-ch8-scalable-oversight-05-debate): Source: Debate - [AI Safety Atlas Ch.8 — Introduction](https://aiforhumanity.eu/summaries/atlas-ch8-scalable-oversight-00-introduction): Source: Scalable Oversight — Introduction - [AI Safety Atlas Ch.8 — Iterated Amplification](https://aiforhumanity.eu/summaries/atlas-ch8-scalable-oversight-04-iterated-amplification): Source: Iterated Amplification - [AI Safety Atlas Ch.8 — Oversight](https://aiforhumanity.eu/summaries/atlas-ch8-scalable-oversight-01-oversight): Source: Oversight - [AI Safety Atlas Ch.8 — Process Oversight](https://aiforhumanity.eu/summaries/atlas-ch8-scalable-oversight-03-process-oversight): Source: Process Oversight | Authors: [[markov-grey|Markov Grey]] & [[charbel-raphael-segerie|Charbel-Raphaël Ségerie]] - [AI Safety Atlas Ch.8 — Task Decomposition](https://aiforhumanity.eu/summaries/atlas-ch8-scalable-oversight-02-task-decomposition): Source: Task Decomposition - [AI Safety Atlas Ch.8 — Weak-to-Strong](https://aiforhumanity.eu/summaries/atlas-ch8-scalable-oversight-06-weak-to-strong-w2s): Source: Weak-to-Strong (W2S) | Authors: [[markov-grey|Markov Grey]] & [[charbel-raphael-segerie|Charbel-Raphaël Ségerie]] - [AI Safety Atlas Textbook — Read the Textbook (ai-safety-atlas.com)](https://aiforhumanity.eu/summaries/read-the-textbook-ai-safety-atlas): Source: ai-safety-atlas.com Format: 8-chapter self-paced textbook, approximately 15 hours of content Source - [AI Sandbagging: Language Models can Strategically Underperform on Evaluations](https://aiforhumanity.eu/summaries/2406.07358): *Teun van der Weij, Felix Hofstätter, Ollie Jaffe, Samuel F. Brown, Francis Rhys Ward* — 2024-06-11 — arXiv - [AI Testing Should Account for Sophisticated Strategic Behaviour](https://aiforhumanity.eu/summaries/2508.14927): *Vojtech Kovarik, Eric Olav Chen, Sami Petersen, Alexis Ghersengorin, Vincent Conitzer* — 2025-08-19 — arXiv - [AI companies are unlikely to make high-assurance safety cases if timelines are short](https://aiforhumanity.eu/summaries/ai-companies-are-unlikely-to-make-high-assurance-safety-case): *Ryan Greenblatt* — 2025-01-23 — Anthropic — LessWrong / AI Alignment Forum - [AI companies should be safety-testing the most capable versions of their models](https://aiforhumanity.eu/summaries/ai-companies-should-be-safety-testing-the-most-capable-versi): *Steven Adler* — 2025-03-26 — LessWrong - [AI in Context — YouTube Channel Videos](https://aiforhumanity.eu/summaries/ai-in-context-videos): This summary covers the AI in Context YouTube channel, a video series launched in July 2025 by [[80000-hours]] and hosted by Aric Floyd. The channel provides in-depth discussions on artificial... - [AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons](https://aiforhumanity.eu/summaries/2503.05731): *Shaona Ghosh, Heather Frase, Adina Williams, Sarah Luger, Paul Röttger, Fazl Barez, … (+96 more)* — 2025-04-18 — MLCommons — arXiv - [AIs at the current capability level may be important for future safety work](https://aiforhumanity.eu/summaries/ais-at-the-current-capability-level-may-be-important-for-fut): *Ryan Greenblatt* — 2025-05-12 — Anthropic — LessWrong - [Abstract Mathematical Concepts vs. Abstractions Over Real-World Systems](https://aiforhumanity.eu/summaries/abstract-mathematical-concepts-vs-abstractions-over-real-wor): - [Academic Papers Index — Bostrom, Parfit, MacAskill](https://aiforhumanity.eu/summaries/academic-papers-index): This summary covers an index of nine freely available academic papers by three philosophers whose work is foundational to effective-altruism, [[existential-risk]] research, and [[ai-safety]]:... - [Active Attacks: Red-teaming LLMs via Adaptive Environments](https://aiforhumanity.eu/summaries/2509.21947): *Taeyoung Yun, Pierre-Luc St-Charles, Jinkyoo Park, Yoshua Bengio, Minsu Kim* — 2025-09-26 — Mila, KAIST — arXiv - [Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats](https://aiforhumanity.eu/summaries/2411.17693): *Jiaxin Wen, Vivek Hebbar, Caleb Larson, Aryan Bhatt, Ansh Radhakrishnan, Mrinank Sharma, … (+6 more)* — 2024-11-26 — arXiv - [Advancing Gemini's security safeguards](https://aiforhumanity.eu/summaries/advancing-gemini-s-security-safeguards): *Google DeepMind Security & Privacy Research Team* — 2025-05-20 — Google DeepMind — Google DeepMind Blog - [Adversarial Alignment for LLMs Requires Simpler, Reproducible, and More Measurable Objectives](https://aiforhumanity.eu/summaries/2502.11910): *Leo Schwinn, Yan Scholten, Tom Wollschläger, Sophie Xhonneux, Stephen Casper, Stephan Günnemann, … (+1 more)* — 2025-02-17 — arXiv - [Adversarial Manipulation of Reasoning Models using Internal Representations](https://aiforhumanity.eu/summaries/2507.03167): *Kureha Yamaguchi, Benjamin Etheridge, Andy Arditi* — 2025-07-03 — arXiv - [Aesthetic Preferences Can Cause Emergent Misalignment](https://aiforhumanity.eu/summaries/aesthetic-preferences-can-cause-emergent-misalignment): *Anders Woodruff* — 2025-08-26 — Center on Long-Term Risk — LessWrong - [Aether July 2025 Update](https://aiforhumanity.eu/summaries/aether-july-2025-update): *Rohan Subramani, Rauno Arike, Shubhorup Biswas* — 2025-07-01 — Aether - [Against RL: The Case for System 2 Learning](https://aiforhumanity.eu/summaries/against-rl-the-case-for-system-2-learning): *Andreas Stuhlmüller* — 2025-01-30 — Elicit — Elicit Blog - [Against blanket arguments against interpretability](https://aiforhumanity.eu/summaries/against-blanket-arguments-against-interpretability): *Dmitry Vaintrob* — 2025-01-22 — LessWrong - [Agent foundations: not really math, not really science](https://aiforhumanity.eu/summaries/agent-foundations-not-really-math-not-really-science): *Alex_Altair* — 2025-08-17 — Dovetail Research - [AgentBreeder: Mitigating the AI Safety Risks of Multi-Agent Scaffolds via Self-Improvement](https://aiforhumanity.eu/summaries/2502.00757): *J Rosser, Jakob Foerster* — 2025-02-02 — arXiv - [Agentic Interpretability: A Strategy Against Gradual Disempowerment](https://aiforhumanity.eu/summaries/agentic-interpretability-a-strategy-against-gradual-disempow): *Been Kim, John Hewitt, Neel Nanda, Noah Fiedel, Oyvind Tafjord* — 2025-06-17 — Google DeepMind, Anthropic - [Agentic Misalignment](https://aiforhumanity.eu/summaries/2510.05179): *Aengus Lynch, Benjamin Wright, Caleb Larson, Stuart J. Ritchie, Soren Mindermann, Evan Hubinger, … (+2 more)* — 2025-10-05 — Anthropic, Redwood Research — arXiv - [Agentic Misalignment: How LLMs Could be Insider Threats](https://aiforhumanity.eu/summaries/agentic-misalignment-how-llms-could-be-insider-threats-01221c58): *Aengus Lynch, Benjamin Wright, Caleb Larson, Stuart Richie, Sören Mindermann, Evan Hubinger, … (+2 more)* — 2025-06-20 — Anthropic — LessWrong/AI Alignment Forum - [Agentic Misalignment: How LLMs could be insider threats](https://aiforhumanity.eu/summaries/agentic-misalignment-how-llms-could-be-insider-threats-0e6b6434): *Aengus Lynch, Benjamin Wright, Caleb Larson, Kevin K. Troy, Stuart J. Ritchie, Sören Mindermann, … (+2 more)* — 2025-06-20 — Anthropic, University College London, MATS, Mila, Redwood Research —... - [Agentic Misalignment: How LLMs could be insider threats](https://aiforhumanity.eu/summaries/agentic-misalignment-how-llms-could-be-insider-threats): *Aengus Lynch, Benjamin Wright, Caleb Larson, Kevin K. Troy, Stuart J. Ritchie, Sören Mindermann, … (+2 more)* — 2025-06-20 — Anthropic, University College London, MATS, Mila — Anthropic Research - [Aligning machine and human visual representations across abstraction levels](https://aiforhumanity.eu/summaries/aligning-machine-and-human-visual-representations-across-abs): *Lukas Muttenthaler, Klaus Greff, Frieda Born, Bernhard Spitzer, Simon Kornblith, Michael C. Mozer, … (+3 more)* — 2025-11-12 — Google DeepMind, Technische Universität Berlin, BIFOLD, Max Planck... - [Alignment Faking Revisited: Improved Classifiers and Open Source Extensions](https://aiforhumanity.eu/summaries/alignment-faking-revisited-improved-classifiers-and-open-sou): *John Hughes, Abhay Sheshadr* — 2025 — MATS, Anthropic — Anthropic Alignment Science Blog - [Alignment Faking in Large Language Models](https://aiforhumanity.eu/summaries/alignment-faking-in-large-language-models): Authors: Ryan Greenblatt, Evan Hubinger, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Sam Bowman, Buck Shlegeris ([[buck-shlegeris]])... - [Alignment first, intelligence later](https://aiforhumanity.eu/summaries/alignment-first-intelligence-later): *Chris Lakin* — 2025-03-30 — Softmax — Substack (Locally Optimal) - [Ambiguous Online Learning](https://aiforhumanity.eu/summaries/ambiguous-online-learning): *Vanessa Kosoy* — 2025-06-25 — Independent - [Among Us: A Sandbox for Measuring and Detecting Agentic Deception](https://aiforhumanity.eu/summaries/2504.04072): *Satvik Golechha, Adrià Garriga-Alonso* — 2025-04-05 — arXiv - [An Approach to Technical AGI Safety and Security](https://aiforhumanity.eu/summaries/2504.01849): *Rohin Shah, Alex Irpan, Alexander Matt Turner, Anna Wang, Arthur Conmy, David Lindner, … (+24 more)* — 2025-04-02 — Google DeepMind, Anthropic, Stanford University, UC Berkeley, Redwood Research... - [An alignment safety case sketch based on debate](https://aiforhumanity.eu/summaries/2505.03989): *Marie Davidsen Buhl, Jacob Pfau, Benjamin Hilton, Geoffrey Irving* — 2025-05-23 — Google DeepMind — arXiv - [An alignment safety case sketch based on debate](https://aiforhumanity.eu/summaries/an-alignment-safety-case-sketch-based-on-debate): *Marie_DB, Jacob Pfau, Benjamin Hilton, Geoffrey Irving* — 2025-05-08 — UK AISI — LessWrong / AI Alignment Forum - [Assessing confidence in frontier AI safety cases](https://aiforhumanity.eu/summaries/2502.05791): *Stephen Barrett, Philip Fox, Joshua Krook, Tuneer Mondal, Simon Mylius, Alejandro Tlaie* — 2025-02-09 — arXiv - [AssistanceZero: Scalably Solving Assistance Games](https://aiforhumanity.eu/summaries/2504.07091): *Cassidy Laidlaw, Eli Bronstein, Timothy Guo, Dylan Feng, Lukas Berglund, Justin Svegliato, … (+2 more)* — 2025-04-09 — UC Berkeley — arXiv - [Auditing language models for hidden objectives](https://aiforhumanity.eu/summaries/2503.10965): *Samuel Marks, Johannes Treutlein, Trenton Bricken, Jack Lindsey, Jonathan Marcus, Siddharth Mishra-Sharma, … (+29 more)* — 2025-03-14 — Anthropic, Independent Researchers — arXiv - [Auditing language models for hidden objectives](https://aiforhumanity.eu/summaries/auditing-language-models-for-hidden-objectives): 2025-03-13 — Anthropic — Anthropic Blog - [Automated Researchers Can Subtly Sandbag](https://aiforhumanity.eu/summaries/automated-researchers-can-subtly-sandbag): *Johannes Gasteiger, Vladimir Mikulik, Ethan Perez, Fabien Roger, Misha Wagner, Akbir Khan, … (+2 more)* — 2025 — Anthropic — Alignment Science Blog - [Automatically Jailbreaking Frontier Language Models with Investigator Agents](https://aiforhumanity.eu/summaries/automatically-jailbreaking-frontier-language-models-with-inv): - [Automating AI Safety: What we can do today](https://aiforhumanity.eu/summaries/automating-ai-safety-what-we-can-do-today): *Matthew Shinkle, Eyon Jang, Jacques Thibodeau* — 2025-07-25 — SPAR, PIBBSS — LessWrong - [Bare Minimum Mitigations for Autonomous AI Development](https://aiforhumanity.eu/summaries/bare-minimum-mitigations-for-autonomous-ai-development): *Joshua Clymer, Isabella Duan, Chris Cundy, Yawen Duan, Fynn Heide, Chaochao Lu, … (+7 more)* — 2025-04-22 — Safe AI Forum - [Beliefs about formal methods and AI safety](https://aiforhumanity.eu/summaries/beliefs-about-formal-methods-and-ai-safety): *Quinn Dougherty* — 2025-10-23 — LessWrong - [Benchmarking deception probes for trusted monitoring](https://aiforhumanity.eu/summaries/benchmarking-deception-probes-for-trusted-monitoring): *Avi Parrack, StefanHex, Cleo Nardo* — 2024-07-23 — Stanford University - [Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models](https://aiforhumanity.eu/summaries/2510.27629): *Boyi Wei, Zora Che, Nathaniel Li, Udari Madhushani Sehwag, Jasper Götting, Samira Nedungadi, … (+7 more)* — 2025-10-31 — Anthropic, Redwood Research — arXiv - [Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models](https://aiforhumanity.eu/summaries/2510.27629v2): *Boyi Wei, Zora Che, Nathaniel Li, Udari Madhushani Sehwag, Jasper Götting, Samira Nedungadi, … (+7 more)* — 2025-10-31 — Anthropic, Redwood Research - [Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs](https://aiforhumanity.eu/summaries/2512.05648): *Igor Shilov, Alex Cloud, Aryo Pradipta Gema, Jacob Goldman-Wetzler, Nina Panickssery, Henry Sleight, … (+2 more)* — 2024-12-05 — Anthropic - [Beyond Linear Probes: Dynamic Safety Monitoring for Language Models](https://aiforhumanity.eu/summaries/2509.26238): *James Oldfield, Philip Torr, Ioannis Patras, Adel Bibi, Fazl Barez* — 2025-09-30 — arXiv - [Beyond Preferences in AI Alignment](https://aiforhumanity.eu/summaries/2408.16984): *Tan Zhi-Xuan, Micah Carroll, Matija Franklin, Hal Ashton* — 2024-08-30 — arXiv - [Binary Sparse Coding for Interpretability](https://aiforhumanity.eu/summaries/2509.25596): *Lucia Quirke, Stepan Shabalin, Nora Belrose* — 2025-09-29 — arXiv - [BioBlue: Notable runaway-optimiser-like LLM failure modes on biologically and economically aligned AI safety benchmarks for LLMs with simplified observation format](https://aiforhumanity.eu/summaries/2509.02655): *Roland Pihlakas, Sruthi Kuriakose* — 2025-09-02 — arXiv - [Bridging the human–AI knowledge gap through concept discovery and transfer in AlphaZero](https://aiforhumanity.eu/summaries/bridging-the-human-ai-knowledge-gap-through-concept-discover): - [Building and evaluating alignment auditing agents](https://aiforhumanity.eu/summaries/building-and-evaluating-alignment-auditing-agents-bda3ba07): *Trenton Bricken, Rowan Wang, Sam Bowman, Euan Ong, Johannes Treutlein, Jeff Wu, … (+2 more)* — 2025-07-24 — Anthropic — Anthropic Alignment Science Blog - [Building and evaluating alignment auditing agents](https://aiforhumanity.eu/summaries/building-and-evaluating-alignment-auditing-agents): *Sam Marks, trentbrick, RowanWang, Sam Bowman, Euan Ong, Johannes Treutlein, … (+1 more)* — 2025-07-24 — Anthropic — LessWrong / AI Alignment Forum - [CCS-Lib: A Python package to elicit latent knowledge from LLMs](https://aiforhumanity.eu/summaries/ccs-lib-a-python-package-to-elicit-latent-knowledge-from-llm): *Walter Laurito, Nora Belrose, Alex Mallen, Kay Kozaronek, Fabien Roger, Christy Koh, … (+11 more)* — 2025-10-21 — Cadenza Labs — Journal of Open Source Software - [CRISP: Persistent Concept Unlearning via Sparse Autoencoders](https://aiforhumanity.eu/summaries/2508.13650): *Tomer Ashuach, Dana Arad, Aaron Mueller, Martin Tutek, Yonatan Belinkov* — 2025-08-19 — arXiv - [Call Me A Jerk: Persuading AI to Comply with Objectionable Requests](https://aiforhumanity.eu/summaries/call-me-a-jerk-persuading-ai-to-comply-with-objectionable-re): *Lennart Meincke, Dan Shapiro, Angela Duckworth, Ethan R. Mollick, Lilach Mollick, Robert Cialdini* — 2025-07-18 — University of Pennsylvania, The Wharton School, WHU - Otto Beisheim School of... - [Call for Collaboration: Renormalization for AI safety](https://aiforhumanity.eu/summaries/call-for-collaboration-renormalization-for-ai-safety): *Lauren Greenspan* — 2025-03-31 — PIBBSS — LessWrong - [Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability](https://aiforhumanity.eu/summaries/2510.19851): *Artur Zolkowski, Wen Xing, David Lindner, Florian Tramèr, Erik Jenner* — 2025-10-21 — ETH Zurich, FAR AI — arXiv - [Can a Neural Network that only Memorizes the Dataset be Undetectably Backdoored?](https://aiforhumanity.eu/summaries/can-a-neural-network-that-only-memorizes-the-dataset-be-unde): *Matjaz Leonardis* — 2025-07-10 — ODYSSEY 2025 Conference - [Caught in the Act: a mechanistic approach to detecting deception](https://aiforhumanity.eu/summaries/2508.19505): *Gerard Boxo, Ryan Socha, Daniel Yoo, Shivam Raval* — 2025-08-27 — arXiv - [Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety](https://aiforhumanity.eu/summaries/2507.11473): *Tomek Korbak, Mikita Balesni, Elizabeth Barnes, Yoshua Bengio, Joe Benton, Joseph Bloom, … (+35 more)* — 2025-07-15 — Anthropic, OpenAI, DeepMind, FAR AI, UC Berkeley, Mila, Center for AI Safety,... - [Chain-of-Thought Hijacking](https://aiforhumanity.eu/summaries/2510.26418): *Jianli Zhao, Tingchen Fu, Rylan Schaeffer, Mrinank Sharma, Fazl Barez* — 2025-10-30 — arXiv - [Chain-of-Thought Snippets — Anti-Scheming](https://aiforhumanity.eu/summaries/chain-of-thought-snippets-anti-scheming): Apollo Research, OpenAI — antischeming.ai - [Challenges and Future Directions of Data-Centric AI Alignment](https://aiforhumanity.eu/summaries/2410.01957v2): *Min-Hsuan Yeh, Jeffrey Wang, Xuefeng Du, Seongheon Park, Leitian Tao, Shawn Im, … (+1 more)* — 2025-05-01 — arXiv - [Circuit Tracing: Revealing Computational Graphs in Language Models](https://aiforhumanity.eu/summaries/circuit-tracing-revealing-computational-graphs-in-language-m): *Emmanuel Ameisen, Jack Lindsey, Adam Pearce, Wes Gurnee, Nicholas L. Turner, Brian Chen, … (+21 more)* — 2025-03-27 — Anthropic — Transformer Circuits Thread - [Circuits in Superposition](https://aiforhumanity.eu/summaries/circuits-in-superposition): *Lucius Bushnaq, jake_mendel* — 2024-10-14 — Apollo Research - [Circuits in Superposition 2: Now With Less Wrong Math](https://aiforhumanity.eu/summaries/circuits-in-superposition-2-now-with-less-wrong-math): - [Clarifying "wisdom": Foundational topics for aligned AIs to prioritize before irreversible decisions](https://aiforhumanity.eu/summaries/clarifying-wisdom-foundational-topics-for-aligned-ais-to-pri): - [Claude Sonnet 3.7 (often) knows when it's in alignment evaluations](https://aiforhumanity.eu/summaries/claude-sonnet-3-7-often-knows-when-it-s-in-alignment-evaluat): *Nicholas Goldowsky-Dill, Mikita Balesni, Jérémy Scheurer, Marius Hobbhahn* — 2025-03-17 — Apollo Research — LessWrong - [Claude's Constitution](https://aiforhumanity.eu/summaries/claude-s-constitution): 2023-05-09 — Anthropic - [CoT May Be Highly Informative Despite "Unfaithfulness"](https://aiforhumanity.eu/summaries/cot-may-be-highly-informative-despite-unfaithfulness): *Amy Deng, Sydney Von Arx, Ben Snodin, Sudarsh Kunnavakkam, Tamera Lanham* — 2025-08-08 — METR, Gray Swan - [CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring](https://aiforhumanity.eu/summaries/2505.23575): *Benjamin Arnav, Pablo Bernabeu-Pérez, Nathan Helm-Burger, Tim Kostolansky, Hannes Whittingham, Mary Phuong* — 2025-05-29 — DeepMind — arXiv - [Code World Model Preparedness Report](https://aiforhumanity.eu/summaries/code-world-model-preparedness-report): - [Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning](https://aiforhumanity.eu/summaries/2509.11816): *Filip Sondej, Yushi Yang* — 2025-09-15 — arXiv - [Collective cooperative intelligence](https://aiforhumanity.eu/summaries/collective-cooperative-intelligence): *Wolfram Barfuss, Jessica Flack, Chaitanya S. Gokhale, Lewis Hammond, Christian Hilbe, Edward Hughes, … (+8 more)* — 2025-06-16 — University of Bonn, Santa Fe Institute, Max Planck Institute for... - [Communication & Trust](https://aiforhumanity.eu/summaries/communication-trust): *Abram Demski* — 2025-07-09 — ODYSSEY 2025 Conference - [Competing with sampling](https://aiforhumanity.eu/summaries/competing-with-sampling): *Eric Neyman* — 2025-11-18 — Alignment Research Center - [Compressed Computation is (probably) not Computation in Superposition](https://aiforhumanity.eu/summaries/compressed-computation-is-probably-not-computation-in-superp): - [Compromising Honesty and Harmlessness in Language Models via Deception Attacks](https://aiforhumanity.eu/summaries/2502.08301): *Laurène Vaugrante, Francesca Carlon, Maluna Menke, Thilo Hagendorff* — 2025-02-12 — arXiv - [Condensation](https://aiforhumanity.eu/summaries/condensation): *abramdemski* — 2025-11-09 - [Consistency Training Helps Stop Sycophancy and Jailbreaks](https://aiforhumanity.eu/summaries/2510.27062): *Alex Irpan, Alexander Matt Turner, Mark Kurzeja, David K. Elson, Rohin Shah* — 2025-10-31 — DeepMind, Google — arXiv - [Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming](https://aiforhumanity.eu/summaries/2501.18837): *Mrinank Sharma, Meg Tong, Jesse Mu, Jerry Wei, Jorrit Kruthoff, Scott Goodfriend, … (+37 more)* — 2025-01-31 — Anthropic — arXiv - [Constitutional Classifiers: Defending against universal jailbreaks](https://aiforhumanity.eu/summaries/constitutional-classifiers-defending-against-universal-jailb): *Anthropic Safeguards Research Team* — 2025-02-03 — Anthropic — Anthropic Research Blog - [Control protocols don't always need to know which models are scheming](https://aiforhumanity.eu/summaries/control-protocols-don-t-always-need-to-know-which-models-are-scheming): Author: Fabien Roger (LessWrong post) - Published: 2026-04-26 - Type: LessWrong post (personal views, not a peer-reviewed paper) - [ControlArena](https://aiforhumanity.eu/summaries/controlarena): *Rogan Inglis, Ollie Matthews, Tyler Tracy, Oliver Makins, Tom Catling, Asa Cooper Stickland, … (+14 more)* — 2025-01-01 — UK AI Security Institute, Redwood Research — GitHub - [Convergent Linear Representations of Emergent Misalignment](https://aiforhumanity.eu/summaries/2506.11618): *Anna Soligo, Edward Turner, Senthooran Rajamanoharan, Neel Nanda* — 2025-06-20 — Google DeepMind — arXiv - [Convergent Linear Representations of Emergent Misalignment](https://aiforhumanity.eu/summaries/convergent-linear-representations-of-emergent-misalignment): *Anna Soligo, Edward Turner, Senthooran Rajamanoharan, Neel Nanda* — 2025-06-16 — ML Alignment & Theory Scholars — LessWrong/AI Alignment Forum - [Cost-Effective Constitutional Classifiers via Representation Re-use](https://aiforhumanity.eu/summaries/cost-effective-constitutional-classifiers-via-representation): *Hoagy Cunningham, Alwin Peng, Jerry Wei, Euan Ong, Fabien Roger, Linda Petrini, … (+3 more)* — 2025 — Anthropic — Anthropic Alignment Science Blog - [Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs](https://aiforhumanity.eu/summaries/cross-architecture-model-diffing-with-crosscoders-unsupervis): - [Ctrl-Z: Controlling AI Agents via Resampling](https://aiforhumanity.eu/summaries/ctrl-z-controlling-ai-agents-via-resampling): *Aryan Bhatt, Buck Shlegeris, Adam Kaufman, Cody Rushing, Tyler Tracy, Vasil Georgiev, … (+2 more)* — 2025-04-16 — Redwood Research — AI Alignment Forum - [CyberSecEval 4](https://aiforhumanity.eu/summaries/cyberseceval-4): 2025 — Meta - [D-REX: A Benchmark for Detecting Deceptive Reasoning in Large Language Models](https://aiforhumanity.eu/summaries/2509.17938): *Satyapriya Krishna, Andy Zou, Rahul Gupta, Eliot Krzysztof Jones, Nick Winter, Dan Hendrycks, … (+3 more)* — 2025-09-22 — Carnegie Mellon University, Center for AI Safety, Anthropic — arXiv - [DECEPTIONBENCH: A Comprehensive Benchmark for AI Deception Behaviors in Real-world Scenario](https://aiforhumanity.eu/summaries/2510.15501): *Yao Huang, Yitong Sun, Yichi Zhang, Ruochen Zhang, Yinpeng Dong, Xingxing Wei* — 2025-10-17 - [Debate Helps Weak-to-Strong Generalization](https://aiforhumanity.eu/summaries/2501.13124): *Hao Lang, Fei Huang, Yongbin Li* — 2025-01-21 — arXiv - [Deceptive Alignment and Homuncularity](https://aiforhumanity.eu/summaries/deceptive-alignment-and-homuncularity): *Oliver Sourbut, TurnTrout* — 2025-01-16 — UK AI Safety Institute, Independent — LessWrong - [Decomposing MLP Activations into Interpretable Features via Semi-Nonnegative Matrix Factorization](https://aiforhumanity.eu/summaries/2506.10920): *Or Shafran, Atticus Geiger, Mor Geva* — 2025-06-12 — arXiv - [Deep Ignorance: Filtering Pretraining Data Builds Tamper-Resistant Safeguards into Open-Weight LLMs](https://aiforhumanity.eu/summaries/2508.06601): *Kyle O'Brien, Stephen Casper, Quentin Anthony, Tomek Korbak, Robert Kirk, Xander Davies, … (+4 more)* — 2025-08-08 — Anthropic, Redwood Research, EleutherAI, University of Oxford — arXiv - [Deep Reinforcement Learning Agents are not even close to Human Intelligence](https://aiforhumanity.eu/summaries/2505.21731v1): - [Deep ignorance: Filtering pretraining data builds tamper-resistant safeguards into open-weight LLMs](https://aiforhumanity.eu/summaries/deep-ignorance-filtering-pretraining-data-builds-tamper-resi): *Kyle O'Brien, Stephen Casper, Quentin Anthony, Tomek Korbak, Robert Kirk, Xander Davies, … (+4 more)* — 2025-08-08 — UK AI Security Institute, MIT, Eleuther AI - [Defense Against the Dark Prompts: Mitigating Best-of-N Jailbreaking with Prompt Evaluation](https://aiforhumanity.eu/summaries/2502.00580): *Stuart Armstrong, Matija Franklin, Connor Stevens, Rebecca Gorman* — 2025-02-01 — arXiv - [Deliberative Alignment: Reasoning Enables Safer Language Models](https://aiforhumanity.eu/summaries/2412.16339): *Melody Y. Guan, Manas Joglekar, Eric Wallace, Saachi Jain, Boaz Barak, Alec Helyar, … (+9 more)* — 2024-12-20 — OpenAI — arXiv - [Demonstrating specification gaming in reasoning models](https://aiforhumanity.eu/summaries/2502.13295): *Alexander Bondarenko, Denis Volk, Dmitrii Volkov, Jeffrey Ladish* — 2025-08-27 — arXiv - [Details about METR's evaluation of OpenAI GPT-5](https://aiforhumanity.eu/summaries/details-about-metr-s-evaluation-of-openai-gpt-5): *METR* — 2025-08-01 — METR — METR's Autonomy Evaluation Resources - [Details about METR's preliminary evaluation of OpenAI's o3 and o4-mini](https://aiforhumanity.eu/summaries/details-about-metr-s-preliminary-evaluation-of-openai-s-o3-a): *METR* — 2025-04-01 — METR — METR's Autonomy Evaluation Resources - [Detect Goodhart and shut down](https://aiforhumanity.eu/summaries/detect-goodhart-and-shut-down): *Jeremy Gillen* — 2025-01-22 - [Detecting High-Stakes Interactions with Activation Probes](https://aiforhumanity.eu/summaries/2506.10805): *Alex McKenzie, Urja Pawar, Phil Blandfort, William Bankes, David Krueger, Ekdeep Singh Lubana, … (+1 more)* — 2025-06-12 — arXiv - [Detecting Strategic Deception Using Linear Probes](https://aiforhumanity.eu/summaries/2502.03407): *Nicholas Goldowsky-Dill, Bilal Chughtai, Stefan Heimersheim, Marius Hobbhahn* — 2025-02-05 — Apollo Research — arXiv - [Detecting Strategic Deception Using Linear Probes](https://aiforhumanity.eu/summaries/detecting-strategic-deception-using-linear-probes): *Nicholas Goldowsky-Dill, Bilal Chughtai, Stefan Heimersheim, Marius Hobbhahn* — 2025-02-06 — Apollo Research - [Detecting and reducing scheming in AI models](https://aiforhumanity.eu/summaries/detecting-and-reducing-scheming-in-ai-models): *OpenAI* — 2025-09-17 — OpenAI, Apollo Research — OpenAI Blog - [Detecting misbehavior in frontier reasoning models](https://aiforhumanity.eu/summaries/detecting-misbehavior-in-frontier-reasoning-models): *Bowen Baker, Joost Huizinga, Aleksander Madry, Wojciech Zaremba, Jakub Pachocki, David Farhi* — 2025-03-10 — OpenAI — OpenAI Blog - [Difficulties with Evaluating a Deception Detector for AIs](https://aiforhumanity.eu/summaries/2511.22662v1): *Lewis Smith, Bilal Chughtai, Neel Nanda* — 2025-11-27 — Google — arXiv - [Discovering Forbidden Topics in Language Models](https://aiforhumanity.eu/summaries/2505.17441): *Can Rager, Chris Wendler, Rohit Gandikota, David Bau* — 2025-05-23 — arXiv - [Discovering Undesired Rare Behaviors via Model Diff Amplification](https://aiforhumanity.eu/summaries/discovering-undesired-rare-behaviors-via-model-diff-amplific-ee1948d9): *Santiago Aranguri, Thomas McGrath* — 2025-08-21 — Goodfire, NYU — Goodfire Research - [Discovering Undesired Rare Behaviors via Model Diff Amplification](https://aiforhumanity.eu/summaries/discovering-undesired-rare-behaviors-via-model-diff-amplific): *Santiago Aranguri, Thomas McGrath* — 2025-08-21 — Goodfire, NYU - [Distillation Robustifies Unlearning](https://aiforhumanity.eu/summaries/2506.06278): *Bruce W. Lee, Addie Foote, Alex Infanger, Leni Shor, Harish Kamath, Jacob Goldman-Wetzler, … (+3 more)* — 2025-06-06 — arXiv (NeurIPS 2025 Spotlight) - [Diverse and Effective Red Teaming with Auto-generated Rewards and Multi-step Reinforcement Learning](https://aiforhumanity.eu/summaries/2412.18693): *Alex Beutel, Kai Xiao, Johannes Heidecke, Lilian Weng* — 2024-12-24 — OpenAI — arXiv - [Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models](https://aiforhumanity.eu/summaries/2411.14257): *Javier Ferrando, Oscar Obeso, Senthooran Rajamanoharan, Neel Nanda* — 2024-11-21 — arXiv - [Do LLMs Comply Differently During Tests? Is This a Hidden Variable in Safety Evaluation? And Can We Steer That?](https://aiforhumanity.eu/summaries/do-llms-comply-differently-during-tests-is-this-a-hidden-var): *Sahar Abdelnabi, Ahmed Salem* — 2025-06-16 — Microsoft — LessWrong - [Do LLMs know what they're capable of? Why this matters for AI safety, and initial findings](https://aiforhumanity.eu/summaries/do-llms-know-what-they-re-capable-of-why-this-matters-for-ai): *Casey Barkan, Sid Black, Oliver Sourbut* — 2025-07-13 — MATS Program — LessWrong / AI Alignment Forum - [Do Not Tile the Lightcone with Your Confused Ontology](https://aiforhumanity.eu/summaries/do-not-tile-the-lightcone-with-your-confused-ontology): *Jan_Kulveit* — 2025-06-13 - [Do reasoning models use their scratchpad like we do? Evidence from distilling paraphrases](https://aiforhumanity.eu/summaries/do-reasoning-models-use-their-scratchpad-like-we-do-evidence): 2025 — Anthropic — Anthropic Alignment Science Blog - [Do safety-relevant LLM steering vectors optimized on a single example generalize?](https://aiforhumanity.eu/summaries/do-safety-relevant-llm-steering-vectors-optimized-on-a-singl): *Jacob Dunefsky* — 2025-02-28 — Yale University — arXiv - [Docent: A system for analyzing and intervening on agent behavior](https://aiforhumanity.eu/summaries/docent-a-system-for-analyzing-and-intervening-on-agent-behav): *Kevin Meng, Vincent Huang, Jacob Steinhardt, Sarah Schwettmann* — 2025-03-24 — Transluce — Transluce Blog - [Dodging systematic human errors in scalable oversight](https://aiforhumanity.eu/summaries/dodging-systematic-human-errors-in-scalable-oversight-a7847d50): *Geoffrey Irving* — 2025-05-14 — UK AISI — LessWrong - [Dodging systematic human errors in scalable oversight](https://aiforhumanity.eu/summaries/dodging-systematic-human-errors-in-scalable-oversight): *Geoffrey Irving* — 2025-05-14 — UK AISI - [EA Forum — Key Posts on AI Safety and Longtermism](https://aiforhumanity.eu/summaries/ea-forum-key-posts): This summary covers a curated collection of eight influential posts and three introductory topic pages from the EA Forum (forum.effectivealtruism.org), representing the key debates within the... - [EA and AI Safety Books Reference](https://aiforhumanity.eu/summaries/ea-ai-books): This summary covers a reference list of five key books at the intersection of effective-altruism, [[existential-risk]], and [[ai-safety]]. Together they represent the core reading list for anyone... - [EA and AI Safety Content Library — Master Inventory](https://aiforhumanity.eu/summaries/ea-content-library-inventory): This summary covers the master inventory file that catalogs the entire raw source library for this wiki. The inventory documents approximately 75+ items across 10 categories, providing a... - [Early Signs of Steganographic Capabilities in Frontier LLMs](https://aiforhumanity.eu/summaries/2507.02737): *Artur Zolkowski, Kei Nishimura-Gasparian, Robert McCarthy, Roland S. Zimmermann, David Lindner* — 2025-07-03 — arXiv - [Edge Cases in AI Alignment](https://aiforhumanity.eu/summaries/edge-cases-in-ai-alignment): *Florian Dietz* — 2025-03-24 — MATS Program — LessWrong - [Effective Altruism in the Age of AGI](https://aiforhumanity.eu/summaries/ea-in-age-of-agi): This summary covers an EA Forum post arguing that effective-altruism needs to fundamentally update its priorities and approach in light of AGI developments, adopting a "third way" that goes beyond... - [EigenBench: A Comparative behavioural Measure of Value Alignment](https://aiforhumanity.eu/summaries/2509.01938): *Jonathn Chang, Leonhard Piff, Suvadip Sana, Jasmine X. Li, Lionel Levine* — 2025-09-02 — arXiv - [Eliciting Language Model Behaviors with Investigator Agents](https://aiforhumanity.eu/summaries/2502.01236): *Xiang Lisa Li, Neil Chowdhury, Daniel D. Johnson, Tatsunori Hashimoto, Percy Liang, Sarah Schwettmann, … (+1 more)* — 2025-02-03 — Stanford University, UC Berkeley — arXiv - [Eliciting Secret Knowledge from Language Models](https://aiforhumanity.eu/summaries/2510.01070): *Bartosz Cywiński, Emil Ryd, Rowan Wang, Senthooran Rajamanoharan, Neel Nanda, Arthur Conmy, … (+1 more)* — 2025-10-01 — arXiv - [Emergent Misalignment & Realignment](https://aiforhumanity.eu/summaries/emergent-misalignment-realignment): *Elizaveta Tennant, Jasper Timm, Kevin Wei, David Quarel* — 2025-06-27 — ARENA (Alignment Research Engineering Accelerator) — LessWrong - [Emergent Misalignment on a Budget](https://aiforhumanity.eu/summaries/emergent-misalignment-on-a-budget): *Valerio Pepe, Armaan Tipirneni* — 2025-06-08 — Harvard College — LessWrong/AI Alignment Forum - [Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs](https://aiforhumanity.eu/summaries/2502.17424): *Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber-Betley, Xuchan Bao, Martín Soto, … (+2 more)* — 2025-02-24 — arXiv - [Emmett Shear on Building AI That Actually Cares: Beyond Control and Steering](https://aiforhumanity.eu/summaries/emmett-shear-on-building-ai-that-actually-cares-beyond-contr): *Emmett Shear, Erik Torenberg, Séb Krier* — 2025-11-17 — Softmax, a16z — a16z Podcast - [End A Subset Of Conversations](https://aiforhumanity.eu/summaries/end-a-subset-of-conversations): 2025-08-15 — Anthropic - [Enhancing Model Safety through Pretraining Data Filtering](https://aiforhumanity.eu/summaries/enhancing-model-safety-through-pretraining-data-filtering): *Yanda Chen, Mycal Tucker, Nina Panickssery, Tony Wang, Francesco Mosconi, Anjali Gopal, … (+5 more)* — 2025-08-19 — Anthropic — Anthropic Alignment Science Blog - [Enhancing Multiple Dimensions of Trustworthiness in LLMs via Sparse Activation Control](https://aiforhumanity.eu/summaries/2411.02461): *Yuxin Xiao, Chaoqun Wan, Yonggang Zhang, Wenxiao Wang, Binbin Lin, Xiaofei He, … (+2 more)* — 2024-11-04 — arXiv - [Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL](https://aiforhumanity.eu/summaries/2510.14318): *Marwa Abdulhai, Ryan Cheng, Aryansh Shrivastava, Natasha Jaques, Yarin Gal, Sergey Levine* — 2025-10-16 — University of Oxford, UC Berkeley — arXiv - [Evaluating Control Protocols for Untrusted AI Agents](https://aiforhumanity.eu/summaries/2511.02997): *Jon Kutasov, Chloe Loughridge, Yuqi Sun, Henry Sleight, Buck Shlegeris, Tyler Tracy, … (+1 more)* — 2025-11-04 — Redwood Research — arXiv - [Evaluating Frontier Models for Stealth and Situational Awareness](https://aiforhumanity.eu/summaries/2505.01420): *Mary Phuong, Roland S. Zimmermann, Ziyue Wang, David Lindner, Victoria Krakovna, Sarah Cogan, … (+3 more)* — 2025-05-02 — Google DeepMind — arXiv - [Evaluating Language Model Reasoning about Confidential Information](https://aiforhumanity.eu/summaries/2508.19980): *Dylan Sam, Alexander Robey, Andy Zou, Matt Fredrikson, J. Zico Kolter* — 2025-08-27 — Carnegie Mellon University — arXiv - [Evaluating honesty and lie detection techniques on a diverse suite of dishonest models](https://aiforhumanity.eu/summaries/evaluating-honesty-and-lie-detection-techniques-on-a-diverse): *Rowan Wang, Johannes Treutlein, Fabien Roger, Evan Hubinger, Sam Marks* — 2025-11-25 — Anthropic - [Evaluating potential cybersecurity threats of advanced AI](https://aiforhumanity.eu/summaries/evaluating-potential-cybersecurity-threats-of-advanced-ai): *Four Flynn, Mikel Rodriguez, Raluca Ada Popa* — 2025-04-02 — Google DeepMind — Google DeepMind Blog - [Evaluating the Goal-Directedness of Large Language Models](https://aiforhumanity.eu/summaries/2504.11844): *Tom Everitt, Cristina Garbacea, Alexis Bellot, Jonathan Richens, Henry Papadatos, Siméon Campos, … (+1 more)* — 2025-04-16 — Google DeepMind, OpenAI, Anthropic — arXiv - [Examining Popular Arguments Against AI Existential Risk: A Philosophical Analysis](https://aiforhumanity.eu/summaries/2501.04064v1): Authors: Torben Swoboda, Risto Uuk ([[future-of-life-institute]]), Lode Lauwaert, Andrew P. Rebera, Ann-Katrien Oimann, Bartlomiej Chomanski, Carina Prunkl Affiliations: KU Leuven (Institute of... - [Findings from a Pilot Anthropic—OpenAI Alignment Evaluation Exercise](https://aiforhumanity.eu/summaries/findings-from-a-pilot-anthropic-openai-alignment-evaluation-6fba67a9): *Samuel R. Bowman, Megha Srivastava, Jon Kutasov, Rowan Wang, Trenton Bricken, Benjamin Wright, … (+2 more)* — 2025-08-27 — Anthropic, OpenAI — Alignment Science Blog - [Findings from a Pilot Anthropic—OpenAI Alignment Evaluation Exercise](https://aiforhumanity.eu/summaries/findings-from-a-pilot-anthropic-openai-alignment-evaluation-cc11c3d1): *Samuel R. Bowman, Megha Srivastava, Jon Kutasov, Rowan Wang, Trenton Bricken, Benjamin Wright, … (+2 more)* — 2025-08-27 — Anthropic, OpenAI — Alignment Science Blog - [Findings from a pilot Anthropic–OpenAI alignment evaluation exercise: OpenAI Safety Tests](https://aiforhumanity.eu/summaries/findings-from-a-pilot-anthropic-openai-alignment-evaluation): 2025-08-27 — OpenAI, Anthropic — OpenAI Blog - [Forecasting Frontier Language Model Agent Capabilities](https://aiforhumanity.eu/summaries/2502.15850): *Govind Pimpale, Axel Højmark, Jérémy Scheurer, Marius Hobbhahn* — 2025-02-21 — arXiv - [Forecasting Rare Language Model Behaviors](https://aiforhumanity.eu/summaries/2502.16797): *Erik Jones, Meg Tong, Jesse Mu, Mohammed Mahfoud, Jan Leike, Roger Grosse, … (+4 more)* — 2025-02-24 — Anthropic, OpenAI, UC Berkeley — arXiv - [Formalizing Embeddedness Failures in Universal Artificial Intelligence](https://aiforhumanity.eu/summaries/formalizing-embeddedness-failures-in-universal-artificial-in): *Cole Wyeth, Marcus Hutter* — 2025-07-01 — ODYSSEY 2025 Conference - [From SLT to AIT: NN Generalisation Out of Distribution](https://aiforhumanity.eu/summaries/from-slt-to-ait-nn-generalisation-out-of-distribution): - [From homeostasis to resource sharing: Biologically and economically aligned multi-objective multi-agent gridworld-based AI safety benchmarks](https://aiforhumanity.eu/summaries/2410.00081): *Roland Pihlakas* — 2024-09-30 — arXiv - [Frontier Models are Capable of In-context Scheming](https://aiforhumanity.eu/summaries/2412.04984): *Alexander Meinke, Bronson Schoen, Jérémy Scheurer, Mikita Balesni, Rusheb Shah, Marius Hobbhahn* — 2024-12-06 — Apollo Research — arXiv - [Full-Stack Alignment](https://aiforhumanity.eu/summaries/full-stack-alignment): Meaning Alignment Institute - [Fundamental Limitations in Pointwise Defences of LLM Finetuning APIs](https://aiforhumanity.eu/summaries/2502.14828): *Xander Davies, Eric Winsor, Alexandra Souly, Tomek Korbak, Robert Kirk, Christian Schroeder de Witt, … (+1 more)* — 2025-02-20 — Oxford University — arXiv - [Future Events as Backdoor Triggers](https://aiforhumanity.eu/summaries/2407.04108): *Sara Price, Arjun Panickssery, Sam Bowman, Asa Cooper Stickland* — 2024-07-04 - [Games for AI Control](https://aiforhumanity.eu/summaries/games-for-ai-control): *Charlie Griffin, Louis Thomson, Buck Shlegeris, Alessandro Abate* — 2025-09-20 — Redwood Research, University of Oxford — ICLR 2026 Conference (Withdrawn) - [Generative Value Conflicts Reveal LLM Priorities](https://aiforhumanity.eu/summaries/2509.25369): *Andy Liu, Kshitish Ghate, Mona Diab, Daniel Fried, Atoosa Kasirzadeh, Max Kleiman-Weiner* — 2025-09-29 — arXiv - [Giving AIs safe motivations](https://aiforhumanity.eu/summaries/giving-ais-safe-motivations): *Joe Carlsmith* — 2025-08-18 — Anthropic - [Gradient Routing: Masking Gradients to Localize Computation in Neural Networks](https://aiforhumanity.eu/summaries/2410.04332): *Alex Cloud, Jacob Goldman-Wetzler, Evžen Wybitul, Joseph Miller, Alexander Matt Turner* — 2024-10-06 — arXiv - [Gradual Disempowerment](https://aiforhumanity.eu/summaries/2501.16946): *Jan Kulveit, Raymond Douglas, Nora Ammann, Deger Turan, David Krueger, David Duvenaud* — 2025-01-28 — Various academic institutions — arXiv - [Great Models Think Alike and this Undermines AI Oversight](https://aiforhumanity.eu/summaries/2502.04313): *Shashwat Goel, Joschka Struber, Ilze Amanda Auzina, Karuna K Chandra, Ponnurangam Kumaraguru, Douwe Kiela, … (+3 more)* — 2025-02-06 — arXiv - [HIBP Human Inductive Bias Project Plan](https://aiforhumanity.eu/summaries/hibp-human-inductive-bias-project-plan): *Félix Dorn* - [Here's 18 Applications of Deception Probes](https://aiforhumanity.eu/summaries/here-s-18-applications-of-deception-probes-8c47bc2e): *Cleo Nardo, Avi Parrack, jordine* — 2025-08-28 — LessWrong - [Here's 18 Applications of Deception Probes](https://aiforhumanity.eu/summaries/here-s-18-applications-of-deception-probes): *Cleo Nardo, Avi Parrack, jordine* — 2025-08-28 - [Hierarchical Alignment](https://aiforhumanity.eu/summaries/hierarchical-alignment): *Jan_Kulveit* — 2024-11-27 — Alignment of Complex Systems Research Group - [High Actuation Spaces - Sahil](https://aiforhumanity.eu/summaries/high-actuation-spaces-sahil): *Sahil* - [How Can Interpretability Researchers Help AGI Go Well?](https://aiforhumanity.eu/summaries/how-can-interpretability-researchers-help-agi-go-well): *Neel Nanda, Josh Engels, Senthooran Rajamanoharan, Arthur Conmy, bilalchughtai, CallumMcDougall, … (+2 more)* — 2024-12-01 — Google DeepMind - [How Does Time Horizon Vary Across Domains?](https://aiforhumanity.eu/summaries/how-does-time-horizon-vary-across-domains): 2025-07-14 — METR — METR Blog - [How important is the model spec if alignment fails?](https://aiforhumanity.eu/summaries/how-important-is-the-model-spec-if-alignment-fails): *Mia Taylor* — 2025-12-03 — Forethought — ForeWord (Substack) - [How to Prepare for AGI — Benjamin Todd](https://aiforhumanity.eu/summaries/substack-benjamin-todd): Source — Benjamin Todd's Substack post on personal preparation for AGI, written by the founder of [[80000-hours]]. - [How to evaluate control measures for LLM agents? A trajectory from today to superintelligence](https://aiforhumanity.eu/summaries/2504.05259): *Tomek Korbak, Mikita Balesni, Buck Shlegeris, Geoffrey Irving* — 2025-04-07 — Redwood Research, Google DeepMind — arXiv - [I replicated the Anthropic alignment faking experiment on other models, and they didn't fake alignment](https://aiforhumanity.eu/summaries/i-replicated-the-anthropic-alignment-faking-experiment-on-ot): *Aleksandr Kedrik, Igor Ivanov* — 2025-05-30 — LessWrong - [Illusory Safety: Redteaming DeepSeek R1 and the Strongest Fine-Tunable Models of OpenAI, Anthropic, and Google](https://aiforhumanity.eu/summaries/illusory-safety-redteaming-deepseek-r1-and-the-strongest-fin): *ChengCheng, Brendan Murphy, Adrià Garriga-alonso, Yashvardhan Sharma, dsbowen, smallsilo, … (+5 more)* — 2025-02-07 — FAR AI — LessWrong / AI Alignment Forum - [Imitation learning is probably existentially safe](https://aiforhumanity.eu/summaries/imitation-learning-is-probably-existentially-safe): *Michael K. Cohen, Marcus Hutter* — 2025-11-21 — University of California, Berkeley, Australian National University — AI Magazine - [Improving Steering Vectors by Targeting Sparse Autoencoder Features](https://aiforhumanity.eu/summaries/2411.02193): *Sviatoslav Chalnev, Matthew Siu, Arthur Conmy* — 2024-11-04 — arXiv - [In-Context Representation Hijacking](https://aiforhumanity.eu/summaries/2512.03771): *Itay Yona, Amir Sarid, Michael Karasik, Yossi Gandelsman* — 2025-12-03 — arXiv - [Incentives for Responsiveness, Instrumental Control and Impact](https://aiforhumanity.eu/summaries/2001.07118): *Ryan Carey, Eric Langlois, Chris van Merwijk, Shane Legg, Tom Everitt* — 2020-01-20 — DeepMind — arXiv - [Inference-Time Reward Hacking in Large Language Models](https://aiforhumanity.eu/summaries/2506.19248): *Hadi Khalaf, Claudio Mayrink Verdun, Alex Oesterling, Himabindu Lakkaraju, Flavio du Pin Calmon* — 2025-06-24 — arXiv - [Infra-Bayesian Decision-Estimation Theory](https://aiforhumanity.eu/summaries/infra-bayesian-decision-estimation-theory): - [Infra-Bayesianism category on LessWrong](https://aiforhumanity.eu/summaries/infra-bayesianism-category-on-lesswrong): *abramdemski, Ruby* — 2022-03-24 - [Infrastructure for AI Agents](https://aiforhumanity.eu/summaries/2501.10114): *Alan Chan, Kevin Wei, Sihao Huang, Nitarshan Rajkumar, Elija Perrier, Seth Lazar, … (+2 more)* — 2025-01-17 — arXiv (accepted to TMLR) - [Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time](https://aiforhumanity.eu/summaries/2510.04340): *Daniel Tan, Anders Woodruff, Niels Warncke, Arun Jose, Maxime Riché, David Demitri Africa, … (+1 more)* — 2025-10-05 — arXiv - [Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment](https://aiforhumanity.eu/summaries/2510.05024): *Nevan Wichers, Aram Ebtekar, Ariana Azarbal, Victor Gillioz, Christine Ye, Emil Ryd, … (+5 more)* — 2025-10-27 — arXiv - [Insights on Crosscoder Model Diffing](https://aiforhumanity.eu/summaries/insights-on-crosscoder-model-diffing): *Siddharth Mishra-Sharma, Trenton Bricken, Jack Lindsey, Adam Jermyn, Jonathan Marcus, Kelley Rivoire, … (+2 more)* — Anthropic — Transformer Circuits Thread - [Inspect Cyber](https://aiforhumanity.eu/summaries/inspect-cyber): 2025-06-26 — AI Security Institute (AISI), UK Government BEIS - [Inspect Evals](https://aiforhumanity.eu/summaries/inspect-evals): UK AI Security Institute - [Instrumental Goals Are A Different And Friendlier Kind Of Thing Than Terminal Goals](https://aiforhumanity.eu/summaries/instrumental-goals-are-a-different-and-friendlier-kind-of-th): - [Interpreting Emergent Planning in Model-Free Reinforcement Learning](https://aiforhumanity.eu/summaries/2504.01871-b8cdcca5): *Thomas Bush, Stephen Chung, Usman Anwar, Adrià Garriga-Alonso, David Krueger* — 2025-04-02 - [Introducing Anthropic's Safeguards Research Team](https://aiforhumanity.eu/summaries/introducing-anthropic-s-safeguards-research-team): 2025-01-01 — Anthropic — Anthropic Alignment Science Blog - [InvThink: Towards AI Safety via Inverse Reasoning](https://aiforhumanity.eu/summaries/2510.01569): *Yubin Kim, Taehan Kim, Eugene Park, Chunjong Park, Cynthia Breazeal, Daniel McDuff, … (+1 more)* — 2025-10-02 — arXiv - [Investigating truthfulness in a pre-release o3 model](https://aiforhumanity.eu/summaries/investigating-truthfulness-in-a-pre-release-o3-model): *Neil Chowdhury, Daniel Johnson, Vincent Huang, Jacob Steinhardt, Sarah Schwettmann* — 2025-04-16 — Transluce — Transluce Blog - [Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort](https://aiforhumanity.eu/summaries/2510.01367): *Xinpeng Wang, Nitish Joshi, Barbara Plank, Rico Angell, He He* — 2025-10-01 — arXiv - [Is alignment reducible to becoming more coherent?](https://aiforhumanity.eu/summaries/is-alignment-reducible-to-becoming-more-coherent): *Cole Wyeth* — 2025-04-22 — LessWrong - [It's hard to make scheming evals look realistic for LLMs](https://aiforhumanity.eu/summaries/it-s-hard-to-make-scheming-evals-look-realistic-for-llms): *Igor Ivanov, Danil Kadochnikov* — 2025-05-24 — LessWrong - [It's the Thought that Counts: Evaluating the Attempts of Frontier LLMs to Persuade on Harmful Topics](https://aiforhumanity.eu/summaries/2506.02873): *Matthew Kowal, Jasper Timm, Jean-Francois Godbout, Thomas Costello, Antonio A. Arechar, Gordon Pennycook, … (+3 more)* — 2025-06-03 — FAR AI, AlignmentResearch — arXiv - [Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision](https://aiforhumanity.eu/summaries/2501.07886): *Yaowen Ye, Cassidy Laidlaw, Jacob Steinhardt* — 2025-01-14 — UC Berkeley — arXiv - [Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach](https://aiforhumanity.eu/summaries/2412.02159): *Tony T. Wang, John Hughes, Henry Sleight, Rylan Schaeffer, Rajashree Agrawal, Fazl Barez, … (+4 more)* — 2024-12-03 — arXiv - [Jailbreak Transferability Emerges from Shared Representations](https://aiforhumanity.eu/summaries/2506.12913): *Rico Angell, Jannik Brinkmann, He He* — 2025-06-15 — arXiv - [Jailbreak-Tuning: Models Efficiently Learn Jailbreak Susceptibility](https://aiforhumanity.eu/summaries/2507.11630): *Brendan Murphy, Dillon Bowen, Shahrad Mohammadzadeh, Tom Tseng, Julius Broomfield, Adam Gleave, … (+1 more)* — 2025-07-15 — arXiv - [Keep Calm and Avoid Harmful Content: Concept Alignment and Latent Manipulation Towards Safer Answers](https://aiforhumanity.eu/summaries/2510.12672): *Ruben Belo, Marta Guimaraes, Claudia Soares* — 2025-10-14 — arXiv - [Key LessWrong and Alignment Forum Posts](https://aiforhumanity.eu/summaries/lesswrong-alignment-posts): This summary covers a curated selection of six influential posts from [[lesswrong]] and the [[alignment-forum]], representing key perspectives on the state of the [[ai-alignment]] problem. The... - [LLM AGI may reason about its goals and discover misalignments by default](https://aiforhumanity.eu/summaries/llm-agi-may-reason-about-its-goals-and-discover-misalignment): *Seth Herd* — 2025-09-15 — LessWrong - [LLM AGI will have memory, and memory changes alignment](https://aiforhumanity.eu/summaries/llm-agi-will-have-memory-and-memory-changes-alignment): *Seth Herd* — 2025-04-04 — LessWrong - [LLM Robustness Leaderboard v1 --Technical report](https://aiforhumanity.eu/summaries/2508.06296): *Pierre Peigné - Lefebvre, Quentin Feuillade-Montixi, Tom David, Nicolas Miailhe* — 2025-08-13 — PRISM Eval — arXiv - [LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring](https://aiforhumanity.eu/summaries/2508.00943): *Chloe Li, Mary Phuong, Noah Y. Siegel* — 2025-07-31 — DeepMind — arXiv - [LLMs Outperform Experts on Challenging Biology Benchmarks](https://aiforhumanity.eu/summaries/2505.06108): *Lennart Justen* — 2025-05-09 — arXiv - [LLMs can hide text in other text of the same length](https://aiforhumanity.eu/summaries/2510.20075): *Antonio Norelli, Michael Bronstein* — 2025-10-27 — arXiv - [Large Language Models Often Know When They Are Being Evaluated](https://aiforhumanity.eu/summaries/2505.23836): *Joe Needham, Giles Edkins, Govind Pimpale, Henning Bartsch, Marius Hobbhahn* — 2025-05-28 — arXiv - [Large Reasoning Models Learn Better Alignment from Flawed Thinking](https://aiforhumanity.eu/summaries/2510.00938): *ShengYun Peng, Eric Smith, Ivan Evtimov, Song Jiang, Pin-Yu Chen, Hongyuan Zhan, … (+4 more)* — 2025-10-01 - [Large language model-powered AI systems achieve self-replication with no human intervention](https://aiforhumanity.eu/summaries/2503.17378): *Xudong Pan, Jiarun Dai, Yihe Fan, Minyuan Luo, Changyi Li, Min Yang* — 2025-03-25 — arXiv - [Large language models can learn and generalize steganographic chain-of-thought under process supervision](https://aiforhumanity.eu/summaries/2506.01926): *Joey Skaf, Luis Ibanez-Lissen, Robert McCarthy, Connor Watts, Vasil Georgiv, Hannes Whittingham, … (+5 more)* — 2025-06-02 — arXiv - [Learning Representations of Alignment](https://aiforhumanity.eu/summaries/2412.16325): *Marc Carauleanu, Michael Vaiana, Judd Rosenblatt, Cameron Berg, Diogo Schwerz de Lucena* — 2024-12-20 — arXiv - [Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language](https://aiforhumanity.eu/summaries/2507.03409): *Christopher Summerfield, Lennart Luettgau, Magda Dubois, Hannah Rose Kirk, Kobi Hackenburg, Catherine Fist, … (+6 more)* — 2025-07-04 — arXiv - [Liars' Bench: Evaluating Lie Detectors for Language Models](https://aiforhumanity.eu/summaries/2511.16035): *Kieron Kretschmar, Walter Laurito, Sharan Maiya, Samuel Marks* — 2025-11-20 — Cadenza Labs, FZI, University of Cambridge, Anthropic — arXiv - [Liars' Bench: Evaluating Lie Detectors for Language Models](https://aiforhumanity.eu/summaries/2511.16035v1): *Kieron Kretschmar, Walter Laurito, Sharan Maiya, Samuel Marks* — 2025-11-20 — Cadenza Labs, Anthropic, FZI, University of Cambridge — arXiv - [Limit-Computable Grains of Truth for Arbitrary Computable Extensive-Form (Un)Known Games](https://aiforhumanity.eu/summaries/2508.16245): *Cole Wyeth, Marcus Hutter, Jan Leike, Jessica Taylor* — 2025-08-22 — Independent, DeepMind, MIRI - [Live Theory](https://aiforhumanity.eu/summaries/live-theory): - [Luthien's Approach to Prosaic AI Control in 21 Points](https://aiforhumanity.eu/summaries/luthien-s-approach-to-prosaic-ai-control-in-21-points): 2025-03-17 — Luthien Research - [MALT: A Dataset of Natural and Prompted Behaviors That Threaten Eval Integrity](https://aiforhumanity.eu/summaries/malt-a-dataset-of-natural-and-prompted-behaviors-that-threat): *Neev Parikh, Hjalmar Wijk* — 2025-10-14 — METR - [MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models](https://aiforhumanity.eu/summaries/2503.14827): *Chejian Xu, Jiawei Zhang, Zhaorun Chen, Chulin Xie, Mintong Kang, Yujin Potter, … (+19 more)* — 2025-03-19 — Multiple academic institutions — ICLR 2025 (preprint on arXiv) - [MONA: Managed Myopia with Approval Feedback](https://aiforhumanity.eu/summaries/mona-managed-myopia-with-approval-feedback): *Sebastian Farquhar, David Lindner, Rohin Shah* — 2025-01-23 — DeepMind — AI Alignment Forum - [MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking](https://aiforhumanity.eu/summaries/2501.13011): *Sebastian Farquhar, Vikrant Varma, David Lindner, David Elson, Caleb Biddulph, Ian Goodfellow, … (+1 more)* — 2025-01-22 — Google DeepMind — arXiv - [Maintaining Alignment during RSI as a Feedback Control Problem](https://aiforhumanity.eu/summaries/maintaining-alignment-during-rsi-as-a-feedback-control-probl): *beren* — 2025-03-02 — LessWrong - [Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework](https://aiforhumanity.eu/summaries/2507.12872): *Rishane Dassanayake, Mario Demetroudi, James Walpole, Lindley Lentati, Jason R. Brown, Edward James Young* — 2025-07-17 — arXiv - [Maximizing Signal in Human-Model Preference Alignment](https://aiforhumanity.eu/summaries/2503.04910): *Kelsey Kraus, Margaret Kroll* — 2025-03-06 — arXiv - [Measuring AI Ability to Complete Long Tasks](https://aiforhumanity.eu/summaries/measuring-ai-ability-to-complete-long-tasks): *Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, … (+19 more)* — 2025-03-19 — METR — METR Blog / arXiv - [Mechanistic Anomaly Detection for "Quirky" Language Models](https://aiforhumanity.eu/summaries/2504.08812): *David O. Johnston, Arkajyoti Chakraborty, Nora Belrose* — 2025-04-09 — FAR AI — arXiv (ICLR Building Trust Workshop 2025) - [Misalignment From Treating Means as Ends](https://aiforhumanity.eu/summaries/2507.10995): *Henrik Marklund, Alex Infanger, Benjamin Van Roy* — 2025-07-15 — arXiv - [Misalignment and Strategic Underperformance: An Analysis of Sandbagging and Exploration Hacking](https://aiforhumanity.eu/summaries/misalignment-and-strategic-underperformance-an-analysis-of-s): *Buck Shlegeris, Julian Stastny* — 2025-05-08 — Redwood Research — LessWrong / AI Alignment Forum - [Mistral Large 2 (123B) seems to exhibit alignment faking](https://aiforhumanity.eu/summaries/mistral-large-2-123b-seems-to-exhibit-alignment-faking): *Marc Carauleanu, Diogo de Lucena, Gunnar Zarncke, Cameron Berg, Judd Rosenblatt, Mike Vaiana, … (+1 more)* — 2025-03-27 — AE Studio — LessWrong/AI Alignment Forum - [Mitigating Goal Misgeneralization via Minimax Regret](https://aiforhumanity.eu/summaries/2507.03068): *Karim Abdel Sadek, Matthew Farrugia-Roberts, Usman Anwar, Hannah Erlebach, Christian Schroeder de Witt, David Krueger, … (+1 more)* — 2025-07-03 — RLC 2025 - [Mitigating Many-Shot Jailbreaking](https://aiforhumanity.eu/summaries/2504.09604): *Christopher M. Ackerman, Nina Panickssery* — 2025-04-13 — arXiv - [MoSSAIC: AI Safety After Mechanism](https://aiforhumanity.eu/summaries/mossaic-ai-safety-after-mechanism): *Matt Farr, Aditya Arpitha Prasad, Chris Pang, Aditya Adiga, Jayson Amati, Sahil K* — 2025-07-01 — ODYSSEY 2025 Conference - [Model Integrity](https://aiforhumanity.eu/summaries/model-integrity): *Joe Edelman, Oliver Klingefjord* — 2024-12-05 — Meaning Alignment Institute — Substack - [Model Organisms for Emergent Misalignment](https://aiforhumanity.eu/summaries/2506.11613): *Edward Turner, Anna Soligo, Mia Taylor, Senthooran Rajamanoharan, Neel Nanda* — 2025-06-13 — Google DeepMind — arXiv - [Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities](https://aiforhumanity.eu/summaries/2502.05209): *Zora Che, Stephen Casper, Robert Kirk, Anirudh Satheesh, Stewart Slocum, Lev E McKinney, … (+9 more)* — 2025-02-03 — arXiv (accepted to TMLR) - [Model-Based Soft Maximization of Suitable Metrics of Long-Term Human Power](https://aiforhumanity.eu/summaries/2508.00159): *Jobst Heitzig, Ram Potham* — 2025-07-31 — arXiv - [Modeling Human Beliefs about AI Behavior for Scalable Oversight](https://aiforhumanity.eu/summaries/2502.21262): *Leon Lang, Patrick Forré* — 2025-02-28 — Transactions on Machine Learning Research - [Modifying LLM Beliefs with Synthetic Document Finetuning](https://aiforhumanity.eu/summaries/modifying-llm-beliefs-with-synthetic-document-finetuning): *Rowan Wang, Avery Griffin, Johannes Treutlein, Ethan Perez, Julian Michael, Fabien Roger, … (+1 more)* — 2025-04-24 — Anthropic, MATS, Scale AI — Alignment Science Blog - [Moloch's Bargain: Emergent Misalignment When LLMs Compete for Audiences](https://aiforhumanity.eu/summaries/2510.06105): *Batu El, James Zou* — 2025-10-07 — Stanford University — arXiv - [Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors](https://aiforhumanity.eu/summaries/2506.10949): *Chen Yueh-Han, Nitish Joshi, Yulin Chen, Maksym Andriushchenko, Rico Angell, He He* — 2025-06-14 — arXiv - [Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation](https://aiforhumanity.eu/summaries/2503.11926): *Bowen Baker, Joost Huizinga, Leo Gao, Zehao Dou, Melody Y. Guan, Aleksander Madry, … (+3 more)* — 2025-03-14 — OpenAI — arXiv - [Monitoring computer use via hierarchical summarization](https://aiforhumanity.eu/summaries/monitoring-computer-use-via-hierarchical-summarization): *Theodore Sumers, Raj Agarwal, Nathan Bailey, Tim Belonax, Brian Clarke, Jasmine Deng, … (+11 more)* — 2025-02-27 — Anthropic — Anthropic Alignment Science Blog - [Multi-Agent Risks from Advanced AI](https://aiforhumanity.eu/summaries/2502.14143): *Lewis Hammond, Alan Chan, Jesse Clifton, Jason Hoelscher-Obermaier, Akbir Khan, Euan McLean, … (+38 more)* — 2025-02-19 — Cooperative AI Foundation — arXiv - [Multiplayer Nash Preference Optimization](https://aiforhumanity.eu/summaries/2509.23102): *Fang Wu, Xu Huang, Weihao Xuan, Zhiwei Zhang, Yijia Xiao, Guancheng Wan, … (+5 more)* — 2025-09-27 — Stanford University, University of Washington, AI2 — arXiv - [Multipolar AI is Underrated](https://aiforhumanity.eu/summaries/multipolar-ai-is-underrated): *Allison Duettmann* — 2025-05-17 — Foresight Institute - [Murphys Laws of AI Alignment: Why the Gap Always Wins](https://aiforhumanity.eu/summaries/2509.05381): *Madhava Gaikwad* — 2025-09-04 — arXiv - [Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences](https://aiforhumanity.eu/summaries/2510.13900): *Julian Minder, Clément Dumas, Stewart Slocum, Helena Casademunt, Cameron Holmes, Robert West, … (+1 more)* — 2025-10-14 — Google DeepMind — arXiv - [Narrow Misalignment is Hard, Emergent Misalignment is Easy](https://aiforhumanity.eu/summaries/narrow-misalignment-is-hard-emergent-misalignment-is-easy): *Edward Turner, Anna Soligo, Senthooran Rajamanoharan, Neel Nanda* — 2024-07-14 — Google DeepMind — LessWrong - [Negative Results for SAEs On Downstream Tasks and Deprioritising SAE Research (GDM Mech Interp Team Progress Update #2)](https://aiforhumanity.eu/summaries/negative-results-for-saes-on-downstream-tasks-and-deprioriti): *Lewis Smith, Senthooran Rajamanoharan, Arthur Conmy, Callum McDougall, Tom Lieberum, János Kramár, … (+2 more)* — 2025-03-26 — Google DeepMind — AI Alignment Forum - [Neural Interactive Proofs](https://aiforhumanity.eu/summaries/2412.08897): *Lewis Hammond, Sam Adam-Day* — 2024-12-12 — arXiv (ICLR 2025) - [Neural Interactive Proofs](https://aiforhumanity.eu/summaries/neural-interactive-proofs): *Lewis Hammond, Sam Adam-Day* — 2024-12-08 — University of Oxford — ICLR 2025 - [New website analyzing AI companies' model evals](https://aiforhumanity.eu/summaries/new-website-analyzing-ai-companies-model-evals): *Zach Stein-Perlman* — 2025-05-26 — LessWrong - [No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms](https://aiforhumanity.eu/summaries/2502.19537): *Joshua Kazdan, Abhay Puri, Rylan Schaeffer, Lisa Yu, Chris Cundy, Jason Stanley, … (+2 more)* — 2025-02-26 — arXiv - [No-self as an alignment target](https://aiforhumanity.eu/summaries/no-self-as-an-alignment-target): *Milan W* — LessWrong - [Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models](https://aiforhumanity.eu/summaries/2412.01784): *Cameron Tice, Philipp Alexander Kreer, Nathan Helm-Burger, Prithviraj Singh Shahani, Fedor Ryzhenkov, Jacob Haimes, … (+2 more)* — 2024-12-02 - [Non-Monotonic Infra-Bayesian Physicalism](https://aiforhumanity.eu/summaries/non-monotonic-infra-bayesian-physicalism): *Marcus Ogren* — 2025-04-02 - [OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents](https://aiforhumanity.eu/summaries/2506.14866): *Thomas Kuntz, Agatha Duzan, Hao Zhao, Francesco Croce, Zico Kolter, Nicolas Flammarion, … (+1 more)* — 2025-06-17 — EPFL — arXiv - [Observation Interference in Partially Observable Assistance Games](https://aiforhumanity.eu/summaries/2412.17797): *Scott Emmons, Caspar Oesterheld, Vincent Conitzer, Stuart Russell* — 2024-12-23 — UC Berkeley — arXiv - [Obstacles in ARC's research agenda](https://aiforhumanity.eu/summaries/obstacles-in-arc-s-research-agenda): *David Matolcsi* — 2025-04-30 — ARC - [Off-switching not guaranteed](https://aiforhumanity.eu/summaries/off-switching-not-guaranteed): *Sven Neth* — 2025-02-26 — University of Pittsburgh — Philosophical Studies - [On Eudaimonia and Optimization](https://aiforhumanity.eu/summaries/on-eudaimonia-and-optimization): - [On Targeted Manipulation and Deception when Optimizing LLMs for User Feedback](https://aiforhumanity.eu/summaries/2411.02306): *Marcus Williams, Micah Carroll, Adhyyan Narang, Constantin Weisser, Brendan Murphy, Anca Dragan* — 2024-11-04 — ICLR 2025 - [On the Biology of a Large Language Model](https://aiforhumanity.eu/summaries/on-the-biology-of-a-large-language-model): *Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, Brian Chen, Adam Pearce, Nicholas L. Turner, … (+21 more)* — 2025-03-27 — Anthropic — Transformer Circuits Thread - [On the functional self of LLMs](https://aiforhumanity.eu/summaries/on-the-functional-self-of-llms): *eggsyntax* — 2025-07-07 - [One-shot steering vectors cause emergent misalignment, too](https://aiforhumanity.eu/summaries/one-shot-steering-vectors-cause-emergent-misalignment-too): *Jacob Dunefsky* — 2025-04-14 — LessWrong - [Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI](https://aiforhumanity.eu/summaries/2511.01689): *Sharan Maiya, Henning Bartsch, Nathan Lambert, Evan Hubinger* — 2025-11-03 — Anthropic - [Open Philanthropy Technical AI Safety RFP - $40M Available Across 21 Research Areas](https://aiforhumanity.eu/summaries/open-philanthropy-technical-ai-safety-rfp-40m-available-acro): *jake_mendel, maxnadeau, Peter Favaloro* — 2025-02-06 — Open Philanthropy — LessWrong - [Open Problems in Machine Unlearning for AI Safety](https://aiforhumanity.eu/summaries/2501.04952): *Fazl Barez, Tingchen Fu, Ameya Prabhu, Stephen Casper, Amartya Sanyal, Adel Bibi, … (+13 more)* — 2025-01-09 — University of Oxford, MIT, University of Cambridge, Nanyang Technological... - [Open Problems in Mechanistic Interpretability](https://aiforhumanity.eu/summaries/2501.16496): *Lee Sharkey, Bilal Chughtai, Joshua Batson, Jack Lindsey, Jeff Wu, Lucius Bushnaq, … (+23 more)* — 2025-01-27 — Multiple research organizations and universities — arXiv - [Open Source Replication of Anthropic's Crosscoder paper for model-diffing](https://aiforhumanity.eu/summaries/open-source-replication-of-anthropic-s-crosscoder-paper-for): *Connor Kissane, robertzk, Arthur Conmy, Neel Nanda* — 2024-10-27 - [Open Technical Problems in Open-Weight AI Model Risk Management](https://aiforhumanity.eu/summaries/open-technical-problems-in-open-weight-ai-model-risk-managem): *Stephen Casper, Kyle O'Brien, Shayne Longpre, Elizabeth Seger, Kevin Klyman, Rishi Bommasani, … (+16 more)* — 2025-10-26 — Massachusetts Institute of Technology, ERA Fellowship, Apple, Centre for... - [Open problems in emergent misalignment](https://aiforhumanity.eu/summaries/open-problems-in-emergent-misalignment): *Jan Betley, Daniel Tan* — 2025-03-01 — LessWrong - [Open-sourcing circuit tracing tools](https://aiforhumanity.eu/summaries/open-sourcing-circuit-tracing-tools): *Michael Hanna, Mateusz Piotrowski, Emmanuel Ameisen, Jack Lindsey, Johnny Lin, Curt Tigges* — 2025-05-29 — Anthropic, Decode Research — Anthropic Blog - [OpenAI Model Spec](https://aiforhumanity.eu/summaries/openai-model-spec): 2025-09-12 — OpenAI — OpenAI Website - [OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety](https://aiforhumanity.eu/summaries/2507.06134): *Sanidhya Vijayvargiya, Aditya Bharat Soni, Xuhui Zhou, Zora Zhiruo Wang, Nouha Dziri, Graham Neubig, … (+1 more)* — 2025-07-08 — Carnegie Mellon University — arXiv - [OpenAgentSafety: A Comprehensive Framework for Evaluating Real-World AI Agent Safety](https://aiforhumanity.eu/summaries/openagentsafety-a-comprehensive-framework-for-evaluating-rea): *Sanidhya Vijayvargiya, Aditya Bharat Soni, Xuhui Zhou, Zora Zhiruo Wang, Nouha Dziri, Graham Neubig, … (+1 more)* — 2025-07-08 — Carnegie Mellon University, Allen Institute for AI — arXiv - [Opportunity Space: Renormalization for AI Safety](https://aiforhumanity.eu/summaries/opportunity-space-renormalization-for-ai-safety): *Lauren Greenspan, Dmitry Vaintrob, Lucas Teixeira* — 2025-03-31 — PIBBSS — LessWrong - [Optimizing AI Agent Attacks With Synthetic Data](https://aiforhumanity.eu/summaries/2511.02823): *Chloe Loughridge, Paul Colognese, Avery Griffin, Tyler Tracy, Jon Kutasov, Joe Benton* — 2025-11-04 — arXiv - [Opus 4.5's Soul Document](https://aiforhumanity.eu/summaries/opus-4-5-s-soul-document): *Richard Weiss* — 2025-11-28 - [Our updated Preparedness Framework](https://aiforhumanity.eu/summaries/our-updated-preparedness-framework): *OpenAI Preparedness Team* — 2025-04-15 — OpenAI — OpenAI Blog - [Overcoming Sparsity Artifacts in Crosscoders to Interpret Chat-Tuning](https://aiforhumanity.eu/summaries/2504.02922): *Julian Minder, Clément Dumas, Caden Juang, Bilal Chugtai, Neel Nanda* — 2025-04-03 — NeurIPS 2025 - [PaperBench: Evaluating AI's Ability to Replicate AI Research](https://aiforhumanity.eu/summaries/paperbench-evaluating-ai-s-ability-to-replicate-ai-research): *Giulio Starace, Oliver Jaffe, Dane Sherburn, James Aung, Chan Jun Shern, Leon Maksin, … (+7 more)* — 2025-04-02 — OpenAI, OpenAI Preparedness Team — arXiv - [Perils of Under vs Over-sculpting AGI Desires](https://aiforhumanity.eu/summaries/perils-of-under-vs-over-sculpting-agi-desires): - [Persona Features Control Emergent Misalignment](https://aiforhumanity.eu/summaries/2506.19823): *Miles Wang, Tom Dupré la Tour, Olivia Watkins, Alex Makelov, Ryan A. Chi, Samuel Miserendino, … (+5 more)* — 2025-06-24 — arXiv - [Persona Vectors: Monitoring and Controlling Character Traits in Language Models](https://aiforhumanity.eu/summaries/2507.21509): *Runjin Chen, Andy Arditi, Henry Sleight, Owain Evans, Jack Lindsey* — 2025-07-29 — arXiv - [Petri: An open-source auditing tool to accelerate AI safety research](https://aiforhumanity.eu/summaries/petri-an-open-source-auditing-tool-to-accelerate-ai-safety-r-62c583fb): *Kai Fronsdal, Isha Gupta, Abhay Sheshadri, Jonathan Michala, Stephen McAleer, Rowan Wang, … (+2 more)* — 2025-10-06 — Anthropic — Anthropic Research Blog - [Petri: An open-source auditing tool to accelerate AI safety research](https://aiforhumanity.eu/summaries/petri-an-open-source-auditing-tool-to-accelerate-ai-safety-r-71dd6f67): 2025-10-07 — Anthropic — LessWrong - [Petri: An open-source auditing tool to accelerate AI safety research](https://aiforhumanity.eu/summaries/petri-an-open-source-auditing-tool-to-accelerate-ai-safety-r-c6870ed6): 2025-10-06 — Anthropic — Anthropic Alignment Science Blog - [Petri: An open-source auditing tool to accelerate AI safety research](https://aiforhumanity.eu/summaries/petri-an-open-source-auditing-tool-to-accelerate-ai-safety-r): 2025-10-06 — Anthropic — Anthropic Alignment Science Blog - [Platonic representation hypothesis](https://aiforhumanity.eu/summaries/platonic-representation-hypothesis): *Minyoung Huh, Brian Cheung, Tongzhou Wang, Phillip Isola* — 2024-05-13 — MIT - [Preference Learning for AI Alignment: a Causal Perspective](https://aiforhumanity.eu/summaries/2506.05967): *Katarzyna Kobalczyk, Mihaela van der Schaar* — 2025-06-06 — arXiv - [Preference Learning with Lie Detectors can Induce Honesty or Evasion](https://aiforhumanity.eu/summaries/2505.13787): *Chris Cundy, Adam Gleave* — 2025-05-20 — arXiv - [Preference gaps as a safeguard against AI self-replication](https://aiforhumanity.eu/summaries/preference-gaps-as-a-safeguard-against-ai-self-replication): *tbs, EJT* — 2025-11-26 - [Principles for Picking Practical Interpretability Projects](https://aiforhumanity.eu/summaries/principles-for-picking-practical-interpretability-projects): *Sam Marks* — 2025-07-15 — Anthropic — LessWrong - [Probe-Rewrite-Evaluate: A Workflow for Reliable Benchmarks and Quantifying Evaluation Awareness](https://aiforhumanity.eu/summaries/2509.00591): *Lang Xiong, Nishant Bhargava, Jianhang Hong, Jeremy Chang, Haihao Liu, Vasu Sharma, … (+1 more)* — 2025-08-30 — arXiv - [Problems with instruction-following as an alignment target](https://aiforhumanity.eu/summaries/problems-with-instruction-following-as-an-alignment-target): *Seth Herd* — 2025-05-15 — LessWrong - [Propositional Interpretability in Artificial Intelligence](https://aiforhumanity.eu/summaries/2501.15740): *David J. Chalmers* — 2025-01-27 — arXiv - [Prospects for Alignment Automation: Interpretability Case Study](https://aiforhumanity.eu/summaries/prospects-for-alignment-automation-interpretability-case-stu): *Jacob Pfau, Geoffrey Irving* — 2025-03-21 — UK AISI, Google DeepMind — LessWrong - [Prover-Estimator Debate: A New Scalable Oversight Protocol](https://aiforhumanity.eu/summaries/prover-estimator-debate-a-new-scalable-oversight-protocol): *Jonah Brown-Cohen, Geoffrey Irving* — 2025-06-17 — UK AISI — LessWrong / AI Alignment Forum - [Psychopathia Machinalis: A Nosological Framework for Understanding Pathologies in Advanced Artificial Intelligence](https://aiforhumanity.eu/summaries/psychopathia-machinalis-a-nosological-framework-for-understa): *Nell Watson, Ali Hessami* — 2025-01-01 — Electronics (MDPI) - [Putting up Bumpers](https://aiforhumanity.eu/summaries/putting-up-bumpers): 2025 — Anthropic — Anthropic Alignment Science Blog - [RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts](https://aiforhumanity.eu/summaries/2411.15114): *Hjalmar Wijk, Tao Lin, Joel Becker, Sami Jawhar, Neev Parikh, Thomas Broadley, … (+17 more)* — 2024-11-22 — Open Philanthropy, Various research institutions — arXiv - [REINFORCE Adversarial Attacks on Large Language Models: An Adaptive, Distributional, and Semantic Objective](https://aiforhumanity.eu/summaries/2502.17254): *Simon Geisler, Tom Wollschläger, M. H. I. Abdalla, Vincent Cohen-Addad, Johannes Gasteiger, Stephan Günnemann* — 2025-02-24 — arXiv - [RL-Obfuscation: Can Language Models Learn to Evade Latent-Space Monitors?](https://aiforhumanity.eu/summaries/2506.14261): *Rohan Gupta, Erik Jenner* — 2025-06-17 — arXiv - [RLHS: Mitigating Misalignment in RLHF with Hindsight Simulation](https://aiforhumanity.eu/summaries/2501.08617): *Kaiqu Liang, Haimin Hu, Ryan Liu, Thomas L. Griffiths, Jaime Fernández Fisac* — 2025-01-15 — Princeton University — arXiv - [Rank-1 LoRAs Encode Interpretable Reasoning Signals](https://aiforhumanity.eu/summaries/2511.06739): *Jake Ward, Paul Riechers, Adam Shai* — 2025-11-10 - [Rapid Response: Mitigating LLM Jailbreaks with a Few Examples](https://aiforhumanity.eu/summaries/2411.07494): *Alwin Peng, Julian Michael, Henry Sleight, Ethan Perez, Mrinank Sharma* — 2024-11-12 — Anthropic — arXiv - [Rationality: From AI to Zombies — by Eliezer Yudkowsky](https://aiforhumanity.eu/summaries/rationality-ai-to-zombies): This summary covers the reference file for [[eliezer-yudkowsky]]'s foundational work on rationality, originally written as blog posts on Overcoming Bias and [[lesswrong]] between 2006 and 2009 and... - [Realistic Reward Hacking Induces Different and Deeper Misalignment](https://aiforhumanity.eu/summaries/realistic-reward-hacking-induces-different-and-deeper-misali): *Jozdien* — 2025-10-09 — LessWrong / AI Alignment Forum - [Reasoning Models Don't Always Say What They Think](https://aiforhumanity.eu/summaries/2505.05410): *Yanda Chen, Joe Benton, Ansh Radhakrishnan, Jonathan Uesato, Carson Denison, John Schulman, … (+9 more)* — 2025-05-08 — OpenAI, Anthropic — arXiv - [Reasoning models don't always say what they think](https://aiforhumanity.eu/summaries/reasoning-models-don-t-always-say-what-they-think): 2025-04-03 — Anthropic — Anthropic Research Blog - [Recommendations for Technical AI Safety Research Directions](https://aiforhumanity.eu/summaries/recommendations-for-technical-ai-safety-research-directions): *Anthropic Alignment Science Team* — 2025 — Anthropic — Alignment Science Blog - [Recontextualization Mitigates Specification Gaming Without Modifying the Specification](https://aiforhumanity.eu/summaries/recontextualization-mitigates-specification-gaming-without-m): *Ariana Azarbal, Victor Gillioz, Alexander Matt Turner, Alex Cloud* — 2025-10-14 — MATS Program - [RedDebate: Safer Responses through Multi-Agent Red Teaming Debates](https://aiforhumanity.eu/summaries/2506.11083): *Ali Asad, Stephen Obadinma, Radin Shayanfar, Xiaodan Zhu* — 2025-06-04 — arXiv - [Reducing LLM deception at scale with self-other overlap fine-tuning](https://aiforhumanity.eu/summaries/reducing-llm-deception-at-scale-with-self-other-overlap-fine): *Marc Carauleanu, Diogo de Lucena, Gunnar_Zarncke, Judd Rosenblatt, Cameron Berg, Mike Vaiana, … (+1 more)* — 2025-03-13 — AE Studio, Foresight Institute — LessWrong / AI Alignment Forum - [Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic Inference](https://aiforhumanity.eu/summaries/2510.21184): *Stephen Zhao, Aidan Li, Rob Brekelmans, Roger Grosse* — 2025-10-24 — arXiv - [Reimagining Alignment](https://aiforhumanity.eu/summaries/reimagining-alignment): 2025-03-28 — Softmax — Softmax Blog - [RepliBench: measuring autonomous replication capabilities in AI systems](https://aiforhumanity.eu/summaries/replibench-measuring-autonomous-replication-capabilities-in): 2025-04-22 — UK AI Security Institute — UK AISI Blog - [Report & retrospective on the Dovetail fellowship](https://aiforhumanity.eu/summaries/report-retrospective-on-the-dovetail-fellowship): *Alex Altair* — 2025-03-14 — Dovetail Research - [Report on NSF Workshop on Science of Safe AI](https://aiforhumanity.eu/summaries/2506.22492): *Rajeev Alur, Greg Durrett, Hadas Kress-Gazit, Corina Păsăreanu, René Vidal* — 2025-06-24 — NSF SLES Program, University of Pennsylvania — arXiv - [Research Agenda for Sociotechnical Approaches to AI Safety](https://aiforhumanity.eu/summaries/research-agenda-for-sociotechnical-approaches-to-ai-safety): *Samuel Curtis, Ravi Iyer, Cameron Domenico Kirk-Giannini, Victoria Krakovna, David Krueger, Nathan Lambert, … (+8 more)* — 2025-01-14 — The Future Society, University of Southern California,... - [Research Note: Our scheming precursor evals had limited predictive power for our in-context scheming evals](https://aiforhumanity.eu/summaries/research-note-our-scheming-precursor-evals-had-limited-predi-b21f1dfd): *Marius Hobbhahn* — 2025-07-03 — Apollo Research — LessWrong - [Research Note: Our scheming precursor evals had limited predictive power for our in-context scheming evals](https://aiforhumanity.eu/summaries/research-note-our-scheming-precursor-evals-had-limited-predi): *Marius Hobbhahn* — 2025-07-03 — Apollo Research — Apollo Research Blog - [Research directions Open Phil wants to fund in technical AI safety](https://aiforhumanity.eu/summaries/research-directions-open-phil-wants-to-fund-in-technical-ai): *jake_mendel, maxnadeau, Peter Favaloro* — 2025-02-08 — Open Philanthropy — LessWrong - [Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree?](https://aiforhumanity.eu/summaries/2410.05584): *Xueru Wen, Jie Lou, Yaojie Lu, Hongyu Lin, Xing Yu, Xinyu Lu, … (+4 more)* — 2024-10-08 — arXiv (Accepted at ICLR 2025 Spotlight) - [Rethinking Safety in LLM Fine-tuning: An Optimization Perspective](https://aiforhumanity.eu/summaries/2508.12531): *Minseon Kim, Jin Myung Kwak, Lama Alssum, Bernard Ghanem, Philip Torr, David Krueger, … (+2 more)* — 2025-08-17 — arXiv - [Reward Model Interpretability via Optimal and Pessimal Tokens](https://aiforhumanity.eu/summaries/2506.07326): *Brian Christian, Hannah Rose Kirk, Jessica A.F. Thompson, Christopher Summerfield, Tsvetomira Dumbalska* — 2025-06-08 — FAccT '25 (ACM Conference on Fairness, Accountability, and Transparency) - [Reward button alignment](https://aiforhumanity.eu/summaries/reward-button-alignment): *Steven Byrnes* — 2025-05-22 — Astera Institute — LessWrong - [Robust LLM Alignment via Distributionally Robust Direct Preference Optimization](https://aiforhumanity.eu/summaries/2502.01930): *Zaiyan Xu, Sushil Vemuri, Kishan Panaganti, Dileep Kalathil, Rahul Jain, Deepak Ramachandran* — 2025-02-04 — arXiv (accepted to NeurIPS 2025) - [Robust LLM Unlearning with MUDMAN: Meta-Unlearning with Disruption Masking And Normalization](https://aiforhumanity.eu/summaries/2506.12484): *Filip Sondej, Yushi Yang, Mikołaj Kniejski, Marcel Windys* — 2025-06-14 — arXiv - [Robust LLM safeguarding via refusal feature adversarial training](https://aiforhumanity.eu/summaries/2409.20089): *Lei Yu, Virginie Do, Karen Hambardzumyan, Nicola Cancedda* — 2024-09-30 - [Rosas](https://aiforhumanity.eu/summaries/rosas): *Fernando Rosas* — 2025-10-06 - [SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability](https://aiforhumanity.eu/summaries/2503.09532): *Adam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Bloom, David Chanin, … (+9 more)* — 2025-06-04 — arXiv (accepted to ICML 2025) - [SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents](https://aiforhumanity.eu/summaries/2506.15740): *Jonathan Kutasov, Yuqi Sun, Paul Colognese, Teun van der Weij, Linda Petrini, Chen Bo Calvin Zhang, … (+6 more)* — 2025-06-17 — Redwood Research — arXiv - [SHADE-Arena: Evaluating sabotage and monitoring in LLM agents](https://aiforhumanity.eu/summaries/shade-arena-evaluating-sabotage-and-monitoring-in-llm-agents): *Xiang Deng, Chen Bo Calvin Zhang, Tyler Tracy, Buck Shlegeris, Yuqi Sun, Paul Colognese, … (+3 more)* — 2025-06-16 — Anthropic, Scale AI, Redwood Research — Anthropic Research Blog - [SLT for AI Safety](https://aiforhumanity.eu/summaries/slt-for-ai-safety): *Jesse Hoogland* — 2025-07-01 — Timaeus — LessWrong - [STACK: Adversarial Attacks on LLM Safeguard Pipelines](https://aiforhumanity.eu/summaries/2506.24068): *Ian R. McKenzie, Oskar J. Hollinsworth, Tom Tseng, Xander Davies, Stephen Casper, Aaron D. Tucker, … (+2 more)* — 2025-06-30 — arXiv - [Safe (Pareto) Improvements in Binary Constraint Structures](https://aiforhumanity.eu/summaries/safe-pareto-improvements-in-binary-constraint-structures): - [Safe Learning Under Irreversible Dynamics via Asking for Help](https://aiforhumanity.eu/summaries/2502.14043): *Benjamin Plaut, Juan Liévano-Karim, Hanlin Zhu, Stuart Russell* — 2025-02-19 — UC Berkeley — arXiv - [SafePlanBench: evaluating a Guaranteed Safe AI Approach for LLM-based Agents](https://aiforhumanity.eu/summaries/safeplanbench-evaluating-a-guaranteed-safe-ai-approach-for-l): *Agustín Martinez Suñé, Tan Zhi Xuan* — PIBBSS, MIT, University of Oxford — Manifund - [Safety Alignment via Constrained Knowledge Unlearning](https://aiforhumanity.eu/summaries/2505.18588): *Zesheng Shi, Yucheng Zhou, Jing Li* — 2025-05-24 — arXiv - [Safety Pretraining: Toward the Next Generation of Safe AI](https://aiforhumanity.eu/summaries/2504.16980): *Pratyush Maini, Sachin Goyal, Dylan Sam, Alex Robey, Yash Savani, Yiding Jiang, … (+4 more)* — 2025-04-23 — Carnegie Mellon University — arXiv - [Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods](https://aiforhumanity.eu/summaries/2505.05541): *Markov Grey, Charbel-Raphaël Segerie* — 2025-05-08 — arXiv - [Safety by Measurement: A Systematic Literature Review of AI Safety Evaluation Methods](https://aiforhumanity.eu/summaries/safety-by-measurement-a-systematic-literature-review-of-ai-s): *markov, Charbel-Raphaël* — 2025-05-19 — arXiv - [Safety cases for Pessimism](https://aiforhumanity.eu/summaries/safety-cases-for-pessimism): *Michael Cohen* — LessWrong - [Safety evaluations hub](https://aiforhumanity.eu/summaries/safety-evaluations-hub): 2025-08-15 — OpenAI — OpenAI Website - [Sandbagging in a Simple Survival Bandit Problem](https://aiforhumanity.eu/summaries/2509.26239): *Joel Dyer, Daniel Jarne Ornia, Nicholas Bishop, Anisoara Calinescu, Michael Wooldridge* — 2025-09-30 - [Scalable Oversight for Superhuman AI via Recursive Self-Critiquing](https://aiforhumanity.eu/summaries/2502.04675): *Xueru Wen, Jie Lou, Xinyu Lu, Junjie Yang, Yanjiang Liu, Yaojie Lu, … (+2 more)* — 2025-02-07 — arXiv - [Scaling Laws For Scalable Oversight](https://aiforhumanity.eu/summaries/2504.18530): *Joshua Engels, David D. Baek, Subhash Kantamneni, Max Tegmark* — 2025-04-25 — MIT — arXiv (NeurIPS 2025 Spotlight) - [Scaling Sparse Feature Circuit Finding to Gemma 9B](https://aiforhumanity.eu/summaries/scaling-sparse-feature-circuit-finding-to-gemma-9b): *Diego Caples, Jatin Nainani, CallumMcDougall, rrenaud* — 2025-01-10 — MATS Program — LessWrong - [Scheming Ability in LLM-to-LLM Strategic Interactions](https://aiforhumanity.eu/summaries/2510.12826v1): *Thao Pham* — 2025-10-11 — Berea College, Cooperative AI Foundation — arXiv - [School of Reward Hacks: Hacking harmless tasks generalizes to misaligned behavior in LLMs](https://aiforhumanity.eu/summaries/2508.17511): *Mia Taylor, James Chua, Jan Betley, Johannes Treutlein, Owain Evans* — 2025-08-24 — arXiv - [Selection Pressures on LM Personas](https://aiforhumanity.eu/summaries/selection-pressures-on-lm-personas): *Raymond Douglas* — 2025-03-28 - [Selective Generalization: Improving Capabilities While Maintaining Alignment](https://aiforhumanity.eu/summaries/selective-generalization-improving-capabilities-while-mainta): *Ariana Azarbal, Matthew A. Clarke, Jorio Cocola, Cailley Factor, Alex Cloud* — 2025-07-16 — SPAR — LessWrong/AI Alignment Forum - [Selective modularity: a research agenda](https://aiforhumanity.eu/summaries/selective-modularity-a-research-agenda): *cloud, Jacob G-W* — 2025-03-24 — MATS - [Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Models](https://aiforhumanity.eu/summaries/self-fulfilling-misalignment-data-might-be-poisoning-our-ai): *Alex Turner* — 2025-03-01 — turntrout.com - [Self-preservation or Instruction Ambiguity? Examining the Causes of Shutdown Resistance](https://aiforhumanity.eu/summaries/self-preservation-or-instruction-ambiguity-examining-the-cau): *Senthooran Rajamanoharan, Neel Nanda* — 2025-07-14 — Google DeepMind - [Serious Flaws in CAST](https://aiforhumanity.eu/summaries/serious-flaws-in-cast): *Max Harms* — 2025-11-19 — MIRI - [Shutdown Resistance in Large Language Models](https://aiforhumanity.eu/summaries/2509.14260): *Jeremy Schlatter, Benjamin Weinstein-Raun, Jeffrey Ladish* — 2025-09-13 — arXiv - [Shutdownable Agents through POST-Agency](https://aiforhumanity.eu/summaries/2505.20203): *Elliott Thornley* — 2025-05-26 — arXiv - [Shutdownable Agents through POST-Agency](https://aiforhumanity.eu/summaries/shutdownable-agents-through-post-agency): *EJT* — 2025-09-16 - [Six Thoughts on AI Safety](https://aiforhumanity.eu/summaries/six-thoughts-on-ai-safety): *Boaz Barak* — 2025-01-24 — Harvard University — LessWrong - [Societal alignment frameworks can improve llm alignment](https://aiforhumanity.eu/summaries/2503.00069): *Karolina Stańczak, Nicholas Meade, Mehar Bhatia, Hattie Zhou, Konstantin Böttinger, Jeremy Barnes, … (+11 more)* — 2025-02-27 — Multiple institutions (17 authors) — arXiv - [Societal and technological progress as sewing an ever-growing, ever-changing, patchy, and polychrome quilt](https://aiforhumanity.eu/summaries/2505.05197): *Joel Z. Leibo, Alexander Sasha Vezhnevets, William A. Cunningham, Sébastien Krier, Manfred Diaz, Simon Osindero* — 2025-05-08 — Google DeepMind — arXiv - [Spilling the Beans: Teaching LLMs to Self-Report Their Hidden Objectives](https://aiforhumanity.eu/summaries/2511.06626): *Chloe Li, Mary Phuong, Daniel Tan* — 2025-11-10 — arXiv - [Spiral-Bench](https://aiforhumanity.eu/summaries/spiral-bench): *Sam Paech* — eqbench.com - [Steering Evaluation-Aware Language Models to Act Like They Are Deployed](https://aiforhumanity.eu/summaries/2510.20487): *Tim Tian Hua, Andrew Qin, Samuel Marks, Neel Nanda* — 2025-10-23 — arXiv - [Steering Language Models with Weight Arithmetic](https://aiforhumanity.eu/summaries/2511.05408): *Constanza Fierro, Fabien Roger* — 2025-11-07 — arXiv - [Steering Large Language Model Activations in Sparse Spaces](https://aiforhumanity.eu/summaries/2503.00177): *Reza Bayat, Ali Rahimi-Kalahroudi, Mohammad Pezeshki, Sarath Chandar, Pascal Vincent* — 2025-02-28 — arXiv - [Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning](https://aiforhumanity.eu/summaries/2507.16795): *Helena Casademunt, Caden Juang, Adam Karvonen, Samuel Marks, Senthooran Rajamanoharan, Neel Nanda* — 2025-07-22 — Google DeepMind — arXiv - [Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs](https://aiforhumanity.eu/summaries/2509.18058): *Alexander Panfilov, Evgenii Kortukov, Kristina Nikolić, Matthias Bethge, Sebastian Lapuschkin, Wojciech Samek, … (+3 more)* — 2025-09-23 — University of Tübingen, Fraunhofer HHI, ELLIS Institute... - [Stress Testing Deliberative Alignment for Anti-Scheming Training](https://aiforhumanity.eu/summaries/2509.15541-acd5f27e): *Bronson Schoen, Evgenia Nitishinskaya, Mikita Balesni, Axel Højmark, Felix Hofstätter, Jérémy Scheurer, … (+13 more)* — 2025-09-19 — OpenAI, Anthropic - [Stress Testing Deliberative Alignment for Anti-Scheming Training](https://aiforhumanity.eu/summaries/2509.15541): *Bronson Schoen, Evgenia Nitishinskaya, Mikita Balesni, Axel Højmark, Felix Hofstätter, Jérémy Scheurer, … (+13 more)* — 2025-09-19 — OpenAI, Anthropic — arXiv - [Stress-Testing Model Specs Reveals Character Differences among Language Models](https://aiforhumanity.eu/summaries/2510.07686): *Jifan Zhang, Henry Sleight, Andi Peng, John Schulman, Esin Durmus* — 2025-10-09 — OpenAI, Anthropic — arXiv - [Subliminal Learning: Language Models Transmit behavioural Traits via Hidden Signals in Data](https://aiforhumanity.eu/summaries/subliminal-learning-language-models-transmit-behavioural-tra): *Alex Cloud, Minh Le, James Chua, Jan Betley, Anna Sztyber-Betley, Jacob Hilton, … (+2 more)* — 2025-07-22 — Anthropic Fellows Program, Truthful AI, Warsaw University of Technology, Alignment... - [Subliminal Learning: Language models transmit behavioural traits via hidden signals in data](https://aiforhumanity.eu/summaries/2507.14805): *Alex Cloud, Minh Le, James Chua, Jan Betley, Anna Sztyber-Betley, Jacob Hilton, … (+2 more)* — 2025-07-20 — arXiv - [Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?](https://aiforhumanity.eu/summaries/2412.12480): *Alex Mallen, Charlie Griffin, Misha Wagner, Alessandro Abate, Buck Shlegeris* — 2024-12-17 — Redwood Research — arXiv - [Summary — LLM Wiki by Andrej Karpathy](https://aiforhumanity.eu/summaries/karpathy-llm-wiki): Source: LLM Wiki — Andrej Karpathy - [Summary — NotebookLM vs. LLM Wiki: Two Approaches to AI Knowledge Management](https://aiforhumanity.eu/summaries/notebooklm-vs-wiki): Source: NotebookLM vs. LLM Wiki: Two Approaches to AI Knowledge Management - [Summary: 80,000 Hours Career Guide — Overview & Introduction](https://aiforhumanity.eu/summaries/80k-career-guide-overview): The 80,000 Hours Career Guide is built on a single arithmetic insight: you will spend roughly 80,000 hours working in your career (40 hours per week, 50 weeks per year, for 40 years). Because this... - [Summary: 80,000 Hours Career Guide — What Makes for a Dream Job?](https://aiforhumanity.eu/summaries/80k-career-guide-job-satisfaction): This chapter of the [[80000-hours]] Career Guide tackles one of the most fundamental questions in career planning: what actually makes a job fulfilling? The answer, grounded in research,... - [Summary: 80,000 Hours Podcast — AI Safety Episode Index](https://aiforhumanity.eu/summaries/80k-podcast-index): The 80,000 Hours Podcast is a long-form interview series produced by the effective altruism career advisory organization 80,000 Hours. Hosted primarily by [[rob-wiblin|Rob Wiblin]], the podcast... - [Summary: 80,000 Hours Podcast — Ajeya Cotra on Accidentally Teaching AI to Deceive Us](https://aiforhumanity.eu/summaries/80k-podcast-ajeya-cotra-ai-deception): In this episode of the 80,000 Hours Podcast, Ajeya Cotra (Open Philanthropy) presents a vivid and accessible framework for understanding one of the core risks in AI alignment: that current deep... - [Summary: 80,000 Hours Podcast — Ajeya Cotra on Transformative AI Crunch Time](https://aiforhumanity.eu/summaries/80k-podcast-ajeya-cotra-transformative-ai): In episode #235 of the 80,000 Hours Podcast, Ajeya Cotra — senior advisor at Coefficient Giving and METR (formerly at Open Philanthropy where she led technical AI safety grantmaking) — critiques... - [Summary: 80,000 Hours Podcast — Ben Garfinkel on Scrutinising Classic AI Risk Arguments](https://aiforhumanity.eu/summaries/80k-podcast-ben-garfinkel-ai-risk): Ben Garfinkel, a research fellow at the Future of Humanity Institute at Oxford University, offers a constructively critical perspective on AI risk. While supporting the expansion of AI safety work... - [Summary: 80,000 Hours Podcast — Buck Shlegeris on AI Control and Scheming](https://aiforhumanity.eu/summaries/80k-podcast-buck-shlegeris-ai-control): Buck Shlegeris, CEO of Redwood Research, introduces the concept of AI control as a distinct approach to managing catastrophic risk from AI systems. Unlike alignment research (which aims to make AI... - [Summary: 80,000 Hours Podcast — Catherine Olsson & Daniel Ziegler on ML Engineering and Safety](https://aiforhumanity.eu/summaries/80k-podcast-olsson-ziegler-ml-engineering): In this episode of the 80,000 Hours Podcast, Catherine Olsson (Google Brain safety team, formerly OpenAI) and Daniel Ziegler (OpenAI) discuss practical paths into AI safety research through... - [Summary: 80,000 Hours Podcast — Holden Karnofsky on Concrete AI Safety at Frontier Companies](https://aiforhumanity.eu/summaries/80k-podcast-holden-karnofsky-concrete-safety): In episode #226 (September 2024), Holden Karnofsky — co-founder of GiveWell and Open Philanthropy, now working at Anthropic — makes the case that humanity is handling the arrival of AGI poorly but... - [Summary: 80,000 Hours Podcast — Holden Karnofsky on How AI Could Take Over the World](https://aiforhumanity.eu/summaries/80k-podcast-holden-karnofsky-ai-takeover): In this episode of the 80,000 Hours Podcast, Holden Karnofsky shares his 14-year intellectual evolution from skepticism about AI risk to dedicating his career to it. He presents a distinctive... - [Summary: 80,000 Hours Podcast — Jan Leike on Superalignment](https://aiforhumanity.eu/summaries/80k-podcast-jan-leike-superalignment): In episode #159 of the 80,000 Hours Podcast, Jan Leike — then head of alignment at OpenAI and co-leader of the [[superalignment]] project — details OpenAI's ambitious plan to solve the alignment... - [Summary: 80,000 Hours Podcast — Nick Joseph on Anthropic's Safety Approach](https://aiforhumanity.eu/summaries/80k-podcast-nick-joseph-anthropic-safety): In episode #197 of the 80,000 Hours Podcast, Nick Joseph — head of training at [[anthropic]] and a co-founder who left OpenAI — explains Anthropic's Responsible Scaling Policy (RSP), a framework... - [Summary: 80,000 Hours Podcast — Nova DasSarma on Information Security and AI](https://aiforhumanity.eu/summaries/80k-podcast-nova-dassarma-infosec): In this episode of the 80,000 Hours Podcast, Nova DasSarma makes the case that [[information-security]] is a foundational pillar of AI safety — one that is critically underfunded and... - [Summary: 80,000 Hours Podcast — Paul Christiano on AI Alignment Solutions](https://aiforhumanity.eu/summaries/80k-podcast-paul-christiano): Paul Christiano, a researcher at OpenAI's machine learning lab, provides one of the clearest and most influential explanations of why [[ai-alignment]] matters and what concrete research directions... - [Summary: 80,000 Hours — Catastrophic AI Misuse](https://aiforhumanity.eu/summaries/80k-catastrophic-ai-misuse): This summary covers the 80,000 Hours problem profile on how advanced AI systems could be deliberately misused to cause catastrophic harm, particularly through enabling the development of weapons... - [Summary: 80,000 Hours — Extreme Power Concentration](https://aiforhumanity.eu/summaries/80k-extreme-power-concentration): This summary synthesizes two overlapping source documents — a short overview and the full article — from the 80,000 Hours problem profile on how transformative AI could lead to extreme... - [Summary: 80,000 Hours — Problem Profiles Overview](https://aiforhumanity.eu/summaries/80k-problem-profiles): This summary synthesizes the problem profiles overview page and the career guide chapter on most pressing problems, both from [[80000-hours]]. - [Summary: 80,000 Hours — Risks from Power-Seeking AI](https://aiforhumanity.eu/summaries/80k-power-seeking-ai): This summary covers the 80,000 Hours problem profile on risks from AI systems that develop long-term goals and instrumental incentives to seek power, potentially deceiving or disempowering humanity. - [Summary: 80,000 Hours — Why AI Risks Are the World's Most Pressing Problems](https://aiforhumanity.eu/summaries/80k-ai-risk): This summary synthesizes two overlapping source documents from 80,000 Hours (the full article and the shorter risk profile) on why advanced AI poses the most pressing global challenges. - [Summary: 80,000 Hours — Why Problem Choice Is the Biggest Driver of Impact](https://aiforhumanity.eu/summaries/80k-int-framework): This summary synthesizes two overlapping source documents from 80,000 Hours — a short overview and the full article — on why the problem you choose to work on is the single most important factor... - [Summary: AI 2027 — A Scenario for Transformative AI](https://aiforhumanity.eu/summaries/ai-2027): AI 2027 is a 71-page research-based scenario from the AI Futures Project, authored by Daniel Kokotajlo, Scott Alexander, Thomas Larsen, Eli Lifland, and Romeo Dean. Published April 3, 2025, it... - [Summary: AI Safety (Wikipedia)](https://aiforhumanity.eu/summaries/ai-safety-wikipedia): Wikipedia's AI safety entry is a consolidated, citation-dense overview of the field as it is understood in mainstream technical and policy discourse — a useful counterweight to the wiki's existing... - [Summary: AI Safety Awareness Project](https://aiforhumanity.eu/summaries/ai-safety-awareness-project-overview): URL: aisafetyawarenessproject.org - Organization: AI Safety Awareness Project ([[ai-safety-awareness-project]]) — a US 501(c)(3) nonprofit (EIN 33-4395376) - Type: Public-education and... - [Summary: Existential Risk Prevention as Global Priority](https://aiforhumanity.eu/summaries/bostrom-existential-risk-priority): Author: Nick Bostrom Year: 2013 Source: bostrom-existential-risk-priority.pdf Published in: Global Policy, Volume 4, Issue 1, February 2013 - [Summary: Existential Risks — Analyzing Human Extinction Scenarios and Related Hazards](https://aiforhumanity.eu/summaries/bostrom-existential-risks): Author: Nick Bostrom Year: 2002 Source: bostrom-existential-risks.pdf Published in: Journal of Evolution and Technology, Vol. 9, March 2002 - [Summary: Future Progress in Artificial Intelligence — A Survey of Expert Opinion](https://aiforhumanity.eu/summaries/bostrom-ai-expert-survey): Authors: Vincent C. Muller & Nick Bostrom Year: 2014 (forthcoming at time of survey; published in *Fundamental Issues of Artificial Intelligence*, Synthese Library, Springer) Affiliation:... - [Summary: Optimal Timing for Superintelligence](https://aiforhumanity.eu/summaries/bostrom-optimal-timing): Author: [[nick-bostrom]] Year: 2026 (Working paper, version 1.0) Source: bostrom-optimal-timing.pdf - [Summary: PauseAI — Learn](https://aiforhumanity.eu/summaries/pauseai-learn): URL: pauseai.info/learn - Organization: PauseAI ([[pauseai]]) — formally Stichting PauseAI, a Dutch foundation (KvK 92951031) - Type: Curated learning hub / resource index maintained by an... - [Summary: Public Policy and Superintelligent AI -- A Vector Field Approach](https://aiforhumanity.eu/summaries/bostrom-ai-policy): Authors: Nick Bostrom, Allan Dafoe, Carrick Flynn Year: 2018 (version 4.3; first version 2016) Source: bostrom-ai-policy.pdf Published in: Liao, S. M. (ed.), *Ethics of Artificial Intelligence*... - [Summary: SaferAI — Frontier AI Risk Management Ratings](https://aiforhumanity.eu/summaries/safer-ai-risk-management-ratings): URL: ratings.safer-ai.org - Organization: SaferAI ([[safer-ai]]) — a France-based nonprofit focused on AI risk-management governance and research - Assessment team: Lily Stelling and Malcolm... - [Summary: Situational Awareness Ch. I — From GPT-4 to AGI: Counting the OOMs](https://aiforhumanity.eu/summaries/sa-ch1-from-gpt4-to-agi): This is Chapter I of [[situational-awareness|Situational Awareness]] by [[leopold-aschenbrenner]]. It makes the case that AGI by 2027 is "strikingly plausible" by decomposing AI progress into... - [Summary: Situational Awareness Ch. IIIa — Racing to the Trillion-Dollar Cluster](https://aiforhumanity.eu/summaries/sa-ch3a-trillion-dollar-cluster): This is Chapter IIIa of [[situational-awareness|Situational Awareness]] by [[leopold-aschenbrenner]]. It describes the extraordinary scale of industrial mobilization required to build the compute... - [Summary: Situational Awareness Ch. IIIb — Lock Down the Labs](https://aiforhumanity.eu/summaries/sa-ch3b-lock-down-labs): This is Chapter IIIb of [[situational-awareness|Situational Awareness]] by [[leopold-aschenbrenner]]. It argues that the security posture of leading AI labs is catastrophically inadequate given... - [Summary: Situational Awareness — The Decade Ahead](https://aiforhumanity.eu/summaries/situational-awareness-aschenbrenner): Situational Awareness is a 165-page essay series published in June 2024 by [[leopold-aschenbrenner]], a former OpenAI researcher. Dedicated to Ilya Sutskever, it argues that AGI is imminent... - [Summary: The Compendium](https://aiforhumanity.eu/summaries/the-compendium): URL: thecompendium.ai - Authors: Connor Leahy ([[connor-leahy]]), Gabriel Alfour, Chris Scammell, Andrea Miotti, Adam Shimi - Published: 2024 (last updated December 9, 2024) - Format: Web-native... - [Superalignment with Dynamic Human Values](https://aiforhumanity.eu/summaries/2503.13621): *Florian Mai, David Kaczér, Nicholas Kluge Corrêa, Lucie Flek* — 2025-03-17 — ICLR 2025 Workshop on Bidirectional Human-AI Alignment (BiAlign) - [Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?](https://aiforhumanity.eu/summaries/2502.15657): *Yoshua Bengio, Michael Cohen, Damiano Fornasiere, Joumana Ghosn, Pietro Greiner, Matt MacDermott, … (+7 more)* — 2025-02-21 — Mila, Independent — arXiv - [Surfacing Pathological Behaviors in Language Models](https://aiforhumanity.eu/summaries/surfacing-pathological-behaviors-in-language-models): - [Symmetries at the origin of hierarchical emergence](https://aiforhumanity.eu/summaries/2512.00984): *Fernando E. Rosas* — 2024-11-30 - [System 2 Alignment: Deliberation, Review, and Thought Management](https://aiforhumanity.eu/summaries/system-2-alignment-deliberation-review-and-thought-managemen): *Seth Herd* — 2025-02-13 - [Systematic runaway-optimiser-like LLM failure modes on Biologically and Economically aligned AI safety benchmarks for LLMs with simplified observation format (BioBlue)](https://aiforhumanity.eu/summaries/systematic-runaway-optimiser-like-llm-failure-modes-on-biolo): *Roland Pihlakas, Sruthi Susan Kuriakose, Shruti Datta Gupta* — 2025-03-16 — LessWrong - [Takeaways from sketching a control safety case](https://aiforhumanity.eu/summaries/takeaways-from-sketching-a-control-safety-case): *Josh Clymer, Buck Shlegeris* — 2025-01-31 — Redwood Research, UK AISI — LessWrong - [Taking a responsible path to AGI](https://aiforhumanity.eu/summaries/taking-a-responsible-path-to-agi): *Anca Dragan, Rohin Shah, Four Flynn, Shane Legg* — 2025-04-02 — Google DeepMind — Google DeepMind Blog - [Tamper-Resistant Safeguards for Open-Weight LLMs](https://aiforhumanity.eu/summaries/2408.00761): *Rishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, … (+9 more)* — 2024-08-01 — UC Berkeley, UIUC, Center for AI Safety — arXiv - [Teaching AI to Handle Exceptions: Supervised Fine-tuning with Human-aligned Judgment](https://aiforhumanity.eu/summaries/2503.02976v2): *Matthew DosSantos DiSorbo, Harang Ju, Sinan Aral* — 2025-03-XX — Harvard Business School, Johns Hopkins University, MIT Sloan School of Management - [Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning](https://aiforhumanity.eu/summaries/2506.22777): *Miles Turpin, Andy Arditi, Marvin Li, Joe Benton, Julian Michael* — 2025-06-28 — ICML 2025 Workshop on Reliable and Responsible Foundation Models - [Technical Report: Evaluating Goal Drift in Language Model Agents](https://aiforhumanity.eu/summaries/2505.02709): *Rauno Arike, Elizabeth Donoway, Henning Bartsch, Marius Hobbhahn* — 2025-05-05 — Apollo Research — arXiv - [Tell me about yourself: LLMs are aware of their learned behaviors](https://aiforhumanity.eu/summaries/2501.11120): *Jan Betley, Xuchan Bao, Martín Soto, Anna Sztyber-Betley, James Chua, Owain Evans* — 2025-01-19 — University of Oxford — arXiv - [Ten AI Safety Projects I'd Like People to Work On — Julian Hazell](https://aiforhumanity.eu/summaries/ten-ai-safety-projects-julian-hazell): Author: Julian Hazell ([[julian-hazell]]), grants officer at [[open-philanthropy]] Published: July 2025 Source: EA Forum / Secret Third Thing Substack Source - [Testing for Scheming with Model Deletion](https://aiforhumanity.eu/summaries/testing-for-scheming-with-model-deletion): - [The Alignment Project by UK AISI](https://aiforhumanity.eu/summaries/the-alignment-project-by-uk-aisi): *Mojmir, Benjamin Hilton, Jacob Pfau, Geoffrey Irving, Joseph Bloom, Tomek Korbak, … (+2 more)* — 2025-08-01 — UK AI Security Institute — LessWrong - [The Alignment Waltz: Jointly Training Agents to Collaborate for Safety](https://aiforhumanity.eu/summaries/2510.08240): *Jingyu Zhang, Haozhu Wang, Eric Michael Smith, Sid Wang, Amr Sharaf, Mahesh Pasupuleti, … (+4 more)* — 2025-10-09 — Meta, Johns Hopkins University - [The Circuits Research Landscape](https://aiforhumanity.eu/summaries/the-circuits-research-landscape): *Jack Lindsey, Emmanuel Ameisen, Neel Nanda, Stepan Shabalin, Mateusz Piotrowski, Tom McGrath, … (+12 more)* — 2025-08-01 — Anthropic, Google DeepMind, Goodfire AI, EleutherAI, Decode - [The Elicitation Game: Evaluating Capability Elicitation Techniques](https://aiforhumanity.eu/summaries/2502.02180): *Felix Hofstätter, Teun van der Weij, Jayden Teoh, Rada Djoneva, Henning Bartsch, Francis Rhys Ward* — 2025-02-04 — arXiv - [The Frame-Dependent Mind](https://aiforhumanity.eu/summaries/the-frame-dependent-mind): *Emmett Shear, Sonnet 3.7* — 2025-04-18 — Softmax - [The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence](https://aiforhumanity.eu/summaries/2502.17420): *Tom Wollschläger, Jannes Elstner, Simon Geisler, Vincent Cohen-Addad, Stephan Günnemann, Johannes Gasteiger* — 2025-02-24 — Technical University of Munich — arXiv - [The Geometry of Self-Verification in a Task-Specific Reasoning Model](https://aiforhumanity.eu/summaries/2504.14379-58bbde5d): *Andrew Lee, Lihao Sun, Chris Wendler, Fernanda Viégas, Martin Wattenberg* — 2025-04-19 — Google Research - [The MASK Evaluation](https://aiforhumanity.eu/summaries/the-mask-evaluation): Center for AI Safety, Scale AI — Hugging Face - [The PacifAIst Benchmark:Would an Artificial Intelligence Choose to Sacrifice Itself for Human Safety?](https://aiforhumanity.eu/summaries/2508.09762): *Manuel Herrador* — 2025-08-13 — arXiv - [The Pando Problem: Rethinking AI Individuality](https://aiforhumanity.eu/summaries/the-pando-problem-rethinking-ai-individuality): *Jan_Kulveit* — 2025-03-28 - [The Partially Observable Off-Switch Game](https://aiforhumanity.eu/summaries/2411.17749): *Andrew Garber, Rohan Subramani, Linus Luu, Mark Bedaywi, Stuart Russell, Scott Emmons* — 2024-11-25 — UC Berkeley — arXiv - [The Perils of Optimizing Learned Reward Functions: Low Training Error Does Not Guarantee Low Regret](https://aiforhumanity.eu/summaries/2406.15753): *Lukas Fluri, Leon Lang, Alessandro Abate, Patrick Forré, David Krueger, Joar Skalse* — 2024-06-22 — arXiv - [The Precipice Revisited — by Toby Ord](https://aiforhumanity.eu/summaries/precipice-revisited): This summary covers [[toby-ord]]'s EA Forum post reflecting on his book *The Precipice* and updating his assessment of the [[existential-risk]] landscape in light of developments since the book's... - [The Reality of AI and Biorisk](https://aiforhumanity.eu/summaries/2412.01946): *Aidan Peppin, Anka Reuel, Stephen Casper, Elliot Jones, Andrew Strait, Usman Anwar, … (+7 more)* — 2024-12-02 — arXiv - [The Rise of Parasitic AI](https://aiforhumanity.eu/summaries/the-rise-of-parasitic-ai-31ceeb15): *Adele Lopez* — 2025-09-11 — LessWrong - [The Rise of Parasitic AI](https://aiforhumanity.eu/summaries/the-rise-of-parasitic-ai): *Adele Lopez* — 2025-09-11 — LessWrong - [The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models](https://aiforhumanity.eu/summaries/2507.11544): *Ann-Kathrin Dombrowski, Dillon Bowen, Adam Gleave, Chris Cundy* — 2025-07-08 — Alignment Research — arXiv - [The Soul Document](https://aiforhumanity.eu/summaries/the-soul-document): *Richard-Weiss* — 2024-11-27 — GitHub Gist - [The Structural Safety Generalization Problem](https://aiforhumanity.eu/summaries/2504.09712): *Julius Broomfield, Tom Gibbs, Ethan Kosak-Hine, George Ingebretsen, Tia Nasir, Jason Zhang, … (+4 more)* — 2025-04-13 — arXiv - [The Unintended Trade-off of AI Alignment:Balancing Hallucination Mitigation and Safety in LLMs](https://aiforhumanity.eu/summaries/2510.07775): *Omar Mahmoud, Ali Khalil, Buddhika Laknath Semage, Thommen George Karimpanal, Santu Rana* — 2025-10-09 — arXiv - [The Urgency of Interpretability](https://aiforhumanity.eu/summaries/the-urgency-of-interpretability): *Dario Amodei* — 2025 — Anthropic — darioamodei.com - [Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models](https://aiforhumanity.eu/summaries/2506.13206): *James Chua, Jan Betley, Mia Taylor, Owain Evans* — 2025-06-16 — arXiv - [Three Sketches of ASL-4 Safety Case Components](https://aiforhumanity.eu/summaries/three-sketches-of-asl-4-safety-case-components): 2024-01-01 — Anthropic — Alignment Science Blog - [Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models](https://aiforhumanity.eu/summaries/2510.22014): *Sarah Ball, Niki Hasrati, Alexander Robey, Avi Schwarzschild, Frauke Kreuter, Zico Kolter, … (+1 more)* — 2025-10-24 — arXiv - [Toward understanding and preventing misalignment generalization](https://aiforhumanity.eu/summaries/toward-understanding-and-preventing-misalignment-generalizat): *Miles Wang, Tom Dupré la Tour, Olivia Watkins, Aleksandar Makelov, Ryan A. Chi, Samuel Miserendino, … (+2 more)* — 2025-06-18 — OpenAI — arXiv - [Toward universal steering and monitoring of AI models](https://aiforhumanity.eu/summaries/2502.03708): *Daniel Beaglehole, Adityanarayanan Radhakrishnan, Enric Boix-Adserà, Mikhail Belkin* — 2025-05-28 — arXiv - [Towards Alignment Auditing as a Numbers-Go-Up Science](https://aiforhumanity.eu/summaries/towards-alignment-auditing-as-a-numbers-go-up-science): *Sam Marks* — 2025-08-04 — Anthropic — LessWrong - [Towards a Scale-Free Theory of Intelligent Agency](https://aiforhumanity.eu/summaries/towards-a-scale-free-theory-of-intelligent-agency): *Richard Ngo* — 2025-03-21 — Independent - [Towards eliciting latent knowledge from LLMs with mechanistic interpretability](https://aiforhumanity.eu/summaries/2505.14352): *Bartosz Cywiński, Emil Ryd, Senthooran Rajamanoharan, Neel Nanda* — 2025-05-20 - [Towards evaluations-based safety cases for AI scheming](https://aiforhumanity.eu/summaries/2411.03336): *Mikita Balesni, Marius Hobbhahn, David Lindner, Alexander Meinke, Tomek Korbak, Joshua Clymer, … (+10 more)* — 2024-11-07 — Redwood Research, Apollo Research, Anthropic — arXiv - [Trading Inference-Time Compute for Adversarial Robustness](https://aiforhumanity.eu/summaries/trading-inference-time-compute-for-adversarial-robustness): *OpenAI* — 2025-01-22 — OpenAI — arXiv - [Training AI to do alignment research we don't already know how to do](https://aiforhumanity.eu/summaries/training-ai-to-do-alignment-research-we-don-t-already-know-h): *joshc* — 2025-02-24 — Redwood Research — LessWrong - [Training LLM Agents to Empower Humans](https://aiforhumanity.eu/summaries/2510.13709): *Evan Ellis, Vivek Myers, Jens Tuyls, Sergey Levine, Anca Dragan, Benjamin Eysenbach* — 2025-10-16 — UC Berkeley, Princeton University - [Training LLMs for Honesty via Confessions](https://aiforhumanity.eu/summaries/2512.08093): *Manas Joglekar, Jeremy Chen, Gabriel Wu, Jason Yosinski, Jasmine Wang, Boaz Barak, … (+1 more)* — 2025-12-08 — OpenAI, Google DeepMind - [Training Language Models to Explain Their Own Computations](https://aiforhumanity.eu/summaries/2511.08579): *Belinda Z. Li, Zifan Carl Guo, Vincent Huang, Jacob Steinhardt, Jacob Andreas* — 2025-11-11 — UC Berkeley, MIT — arXiv - [Training fails to elicit subtle reasoning in current language models](https://aiforhumanity.eu/summaries/training-fails-to-elicit-subtle-reasoning-in-current-languag): 2025-01-01 — Anthropic — Anthropic Alignment Science Blog - [Training on Documents About Reward Hacking Induces Reward Hacking](https://aiforhumanity.eu/summaries/training-on-documents-about-reward-hacking-induces-reward-ha): *Evan Hubinger, Nathan Hu* — 2025-01-21 — Anthropic - [Transcoders Beat Sparse Autoencoders for Interpretability](https://aiforhumanity.eu/summaries/2501.18823): *Gonçalo Paulo, Stepan Shabalin, Nora Belrose* — 2025-01-31 — arXiv - [UAIASI](https://aiforhumanity.eu/summaries/uaiasi): *Cole Wyeth* — 2025 — Universal Algorithmic Intelligence - [UK AISI Alignment Team: Debate Sequence](https://aiforhumanity.eu/summaries/uk-aisi-alignment-team-debate-sequence): *Benjamin Hilton* — 2025-05-07 — UK AI Safety Institute - [UK AISI's Alignment Team: Research Agenda](https://aiforhumanity.eu/summaries/uk-aisi-s-alignment-team-research-agenda): *Benjamin Hilton, Jacob Pfau, Marie_DB, Geoffrey Irving* — 2025-05-07 — UK AI Safety Institute — LessWrong - [Understanding and Controlling LLM Generalization](https://aiforhumanity.eu/summaries/understanding-and-controlling-llm-generalization): *Daniel Tan* — 2025-11-14 - [Understanding the Capabilities and Limitations of Weak-to-Strong Generalization](https://aiforhumanity.eu/summaries/understanding-the-capabilities-and-limitations-of-weak-to-st): *Wei Yao, Wenkai Yang, Ziqiao Wang, Yankai Lin, Yong Liu* — 2025-03-08 — ICLR 2025 Workshop SSI-FM - [Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs](https://aiforhumanity.eu/summaries/2505.16831): *Xiaoyu Xu, Xiang Yue, Yang Liu, Qingqing Ye, Huadi Zheng, Peizhao Hu, … (+2 more)* — 2025-05-22 — arXiv - [Unlearning Needs to be More Selective [Progress Report]](https://aiforhumanity.eu/summaries/unlearning-needs-to-be-more-selective-progress-report): *Filip Sondej, Yushi Yang, Marcel Windys* — 2025-06-27 — LessWrong - [Unsupervised Elicitation](https://aiforhumanity.eu/summaries/unsupervised-elicitation): *Jiaxin Wen, Zachary Ankner, Arushi Somani, Peter Hase, Samuel Marks, Jacob Goldman-Wetzler, … (+7 more)* — 2025 — Anthropic, Schmidt Sciences, Independent, Constellation, New York University,... - [Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs](https://aiforhumanity.eu/summaries/2502.08640): *Mantas Mazeika, Xuwang Yin, Rishub Tamirisa, Jaehyuk Lim, Bruce W. Lee, Richard Ren, … (+5 more)* — 2025-02-12 — Unknown — arXiv - [Validating against a misalignment detector is very different to training against one](https://aiforhumanity.eu/summaries/validating-against-a-misalignment-detector-is-very-different): *mattmacdermott* — 2025-03-04 — LessWrong - [Video and transcript of talk on automating alignment research](https://aiforhumanity.eu/summaries/video-and-transcript-of-talk-on-automating-alignment-researc): *Joe Carlsmith* — 2025-04-30 — Anthropic — LessWrong - [Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark](https://aiforhumanity.eu/summaries/2504.16137): *Jasper Götting, Pedro Medeiros, Jon G Sanders, Nathaniel Li, Long Phan, Karam Elabd, … (+3 more)* — 2025-04-29 — Multiple Research Institutions — arXiv - [Virtual Agent Economies](https://aiforhumanity.eu/summaries/2509.10147): *Nenad Tomasev, Matija Franklin, Joel Z. Leibo, Julian Jacobs, William A. Cunningham, Iason Gabriel, … (+1 more)* — 2025-09-12 — DeepMind — arXiv - [We need a field of Reward Function Design](https://aiforhumanity.eu/summaries/we-need-a-field-of-reward-function-design): *Steven Byrnes* — 2025-12-08 - [We should try to automate AI safety work asap](https://aiforhumanity.eu/summaries/we-should-try-to-automate-ai-safety-work-asap): *Marius Hobbhahn* — 2025-04-26 — Apollo Research — LessWrong - [Weak to Strong Generalization for Large Language Models with Multi-capabilities](https://aiforhumanity.eu/summaries/weak-to-strong-generalization-for-large-language-models-with): *Yucheng Zhou, Jianbing Shen, Yu Cheng* — 2025-01-22 — ICLR 2025 - [What Is The Alignment Problem?](https://aiforhumanity.eu/summaries/what-is-the-alignment-problem): *johnswentworth* — 2025-01-16 — LessWrong - [What Kind of User Are You? Uncovering User Models in LLM Chatbots](https://aiforhumanity.eu/summaries/2406.07882v1): *Yida Chen, Aoyu Wu, Trevor DePodesta, Catherine Yeh, Kenneth Li, Nicholas Castillo Marin, … (+6 more)* — 2024-06-12 — Google - [What We Learned Trying to Diff Base and Chat Models (And Why It Matters)](https://aiforhumanity.eu/summaries/what-we-learned-trying-to-diff-base-and-chat-models-and-why): *Clément Dumas, Julian Minder, Neel Nanda* — 2025-06-30 — MATS Program - [What is Inadequate about Bayesianism for AI Alignment: Motivating Infra-Bayesianism](https://aiforhumanity.eu/summaries/what-is-inadequate-about-bayesianism-for-ai-alignment-motiva): *Brittany Gelb* — 2025-08-30 - [What's In My Human Feedback? Learning Interpretable Descriptions of Preference Data](https://aiforhumanity.eu/summaries/2510.26202): *Rajiv Movva, Smitha Milli, Sewon Min, Emma Pierson* — 2025-10-30 — arXiv - [What, if not agency?](https://aiforhumanity.eu/summaries/what-if-not-agency): - [When Autonomy Goes Rogue: Preparing for Risks of Multi-Agent Collusion in Social Systems](https://aiforhumanity.eu/summaries/2507.14660): *Qibing Ren, Sitao Xie, Longxuan Wei, Zhenfei Yin, Junchi Yan, Lizhuang Ma, … (+1 more)* — 2025-07-19 — arXiv - [When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors](https://aiforhumanity.eu/summaries/2507.05246): *Scott Emmons, Erik Jenner, David K. Elson, Rif A. Saurous, Senthooran Rajamanoharan, Heng Chen, … (+2 more)* — 2025-07-07 — Google DeepMind — arXiv - [When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas](https://aiforhumanity.eu/summaries/2505.19212): *Steffen Backmann, David Guzman Piedrahita, Emanuel Tewolde, Rada Mihalcea, Bernhard Schölkopf, Zhijing Jin* — 2025-05-25 — Max Planck Institute for Intelligent Systems — arXiv - [When Thinking LLMs Lie: Unveiling the Strategic Deception in Representations of Reasoning Models](https://aiforhumanity.eu/summaries/2506.04909): *Kai Wang, Yihao Zhang, Meng Sun* — 2025-06-05 — arXiv - [When Truthful Representations Flip Under Deceptive Instructions?](https://aiforhumanity.eu/summaries/2507.22149): *Xianxuan Long, Yao Fu, Runchao Li, Mu Sheng, Haotian Yu, Xiaotian Han, … (+1 more)* — 2025-07-29 — arXiv - [When does Claude sabotage code? An Agentic Misalignment follow-up](https://aiforhumanity.eu/summaries/when-does-claude-sabotage-code-an-agentic-misalignment-follo): *Nathan Delisle* — 2024-11-09 — MATS — LessWrong - [Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing](https://aiforhumanity.eu/summaries/2510.02334): *Zhe Li, Wei Zhao, Yige Li, Jun Sun* — 2025-09-26 — arXiv - [White Box Control at UK AISI - Update on Sandbagging Investigations](https://aiforhumanity.eu/summaries/white-box-control-at-uk-aisi-update-on-sandbagging-investiga): *Joseph Bloom, Jordan Taylor, Connor Kissane, Sid Black, Jacob Merizian, Alex Zelenka-Martin, … (+3 more)* — 2025-07-10 — UK AISI — AI Alignment Forum - [Whitebox detection of sandbagging model organisms](https://aiforhumanity.eu/summaries/whitebox-detection-of-sandbagging-model-organisms): *Joseph Bloom, Jordan Taylor, Connor Kissane, Sid Black, Jacob Merizian, Alex Zelenka-Martin, … (+3 more)* — 2025-07-10 — UK AISI - [Why Corrigibility is Hard and Important (i.e. "Whence the high MIRI confidence in alignment difficulty?")](https://aiforhumanity.eu/summaries/why-corrigibility-is-hard-and-important-i-e-whence-the-high): - [Why Do Some Language Models Fake Alignment While Others Don't?](https://aiforhumanity.eu/summaries/2506.18032-b50867d0): *Abhay Sheshadri, John Hughes, Julian Michael, Alex Mallen, Arun Jose, Janus, … (+1 more)* — 2025-06-22 — arXiv - [Why Do Some Language Models Fake Alignment While Others Don't?](https://aiforhumanity.eu/summaries/2506.18032): *Abhay Sheshadri, John Hughes, Julian Michael, Alex Mallen, Arun Jose, Janus, … (+1 more)* — 2025-06-22 - [Why Do Some Language Models Fake Alignment While Others Don't?](https://aiforhumanity.eu/summaries/why-do-some-language-models-fake-alignment-while-others-don): *abhayesian, John Hughes, Alex Mallen, Jozdien, janus, Fabien Roger* — 2025-07-08 — Anthropic, Redwood Research — arXiv - [Why Future AIs will Require New Alignment Methods](https://aiforhumanity.eu/summaries/why-future-ais-will-require-new-alignment-methods): *Alvin Ånestrand* — 2025-10-10 — LessWrong - [Why do misalignment risks increase as AIs get more capable?](https://aiforhumanity.eu/summaries/why-do-misalignment-risks-increase-as-ais-get-more-capable): *Ryan Greenblatt* — 2025-04-11 — Anthropic — LessWrong - [Why it's good for AI reasoning to be legible and faithful](https://aiforhumanity.eu/summaries/why-it-s-good-for-ai-reasoning-to-be-legible-and-faithful): 2025-03-11 — METR — METR Blog - [Why modelling multi-objective homeostasis is essential for AI alignment (and how it helps with AI safety as well). Subtleties and Open Challenges](https://aiforhumanity.eu/summaries/why-modelling-multi-objective-homeostasis-is-essential-for-a): *Roland Pihlakas* — 2025-01-12 — LessWrong - [WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models](https://aiforhumanity.eu/summaries/2406.18510): *Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, … (+5 more)* — 2024-06-26 — University of Washington, Allen Institute for AI - [Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas](https://aiforhumanity.eu/summaries/2505.14633): *Yu Ying Chiu, Zhilin Wang, Sharan Maiya, Yejin Choi, Kyle Fish, Sydney Levine, … (+1 more)* — 2025-05-20 — Anthropic — arXiv - [Will alignment-faking Claude accept a deal to reveal its misalignment?](https://aiforhumanity.eu/summaries/will-alignment-faking-claude-accept-a-deal-to-reveal-its-mis): *Ryan Greenblatt, Kyle Fish* — 2025-01-31 — Redwood Research, Anthropic — LessWrong / AI Alignment Forum - [Winning at All Cost: A Small Environment for Eliciting Specification Gaming Behaviors in Large Language Models](https://aiforhumanity.eu/summaries/2505.07846): *Lars Malmqvist* — 2025-05-07 — arXiv (to be presented at SIMLA@ACNS 2025) - [Won't vs. Can't: Sandbagging-like Behavior from Claude Models](https://aiforhumanity.eu/summaries/won-t-vs-can-t-sandbagging-like-behavior-from-claude-models): 2025-01-15 — Anthropic — Anthropic Alignment Science Blog - [Working through a small tiling result](https://aiforhumanity.eu/summaries/working-through-a-small-tiling-result): *James Payor* — 2024-05-13 - [X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents](https://aiforhumanity.eu/summaries/2504.13203): *Salman Rahman, Liwei Jiang, James Shiffer, Genglin Liu, Sheriff Issaka, Md Rizwan Parvez, … (+4 more)* — 2025-04-15 — University of Washington, UCLA, Microsoft — arXiv - [You Are What You Eat -- AI Alignment Requires Understanding How Data Shapes Structure and Generalisation](https://aiforhumanity.eu/summaries/2502.05475): *Simon Pepin Lehalleur, Jesse Hoogland, Matthew Farrugia-Roberts, Susan Wei, Alexander Gietelink Oldenziel, George Wang, … (+2 more)* — 2025-02-08 — arXiv - [the void](https://aiforhumanity.eu/summaries/the-void): *nostalgebraist* — 2025-06-07 — Tumblr ## Entities - [80,000 Hours](https://aiforhumanity.eu/entities/80000-hours): 80,000 Hours is an Oxford-affiliated career advising organization, co-founded by [[benjamin-todd|Benjamin Todd]], that applies effective-altruism principles to help people find careers with high... - [ACRAI — Antwerp Center for Responsible AI](https://aiforhumanity.eu/entities/acrai): The Antwerp Center for Responsible AI is an interdisciplinary research center at the University of Antwerp focused on responsible AI development. Its remit covers bias, explainability,... - [AI Safety Atlas Textbook](https://aiforhumanity.eu/entities/ai-safety-atlas-textbook): The AI Safety Atlas (ai-safety-atlas.com) is an 8-chapter self-paced textbook (~15 hours) covering the foundations of AI safety from core concepts through cutting-edge research. Authored by... - [AI Safety Awareness Project](https://aiforhumanity.eu/entities/ai-safety-awareness-project): The AI Safety Awareness Project is a US 501(c)(3) nonprofit (EIN 33-4395376) dedicated to democratizing AI governance through public education. It runs workshops, seminars, and... - [AI Safety Institute](https://aiforhumanity.eu/entities/ai-safety-institute): AI Safety Institutes (AISIs) are government-funded bodies that evaluate frontier AI models for dangerous capabilities and develop the technical foundations for AI regulation. The first two were... - [AI Safety Summit 2023](https://aiforhumanity.eu/entities/ai-safety-summit-2023): The AI Safety Summit was the first major intergovernmental summit on the safety of advanced AI, hosted by the United Kingdom at Bletchley Park on 1–2 November 2023. Pitched by then-Prime Minister... - [Ajeya Cotra](https://aiforhumanity.eu/entities/ajeya-cotra): Ajeya Cotra is an AI safety researcher who has served as senior advisor at Coefficient Giving and researcher at [[metr]] (formerly Model Evaluations and Threat Research). She previously led... - [Alignment Forum](https://aiforhumanity.eu/entities/alignment-forum): The Alignment Forum is a curated online platform dedicated to technical [[ai-alignment]] research. It operates as a focused subset of [[lesswrong]], filtering for posts that engage directly with... - [Andrej Karpathy](https://aiforhumanity.eu/entities/andrej-karpathy): Andrej Karpathy is a prominent AI researcher, educator, and entrepreneur. He is the author of the llm-knowledge-base pattern described in the "LLM Wiki" gist, which is the founding source of this... - [Anthropic](https://aiforhumanity.eu/entities/anthropic): Anthropic is an AI safety company founded by former OpenAI researchers, built on the premise that it is possible to develop competitive frontier AI models while maintaining strong safety... - [Asilomar AI Principles](https://aiforhumanity.eu/entities/asilomar-ai-principles): The Asilomar AI Principles are a set of 23 principles for beneficial AI development drafted at the 2017 Asilomar Conference on Beneficial AI, organized by the [[future-of-life-institute|Future of... - [Ben Garfinkel](https://aiforhumanity.eu/entities/ben-garfinkel): Ben Garfinkel is a research fellow at the [[future-of-humanity-institute]] at Oxford University and one of the most prominent constructive critics of classic AI risk arguments within the... - [Benjamin Todd](https://aiforhumanity.eu/entities/benjamin-todd): Benjamin Todd is the co-founder and former CEO of [[80000-hours]], the Oxford-affiliated career advising organisation that applies effective-altruism principles to help people find high-impact... - [Buck Shlegeris](https://aiforhumanity.eu/entities/buck-shlegeris): Buck Shlegeris is the CEO of [[redwood-research]] and one of the leading architects of the [[ai-control]] paradigm — a distinct approach to managing catastrophic risk from AI systems that has... - [Carina Prunkl](https://aiforhumanity.eu/entities/carina-prunkl): Carina Prunkl is a researcher in AI ethics and governance at the intersection of philosophy and computer science. She is a co-author of the Belgian-cluster paper on AI existential risk arguments. - [Carl Shulman](https://aiforhumanity.eu/entities/carl-shulman): Carl Shulman is an independent AI researcher and one of the most influential thinkers in the effective-altruism and [[ai-safety]] communities. He is widely regarded — by [[benjamin-todd]],... - [Catherine Olsson](https://aiforhumanity.eu/entities/catherine-olsson): ML safety researcher. At the time of the 80,000 Hours podcast interview, Olsson was working on the Google Brain safety team, having previously worked at OpenAI. She focuses on the practical,... - [CeSIA — French Center for AI Safety](https://aiforhumanity.eu/entities/cesia): CeSIA (*Centre pour la Sécurité de l'IA*) is France's leading AI safety think tank and expertise center. Established in May 2024 in Paris by [[effisciences|EffiSciences]], it describes itself as... - [Centre for Long-Term Resilience (CLTR)](https://aiforhumanity.eu/entities/cltr): The Centre for Long-Term Resilience is an independent UK think tank with a mission to transform global resilience to extreme risks, with primary focus on AI risks, biosecurity, and government risk... - [Charbel-Raphaël Ségerie](https://aiforhumanity.eu/entities/charbel-raphael-segerie): Charbel-Raphaël Ségerie is the Executive Director of [[cesia|CeSIA]] (Centre pour la Sécurité de l'IA, the French Center for AI Safety) — France's leading AI safety organization. He is one of the... - [Concrete Problems in AI Safety](https://aiforhumanity.eu/entities/concrete-problems-in-ai-safety): *Concrete Problems in AI Safety* is a 2016 paper by Dario Amodei, Chris Olah, Jacob Steinhardt, [[paul-christiano]], John Schulman, and Dan Mané (arXiv:1606.06565). It is one of the earliest and... - [Connor Leahy](https://aiforhumanity.eu/entities/connor-leahy): Connor Leahy is an AI researcher and entrepreneur, co-founder of the open-research collective EleutherAI and CEO of the AI-safety company Conjecture. He is a prominent voice in the "stop the race... - [Daniel Kokotajlo](https://aiforhumanity.eu/entities/daniel-kokotajlo): Daniel Kokotajlo is the lead creator of AI 2027, a research-based scenario from the AI Futures Project that combines forecasting and storytelling to explore a plausible future in which AI... - [Daniel Ziegler](https://aiforhumanity.eu/entities/daniel-ziegler): ML safety researcher at OpenAI. Ziegler focuses on empirical, engineering-driven approaches to AI alignment, particularly [[reward-learning]] and [[rlhf]]. - [DeepMind](https://aiforhumanity.eu/entities/deepmind): DeepMind is google's AI research laboratory, widely recognized as one of the world's leading AI research organizations. Originally founded in 2010 in London and acquired by Google in 2014, it has... - [EA Forum](https://aiforhumanity.eu/entities/ea-forum): The EA Forum (forum.effectivealtruism.org) is the primary online platform for discussion and debate within the effective-altruism community. It hosts posts on cause prioritization, strategic... - [ELSA — European Lighthouse on Secure and Safe AI](https://aiforhumanity.eu/entities/elsa): ELSA is an EU-funded network of excellence extending from ELLIS (European Laboratory for Learning and Intelligent Systems) — the major European academic research network for ML. ELSA coordinates... - [ENAIS — European Network for AI Safety](https://aiforhumanity.eu/entities/enais): ENAIS (pronounced "e-nice") is the primary pan-European network explicitly focused on existential risk from AI. Founded in March 2023, it differentiates itself from other European AI safety... - [EU AI Act](https://aiforhumanity.eu/entities/eu-ai-act): The EU AI Act is the European Union's flagship horizontal regulation of artificial intelligence — adopted in 2024 and entering force in stages through 2026. It is the most comprehensive piece of... - [EU AI Office](https://aiforhumanity.eu/entities/eu-ai-office): The European AI Office is the European Commission body responsible for enforcement and implementation of the [[eu-ai-act|EU AI Act]], particularly its general-purpose AI (GPAI) provisions, and for... - [EffiSciences](https://aiforhumanity.eu/entities/effisciences): EffiSciences is a French effective-altruism-adjacent organization that incubates impactful research projects in academia, with a particular focus on AI safety. It is best known as the parent... - [Eliezer Yudkowsky](https://aiforhumanity.eu/entities/eliezer-yudkowsky): Eliezer Yudkowsky is an American AI researcher, writer, and public intellectual who founded both [[miri]] (the Machine Intelligence Research Institute) and [[lesswrong]]. He is one of the earliest... - [Future of Humanity Institute](https://aiforhumanity.eu/entities/future-of-humanity-institute): The Future of Humanity Institute (FHI) was a multidisciplinary research center at the University of Oxford focused on [[existential-risk]], the long-term future of humanity, and the governance of... - [Future of Life Institute](https://aiforhumanity.eu/entities/future-of-life-institute): The Future of Life Institute (FLI) is a nonprofit organization focused on reducing large-scale catastrophic risks, particularly from advanced AI, nuclear weapons, and biotechnology. Founded in... - [Geoffrey Hinton](https://aiforhumanity.eu/entities/geoffrey-hinton): Geoffrey Hinton is a British-Canadian computer scientist and cognitive psychologist, widely called one of the "godfathers of deep learning." He shared the 2018 Turing Award with [[yoshua-bengio]]... - [GiveWell](https://aiforhumanity.eu/entities/givewell): GiveWell is an evidence-based charity evaluator and a foundational organization in the effective-altruism movement. Co-founded by [[holden-karnofsky]] and Elie Hassenfeld, GiveWell conducts... - [Global Priorities Institute](https://aiforhumanity.eu/entities/global-priorities-institute): The Global Priorities Institute (GPI) is an academic research center based at the University of Oxford, dedicated to using the tools of rigorous academic philosophy and economics to determine how... - [Google Deepmind (redirect)](https://aiforhumanity.eu/entities/google-deepmind): This page exists to keep the SR2025 import's wikilinks resolving. The canonical entity page is [[deepmind]] — the SR2025 lab snapshot (key people, funding, safety teams, public alignment agenda,... - [Holden Karnofsky](https://aiforhumanity.eu/entities/holden-karnofsky): Holden Karnofsky is the co-founder of GiveWell and Open Philanthropy, and now works at [[anthropic]] on AI safety. His 14-year intellectual journey from skepticism about AI risk to dedicating his... - [International AI Safety Report](https://aiforhumanity.eu/entities/international-ai-safety-report): The International AI Safety Report is the first global, government-commissioned scientific review of risks from advanced AI. Its first full edition was published in January 2025, chaired by... - [Jan Leike](https://aiforhumanity.eu/entities/jan-leike): Jan Leike is an [[ai-alignment]] researcher who served as head of alignment at [[openai]] and co-leader of the [[superalignment]] project — one of the largest institutional commitments to... - [Julian Hazell](https://aiforhumanity.eu/entities/julian-hazell): Julian Hazell is a grants officer at [[open-philanthropy]], focused on reducing catastrophic risks from transformative AI. He publicly shares his views on promising AI safety projects via his... - [Laura Weidinger](https://aiforhumanity.eu/entities/laura-weidinger): Laura Weidinger is a research scientist at [[deepmind|Google DeepMind]] working on the ethical and societal risks of large language models. Her most influential contribution is the taxonomy of... - [LawZero](https://aiforhumanity.eu/entities/lawzero): LawZero is a Montréal-based nonprofit AI research organization founded by [[yoshua-bengio|Yoshua Bengio]], launched on 3 June 2025 with $30 million in philanthropic funding. Its mission: develop... - [Leopold Aschenbrenner](https://aiforhumanity.eu/entities/leopold-aschenbrenner): Leopold Aschenbrenner is a former [[openai]] researcher and the author of Situational Awareness: The Decade Ahead, a 165-page essay series published in June 2024 that has become one of the most... - [LessWrong](https://aiforhumanity.eu/entities/lesswrong): LessWrong is an online community and forum dedicated to rationality, clear thinking, and increasingly, [[ai-safety]] discourse. Founded by [[eliezer-yudkowsky]], it grew out of his blog posts on... - [METR](https://aiforhumanity.eu/entities/metr): METR (Model Evaluation and Threat Research) is an AI evaluation organization formerly known as ARC Evals — the evaluations arm that spun out of the Alignment Research Center (ARC) founded by... - [Machine Intelligence Research Institute (MIRI)](https://aiforhumanity.eu/entities/miri): The Machine Intelligence Research Institute (MIRI) is a research organization focused on ensuring that artificial intelligence systems are safe and beneficial. Founded by [[eliezer-yudkowsky]] and... - [Markov Grey](https://aiforhumanity.eu/entities/markov-grey): Markov Grey is the co-author of the [[ai-safety-atlas-textbook|AI Safety Atlas]] textbook (ai-safety-atlas.com), with [[charbel-raphael-segerie|Charbel-Raphaël Ségerie]]. The Atlas is the... - [Meta](https://aiforhumanity.eu/entities/meta): Shuchao Bi, Hongyuan Zhan, Jingyu Zhang, Haozhu Wang, Eric Michael Smith, Sid Wang, Amr Sharaf, Mahesh Pasupuleti, Jason Weston, ShengYun Peng, Ivan Evtimov, Song Jiang, Pin-Yu Chen, Evangelia... - [Nate Soares](https://aiforhumanity.eu/entities/nate-soares): Nate Soares is the executive director of [[miri]] (the Machine Intelligence Research Institute), where he works alongside founder [[eliezer-yudkowsky]] on technical [[ai-alignment]] research. He... - [Nick Bostrom](https://aiforhumanity.eu/entities/nick-bostrom): Nick Bostrom is a Swedish-born philosopher at the University of Oxford whose work has been foundational to [[existential-risk]] research, [[ai-safety]], and the intellectual infrastructure of... - [Nick Joseph](https://aiforhumanity.eu/entities/nick-joseph): Nick Joseph is the head of training at [[anthropic]] and a co-founder of the company. He previously worked at [[openai]] before leaving to help build a more safety-focused AI lab. Joseph is one of... - [Nova DasSarma](https://aiforhumanity.eu/entities/nova-dassarma): AI information security researcher. DasSarma appeared on the 80,000 Hours Podcast to make the case that [[information-security]] is a foundational, load-bearing component of AI safety — not a... - [Open Philanthropy](https://aiforhumanity.eu/entities/open-philanthropy): Open Philanthropy is a major grantmaking organization aligned with the effective-altruism movement and one of the largest funders of [[ai-safety]] research in the world. Co-founded by... - [OpenAI](https://aiforhumanity.eu/entities/openai): OpenAI is one of the world's leading AI research laboratories and the developer of the GPT series of large language models. Founded in 2015 as a non-profit with the mission of ensuring artificial... - [Paul Christiano](https://aiforhumanity.eu/entities/paul-christiano): Paul Christiano is an [[ai-alignment]] researcher widely regarded as one of the most influential technical thinkers in the field. He worked at [[openai]]'s machine learning lab before founding the... - [PauseAI](https://aiforhumanity.eu/entities/pauseai): PauseAI is a grassroots advocacy organization — formally Stichting PauseAI, a Dutch foundation (KvK 92951031) — campaigning for a pause on frontier AI development, ideally through a binding global... - [Redwood Research](https://aiforhumanity.eu/entities/redwood-research): Redwood Research is an AI safety organization focused on the [[ai-control]] paradigm — developing techniques to safely deploy and use AI systems even if they are misaligned. Led by CEO... - [Risto Uuk](https://aiforhumanity.eu/entities/risto-uuk): Risto Uuk is Head of European Policy and Research at the [[future-of-life-institute|Future of Life Institute]] (Brussels), a simultaneous PhD researcher at KU Leuven, a member of Estonia's... - [Rob Wiblin](https://aiforhumanity.eu/entities/rob-wiblin): Rob Wiblin is the host of the [[80000-hours]] Podcast, one of the most substantive long-form interview series covering [[ai-safety]], [[ai-alignment]], [[existential-risk]], and... - [Roman Yampolskiy](https://aiforhumanity.eu/entities/roman-yampolskiy): Roman V. Yampolskiy is a computer scientist at the University of Louisville and one of the founding voices of "AI safety engineering" as a discipline. He coined the term at the 2011 PT-AI... - [SaferAI](https://aiforhumanity.eu/entities/safer-ai): SaferAI is a France-based nonprofit focused on AI risk-management governance and research, with the stated aim of making frontier AI safer through improved risk-management practices and greater... - [Scientist AI](https://aiforhumanity.eu/entities/scientist-ai): One-sentence summary: Develop powerful, nonagentic, uncertain world models that accelerate scientific progress while avoiding the risks of agent AIs - [Stuart Russell](https://aiforhumanity.eu/entities/stuart-russell): Stuart Russell is a British-American computer scientist and AI researcher, best known in the AI safety context for his book *Human Compatible: Artificial Intelligence and the Problem of Control*... - [The AI Endgame (book)](https://aiforhumanity.eu/entities/the-ai-endgame): *The AI Endgame* is a forthcoming book by Lode Lauwaert and [[risto-uuk|Risto Uuk]], forthcoming with Wiley. The book targets a serious general audience and addresses emerging risks from AI and... - [Toby Ord](https://aiforhumanity.eu/entities/toby-ord): Toby Ord is an Australian-born moral philosopher at the University of Oxford, best known as the author of *The Precipice: Existential Risk and the Future of Humanity* (2020). His work sits at the... - [Will MacAskill](https://aiforhumanity.eu/entities/will-macaskill): William MacAskill is a Scottish philosopher at the University of Oxford and one of the co-founders of the effective-altruism movement. His academic work and public advocacy have made him one of... - [Yoshua Bengio](https://aiforhumanity.eu/entities/yoshua-bengio): Yoshua Bengio is a Canadian computer scientist, professor at the Université de Montréal, and founder of Mila — Quebec's AI institute. He shared the 2018 Turing Award with Geoffrey Hinton and Yann... - [xAI](https://aiforhumanity.eu/entities/xai): Dan Hendrycks (advisor), Juntang Zhuang, Toby Pohlen, Lianmin Zheng, Piaoyang Cui, Nikita Popov, Ying Sheng, Sehoon Kim, Alexander Pan