Anthropic Fellows Program — AI Safety
About Anthropic
.
Anthropic Fellows posting
Anthropic Fellows Program overview
public output
What to expect
- 4 months of full-time research
- Direct mentorship from Anthropic researchers
- Access to a shared workspace (in either Berkeley, California or London, UK)
- Connection to the broader AI safety and security research community
- Weekly stipend of 3,850 USD / 2,310 GBP / 4,300 CAD + benefits (these vary by country)
- Funding for compute (~$15k/month) and other research expenses
Interview process
We encourage you to apply even if you do not believe you meet every single qualification.
Compensation
Fellows workstreams
Due to the success of the Anthropic Fellows for AI Safety Research program, we are now expanding it across teams at Anthropic . We expect there to be significant overlap in the types of skills and responsibilities across the roles and will by default consider candidates for all the workstreams. Some of the workstreams may include unique assessment steps; we therefore ask you for workstream preferences in the application. You can see an overview of the current workstreams below:
- AI Safety Fellows
- AI Security Fellows
- ML Systems & Performance Fellows
- Reinforcement Learning Fellows
- Economics & Societal Impacts Fellows
This page is specific to one of the Anthropic Fellows Workstreams, see also the main
Anthropic Fellows posting
Across the workstreams, you may be a good fit if you
- Are motivated by making sure AI is safe and beneficial for society as a whole
- Are excited to transition into empirical AI research and would be interested in a full-time role at Anthropic
- Have a strong technical background in computer science, mathematics, or physics
- Thrive in fast-paced, collaborative environments
- Can implement ideas quickly and communicate clearly
Strong candidates may also have
- Strong background in a discipline relevant to a specific Fellows workstream (e.g. economics, social sciences, or cybersecurity)
- Experience in areas of research or engineering related to their workstream
Candidates must be
- Fluent in Python programming
- Available to work full-time on the Fellows program
Mentors, research areas, & past projects
Fellows will undergo a project selection & mentor matching process. Potential mentors include:
- Sam Bowman
- Sara Price
- Alex Tamkin
- Nina Panickssery
- Trenton Bricken
- Logan Graham
- Jascha Sohl-Dickstein
- Joe Benton
- Collin Burns
- Fabien Roger
- Samuel Marks
- Kyle Fish
- Ethan Perez
Our mentors will lead projects in select AI safety research areas, such as:
•
Scalable Oversight
Adversarial Robustness and AI Control
Model Organisms
Model Internals / Mechanistic Interpretability
AI Welfare
Improving our understanding of potential AI welfare and developing related evaluations and mitigations.
On our Alignment Science and Frontier Red Team blogs, you can read about past projects, including:
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Data: Alex Cloud and Minh Le, et al., mentors including Samuel Marks and Owain Evans
- Open-source circuits: Michael Hanna and Mateusz Piotrowski with mentorship from Emmanuel Ameisen and Jack Lindsey
For a full list of representative projects for each area, please see these blog posts: Introducing the Anthropic Fellows Program for AI Safety Research, Recommendations for Technical AI Safety Research Directions.
Unique candidate criteria
You might be a particularly great fit for this workstream if you:
- Are motivated by reducing catastrophic risks from advanced AI systems
- Have experience with empirical ML research projects
- Have experience working with large language models
- Have experience in one of the research areas mentioned above
- Have a track record of open-source contributions
Logistics Requirements
Workspace Locations
We are also open to remote fellows in the UK, US, or Canada
Visa Sponsorship
not
Program Duration
Please note
Apply here
NOT
Minimum education
Required field of study
Minimum years of experience
Location-based hybrid policy
Visa sponsorship
We encourage you to apply even if you do not believe you meet every single qualification.
Your safety matters to us.
How we're different
Come work with us!
Guidance on Candidates' AI Usage
Not included in the source posting: what you'll do.
Skills
Who can apply
The employer didn't state any visa, work authorization, citizenship or clearance requirements in this posting. Confirm with the employer before applying.
Read automatically from the employer's posting text. Always confirm with the employer — requirements can change after a job is published.