Mechanistic interpretability theory
Research Scientist
TLDR: We're looking for a research scientist to collaborate with our autoresearch framework: understanding and identifying promising outputs, and turning them into precise research questions and experiments. This is a (modestly) paid role on a three-month project.
About us
Verified Mechanisms is an independent, grant-funded research project developing theoretical foundations for mechanistic interpretability. We study how LLM agents can collaborate to discover and formally verify theorems about the internal computations of transformers, beginning with a simple pilot question: how many attention heads are needed to represent a Boolean function?
Recent advances in language models for mathematical reasoning and formal theorem proving may make it possible to produce rigorous results at a lower cost. This project aims to produce foundational results for empirical interpretability by leveraging the growing mathematical abilities of frontier models, while identifying multi-agent harnesses that make such theoretical research faster and more reliable.
About the role
- Read and interpret the outputs from autoresearch runs, identifying promising conjectures, hidden assumptions and useful directions.
- Turn promising outputs into precise research questions and experiments.
- Work with the research engineer to improve the autoresearch workflow.
- Contribute towards public writeups.
The role is 10–15 hours per week, including a weekly meeting.
The autoresearch pipeline results typically vary in scope and significance, so a large part of this role is research judgment: identifying the small subset that matters most.
You will be a good fit if you have
- A strong mathematical background, including experience writing proofs in an (applied) mathematical setting.
- Proficiency in using LLMs and coding agents as part of your daily workflow.
- The ability to critically evaluate LLM-generated research outputs, identify promising conjectures and arguments, and pick out the substantive insights.
- Strong scientific communication, with the ability to explain complicated ideas in an understandable way.
Nice to have
- Proficiency in Python and Git, including managing worktrees created by parallel runs within the same project.
- Familiarity with mechanistic interpretability, or other relevant areas of mathematical machine learning.
- Prior experience with Lean or another proof assistant.
Perks
- A stipend of $1,000 per month, or more if the project receives further donations.
- Claude Max and ChatGPT Pro subscriptions.
Timeline
- A short application form, due August 10.
- A take-home screening task, due August 20.
- A final interview, scheduled over the following 4–5 days.
We plan to wrap up the process by the end of August, so that we can get started by September 1.