Harness design
Research Engineer
TLDR: We're looking for a research engineer to build an orchestration harness that enables several frontier models to collaborate on formulating and proving theorems that are formally verified in Lean 4. This is a (modestly) paid role on a three-month project.
About us
Verified Mechanisms is an independent, grant-funded research project developing theoretical foundations for mechanistic interpretability. We study how LLM agents can collaborate to discover and formally verify theorems about the internal computations of transformers, beginning with a simple pilot question: how many attention heads are needed to represent a Boolean function?
Recent advances in language models for mathematical reasoning and formal theorem proving may make it possible to produce rigorous results at a lower cost. This project aims to produce foundational results for empirical interpretability by leveraging the growing mathematical abilities of frontier models, while identifying multi-agent harnesses that make such theoretical research faster and more reliable.
We would like to understand how models can collaborate effectively: which division of labour actually moves a proof forward. Much of the work is figuring out the right harness for these problems, through rigorous experimentation.
About the role
- Design the framework that coordinates agents to do the math and formally verify the results.
- Build benchmarks and compare harnesses based on verified mathematical progress.
- Build Inspect logging over the autoresearch runs, so that the runs are easy to audit.
The role is 10–15 hours per week, including a weekly meeting.
You will be a good fit if you have
- Experience building and running agentic workflows on cloud infrastructure.
- Experience designing controlled evaluations on multi-agent workflows.
- Proficiency with Python and Git, including reading an unfamiliar codebase and making sense of it.
- Proficiency in using LLMs and coding agents as part of your daily workflow.
Nice to have
- Experience designing effective loops for agents like Claude Code or Codex.
- Prior experience with Lean or another proof assistant.
Perks
- A stipend of $1,000 per month, or more if the project receives further donations.
- Claude Max and ChatGPT Pro subscriptions.
Timeline
- A short application form, due August 10.
- A take-home screening task, due August 20.
- A final interview, scheduled over the following 4–5 days.
We plan to wrap up the process by the end of August, so that we can get started by September 1.