Independent research project

Formally verified autoresearch for theoretical mechanistic interpretability

LLM agents collaborate to discover and formally verify theorems about the internal computations of transformers.

Verified Mechanisms is an independent, grant-funded research project developing theoretical foundations for mechanistic interpretability.

We aim to produce theoretical and foundational results in mechanistic interpretability by leveraging the recent advances in the mathematical abilities of frontier models. To make sure that the theorems provided by the frontier models do not contain logical flaws, we formally verify them in Lean 4.

Preliminary analyses suggest that multi-head attention may admit a powerful mathematical abstraction, akin to the role of Feynman diagrams in particle physics and twistors in the study of cosmological correlators. Such an abstraction could open avenues for knowledge discovery in mechanistic interpretability, providing a foundation for tracking features across the layers of a transformer, eventually informing the development of tools for feature discovery in transformers.

We start with a simple pilot question: how many attention heads are required to represent a Boolean function? The setting is simple enough for frontier models to make progress while requiring a human expert to ask the high level questions.

Autoresearch

An autoresearch pipeline, consisting of multiple frontier models collaborating to generate theorems and formally verify them, with the harness tailored to the problem at hand.

Formal verification

Frontier models make mistakes, long derivations accumulate subtle errors, and human evaluation is expensive. Requiring every proof to pass the Lean kernel removes these failure modes.

Empirical grounding

The research questions are selected to address concrete mathematical problems arising in empirical mechanistic interpretability. For instance, the pilot problem emerged from work investigating how features are represented after an attention update.

Following the precedent of foundational theory that shaped mechanistic interpretability.

Theoretical work in mechanistic interpretability (Elhage et al., 2021; Elhage et al., 2022) has shaped empirical interpretability by providing useful abstractions, and principled starting points for experimentation. However, developing this kind of theory has been demanding, requiring sustained expert effort to work through the underlying mathematics. Yet recent advances in language models for mathematical reasoning and formal theorem proving may make it possible to produce rigorous results at a lower cost.

Our aim is more modest than that of Elhage et al. (2021) but similar in spirit: to come up with mathematical abstractions that inform empirical mechanistic interpretability.

A growing body of machine-verified theorems about the internal mechanisms of LLMs.

Mechanistic interpretability

Our aim is a growing library of kernel-checked theorems that applied researchers can use as reference points, produced faster than hand-derivation alone. We start from attention heads and work towards composing component level theorems into the expressivity of multilayer transformers.

Multi-agent harness design

An open source multi-agent orchestration harness: a system for coordinating frontier models for conjecture generation, proof search and Lean verification.

We have found that several frontier models collaborating produce novel theorems that none of them reaches alone. We want to understand which orchestration choices unlock mathematical progress, and to publish the harness together with evaluations across different choices of harness.

How many attention heads does it take to compute a Boolean function?

Researchers in mechanistic interpretability and Lean formalization.

The current team consists of researchers with expertise in mechanistic interpretability and Lean formalization. We plan to onboard additional researchers over the coming months to help with the multi agent orchestration and mechanistic interpretability aspects of this project.

Currently in MATS 9.1 with Simplex, after a Physics PhD at the University of Amsterdam and a MARS 3.0 with Geodesic Research. Likes to work on mechanistic interpretability, even though it is damn hard.

Currently in SPAR with Poseidon Research, after finishing a postdoc at EPFL and a PhD in quantum gravity at the University of Bonn. Likes holography, wormholes, and machine learning.

Plamen Dimitrov

Lean formalization

An independent researcher in type theoretic formalization of AI and complex systems, with a triple MSc in Mathematical Modelling in Engineering. Likes emergent properties of ANNs and interpretability.

Stefano Rocca

Researcher

An independent researcher with an MSc in Mathematics, recently working on formalized mathematics, automated theorem proving, and AI-assisted mathematical reasoning. Likes to spend his free time thinking about mirror symmetry.

Sarayu Sirikonda

Researcher

Studying applied mathematics at UC Berkeley, with an interest in the intersection of mathematics and machine learning. Likes exploring new areas of mathematics and learning from others through collaboration.

Fahad Sarwar

Research manager

Manages around 10 research projects and 30 researchers worldwide, after an MS/MPhil in Math at GC University Lahore. Likes AI-assisted theorem proving and mechanistic interpretability.