AI Alignment Research

My current AI safety research is supported by a grant from BlueDot Impact.

Projects

emmy1

Evaluation-invariant measurement for alignment of multi-agent systems (github)


llms en garde: ain’t misbehavin’?2

Do LLMs misbehave less when they’re en garde? (github)


liar liar

Is the model scheming to deceive… or is it just wrong? (github)


activation tomography

Natural Language Autoencoders as measurement instruments for AI safety (github)


paper chase

Multi-agent simulation of a scientific publishing ecosystem (github)


Technical AI safety and alignment research interests

I’m interested in developing construct-valid instruments for measuring properties of AI systems. Active projects include emergent properties of AI collectives, measuring an eval-awareness discount (safety evaluation adjustment when a model is aware it’s being tested), and nascent work on measurement in AI control evaluations (monitor calibration, cross-environment validity).

I’m broadly interested in technical AI safety across domains: evaluations and measurement methodology, control, adversarial evaluation, scalable oversight, multi-agent safety, model organisms, mechanistic interpretability.


  1. Name inspired by Emmy Noether (1882–1935), who did foundational work connecting symmetries to invariants. ↩︎

  2. A jazz standard, here performed by Sarah Vaughan (feat. Miles Davis): Ain’t Misbehavin’ ↩︎