about
I'm a final-year PhD student at the University of Oxford, supervised by Prof. Jakob Foerster, researching how to build autonomous and open-ended AI scientists. I'm currently on an internship at Inherent, after more than a year at Meta Superintelligence Labs on their AI Research Agents team.
Previously: a Computer Science degree from Oxford (second in my cohort), then two startups — Halo, a biotech, and genei.io, Y-Combinator backed and now Coloop.ai.
news
- AIRA₃, the research agent from my former team at Meta, won Gold in NVIDIA’s Kaggle reasoning challenge, placing 8th of ~4,000 teams.
- Training AI Scientists to Replicate Research released with Inherent, introducing Faraday and the Replica benchmark. Covered by TechCrunch, Import AI, Hugging Face Journal Club and the Cognitive Revolution podcast.
- AI Research Preference Models released. Discussed on Hugging Face Journal Club, and on X by Bassel Al Omari, Jason Weston and Ed Hughes.
- Started at Inherent.
- LLMs Encode Their Failures accepted to COLM 2026.
- Left Meta after more than a year on the AI Research Agents team.
- Presented at the Meta AI agents community meetup.
- Two papers accepted to ICML 2026: Compute as Teacher and Goal-Conditioned Agents that Learn Everything All at Once.
- Co-authored AIRA₂: Overcoming Bottlenecks in AI Research Agents, released by Meta.
- AIRS-Bench released.
- Presented at NeurIPS 2025.
- Spoke to 6th form students at Michaela Community School in London about startups and LLMs.
- “Measuring What Matters” featured in Guardian and NBC News.
- Spoke at London Open Endedness Conference.
- Three papers accepted to NeurIPS 2025 main track!
- Meta released the Automated LLM Speedrunning Benchmark.
- Presented LILO at Oxford Internet Institute OxRML Group.
- Started at Meta as part-time Research Engineer.
- Presented at Oxford LLM Workshop.
- Learning to Reason at Pre-training Scale accepted to NeurIPS 2024 Language Gamification Workshop.
- Presented at Meta Open Innovation Community Workshop.
- Started PhD with Prof. Jakob Foerster.
- Started at the AIMS CDT, University of Oxford.
why ai scientists?
My interest in AI Scientists sprang from my fascination with scientific discovery – how over the last 500 years mankind developed the scientific method and began to make discoveries at an exponentially increasing rate. But I’m equally excited about the next 500 years: I believe solving open science problems is a genuinely good and responsible reason to build powerful AI, as opposed to, say, advertising tools or addictive online videos.
My working model of open-ended scientific discovery is a distributed multi-agent system in which agents:
- Create tasks they believe will unlock progress toward a goal.
- Use curricula to prioritise work on promising areas.
- Tackle tasks to learn and produce insights.
- Reward insights by the extent to which they unlock or amplify progress elsewhere. This can be a challenging task in itself, often revealing long-delayed, unexpected connections.
Research in this area therefore combines several fields, including LLMs (of course), open endedness, curriculum design, recursive model self-improvement, reinforcement learning and evolutionary search.
selected publications
AI Research Preference Models2026
Research agents can propose far more experiments than they can afford to run. Preference models predict which candidates are worth the GPU time, reaching the unguided agent’s 24-hour performance in roughly 15 hours and setting new state of the art on two AIRS-Bench tasks.
AIRS-Bench: A Suite of Tasks for Frontier AI Research Science Agents2026
Twenty research tasks drawn from state-of-the-art ML papers, spanning language modelling, mathematics, bioinformatics and time series, and assessing the full research lifecycle with no baseline code given. Frontier agents beat human SOTA on four tasks and fall short on the other sixteen.
LILO: Learning to Reason at the Frontier of Learnability2025
An online curriculum for training LLMs with RL, prioritising high-variance questions and yielding large gains in training efficiency across RL algorithms and reasoning benchmarks.
full publication history → google scholar