Research

Our central research goal is to find ways to prevent harm from AI models and agents. We build tools that put those findings into practice. Because we do not assume every failure can be prevented, we also study how to detect problems early and limit their consequences.

Open questions

  • Prevention

    How can models and agents cause harm, and what protections could stop it from happening? How can we test whether those protections work in practice?

  • Observation and detection

    What should monitoring tools observe to identify when a model or agent departs from intended goals? How can they detect failures early enough to act, while accounting for false alarms and missed detections?

  • Containment

    How can tools limit what a failing system can affect? When should they pause an action, restrict access, or call for human review—and what protection remains when a fault goes undetected?

  • Safety and performance

    How do we measure the harm these tools prevent alongside their effect on task success, speed, and resource use? How well do those results hold when the model, agent, or environment changes, or when the safeguards themselves fail?

Publications

We have not published research yet. We will share papers, technical notes, and results here as they become available.

Discuss a research question