Housekeeping

“Harlequin’s Carnival” by Juan Miró. I find the chaos and mind-bending nature of surrealist art both whimsical & relatable.

My desire for this newsletter is to capture my actual voice and read as my honest and original thoughts, provide tactical frameworks and learnings in areas of expertise, and sometimes provide a review of current research, trends, and possibly even some interviews with dear friends & incredible practitioners in the space. Therefore, I will never use AI models to write, only review & provide me with critical feedback, and generate the occasional diagram or image. When I do use AI, I will provide a disclosure. You’re here for my voice and writing, and I very much intend to provide that.

Why This Topic

I’ve been working on risks from frontier AI models for ~90% of my career, and the prospect of even capturing the complexity of this space in this post is daunting. I like to start with the baseline of: I know many things at a high-level, a couple of things at a much deeper level, and everything else is a learning sprint. As a safety generalist working on risk for model and product launches on insane and often, uncomfortable timelines, you’re forced to be a pragmatist. In order to combat this internal fatigue and move at the pace that is required in the current “AI race”, below is how I personally risk triage.

(*very important note to get out of the way: that this does NOT represent the risk view/posture/stance/process of any current or future employers of mine, please & thank you!*)

Risk Triage

At a high level, risk triage means that the flow below will capture the essence of the risks you are concerned about, help document, understand and mitigate (some) risks, and provide a few of what’s occurring in the wild via post-deployment monitoring. Pieces of this are covered across known published frameworks such as NIST’s AI RMF, the establishment of “frontier safety cases” taking inspiration from highly regulated industries such as nuclear and aviation, and approaching safety from a sociotechnical lens to build evaluations that capture contextual, real world harms.

My Flow: capabilities/context → theoretical risks → mitigations → residual risk → deployment → post-deployment monitoring

what novel capabilities does the model have? how was it designed and how will it be deployed?

How was your model trained (is the data licensed, does it contain personal information, how was it filtered), how was it designed and does it present a new or increased capability? Let’s say your new model is quite good at STEM applications as it was fine-tuned on that data—how does this translate to dual-use capabilities in CBRNe? Is there a new modality you are releasing, which ultimately needs its own safety protocols and coverage outside of any text-based protections? What do you envision to be the critical use cases of the model and who is the intended audience? If the model has been designed for educational deployment use cases, then naturally I’d imagine it would appeal to students and under 18 (U18) demographics, for example. A trusted tester launch with 10 users should not be treated the same as an enterprise launch with 1k customers. This is the kind of stream of consciousness that guides my thinking through this first step.

Identify risks (what bad things could happen?)

Next, the risks identified should always be in context of the model’s capabilities, particularly as model capabilities get more autonomous and we lose sight of how they may directly or indirectly impact users. There is plenty of scholarship, blogs, arxiv papers here you could consult to understand what types of risks should be top of mind for a frontier model. The MIT Risk Initiative is doing some notable work in this space to document various risk and sub-risk categories, which is a great place to start. Ultimately, without understanding the model’s core technical functionality and how it will be used, you cannot contextualize risk.

If I were to provide a laundry list of risks I frequently think about with respect to more capable frontier (agentic and multimodal) models it would include: cyber, child safety, affective/well-being risks, security risks (prompt injection/data poisoning/distillation), CBRNe, loss of autonomy, epistemic risks (systemic and individual risks to the future of knowledge, paper of interest/podcast of interest), and fairness and allocative harms (which is think are understudied in the context of frontier AI models).

Assess risk (how likely? how severe?)

Severity x likelihood is a very common framework to assess risk. In application it is often imperfect, as severity can be extremely subjective and likelihood can be difficult to both capture and forecast until you monitor after deployment (especially for more novel or rarer risks). Sometimes, you care even less about likelihood when the risk is severe enough or if we think it could unleash catastrophic harm (cyber/bio are often cited examples). Still, I find this framework to be useful to organize my thinking in a pinch when I need a pragmatic approach in assessing a risk. Below are some common components to consider in each area.

Components of Severity: violation of privacy, dignity or human rights, regulatory/legal risk with liability considerations, affective/mental health harm, how reversible the harm is (if it is difficult to reverse, it is much more severe), impacts to financial or economic well-being, harm to minors/minor safety.
Components of Likelihood: # of users/daily active users (DAU), user base (is your model/product appealing to vulnerable populations, children, malign actors?), empirical understanding of risk (how common are incidents in the wild, what does literature show us), observed model behavior in safety evals, mitigations coverage for a risk (are there classification systems, refusals, etc that will cover the risk?)

Documenting which components of severity x likelihood went into your assessment is critical to capturing your risk understanding and highlighting any gaps.

Design mitigations (how do we reduce risk?)

Without mitigations as an intervention, frontier models cannot be deployed safely. Ideally, these are being explored earlier in the model development process, to understand where and if post-training interventions (RLAIF, RLHF, fine-tuning) can help with risk coverage proactively. Usually this is a “swiss-cheese” approach, in that there will be holes and the goal is to layer mitigations at the model level, on top of the model, and in product surfaces (at the user/account level) and build defense-in-depth when you can, especially for trickier areas like cyber. Often times mitigations are prioritized for P0 risks that can easily be identified, assessed, and defined, such as CSAM images for which industry standard is to use hash matching, novel classifiers, and text-based interventions. This can leave areas that don’t fall in a “P0” and are more challenging to define, measure, and mitigate against, behind. This is something I worry about all the time across the AI safety space.

Monitor for realized harms (are they actually occurring post-launch?)

Monitoring for realized harms is so so important but often feels like the “middle child” of the AI safety pipeline (in that parents only worry about the middle child when something goes visibly wrong). But monitoring is how we can build a full loop of understanding, build improved data flywheels for better classification/mitigations, and rapidly recalibrate our risk assessments for a given model. Without monitoring, a lot of the previous work we did to get here can get incredibly stale and not grounded in real world harms and consequences.

**if this type of content interests you, subscribe here :)

Bilva