I like the embedded-evaluator proposal. Coming from chip and autonomous-vehicle verification, I’d suggest one concrete addition: An evolving coverage map connecting safety claims to evidence.
For each relevant configuration (during training, internal use, and release), record which requirements and situations were tested, what failed, what remains unchecked, and where the checking methods themselves are weak. Also record whether apparent alignment survives further capabilities training and generalizes to situations withheld from alignment training.
The map may be used to guide improving alignment, not just measuring it - for example, by systematically generating alignment stories / training cases across relevant situations. When a problem is found, identify the broader failure class, strengthen alignment across that class, and re-evaluate (including after further capabilities training).
A coverage map cannot establish that all important risks have been identified: Searching for missing dimensions and checkers is part of the work. But it can make the scope and limitations of the evidence inspectable, and help prioritize how to use the time pacing buys us. I discuss this approach in V&V takes on “Pacing the frontier”.
I like the embedded-evaluator proposal. Coming from chip and autonomous-vehicle verification, I’d suggest one concrete addition: An evolving coverage map connecting safety claims to evidence.
For each relevant configuration (during training, internal use, and release), record which requirements and situations were tested, what failed, what remains unchecked, and where the checking methods themselves are weak. Also record whether apparent alignment survives further capabilities training and generalizes to situations withheld from alignment training.
The map may be used to guide improving alignment, not just measuring it - for example, by systematically generating alignment stories / training cases across relevant situations. When a problem is found, identify the broader failure class, strengthen alignment across that class, and re-evaluate (including after further capabilities training).
A coverage map cannot establish that all important risks have been identified: Searching for missing dimensions and checkers is part of the work. But it can make the scope and limitations of the evidence inspectable, and help prioritize how to use the time pacing buys us. I discuss this approach in V&V takes on “Pacing the frontier”.