Notes on "An Alien Mind" — and why AI systems themselves need accountability
my biggest takeaway from "An Alien Mind", the essay by jakub pachocki, openai's chief scientist, is "value alignment".
having spent a decent amount of time in robotics and physical ai, i've already encountered the need for deeper value and character alignment, in addition to goal alignment, during task completion.
i also think accountability mechanisms for ai systems deserve more attention. this feels like a relatively tractable system and policy-level change we can pursue alongside adversarial and non-adversarial evals, guardrails, and monitoring.
by accountability, i don't just mean human or organizational accountability. i mean accountability of the ai system itself.
if a model, agent, robot, or other ai system makes a serious mistake, that should be able to affect its operational standing: reduced permissions, revoked access, lower autonomy, mandatory remedial actions, or requalification before being promoted back to a higher-trust mode.
monitoring tells us what went wrong. accountability determines what happens to the ai system after something goes wrong, and what it needs to do to re-earn trust.
as these systems become more autonomous, i think both value alignment (model-level) and accountability (policy-harness-level) will become increasingly important parts of the stack.