I’m looking for people, advice, critiques, and funding to build a research program on value generalisation – the ability of an AI to correctly extend human values and preferences to situations neither it nor we have seen before. My ongoing research has become convinced that this is necessary if we want to get aligned AIs that operate in the human interest. This would be a focused research organisation or a commercial venture. I’m leaning towards commercial, because alignment techniques confined to academic papers get ignored – or worse, mined for capability-relevant parts while the alignment component is discarded. This post is the research program’s summary. The technical case is in the next post , and one exciting consequence – AIs whose alignment grows with their capabilities – is in the post after that . Without value generalisation, AI can't be reliable: it lacks that capability Nothing technical stands in the way of you handing an AI assistant full control of your devices and accounts today. And I’m not saying using an app or harness that has been designed for these tasks. I mean hand an LLM your passwords, email, social media, and bank access along with a little note stating what you want, plugging inputs and outputs via APIs, and letting it go wild. The capacity to do this exists. What doesn’t is the trust. And the trust is missing for a good reason: today’s AIs cannot be relied on to understand your interests in situations that weren’t covered – explicitly or implic…

Full article content could not be extracted automatically. Read the original below.