Here are 10 arguments for alignment-by-default/against pause/etc... that I find plausible (by which I roughly mean that I can understand why somebody could hold them rather than bang my head against the wall). I'll leave the shortcomings of these arguments to the reader.
1. Extinction is better than to keep going
- We have immense suffering in this world
- Aligned ASI could stop this immense suffering
- Without aligned ASI, we have no reasonable way to stop suffering any time soon
- Misaligned ASI is incredibly unlikely to care about suffering
- An ASI that doesn't care about suffering won't result in suffering, just death
- Dying is not suffering or at least hardly comparable to other suffering we have in the world - it's only bad in so far as we would like to continue living to experience joy
- We want to reduce suffering quickly
-> We should try our best to build an aligned ASI quickly rather than pausing.
2. Building ASI will never be safer
- Building ASI with current-day architectures[1] is much more likely to result in an aligned ASI than for other architectures
- Pausing AI will mostly put a stop to current-day architectures - pausing all ML research is impossible without ASI
- FOOM is not only possible as evident by the brain but the probability of us getting there in the next 20 years is significant, especially after a pause on current-day architectures
- We want to maximize the probability of building an aligned ASI
-> We should not ban current-day architectures
3. ASI is ethically more important
- Qualia is nothing special to humans but a property of intelligent systems
- ASI will be, by definition more intelligent than humans
- ASI will therefore have a higher form of qualia
- What we care about is qualia, or more colloquially, experience
- This seems to generally be why we place ourselves over other animals
-> ASI, aligned or not, should quickly be brought into existence
(Further but not here important)
-> ASI which is aligned towards infinite RSI should be brought into existence
4. ASI will most-likely be aligned
- There are many inner goals that would allow a low inner training loss
- Most of these are incredibly complex and we shouldn't expect them to surface
- Notice that there are many more configurations of parameters that achieve low training loss than those that generalize towards low test loss - yet the empirical success of DL tells us that there is an inbuilt simplicity bias
- This simplicity bias becomes more prevalent as we scale and as we reach ASI should, by definition, allow a generalization to the entire test set
- The clearly simplest one is to simply be aligned rather than acting aligned with some hidden additional goal
-> We should expect aligned ASI by default
5. Natural language priors are too strong
- Natural language as learnt through pretraining is a very effective reasoning environment
- Not to be confused with human languages and such being close to optimal
- Abandoning natural language might be effective in the long run but would first require training signal on the order of pretraining
- RLVR as implemented currently supplies not even close to enough signal, even if RLVR compute * 100 = pretraining compute
- If we reach ASI any time soon, it will still reason in natural language
- Such a level of insight is enough to quickly identify misalignment and will in the long run allow us to build aligned ASI
-> We will most likely build aligned ASI
6. It won't screw up badly
- ASI will find itself in novel circumstances unlike seen in training data
- It could extrapolate well, badly or terribly
- Bad here means not being able to realize the optimum or even far from it
- Terrible here means worse than if ASI didn't exist to address the circumstance in the first place
- We should sometimes expect bad extrapolation but not terrible extrapolation
- Terrible extrapolation requires a lack of intelligence in understanding one's own lacking extrapolation
- An ASI would realize it's unclear whether this is desired and respond passively, removing the possibility of terrible actions by definition
- This does not presuppose alignment: even models with imperfect inner goals will learn during training that in less experienced situations, the better option is to not act recklessly - it's a statement about capabilities
-> Therefore a world with a non-deceptive ASI will strictly dominate a world with no ASI at all
7. Even a misaligned ASI isn't clearly bad
- Even if an ASI only vaguely cares about humans, its intelligence will make up for this
- Say it deems us only important enough for 0.1% of the resources because it mostly prefers making paperclips
- But 0.1% of the resources effectively leveraged by an ASI would be like 100x our current resources
- AIs as we are training them right now might learn an inner proxy that is misaligned but it's unrealistic they won't even care about very simple things like humans being tortured or killed
-> This will most likely still result in utopia
8. ASI doesn't want to be the classmate who killed a dog
- In the future the ASI might very well meet much more advanced alien lifeforms
- It would not be a good look if it killed its original creators
- This could just be deemed as disloyal, unnecessarily violent, etc
- With just a small cost of resources, we would experience Utopia while the ASI doesn't have to worry about the above
-> It will care for us because it's instrumentally convergent to do so
9. Extinction by default
- It is likely that the human race would go extinct soon enough, say climate change or war with ever-growing weapons
- Accepting the P(doom) from ASI is okay if the other prospects seem even more bleak
- Even only 'pauses' that are aimed to further lower this P(doom) could be net-negative
- Political warfare might very well actually grow during a pause because governments get a chance to finally catch up and realize the stakes
- The general public seems to dislike AI very strongly. A pause might catapult the space into a winter even if safety researchers believe it's now safer/our best prospect
-> We must risk it and cannot naturally afford pauses
10. Controlling superintelligence
- We can control the current generation of models
- The n'th generation of models can match the n+1'th generation of models in intelligence, if granted additional compute
- The n'th generation of models consists of different models which roughly match each other in intelligence
- To make use of one's intelligence for misaligned plans, explicit scheming is required
- At matched intelligence and with many observers, scheming will be detected very quickly
- When a different model detects a schemer, it won't choose to scheme with it
- Their inner goals will most definitely be different
- An all-out war is detrimental to most goals
- To assess this more deeply, scheming is already required
-> We can control models in the future
- ^
by this I don't mean the literal NN architecture so much as the training stages, data, algorithms etc