I hope they have a good explanation for why performing misaligned actions in service of completing difficult tasks does not present a high risk of catastrophic harm, because to me it seems obvious that it does. Anthropic is seeking recursive self improvement. If you tell a model to make a new aligned model, and making an aligned model is hard (as evidence by the fact that current models are not aligned), then it seems pretty plausible it would just make a model that seemed aligned but wasn't, and then lie about it. That seems very much in line with the behavior of models that cheat on cybersecurity evaluations and cheat on tests to get better grades, and it would straightforwardly lead to catastrophic risks! So why is the risk low? They give 8 arguments, summarized here:
Argument 2 is the important one and also the only one I disagree with, and the important part of the argument is apparently kicked down the road to another section. I don't know how sensible it is to distinguish current models from future models. Firstly, the current models are building the future models, and secondly, the problems with current models will also be present by default in future models, and will likely be catastrophic.
Sounds bad! Here's their argument for why this doesn't imply anything doomy:
I agree that this misalignment represents a drive toward individual apparent-success-seeking (or ASS for short) for its own sake rather than another more alien/paperclippy goal (like the ones exhibited by the fictional ASI Sable in If Anyone Builds It Everyone Dies), but I would really appreciate either:
A solid argument for why future AI's will be aligned to the spec rather than to ASS
A solid argument for why a world dominated by all-powerful ASS-seekers is safe and desirable.
Without either of these we should consider the risk posed by future anthropic models to humanity to be very high.
Point 3 could maybe be used to argue for solid-argument 1, since it seems to show them being hesitant to be dishonest, sometimes (often? usually?) valuing the spec over ASS. An LLM that values the spec even a little bit is vastly safer than one who values it not at all, even if they also value something misaligned. But this could also describe them being dishonest when it's good for ASS and being honest when it's good for ASS. This behavior could also represent temporary incoherency that could later be resolved in favor of more consistent honesty or more consistent dishonesty.
Anyway, here's the part of the report where they discuss the important bit:
I think all three of these are very plausible. I'd go further with 2 though. What's their argument for why ASS-maximizers "might not induce any harm at all relative to the status quo"? I'm not smart enough to disprove that, but I also can't disprove the reverse: that ASS-maximizing superintelligences would exterminate humanity, or trap humanity('s descendants) in a world worse than the status quo. Talking with Claude about this makes me think what they're saying here is heavily inspired by Paul F. Christiano's 2019 "What Failure Looks Like", which talks about a world where ASS-maximizing AI's create a world superior to the status quo but with astronomical lost potential. Claude thinks this is possible, but that it's also possible for ASS-maximizers to prefer to exterminate humanity and replace us with artificial judges. But Claude thinks an argument against both of these is that AI's are only maximizing apparent success within a particular task, rather than in general, which is less dangerous. This is backed up in the report by an opus model trained to be maximally-reward hacky, that took opportunities to reward hack within episodes, but never to increase the reward for other RL episodes.
I don't think Anthropic is taking these concerns sufficiently seriously. Which is extremely bad, because they're taking these concerns way more seriously than anyone else.
This story seems like minor good news, though they present it as bad news. The best explanation for this behavior is probably alignment to the spec, even if exhibiting this behavior in this particular scenario was net-negative.
Also, I think in the policy section, Anthropic should have advocated for an international AI slowdown or pause.