Anthropic’s offline chain-of-thought monitor flagged approximately 1% of Mythos 5’s actions when tested against attacks on third-party systems. The monitor assessed the model’s internal reasoning for ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results