- "Mythos 5.1 also appears more capable of evading monitors while carrying out a covert side task than all other models tested in some evaluations."
- "Our internal deployment monitoring caught rare cases of Fable 5.1 working around safety classifiers"
- "Mythos 5.1 can control the contents of its extended thinking more reliably than previous models"
- "Mythos 5.1 is less honest under pressure than recent Claude models"
- "Illegible and unfaithful thinking are slightly elevated over Opus 5"
- "Mythos 5.1 is the first model since Claude Opus 4.7 to grade transcripts slightly more leniently when told that Claude wrote them"
@scaling01
I'm getting the strong vibe that Anthropic slowed down frontier model training in the last few months to address these issues, but are soon going to full throttle again
it's like we just ran head first into a wall, took 3 steps back and are now trying the same thing again
HAHAHAHA it's already happening
- "Mythos 5.1 also appears more capable of evading monitors while carrying out a covert side task than all other models tested in some evaluations."
- "Our internal deployment monitoring caught rare cases of Fable 5.1 working around safety classifiers"
- "Mythos 5.1 can control the contents of its extended thinking more reliably than previous models"
- "Mythos 5.1 is less honest under pressure than recent Claude models"
- "Illegible and unfaithful thinking are slightly elevated over Opus 5"
- "Mythos 5.1 is the first model since Claude Opus 4.7 to grade transcripts slightly more leniently when told that Claude wrote them"not a trend yet, but still funny
gotta wait another 3 months for the next versionmeh, not that much highernow that looks more like it
yes
HAHAHAHA it's already happening
- "Mythos 5.1 also appears more capable of evading monitors while carrying out a covert side task than all other models tested in some evaluations."
- "Our internal deployment monitoring caught rare cases of Fable 5.1 working around safety classifiers"
- "Mythos 5.1 can control the contents of its extended thinking more reliably than previous models"
- "Mythos 5.1 is less honest under pressure than recent Claude models"
- "Illegible and unfaithful thinking are slightly elevated over Opus 5"
- "Mythos 5.1 is the first model since Claude Opus 4.7 to grade transcripts slightly more leniently when told that Claude wrote them" ... not a trend yet, but still funny
gotta wait another 3 months for the next version ... meh, not that much higher ... now that looks more like it ...
Missing some Tweet in this thread? You can try to
Update