Computer Using First Step

XBOW tests Anthropic's Mythos Preview for offensive security

Anthropic's Mythos Preview was highly effective at finding vulnerability candidates, especially when analyzing source code.

21h

The end result could be models trained by models to achieve goals set by models, whose safety is verified only by models.

Some results have been hidden because they may be inaccessible to you