Bring your toughest prompts and start voting. Scores coming soon.
In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology.
@claudeai Fable 5.1 is also available in: Code Arena: WebDev, Text, Vision, and Document.
@claudeai
We’re introducing Claude Fable 5.1 and Claude Mythos 5.1.
They're the world’s most advanced models for coding and knowledge work. https://t.co/8P9PSrWPi3
Fable 5.1 by @AnthropicAI is now in the Arena!
Bring your toughest prompts and start voting. Scores coming soon.
In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology.
@claudeai Fable 5.1 is also available in: Code Arena: WebDev, Text, Vision, and Document.Test Fable 5.1 in Battle Mode and Agent Mode at:
yes
Fable 5.1 by @AnthropicAI is now in the Arena!
Bring your toughest prompts and start voting. Scores coming soon.
In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology.
@claudeai Fable 5.1 is also available in: Code Arena: WebDev, Text, Vision, and Document. ... Test Fable 5.1 in Battle Mode and Agent Mode at:
Missing some Tweet in this thread? You can try to
Update