Best for
- Writing
Claude Fable 5.1
Claude Sonnet 5
Claude Opus 5.5
- Conversation
Gemini 3.8 Flash
Claude Sonnet 5
Claude Opus 5.5
- Coding
Kimi K3
Qwen3.8 Max
Claude Opus 5.5
- Reasoning & math
GPT-6 Astra
Kimi K3
Claude Fable 5.1
- Web research
Claude Opus 5.5
Claude Fable 5.1
GPT-6 Astra
- Data & analysis
Claude Fable 5.1
GPT-6 Astra
Claude Opus 5.5
- Long documents
GPT-6 Astra
Muse Spark 1.3
Kimi K3
- Images & PDFs
Gemini 3.8 Flash
Gemini 3.1 Pro (Preview)
Claude Fable 5.1
- Languages
Gemini 3.8 Flash
Gemini 3.1 Pro (Preview)
Qwen3.8 Max
- Quick replies
GPT-5.6 Luna
Gemini 3.8 Flash
DeepSeek V4 Pro
Models
GPT-6 Astra
OpenAI · Sep 2026- Reasoning & math
- Long documents
- Data & analysis
- Computer and browser use
- Long-horizon agents
- Scientific reasoning
- Hard math
- Agentic coding
- Deep web research
- AA Index
- 53
- Arena Elo
- 1480
- GPQA
- 96%
- HLE+tools
- 57.2%
- Raise reasoning effort for hard problems
- Use fast mode (2x price) for latency
- Cache long prompts: $1 per 1M cached
FinCraftly routes coding & reasoning & math hereLinkGPT-5.6 Luna
OpenAI · Jul 2026- Quick replies
- High-volume automation
- Classification and extraction
- Chat support bots
- Cheap sub-agents
- Summarization
- AA Index
- 37
- GPQA
- 92.3%
- SWE Pro
- 62.7%
- TB 2.1
- 84.7%
- Use for bulk, cost-sensitive jobs
- Set reasoning effort none for speed
- Escalate hard cases to Astra
FinCraftly routes quick replies & conversation hereLinkClaude Opus 5.5
Anthropic · Sep 2026- Web research
- Writing
- Conversation
- Agentic coding
- Long-running agents
- Professional knowledge work
- Long-form writing
- Computer use
- Document analysis
- HLE+tools
- 67.7%
- TB 4.0
- 66.4%
- Default pick for most agent workloads
- Use xhigh effort for hardest tasks
- Fast mode: 2x price, faster output
FinCraftly routes writing & conversation hereLinkClaude Fable 5.1
Anthropic · Sep 2026- Writing
- Data & analysis
- Web research
- Multi-hour agent runs
- Scientific research
- Hard reasoning
- Agentic coding
- Complex analysis
- AA Index
- 53
- Arena Elo
- 1498
- HLE+tools
- 65.6%
- TB 4.0
- 55.8%
- Reserve for the hardest, longest tasks
- Cache reads $0.25/1M cut agent costs
- Try Opus 5.5 first; similar on most work
FinCraftly routes writing & coding hereLinkClaude Sonnet 5
Anthropic · Jun 2026- Writing
- Conversation
- Everyday coding
- Business writing
- Customer-facing chat
- Computer use
- Agent workflows
- AA Index
- 38
- HLE+tools
- 57.4%
- HLE
- 43.2%
- SWE Pro
- 63.2%
- Best value Claude for daily work
- Lower effort for faster replies
- Upgrade to Opus 5.5 for long agents
FinCraftly routes writing & conversation hereLinkGemini 3.8 Flash
Google · Sep 2026- Conversation
- Images & PDFs
- Languages
- Fast agentic coding
- Multimodal understanding
- Video and audio input
- High-volume apps
- Chart reasoning
- AA Index
- 41
- Arena Elo
- 1493
- HLE
- 45.4%
- SWE Pro
- 61.6%
- Very fast; good default for agents
- Use high thinking for hard tasks
- Price doubles on 2027-01-01
FinCraftly routes conversation & long documents hereLinkGemini 3.1 Pro (Preview)
Google · Feb 2026- Images & PDFs
- Languages
- Multimodal analysis
- Long documents
- Visual reasoning
- Creative coding
- Research synthesis
- AA Index
- 30
- Arena Elo
- 1487
- GPQA
- 94.3%
- HLE+tools
- 51.4%
- Keep prompts under 200K for base price
- Strong at video and image inputs
- Try 3.8 Flash first for speed
FinCraftly routes long documents & images & pdfs hereLinkGrok 4.7
xAI · Sep 2026- Conversation
- Coding
- Coding
- Engineering problems
- Professional knowledge work
- Real-time X context
- Cost-efficient reasoning
- AA Index
- 46
- TB 4.0
- 38%
- Pricing doubles above 200K prompt tokens
- Use xhigh effort on hard tasks
- Fast variant: 2x price, 2x speed
FinCraftly routes conversation & coding hereLinkMuse Spark 1.3
Meta · Sep 2026- Long documents
- Multi-agent workflows
- Agentic coding
- Long-context recall
- Multimodal input
- Low-cost frontier work
- AA Index
- 48
- Arena Elo
- 1493
- GPQA
- 94%
- TB 2.1
- 88.8%
- Very low cost per task
- Cached input only $0.15/1M
- Access via Meta Model API
DeepSeek V4 Pro
DeepSeek · Apr 2026- Quick replies
- Math and reasoning
- Competitive coding
- Self-hosting
- Budget API use
- Long documents
- AA Index
- 36
- GPQA
- 90.1%
- HLE+tools
- 48.2%
- SWE Pro
- 55.4%
- Off-peak hours cost half
- Weights are MIT-licensed
- No image input; use Flash for vision
FinCraftly routes quick replies & coding hereLinkQwen3.8 Max
Alibaba · Sep 2026- Coding
- Languages
- Multi-step coding projects
- Tool orchestration
- Document parsing
- Chart reasoning
- Multilingual tasks
- AA Index
- 45
- Arena Elo
- 1481
- GPQA
- 92.6%
- SWE Pro
- 67.7%
- Good for long autonomous coding runs
- Use implicit caching to cut cost
- Strong for Chinese and multilingual
FinCraftly routes coding & long documents hereLinkKimi K3
Moonshot AI · Jul 2026- Coding
- Reasoning & math
- Long documents
- Agentic web research
- Large-repo coding
- Frontend and UI code
- Knowledge work
- Self-hosting
- AA Index
- 44
- Arena Elo
- 1485
- GPQA
- 93.5%
- HLE+tools
- 56%
- Top open model for browsing agents
- Cache hits cost $0.30/1M
- Self-host with vLLM or SGLang
FinCraftly routes coding & reasoning & math hereLinkGLM-5.3
Z.ai · Aug 2026- Conversation
- Coding
- Agentic coding
- Automation workflows
- Security analysis
- Self-hosting
- Budget reasoning
- AA Index
- 45
- Arena Elo
- 1483
- Reasoning is always on
- Flash variant for cheap bulk work
- Check custom licence before hosting
Mistral Medium 3.5
Mistral AI · May 2026- Languages
- Writing
- European data residency
- Self-hosted agents
- Coding with Vibe
- Structured outputs
- Multilingual business
- AA Index
- 14
- SWE Verified
- 77.6%
- Runs on as few as four GPUs
- Set reasoning effort per request
- Good fit for EU compliance needs
FinCraftly routes languages & writing hereLinkCommand A+
Cohere · May 2026- Languages
- Quick replies
- Enterprise RAG
- Private deployment
- Multilingual (48 languages)
- Document processing
- Tool use
- AA Index
- 13
- Runs on 2x H100 or 1x B200
- Apache 2.0: fully self-hostable
- Pair with Cohere Rerank for RAG
Benchmarks
| Model | GPQA | HLE | HLE+tools | SWE Pro | SWE Verified | TB 4.0 | TB 2.1 | ARC-AGI-2 | Arena Elo | AA Index |
|---|---|---|---|---|---|---|---|---|---|---|
| — | — | 65.6% | — | — | 55.8% | — | 90% | 1498 | 53 | |
| 96% | — | 57.2% | — | — | 57.9% | — | 95% | 1480 | 53 | |
| 94% | — | — | — | — | — | 88.8% | — | 1493 | 48 | |
| — | — | — | — | — | 38% | — | — | — | 46 | |
| — | — | — | — | — | — | — | — | 1483 | 45 | |
| 92.6% | — | — | 67.7% | — | — | 86.6% | — | 1481 | 45 | |
| 93.5% | 43.5% | 56% | — | — | — | 88.3% | — | 1485 | 44 | |
| — | 45.4% | — | 61.6% | — | — | 90.8% | — | 1493 | 41 | |
| — | 43.2% | 57.4% | 63.2% | — | — | 80.4% | — | — | 38 | |
| 92.3% | — | — | 62.7% | — | — | 84.7% | — | — | 37 | |
| 90.1% | — | 48.2% | 55.4% | 80.6% | — | — | — | — | 36 | |
| 94.3% | 44.4% | 51.4% | 54.2% | 80.6% | — | — | 77.1% | 1487 | 30 | |
| — | — | — | — | 77.6% | — | — | — | — | 14 | |
| — | — | — | — | — | — | — | — | — | 13 | |
| — | — | 67.7% | — | — | 66.4% | — | — | — | — |
Sources
GPT-6 Astra
GPT-5.6 Luna
Claude Fable 5.1
Claude Sonnet 5
Gemini 3.8 Flash
Gemini 3.1 Pro (Preview)
Muse Spark 1.3
DeepSeek V4 Pro
Qwen3.8 Max
Kimi K3
GLM-5.3
Mistral Medium 3.5